Virtual Cell wrangler @genbioai | PhD @CMUCompBio | Creator contextualized.ml | context-adaptive models, disease simulators, personalized medicine

San Francisco, CA
Honored to share a major thread of my PhD research, out now in PNAS. We address a core issue with how models are used for scientific discovery. Models are so important that they define the entire scientific process... 1/n
8
45
319
62,168
Caleb Ellington retweeted
Today, Sculpta releases the first demonstration of bioorthogonal barcoding ("bobcoding"), the process by which we: 1) covalently attach oligonucleotide barcodes to RNA at multiple internal positions with two-step chemistry, and 2) perform a multiplexed reverse transcription reaction to generate barcoded cDNA libraries with >99% barcoding accuracy. To the best of our knowledge, it is the first time ever that either of these steps have ever been performed by anyone. No one has ever broken the barrier of multiplexing prior to any enzymatic step in library preparation, driving simplicity and higher data quality. Instead of other methods that add 0 or 1 barcode, we add one barcode every ~300bp, adding redundancy and fidelity to RNA measurements. In our first application described in our pre-print below, we used bobcodes to create the first RNA isoform-resolved drug screening platform. It beats Novartis' DRUG-seq platform by: - generating full length transcript capture of RNA isoforms, instead of just 3' end counting - capturing 25-fold more RNA splicing events - reducing barcode swapping by 10-fold - using 18 fewer PCR cycles (~250,000 less amplification) - reducing sample-to-sample variability in gene expression measurements - eliminating costly and cumbersome library fragmentation/tagmentation steps completely - reducing workflow complexity and number of steps - reducing overall protocol duration by ~25% Though our chemical barcoding method improved transcriptomic data quality while also being simpler and faster than all existing methods, the implications of this work go well beyond this particular use case: The gate to greater applications of AI in transcriptomics is not a lack of compute. Nor is it a lack of data volume. The elephant in the room has always been our limited ability to faithfully and accurately measure cellular RNAs. Sculpta now has line of sight to build what we call the first ground truth transcriptomics platform [for single cell, spatial and more] for the future of biological research, drug development and AI-enabled discoveries for the betterment of human health and longevity. [PS Until the kind folks at bioRxiv get through the apparent backlog of AI slop submissions, Sculpta will host the pre-print PDF on our website. You can sign up for product offerings and other updates at this link too] sculpta.bio/preprint
13
46
203
31,459
Caleb Ellington retweeted
If every step in a scRNA-seq experiment loses just 1% of the signal, you lose almost 40% of the total signal after 50 steps. Noise and poor reproducibility are common problems when working with this kind of data. Yet, as computational biologists, we often treat data collection as a black box and start our work from the count matrix. I wrote a post about what can go wrong before that matrix reaches us, from sample selection and tissue handling all the way to sequencing and read counting. I hope it helps other computational biologists better understand where their data came from. Post: euxhenh.com/blog/scrnaseq-wh…
1
5
301
Can I make Fable fall back to Opus 4.6
1
4
435
I would like to let Fable run overnight but there's like a 3/4 chance it releases the chaos gremlin
1
55
Caleb Ellington retweeted
We are proud to release the world's first virtual cell -- a multi-modal, multi-scale, dynamic, and stateful world model of the cell, the first step toward an AI-driven Digital Organism (AIDO).
Today we're releasing a preview of AIDO Cell, a world model of the human cell. It simulates one cell across DNA, RNA, protein, through regulatory networks and whole-cell behavior. Not four models handing results to each other. One system.
86
445
3,808
408,636
Caleb Ellington retweeted
1/ For over a decade, my lab and others have shown you can predict 3D structure of proteins, it’s interactions, which mutations may cause disease and even how a virus evolves - all from evolutionary sequence data 2/ Whilst powerful - our approaches are still missing what we really need: how natural variation or design ripples through transcription, translation, complex assembly, and the behavior of the cell/tissue organism it's sitting in. Call that mechanism, etiology, pathophysiology, but it’s where we need new models. 3/ GenBio AI (where I am an advisor) is previewing AIDO Cell today: genbio.ai/aido-cell-simulato… one model spanning DNA, RNA, protein, and whole-cell behavior, so a variant's consequences can be traced forward instead of just scored. The technical report here nice is a way to release the progress.
4
31
176
25,892
GenBio was founded with the audacious goal of simulating all of biology. The natural starting point for biological simulation is the cell: the smallest unit of life. Today we're sharing a preview of our first version of that simulator, AIDO Cell.
Today we're releasing a preview of AIDO Cell, a world model of the human cell. It simulates one cell across DNA, RNA, protein, through regulatory networks and whole-cell behavior. Not four models handing results to each other. One system.
2
12
67
10,207
Caleb Ellington retweeted
Can we use human language to predict novel bioactive molecules that work in vivo? In our new paper, we trained a deep learning model on a dataset derived from PubMed text and used it to discover novel antibiotic co-inhibitors that reduced infection burden, even in mouse models!
2
8
26
4,428
Caleb Ellington retweeted
Biology is noisy enough that "predicting the mean" is often a competitive baseline. I put together a practical catalog of trivial baselines for expression prediction, perturbation modeling, GRNs, and more. euxhenh.com/blog/guide-trivi…
2
9
622
Caleb Ellington retweeted
Replying to @chhaviyadav_
Soccer (2 games/wk in my 25-and-over league), gym 2x/wk, yoga 2x/wk, run to work 3-4 directions/wk (8k each way), bike to work 3-4 directions/wk (8k each way), longer run or bike on weekend.
27
43
786
139,996
Caleb Ellington retweeted
Statistical estimation is being absorbed into context. In a new 90-page review, we provide a unified view of statistical and foundation models through the lens of context-adaptive inference! ...
3
3
13
1,757
Caleb Ellington retweeted
[Reviewer Request] I need a few emergency reviewers for MLCB 2026. Please let me know if you have capacity to help over the next day or two. DM me with your contact email! Please RT!
13
29
45
15,583
Caleb Ellington retweeted
Some problems (protein folding) are a good fit for simulation, while others (lit review) are good for LLMs—the toughest problems, however, span both categories. Our latest Rowan post examines one such case: NMR structure elucidation by AI agents. rowansci.com/blog/simulation…
4
9
59
4,403
Statistical estimation is being absorbed into context. In a new 90-page review, we provide a unified view of statistical and foundation models through the lens of context-adaptive inference! ...
3
3
13
1,757
Context-adaptive inference is doing to statistical modeling what LLMs did to software. What this looks like in the data domain outside of natural/formal languages is an open question, but the evolution of estimation might look similar to @karpathy's evolution of software
1
2
68
Joint work with Yue Yao, Jingyun Jia, Baiheng Chen, Dong Liu, Rikhil Rao, @JiaqiWang_, Samuel Wales-McGrath, Yixin Yang, Zhiyuan Li, @ericxing, and @ben_lengerich.
1
88