The AI world has been obsessed with scaling laws: more data, more compute, bigger models, better results. It’s a pattern that has reshaped everything from natural language processing to protein folding. But for the complex, messy world of human biology, specifically, understanding how genes and cells interact under the influence of drugs, that promise has largely remained just that: a promise. Now, Tahoe Therapeutics, a biotech firm formerly known as Vevo Therapeutics, is pulling back the curtain on Tahoe-x1 single cell (Tx1), a 3-billion-parameter foundation model designed to learn "unified representations" of genes, cells, and drugs.
Tx1 is a bold attempt to bring the scaling revolution directly to the heart of cancer research and drug discovery, promising state-of-the-art performance across critical single-cell biology benchmarks.
For years, two formidable barriers have prevented AI from truly unlocking the secrets of systems biology. First, the sheer lack of large, diverse single-cell data. Second, the absence of compute-efficient models capable of handling the astronomical parameter counts needed for meaningful exploration. Tahoe Therapeutics has been systematically dismantling these obstacles.
Their initial salvo, Tahoe-100M, tackled the data problem head-on. It’s the largest perturbation dataset ever assembled, comprising 100 million single cells across 50 cancer models and 1,100 drug perturbations. The dataset has seen nearly 200,000 downloads in just a few months, a testament to its immediate utility and the hunger for such resources in the biological AI community.
Now, with Tahoe-x1 single cell, the focus shifts to the compute challenge. Tx1 is not only the first billion-parameter foundation model trained on this kind of rich, perturbation-driven single-cell data, but it’s also remarkably efficient. Tahoe Therapeutics claims it's 3 to 30 times more compute-efficient than previous cell-state models, pushing the boundaries of what's feasible at this scale. Crucially, it’s fully open-source, with open weights, training, and evaluation code available on Hugging Face and GitHub.
