Lila Sciences Aims to Build AI Science Factories
Lila Sciences CTO Andrew Beam and co-founder Rafa Gómez-Bombarelli discuss their vision for "AI Science Factories" that leverage experiments as a data source for scaling AI in science.
7 min read

Visual TL;DR
traditional data sources (internet) are finite, like 'fossil fuel' already 'fracked'
From the article 4 mentionsIn the rapidly evolving AI landscape, the race for more capable models often centers on scaling compute, parameters, and data.
Andrew Beam and Rafa Gómez-Bombarelli propose a new approach for AI in science
From the article 6 mentionsLila Sciences' vision extends beyond biotech, encompassing chemistry and materials science, with the ultimate aim of creating a core reasoning LLM-based model.
leveraging scientific experiments as an infinite token generator for AI models
From the article 9+ mentionsThey believe that the next frontier in AI lies not just in processing existing data, but in generating new, high-quality data through scientific experimentation, effectively creating "AI Science Factories.”
moving beyond processing existing data to actively creating high-quality experimental data
From the article 9+ mentionsThis iterative process, where the model learns from its experimental outcomes, is seen as a way to generate truly valuable, incremental data.
envisioning a 'PCI bus' for labs, integrating AI with human scientific workflows
From the articleThis human-AI collaboration is crucial for handling tasks that are difficult to automate, such as removing a cap from a test tube, highlighting a pragmatic approach to automation.
science itself can continuously produce new, valuable data for training AI models
From the articleOn the Latent Space Science podcast, Anya Beam, CTO of Lila Sciences, and Rafa Gómez-Bombarelli, co-founder and Chief Scientific Officer for Physical Sciences, outlined their ambitious thesis: science itself can serve as an infinite token generator for training AI models at scale.
enabling more capable models by continuously feeding them novel, experimental data
From the article 9+ mentionsBeam described it as "rows of server racks, as densely packed as possible, and also as energy efficient as possible." This integrated vision aims to generate diverse data across modalities that can be validated in the lab, effectively adding a new scaling axis for data generation.
Contents(7)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.