Kenny Workman on Building Verifiable AI Evals for Biology

LatchBio CTO Kenny Workman explains how verifiable evaluation frameworks and multi-omics benchmarks drive AI agent progress in biological research.

Kenny Workman presenting LatchBio AI benchmarks for biology at AI Engineer World's Fair
Kenny Workman presenting LatchBio's evaluation frameworks for biological AI agents.· AI Engineer
Visual TL;DR
Biological Data ExplosionDriver
modern techniques generate massive, complex datasets at an unprecedented rate
From the article 3 mentionsBiological research produces vast volumes of complex data every day.
Kenny Workman / LatchBioCore
co-founder and CTO building data infrastructure for biotech and pharma
From the article 2 mentionsAt the AI Engineer World's Fair, Kenny Workman, co-founder and CTO of LatchBio, detailed how his team builds verifiable evaluation environments to train and test AI agents in life sciences.
Verifiable AI EvalsCore
frameworks to train and test AI agents in life sciences research
From the article 3 mentions"Just like code provided a verifiable substrate for complex software tasks that are not inherently verifiable, data analysis might do the same thing in bio," Workman noted.
SpatialBenchContext
a specific benchmarking framework for measuring agent performance on tasks
From the articleTo solve this, LatchBio built SpatialBench, an evaluation suite containing 146 verifiable problems derived from real spatial biology workflows.
Multi-Omics BenchmarksContext
long-horizon benchmarks for complex tasks across genomics and drug discovery
From the article 5 mentionsBy breaking scientific research down into data processing steps, developers can systematically benchmark model performance.
Improved AI ReasoningEffect
From the article 2 mentionsWorkman explained how treating biological data analysis like code execution creates a natural path to improve AI reasoning across genomics and drug discovery.
Advance Biological ResearchOutcome
driving AI agent progress in biological research and drug discovery
From the article 2 mentionsOver five years, LatchBio evolved from an enterprise data management provider into a specialized research lab for biological AI agents.
Address BiosecurityContext
considering model refusals and biosecurity implications in AI development
From the article 2 mentionsAs AI capabilities expand, assessing safety and biosecurity risks becomes critical.
Contents(6)

Biological research produces vast volumes of complex data every day. At the AI Engineer World's Fair, Kenny Workman, co-founder and CTO of LatchBio, detailed how his team builds verifiable evaluation environments to train and test AI agents in life sciences. Workman explained how treating biological data analysis like code execution creates a natural path to improve AI reasoning across genomics and drug discovery.

Who Is Kenny Workman

Kenny Workman co-founded LatchBio out of UC Berkeley to build data infrastructure for biotechnology and pharmaceutical companies. Over five years, LatchBio evolved from an enterprise data management provider into a specialized research lab for biological AI agents. The company develops benchmarking frameworks that major AI research teams now use to measure agent performance on complex scientific tasks.

The Biological Data Explosion

Modern biological techniques generate massive datasets at an unprecedented rate. Single-cell experiments yield between two and six terabytes per run. Spatial biology runs can generate seven terabytes of raw image data. These numbers routinely exceed what individual researchers can store on consumer computers.

Workman argued that data analysis serves as an executable foundation for AI agents in biology. "Just like code provided a verifiable substrate for complex software tasks that are not inherently verifiable, data analysis might do the same thing in bio," Workman noted. By breaking scientific research down into data processing steps, developers can systematically benchmark model performance.

Building SpatialBench and Verifiable Evals

When LatchBio began testing coding models on biological analysis tasks, frontier models frequently struggled. They lacked the ability to combine programming, data analysis, and domain reasoning. Existing benchmarks mostly tested static question answering rather than real experimental workflows.

To solve this, LatchBio built SpatialBench, an evaluation suite containing 146 verifiable problems derived from real spatial biology workflows. The team used deterministic Python functions as graders. Through human verification, LatchBio discovered that many initial task prompts contained ambiguity or relied on arbitrary quality control thresholds. Refining these tasks ensured that evaluation results remained durable across valid alternative analysis paths.

Long Horizon Benchmarks and Multi-Omics

LatchBio expanded its benchmark suite beyond short data processing steps. The team built SpatialBench-Long to simulate multi-step workflows that mirror entire paper results or commercial drug program decisions. None of the evaluated models solved these extended tasks initially, highlighting clear targets for future post-training.

LatchBio extended this methodology across other biological domains. The lab released benchmarks for single-cell biology, epigenomics, and preclinical pharmacology for small molecules. These tools evaluate how AI agents interpret complex experimental designs and prior scientific literature.

Addressing Biosecurity and Model Refusals

As AI capabilities expand, assessing safety and biosecurity risks becomes critical. LatchBio formed a dedicated biosecurity team following its acquisition of Twenty Two. In collaboration with American Wetware and Aclid, LatchBio created BioSecBench-Refusal to evaluate model refusal behavior.

The benchmark tests models against both routine scientific questions and red-team queries designed to obscure dangerous requests. The research revealed that models refuse harmless routine tasks far more often than red-team tasks, signaling a need for better evaluation standards in scientific safety filters.

StartupHub.ai Market Context

LatchBio continues to bridge frontier AI development with practical life sciences tools. StartupHub.ai data shows LatchBio holds a score of 71/100, with verified financials confirming $80M raised in a 2023 Series A round. LatchBio tracks closely alongside peers evaluated in the sector, including OpenAI at 84/100, Alphabet Inc. (NASDAQ:GOOGL) at 73/100, Perplexity AI at 71/100, Lucidworks at 51/100, and matey at 50/100, according to StartupHub.ai data.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.