Kenny Workman on Building Verifiable AI Evals for Biology

LatchBio CTO Kenny Workman explains how verifiable evaluation frameworks and multi-omics benchmarks drive AI agent progress in biological research.

8 min read
Kenny Workman presenting LatchBio AI benchmarks for biology at AI Engineer World's Fair
Kenny Workman presenting LatchBio's evaluation frameworks for biological AI agents.· AI Engineer

Visual TL;DR. Biological Data Explosion drives need Kenny Workman / LatchBio. Kenny Workman / LatchBio develops Verifiable AI Evals. Verifiable AI Evals includes SpatialBench. Verifiable AI Evals extends to Multi-Omics Benchmarks. Verifiable AI Evals enables Improved AI Reasoning. Improved AI Reasoning leads to Advance Biological Research. Advance Biological Research considers Address Biosecurity.

  1. Biological Data Explosion: modern techniques generate massive, complex datasets at an unprecedented rate
  2. Kenny Workman / LatchBio: co-founder and CTO building data infrastructure for biotech and pharma
  3. Verifiable AI Evals: frameworks to train and test AI agents in life sciences research
  4. SpatialBench: a specific benchmarking framework for measuring agent performance on tasks
  5. Multi-Omics Benchmarks: long-horizon benchmarks for complex tasks across genomics and drug discovery
  6. Improved AI Reasoning: treating biological data analysis like code execution improves AI reasoning
  7. Advance Biological Research: driving AI agent progress in biological research and drug discovery
  8. Address Biosecurity: considering model refusals and biosecurity implications in AI development
Visual TL;DR
Visual TL;DR, startuphub.ai Biological Data Explosion drives need Kenny Workman / LatchBio. Kenny Workman / LatchBio develops Verifiable AI Evals. Verifiable AI Evals enables Improved AI Reasoning. Improved AI Reasoning leads to Advance Biological Research drives need develops enables leads to Biological Data Explosion Kenny Workman / LatchBio Verifiable AI Evals Improved AI Reasoning Advance Biological Research From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Biological Data Explosion drives need Kenny Workman / LatchBio. Kenny Workman / LatchBio develops Verifiable AI Evals. Verifiable AI Evals enables Improved AI Reasoning. Improved AI Reasoning leads to Advance Biological Research drives need develops enables leads to Biological DataExplosion Kenny Workman /LatchBio Verifiable AIEvals Improved AIReasoning AdvanceBiological… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Biological Data Explosion drives need Kenny Workman / LatchBio. Kenny Workman / LatchBio develops Verifiable AI Evals. Verifiable AI Evals enables Improved AI Reasoning. Improved AI Reasoning leads to Advance Biological Research drives need develops enables leads to Biological Data Explosion modern techniques generate massive,complex datasets at an unprecedented rate Kenny Workman / LatchBio co-founder and CTO building datainfrastructure for biotech and pharma Verifiable AI Evals frameworks to train and test AI agents inlife sciences research Improved AI Reasoning treating biological data analysis likecode execution improves AI reasoning Advance Biological Research driving AI agent progress in biologicalresearch and drug discovery From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Biological Data Explosion drives need Kenny Workman / LatchBio. Kenny Workman / LatchBio develops Verifiable AI Evals. Verifiable AI Evals enables Improved AI Reasoning. Improved AI Reasoning leads to Advance Biological Research drives need develops enables leads to Biological DataExplosion modern techniquesgenerate massive,complex datasets at… Kenny Workman /LatchBio co-founder and CTObuilding datainfrastructure for… Verifiable AIEvals frameworks to trainand test AI agentsin life sciences… Improved AIReasoning treating biologicaldata analysis likecode execution… AdvanceBiological… driving AI agentprogress inbiological research… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Biological Data Explosion drives need Kenny Workman / LatchBio. Kenny Workman / LatchBio develops Verifiable AI Evals. Verifiable AI Evals includes SpatialBench. Verifiable AI Evals extends to Multi-Omics Benchmarks. Verifiable AI Evals enables Improved AI Reasoning. Improved AI Reasoning leads to Advance Biological Research. Advance Biological Research considers Address Biosecurity drives need develops includes extends to enables leads to considers Biological Data Explosion modern techniques generate massive,complex datasets at an unprecedented rate Kenny Workman / LatchBio co-founder and CTO building datainfrastructure for biotech and pharma Verifiable AI Evals frameworks to train and test AI agents inlife sciences research SpatialBench a specific benchmarking framework formeasuring agent performance on tasks Multi-Omics Benchmarks long-horizon benchmarks for complex tasksacross genomics and drug discovery Improved AI Reasoning treating biological data analysis likecode execution improves AI reasoning Advance Biological Research driving AI agent progress in biologicalresearch and drug discovery Address Biosecurity considering model refusals and biosecurityimplications in AI development From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Biological Data Explosion drives need Kenny Workman / LatchBio. Kenny Workman / LatchBio develops Verifiable AI Evals. Verifiable AI Evals includes SpatialBench. Verifiable AI Evals extends to Multi-Omics Benchmarks. Verifiable AI Evals enables Improved AI Reasoning. Improved AI Reasoning leads to Advance Biological Research. Advance Biological Research considers Address Biosecurity drives need develops includes extends to enables leads to considers Biological DataExplosion modern techniquesgenerate massive,complex datasets at… Kenny Workman /LatchBio co-founder and CTObuilding datainfrastructure for… Verifiable AIEvals frameworks to trainand test AI agentsin life sciences… SpatialBench a specificbenchmarkingframework for… Multi-OmicsBenchmarks long-horizonbenchmarks forcomplex tasks… Improved AIReasoning treating biologicaldata analysis likecode execution… AdvanceBiological… driving AI agentprogress inbiological research… AddressBiosecurity considering modelrefusals andbiosecurity… From startuphub.ai · The publishers behind this format

Biological research produces vast volumes of complex data every day. At the AI Engineer World's Fair, Kenny Workman, co-founder and CTO of LatchBio, detailed how his team builds verifiable evaluation environments to train and test AI agents in life sciences. Workman explained how treating biological data analysis like code execution creates a natural path to improve AI reasoning across genomics and drug discovery.

Kenny Workman on Building Verifiable AI Evals for Biology - AI Engineer
Kenny Workman on Building Verifiable AI Evals for Biology — from AI Engineer

Who Is Kenny Workman

Kenny Workman co-founded LatchBio out of UC Berkeley to build data infrastructure for biotechnology and pharmaceutical companies. Over five years, LatchBio evolved from an enterprise data management provider into a specialized research lab for biological AI agents. The company develops benchmarking frameworks that major AI research teams now use to measure agent performance on complex scientific tasks.

The Biological Data Explosion

Modern biological techniques generate massive datasets at an unprecedented rate. Single-cell experiments yield between two and six terabytes per run. Spatial biology runs can generate seven terabytes of raw image data. These numbers routinely exceed what individual researchers can store on consumer computers.

Workman argued that data analysis serves as an executable foundation for AI agents in biology. "Just like code provided a verifiable substrate for complex software tasks that are not inherently verifiable, data analysis might do the same thing in bio," Workman noted. By breaking scientific research down into data processing steps, developers can systematically benchmark model performance.

Building SpatialBench and Verifiable Evals

When LatchBio began testing coding models on biological analysis tasks, frontier models frequently struggled. They lacked the ability to combine programming, data analysis, and domain reasoning. Existing benchmarks mostly tested static question answering rather than real experimental workflows.

To solve this, LatchBio built SpatialBench, an evaluation suite containing 146 verifiable problems derived from real spatial biology workflows. The team used deterministic Python functions as graders. Through human verification, LatchBio discovered that many initial task prompts contained ambiguity or relied on arbitrary quality control thresholds. Refining these tasks ensured that evaluation results remained durable across valid alternative analysis paths.

Long Horizon Benchmarks and Multi-Omics

LatchBio expanded its benchmark suite beyond short data processing steps. The team built SpatialBench-Long to simulate multi-step workflows that mirror entire paper results or commercial drug program decisions. None of the evaluated models solved these extended tasks initially, highlighting clear targets for future post-training.

LatchBio extended this methodology across other biological domains. The lab released benchmarks for single-cell biology, epigenomics, and preclinical pharmacology for small molecules. These tools evaluate how AI agents interpret complex experimental designs and prior scientific literature.

Addressing Biosecurity and Model Refusals

As AI capabilities expand, assessing safety and biosecurity risks becomes critical. LatchBio formed a dedicated biosecurity team following its acquisition of Twenty Two. In collaboration with American Wetware and Aclid, LatchBio created BioSecBench-Refusal to evaluate model refusal behavior.

The benchmark tests models against both routine scientific questions and red-team queries designed to obscure dangerous requests. The research revealed that models refuse harmless routine tasks far more often than red-team tasks, signaling a need for better evaluation standards in scientific safety filters.

StartupHub.ai Market Context

LatchBio continues to bridge frontier AI development with practical life sciences tools. StartupHub.ai data shows LatchBio holds a score of 71/100, with verified financials confirming $80M raised in a 2023 Series A round. LatchBio tracks closely alongside peers evaluated in the sector, including OpenAI at 84/100, Alphabet Inc. (NASDAQ:GOOGL) at 73/100, Perplexity AI at 71/100, Lucidworks at 51/100, and matey at 50/100, according to StartupHub.ai data.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.