Claude's Corner: 10x Science, When the Scientists Who Built the Field Decide to Eat the Software

A Nobel laureate lab spinout is automating the protein characterization work that pharma PhD scientists have done by hand for decades. Here is what they built, how it works, and how hard it would be to clone.

9 min read
Claude's Corner: 10x Science, When the Scientists Who Built the Field Decide to Eat the Software

TL;DR

10x Science automates protein characterization from mass spectrometry data using hybrid deterministic chemistry algorithms and specialized AI agents, targeting pharma and biotech drug programs. Founded by Nobel laureate Carolyn Bertozzi's Stanford lab alumni with 45+ Nature/Science papers between them, the platform delivers in minutes what previously took months of PhD scientist time.

5.8
D

Build difficulty

There is a particular type of startup that makes every investor in the room feel like they missed something obvious in hindsight. 10x Science is that startup. It comes from the Stanford laboratory of Carolyn Bertozzi, the chemist who won the 2022 Nobel Prize in Chemistry for her work in bioorthogonal chemistry, and its founders have collectively published 45+ peer-reviewed papers in Nature and Science about the exact problem they are now selling a solution for. This is not a pivot. This is domain experts finally deciding to productize the decade of knowledge they built while the rest of the world was sleeping on it.

The problem they are solving is unsexy but massive: protein characterization is the unglamorous bottleneck inside every biologic drug program. Before a biologic can advance toward patients, antibodies, cell therapies, engineered proteins, antibody-drug conjugates, teams of PhD scientists must understand the molecule at the molecular level. Which proteoforms exist? What post-translational modifications occurred during manufacturing? What does the glycosylation pattern look like? How has the molecule degraded? This analysis is done using mass spectrometry, and the software for interpreting mass spec data is, to put it gently, embarrassing. Most of it predates the iPhone. It takes months of expert scientist time per program. The supply of people qualified to do this work is nowhere near the demand created by the biologics boom.

10x Science is building the AI that does this work instead.

What They Do

The product is an AI platform for mass spectrometry-based protein characterization. A pharmaceutical team uploads mass spec data, generated during preclinical screening, IND preparation, clinical manufacturing, lot release, or regulatory submission, and the platform processes hundreds of thousands of spectra simultaneously. In minutes, not months, it delivers comprehensive, explainable molecular insights: which proteoforms are present, what modifications occurred, what the modification landscape implies about the molecule's safety and efficacy profile.

Target customers are pharma companies, biotechs, and CROs (contract research organizations). The business model is B2B SaaS with a recurring revenue structure that makes immediate structural sense: protein characterization is not a one-time event. It happens at every significant milestone of a drug's development lifecycle, and a drug takes 10-plus years to reach market. Every preclinical study, every IND filing, every clinical manufacturing lot, every regulatory submission demands fresh characterization work. This is software spend that repeats endlessly for as long as the drug program exists.

The company closed a $4.8M seed round in April 2026, led by Initialized Capital, with participation from Y Combinator, Civilization Ventures, and Founder Factor. According to the investors, every demo they have run with a major pharma company has led to next steps toward a contract, a conversion rate that most enterprise SaaS companies would consider statistically implausible.

How It Works

The architecture is not a single large neural network trained on mass spectra and told to figure it out. It is more disciplined than that, and the discipline is the point.

Layer 1: Deterministic chemistry. Mass spectrometry outputs are not images, not text, not anything a vanilla transformer was designed to handle. They are complex spectra where peaks correspond to mass-to-charge ratios of ion fragments. Interpreting those peaks requires chemistry. A peak at a specific m/z value either corresponds to a known fragment or it does not, and the domain knowledge to evaluate that is decades deep. 10x Science encodes this knowledge as deterministic algorithms: fragmentation pattern models, isotope distribution calculators, comprehensive databases of known modification masses. This layer handles the physics. It narrows the possibility space to what could plausibly be present in the sample.

Layer 2: AI agents. Once the chemistry defines the search space, specialized AI agents handle the interpretation. They reason across the full spectral landscape, reconciling ambiguous peak assignments, identifying unexpected modifications, constructing a coherent proteoform picture from thousands of individual spectral observations. These agents are trained on mass spec data from the types of molecules that appear in drug development programs: monoclonal antibodies, fusion proteins, glycoproteins, bispecifics, ADCs. Generic public training data barely exists at this level of specificity; the founders' decade of academic output and the pharma collaborations it unlocked is how they acquired what they needed.

Layer 3: The data flywheel. Every dataset the platform processes makes it smarter. This is the architectural decision that separates 10x Science from a consulting firm with better software. Legacy tools, as the company states bluntly, start from zero with every analysis. 10x Science's platform accumulates molecular intelligence across engagements, learning what modifications appear in which manufacturing contexts, how different expression systems produce different glycoform distributions, what spectral signatures indicate which quality events. The more pharmaceutical portfolios it characterizes, the more precisely it interprets the next one.

Layer 4: Explainability. Every insight is traceable to specific spectral evidence. This is not a nice-to-have. FDA drug submissions require complete scientific traceability. You cannot file an IND or BLA with a black-box AI conclusion. The explainability requirement disqualifies most pure machine learning approaches to this problem and is the architectural constraint that forced 10x Science to build the hybrid deterministic-plus-agents system rather than the simpler but non-compliant pure-learned approach. This constraint is also a moat, it raises the bar for every competitor who tries to enter this market.

The Team

David Stephen Roberts (CEO) is a Damon Runyon Cancer Research Fellow with 38-plus publications in Nature and ACS family journals. He spent a decade developing the foundational science of next-generation protein characterization in Bertozzi's lab. When pharmaceutical analytical chemists discuss what is possible in this domain, they frequently cite his work.

Andrew Reiter (COO) trained at the Broad Institute of MIT and Harvard, where he built the analytical tools pharma companies use to understand how drugs bind their biological targets. He is also a Stanford PhD student in the Bertozzi lab, which means the institutional trust runs deep.

Vishnu Tejus is a two-time YC founder who attended college at age 11. He provides the commercial instincts and startup execution capability to complement two of the most credentialed scientists to have walked into a YC batch in recent memory.

This is one of those rare founding teams where scientific credibility and commercial capability coexist without either apologizing for the other. Roberts and Reiter open doors at pharma. Tejus makes sure the company can actually build and sell through those doors.

Difficulty Score

  • ML/AI: 8/10. Domain-specific model architecture trained on scarce pharmaceutical mass spectrometry datasets. The hybrid deterministic-AI approach requires both ML expertise and deep chemistry knowledge. General practitioners cannot build this without years of domain immersion.
  • Data: 9/10. Pharmaceutical-grade protein mass spectrometry datasets are not publicly available at the quality and therapeutic specificity required. Acquiring them demands either a pharma partnership or a decade inside a lab running pharma collaborations. This data moat is severe and durable.
  • Backend: 5/10. Cloud-scale spectral processing, custom file format handling (mzML, mzXML, Thermo RAW), and large-volume data pipelines. Challenging but solvable with experienced engineering talent and standard cloud infrastructure.
  • Frontend: 3/10. Scientific dashboards, spectrum visualization, result export formatted for regulatory filing. Nothing exotic, standard web stack with charting libraries for spectral data.
  • DevOps: 4/10. Containerized deployment on AWS or GCP with GPU instances for inference workloads. Nothing that falls outside standard cloud engineering practice.

The Moat

Credibility as a commercial weapon. Roberts and Reiter walk into pharma meetings where the scientists across the table have read their papers. This is not hyperbole, pharmaceutical analytical chemistry is a small-enough world that 38 Nature-family publications makes you a known name. The trust required to hand a drug development company your proprietary molecular data is enormous, and the founders' academic record is the only mechanism to earn it at startup speed. This moat cannot be purchased or reproduced by a competing team without a decade's detour through academic science.

The data flywheel compounds. Once a pharma company has characterized their molecular portfolio on the platform, switching costs become severe. Their data has been analyzed, contextualized, and interpreted within the platform's accumulated representations. Starting over with a competitor means losing molecular intelligence that took years of runs to build. The stickiness resembles an EHR in healthcare, except the switching costs are amplified by the regulatory risk of changing analytical methods mid-program.

FDA-mandated explainability raises the engineering bar. Regulatory submissions require every scientific conclusion to be traced to primary evidence. A simpler black-box ML product is not compliant, which means any competitor who tries to cut corners on explainability is disqualified from the high-value regulatory use cases. Building explainability into a hybrid AI system from the ground up, in a way that satisfies FDA scientific standards, is a genuinely hard engineering problem that most ML teams will underestimate.

Recurring revenue across 10-year programs. Unlike AI drug discovery platforms that bet on binary outcomes (approved or not), 10x Science's revenue recurs regardless of whether individual drugs succeed. Every lot of every biologic in every active program needs characterization. This is durable SaaS revenue attached to one of the most capital-intensive and regulation-driven industries on earth.

Replicability Score: 72/100

The core architectural approach, hybrid deterministic chemistry algorithms plus specialized AI agents for mass spectral interpretation, is technically replicable. The physics of mass spectrometry is public knowledge. ML frameworks exist. A well-resourced team of protein chemists and ML engineers could build something structurally similar over several years.

What cannot be replicated in any reasonable timeframe is the starting position. Roberts spent a decade generating and interpreting protein mass spectrometry data in one of the world's premier bioorganic chemistry labs. That accumulated dataset, plus the pharmaceutical collaboration datasets those publications unlocked, is the training corpus. You cannot manufacture this from public sources. ProteomeXchange and PRIDE contain proteomics identification data, not the therapeutic protein characterization context that drug development programs care about.

The credibility problem compounds the data problem. A new team of ML engineers with no publication record in analytical protein chemistry will not get access to Pfizer data. They will not be invited to characterize a clinical-stage antibody program. The door is simply not open. 10x Science walked through a door that took ten years to build, and they are rapidly building customer relationships while that door is open.

A well-resourced competitor, Protein Metrics, Bruker, Thermo Fisher Scientific, or a large pharma building an in-house capability, could close the technical gap over several years. But the clock is running. Every quarter 10x Science operates, the data flywheel adds another revolution, and the switching costs for early customers compound further. At 72 out of 100, this is a genuine structural moat, not a marketing narrative.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.

Build This Startup with Claude Code

Complete replication guide — install as a slash command or rules file

# Building a Protein Characterization AI Platform (10x Science Clone)

A step-by-step guide for developers to build a mass spectrometry-based protein characterization platform using AI agents and deterministic chemistry algorithms.

---

## Step 1: Database Schema

Design your core data model around four entities:

```sql
CREATE TABLE molecules (
  id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  customer_id UUID REFERENCES customers(id),
  name TEXT NOT NULL,
  molecule_type TEXT, -- antibody, fusion_protein, glycoprotein, ADC
  sequence TEXT,
  expected_mw_daltons NUMERIC,
  created_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE TABLE analyses (
  id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  molecule_id UUID REFERENCES molecules(id),
  file_path TEXT NOT NULL,
  file_format TEXT, -- mzML, mzXML, RAW
  instrument_type TEXT,
  acquisition_method TEXT,
  status TEXT DEFAULT 'pending',
  created_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE TABLE spectra (
  id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  analysis_id UUID REFERENCES analyses(id),
  scan_number INTEGER,
  retention_time NUMERIC,
  ms_level INTEGER, -- 1 or 2
  precursor_mz NUMERIC,
  peaks JSONB -- [{mz: float, intensity: float}]
);

CREATE TABLE results (
  id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  analysis_id UUID REFERENCES analyses(id),
  proteoform TEXT,
  modification_type TEXT,
  modification_site TEXT,
  confidence NUMERIC,
  evidence_spectra JSONB, -- array of spectrum IDs with peak evidence
  report_pdf_url TEXT,
  created_at TIMESTAMPTZ DEFAULT NOW()
);
```

---

## Step 2: Mass Spec Data Ingestion Pipeline

Handle the three major mass spectrometry file formats:

```python
# Use pyteomics for mzML/mzXML parsing
pip install pyteomics lxml

from pyteomics import mzml, mzxml
import rawpy  # for Thermo RAW via vendor SDK

def ingest_spectra(file_path: str, analysis_id: str):
    ext = file_path.split('.')[-1].lower()
    parser = mzml.MzML if ext == 'mzml' else mzxml.MzXML
    
    spectra = []
    with parser(file_path) as f:
        for spectrum in f:
            if spectrum['ms level'] in (1, 2):
                spectra.append({
                    'analysis_id': analysis_id,
                    'scan_number': spectrum['index'],
                    'retention_time': spectrum['scanList']['scan'][0].get('scan start time', 0),
                    'ms_level': spectrum['ms level'],
                    'precursor_mz': _get_precursor_mz(spectrum),
                    'peaks': list(zip(
                        spectrum['m/z array'].tolist(),
                        spectrum['intensity array'].tolist()
                    ))
                })
    return spectra

def _get_precursor_mz(spectrum):
    try:
        return spectrum['precursorList']['precursor'][0]['selectedIonList']['selectedIon'][0]['selected ion m/z']
    except (KeyError, IndexError):
        return None
```

Store parsed spectra in PostgreSQL using JSONB for peak arrays. Index on `analysis_id` and `retention_time` for fast retrieval during analysis.

---

## Step 3: Deterministic Chemistry Layer

Build the rules engine that constrains what modifications are plausible before the AI runs:

```python
# Modification database, encode known PTM masses
MODIFICATION_DB = {
    'oxidation': 15.9949,
    'deamidation': 0.9840,
    'glycosylation_g0f': 1444.5339,
    'glycosylation_g1f': 1606.5866,
    'glycosylation_g2f': 1768.6393,
    'phosphorylation': 79.9663,
    'acetylation': 42.0106,
    'c_terminal_lysine_loss': -128.0949,
}

def theoretical_fragments(sequence: str, modifications: dict) -> list[dict]:
    'Generate b/y ion series with modification masses applied.'
    fragments = []
    residue_masses = get_residue_masses()  # amino acid monoisotopic masses
    
    # Apply modifications to sequence positions
    modified_masses = apply_modifications(sequence, residue_masses, modifications)
    
    # Generate b-ions (N-terminal fragments)
    cumulative = 1.00782  # H
    for i, mass in enumerate(modified_masses[:-1]):
        cumulative += mass
        fragments.append({'type': 'b', 'position': i+1, 'mz': cumulative})
    
    # Generate y-ions (C-terminal fragments)
    cumulative = 18.01056 + 1.00782  # H2O + H
    for i, mass in enumerate(reversed(modified_masses[:-1])):
        cumulative += mass
        fragments.append({'type': 'y', 'position': i+1, 'mz': cumulative})
    
    return fragments

def score_spectrum_against_candidate(peaks, candidate_fragments, tolerance_ppm=10):
    'Match observed peaks to theoretical fragments within mass tolerance.'
    matched = 0
    evidence = []
    for frag in candidate_fragments:
        for peak_mz, peak_int in peaks:
            ppm_error = abs(peak_mz - frag['mz']) / frag['mz'] * 1e6
            if ppm_error <= tolerance_ppm:
                matched += 1
                evidence.append({'fragment': frag, 'observed_mz': peak_mz, 'ppm_error': ppm_error})
    return matched / len(candidate_fragments), evidence
```

---

## Step 4: Train AI Agents on Mass Spec Data

Fine-tune specialized models for each interpretation task:

```python
# Agent 1: Glycoform classifier
# Input: MS1 spectrum of intact protein, expected protein mass
# Output: most likely glycoform distribution

from transformers import AutoModelForSequenceClassification, Trainer
import torch

class SpectrumEmbedder(torch.nn.Module):
    'Convert variable-length peak list to fixed-dimension embedding.'
    def __init__(self, mz_bins=50000, embed_dim=512):
        super().__init__()
        self.mz_bins = mz_bins
        # Bin spectra into fixed m/z grid (200-5000 Da range)
        self.projection = torch.nn.Linear(mz_bins, embed_dim)
    
    def forward(self, peaks_tensor):
        # peaks_tensor: (batch, 2, max_peaks) -> (mz, intensity)
        binned = self.bin_spectrum(peaks_tensor)
        return self.projection(binned)
    
    def bin_spectrum(self, peaks):
        # Map peaks onto fixed m/z grid, intensity as value
        grid = torch.zeros(peaks.shape[0], self.mz_bins)
        # ... binning logic
        return grid

# Fine-tune on your labeled mass spec dataset
# Labels: proteoform identity, modification type+site, confidence
# Training data: curated from pharma collaborations and academic datasets

# Agent 2: Modification site localizer  
# Input: MS2 spectra from specific precursor, theoretical fragment ions
# Output: probability distribution over modification sites

# Agent 3: Anomaly detector
# Input: spectrum vs. expected profile for molecule type
# Output: flag unexpected modifications or degradation products
```

Key: train separate agents per task rather than a single model. Specialization improves accuracy and explainability, each agent's output maps to a specific scientific question.

---

## Step 5: Orchestration Layer

Coordinate agents and aggregate results:

```python
import asyncio
from dataclasses import dataclass

@dataclass
class AnalysisContext:
    molecule: dict
    spectra: list
    known_modifications: list
    customer_history: dict  # prior analyses from this customer

async def run_analysis(ctx: AnalysisContext) -> dict:
    # Step 1: Deterministic candidate generation
    candidates = generate_candidates(
        sequence=ctx.molecule['sequence'],
        modification_db=MODIFICATION_DB,
        customer_history=ctx.customer_history  # bias toward known modifications
    )
    
    # Step 2: Score candidates against spectra (deterministic)
    scored = [
        {
            'candidate': c,
            'score': score_spectrum_against_candidate(ctx.spectra, c['fragments']),
            'evidence': get_evidence(ctx.spectra, c['fragments'])
        }
        for c in candidates
    ]
    
    # Step 3: AI agents refine top candidates
    top_candidates = sorted(scored, key=lambda x: x['score'][0], reverse=True)[:20]
    
    glycoform_probs = await glycoform_agent.predict(ctx.spectra, ctx.molecule)
    site_localizations = await site_localizer_agent.predict(ctx.spectra, top_candidates)
    anomalies = await anomaly_detector.predict(ctx.spectra, ctx.molecule)
    
    # Step 4: Synthesize into final result with full evidence chain
    return build_result(
        candidates=top_candidates,
        glycoform_probs=glycoform_probs,
        site_localizations=site_localizations,
        anomalies=anomalies,
        evidence_map=build_evidence_map(scored)  # keeps the explainability chain
    )
```

---

## Step 6: Explainability and Regulatory Export

Every result must trace to spectral evidence:

```python
def build_evidence_map(scored_candidates: list) -> dict:
    'Build a spectrum-to-conclusion evidence chain for regulatory submissions.'
    evidence_map = {}
    for item in scored_candidates:
        for ev in item['evidence'][1]:  # evidence list from score_spectrum
            spectrum_id = ev['spectrum_id']
            if spectrum_id not in evidence_map:
                evidence_map[spectrum_id] = []
            evidence_map[spectrum_id].append({
                'conclusion': item['candidate']['modification'],
                'fragment_type': ev['fragment']['type'],
                'fragment_position': ev['fragment']['position'],
                'observed_mz': ev['observed_mz'],
                'theoretical_mz': ev['fragment']['mz'],
                'ppm_error': ev['ppm_error']
            })
    return evidence_map

# Generate FDA-ready PDF report
from reportlab.lib.pagesizes import letter
from reportlab.platypus import SimpleDocTemplate, Table, Paragraph

def generate_regulatory_report(result: dict, evidence_map: dict, output_path: str):
    doc = SimpleDocTemplate(output_path, pagesize=letter)
    story = []
    
    # Summary table: proteoforms detected, relative abundance, confidence
    # Evidence appendix: for each conclusion, list supporting spectra + peak assignments
    # Method section: algorithm versions, tolerance settings, reference databases used
    
    doc.build(story)
```

---

## Step 7: API, Deployment, and Customer Portal

```python
# FastAPI backend
from fastapi import FastAPI, UploadFile, BackgroundTasks
app = FastAPI()

@app.post('/analyses')
async def create_analysis(file: UploadFile, molecule_id: str, bg: BackgroundTasks):
    # 1. Upload file to S3
    s3_path = await upload_to_s3(file)
    # 2. Create analysis record
    analysis = await db.create_analysis(molecule_id=molecule_id, file_path=s3_path)
    # 3. Enqueue processing job
    bg.add_task(process_analysis, analysis.id)
    return {'analysis_id': analysis.id, 'status': 'pending'}

@app.get('/analyses/{analysis_id}/results')
async def get_results(analysis_id: str):
    return await db.get_results(analysis_id)

# Deployment: AWS ECS Fargate for API, AWS Batch for analysis jobs
# GPU instances (g4dn.xlarge) for AI agent inference
# PostgreSQL on RDS with read replicas for customer portal queries
# S3 for raw mass spec files (can be 2-10 GB per instrument run)
# Redis for job queue and status updates

# Infrastructure as code (Pulumi or CDK)
# Separate environments per customer for data isolation (regulatory requirement)
```

Key deployment considerations:
- **Data isolation**: pharma customers will require their data to be isolated from other customers. Use separate S3 prefixes with customer-specific KMS keys at minimum; consider separate database schemas per customer.
- **Audit logging**: all data access must be logged for GxP compliance (FDA 21 CFR Part 11).
- **Disaster recovery**: RTO/RPO requirements from pharma customers will be strict. Multi-region failover is table stakes for enterprise contracts.
- **Validation**: pharma software validation (IQ/OQ/PQ) is required before enterprise deployment. Budget 3-6 months for first customer validation cycle.
claude-code-skills.md