AI Data is the New Bottleneck, Not Models

AI experts at YC Data Club reveal data quality, not model architecture, is the key bottleneck in AI development, emphasizing the need for expert supervision and innovative data strategies.

AI experts discuss data as the bottleneck in AI development at a YC Data Club session.
YC
Visual TL;DR
YC Data Club InsightsContext
industry leaders highlighted this shift in AI landscape at a recent YC Data Club session
From the article 2 mentionsThe overarching message from the YC Data Club session was clear: as AI capabilities advance, the ability to source, curate, and effectively utilize high-quality data is paramount.
AI Development BottleneckDriver
data quality, not model architecture, is the key bottleneck in AI development
From the article 4 mentionsIn a recent YC Data Club session, industry leaders highlighted a significant shift in the AI landscape: data, not models, has become the primary bottleneck for progress.
Data Quality ImperativeContext
meticulously curated datasets and sophisticated evaluation environments are critical for effective AI
From the articleFrancois Chaudard, a PhD student and visiting partner at YC, set the stage by explaining how the focus in AI development has moved from architectural innovation to data quality.
Data-Centric ShiftOutcome
From the article 4 mentionsFast forward to today, and the market capitalization created by data-centric businesses has ballooned into the hundreds of billions.
Expert Supervision NeededDriver
emphasizing the need for expert supervision and innovative data strategies for AI progress
From the article 2 mentionsThe focus on data-centric AI, expert supervision, and the development of sophisticated evaluation benchmarks will be critical for the continued progress and responsible deployment of AI systems across various domains.
Data is DifferentiatorCore
the data itself is the key differentiator, not just a commodity, for AI systems
From the article 9+ mentionsThis dramatic shift in perspective, the speakers emphasized, is driven by a fundamental realization: the data itself is the key differentiator.
Effective AI SystemsEffect
building truly effective AI systems requires focus on data quality and evaluation
From the article 5 mentionsThe conversation, featuring experts from leading AI research labs and startups, underscored the critical role of meticulously curated datasets and sophisticated evaluation environments in building truly effective AI systems.
Contents(4)

In a recent YC Data Club session, industry leaders highlighted a significant shift in the AI landscape: data, not models, has become the primary bottleneck for progress. The conversation, featuring experts from leading AI research labs and startups, underscored the critical role of meticulously curated datasets and sophisticated evaluation environments in building truly effective AI systems.

The Data Imperative in AI

The session opened with a provocative statement: in 2016, the prevailing notion was that data was a commodity, and companies like Scale AI were viewed with skepticism by many VCs regarding their terminal value. Fast forward to today, and the market capitalization created by data-centric businesses has ballooned into the hundreds of billions. This dramatic shift in perspective, the speakers emphasized, is driven by a fundamental realization: the data itself is the key differentiator.

The full discussion can be found on YC's YouTube channel.

Going In Deep On Data | YC Paper Club - YC
Going In Deep On Data | YC Paper Club, from YC

Francois Chaudard, a PhD student and visiting partner at YC, set the stage by explaining how the focus in AI development has moved from architectural innovation to data quality. He cited a common interview question for deep learning roles: after achieving an 85% F1 score, what's next? The incorrect answer, he noted, is to tinker with architecture or hyperparameters. The correct answer, he stressed, is to "look at the data." This involves meticulously classifying false positives and negatives, identifying Pareto buckets of issues, and understanding the root causes, often stemming from the data itself, such as fogged-up refrigerators or occlusions in image recognition tasks.

From PhD to Production: The Data Shift

The transition from academic research to production environments marks a significant change in the allocation of effort. While PhDs might spend 95% of their time on architecture, companies in production flip this, dedicating 95% of their effort to data. This is particularly true now that architectures like transformers are highly effective. Chaudard illustrated this with a pie chart showing the stark difference in focus: PhDs spend 5% on data and 95% on models, while Tesla, for instance, spends 75% on data and 25% on models in production.

The challenge lies in the inherent difference between clean, curated training distributions (like ImageNet) and the messy, unpredictable "test distribution" encountered in the real world. When models encounter data points outside their training distribution, they often fail. To automate the economy, Chaudard argued, we need "both expert data and expert RL environments."

Scaling Expertise: The Snorkel Approach

Vincent Sunn Chen, VP and Founding Team Member at Snorkel.ai, elaborated on the concept of "scaling expert supervision." Snorkel’s core thesis is that leveraging the judgment and knowledge of domain experts, doctors, clinicians, journalists, is the real bottleneck in building effective datasets. Their research focuses on turning this expertise into software, allowing for programmatic labeling, adaptability to changing specifications, and auditability.

Chen distinguished between "Data 1.0" (basic labeling, taking about 30 seconds per data point) and "Data 2.0" (expert agentic environments, requiring 3-30+ hours per data point). He highlighted the limitations of manual labeling: it's not scalable, not robust to noise (requiring redundancy), lacks adaptability to schema changes, and offers no provenance for the labeling decisions. These challenges are amplified in expert domains, making manual labeling practically intractable.

Snorkel's approach, termed "data programming," encodes expert knowledge into labeling functions. These functions, though potentially inaccurate or overlapping, can be combined and denoised to create high-quality datasets. Chen explained the label model's intuition: estimating ground truth from noisy, overlapping signals without any ground truth data, akin to grading student papers without an answer key.

Benchmarks and the Future of Data

Shayne Longpre, an MIT PhD student who recently joined Anthropic, discussed the importance of benchmarks, particularly in the context of multilingual models. He pointed out that most scaling law research is English-centric, neglecting the rest of the world. The scarcity of data for non-English languages presents a significant challenge, and Longpre’s work focuses on measuring cross-lingual transfer synergies and interferences.

Volodymyr Kuleshov, co-founder of Inception Labs and Cornell professor, presented on diffusion language models, highlighting their breakthrough speed (over 1000 tokens per second) and efficiency. He emphasized that while algorithms are crucial, high-quality data is equally important. Inception Labs’ "Data Forge" system synthesizes realistic RL environments based on real-world data to train and evaluate these models, addressing the limitations of existing benchmarks like Towbench, which he argues are often over-benchmarked and don't capture real-world complexity.

The overarching message from the YC Data Club session was clear: as AI capabilities advance, the ability to source, curate, and effectively utilize high-quality data is paramount. The focus on data-centric AI, expert supervision, and the development of sophisticated evaluation benchmarks will be critical for the continued progress and responsible deployment of AI systems across various domains.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.