Sean Cai on the State of AI Data Markets

Sean Cai dissects the AI data market, highlighting fragmentation, the importance of 'type one' data, and the pitfalls of current benchmarks.

8 min read
Sean Cai presenting on the State of Data at AI Engineer World's Fair
AI Engineer

Visual TL;DR. AI Data Market faces Market Fragmentation. Market Fragmentation leads to Specialists Outcompete. Beyond Basic Labeling focus on Type One Data. Type One Data enables True Expertise. Market Fragmentation exacerbates Verifiability Bottleneck. True Expertise drives Future: Custom Data. Verifiability Bottleneck requires Future: Custom Data.

  1. AI Data Market: Sean Cai critiques the current state of AI data markets at AI Engineer World's Fair
  2. Market Fragmentation: shift from vertically integrated giants to a fragmented landscape of specialist providers
  3. Specialists Outcompete: specialists excel in talent sourcing, environment building, rewards, and evaluations
  4. Type One Data: pure capture of real workflows like GitHub commits, crucial for true expertise
  5. Beyond Basic Labeling: focus on basic data labeling is the least interesting part of AI development
  6. True Expertise: data transitions models from generalist competence to true domain expertise
  7. Verifiability Bottleneck: current benchmarks are problematic, hindering reliable data quality assessment
  8. Future: Custom Data: future of data lies in enterprise and highly customized, specialized datasets
Visual TL;DR
Visual TL;DR, startuphub.ai Type One Data enables True Expertise. True Expertise drives Future: Custom Data enables drives Market Fragmentation Type One Data True Expertise Future: Custom Data From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Type One Data enables True Expertise. True Expertise drives Future: Custom Data enables drives MarketFragmentation Type One Data True Expertise Future: CustomData From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Type One Data enables True Expertise. True Expertise drives Future: Custom Data enables drives Market Fragmentation shift from vertically integrated giants toa fragmented landscape of specialistproviders Type One Data pure capture of real workflows like GitHubcommits, crucial for true expertise True Expertise data transitions models from generalistcompetence to true domain expertise Future: Custom Data future of data lies in enterprise andhighly customized, specialized datasets From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Type One Data enables True Expertise. True Expertise drives Future: Custom Data enables drives MarketFragmentation shift fromverticallyintegrated giants… Type One Data pure capture ofreal workflows likeGitHub commits,… True Expertise data transitionsmodels fromgeneralist… Future: CustomData future of data liesin enterprise andhighly customized,… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Data Market faces Market Fragmentation. Market Fragmentation leads to Specialists Outcompete. Beyond Basic Labeling focus on Type One Data. Type One Data enables True Expertise. Market Fragmentation exacerbates Verifiability Bottleneck. True Expertise drives Future: Custom Data. Verifiability Bottleneck requires Future: Custom Data faces leads to focus on enables exacerbates drives requires AI Data Market Sean Cai critiques the current state of AIdata markets at AI Engineer World's Fair Market Fragmentation shift from vertically integrated giants toa fragmented landscape of specialistproviders Specialists Outcompete specialists excel in talent sourcing,environment building, rewards, andevaluations Type One Data pure capture of real workflows like GitHubcommits, crucial for true expertise Beyond Basic Labeling focus on basic data labeling is the leastinteresting part of AI development True Expertise data transitions models from generalistcompetence to true domain expertise Verifiability Bottleneck current benchmarks are problematic,hindering reliable data quality assessment Future: Custom Data future of data lies in enterprise andhighly customized, specialized datasets From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Data Market faces Market Fragmentation. Market Fragmentation leads to Specialists Outcompete. Beyond Basic Labeling focus on Type One Data. Type One Data enables True Expertise. Market Fragmentation exacerbates Verifiability Bottleneck. True Expertise drives Future: Custom Data. Verifiability Bottleneck requires Future: Custom Data faces leads to focus on enables exacerbates drives requires AI Data Market Sean Cai critiquesthe current stateof AI data markets… MarketFragmentation shift fromverticallyintegrated giants… SpecialistsOutcompete specialists excelin talent sourcing,environment… Type One Data pure capture ofreal workflows likeGitHub commits,… Beyond BasicLabeling focus on basic datalabeling is theleast interesting… True Expertise data transitionsmodels fromgeneralist… VerifiabilityBottleneck current benchmarksare problematic,hindering reliable… Future: CustomData future of data liesin enterprise andhighly customized,… From startuphub.ai · The publishers behind this format

Sean Cai, an independent researcher and speaker, delivered a compelling talk on the "State of Data" at the AI Engineer World's Fair, offering a critical look at the AI data market. Cai argued that the focus on basic data labeling is the "least interesting part" of AI development, with the real value lying in data that transitions models from generalist competence to true expertise.

Sean Cai on the State of AI Data Markets - AI Engineer
Sean Cai on the State of AI Data Markets — from AI Engineer

The Evolving Data Market: Fragmentation and Quality

Cai highlighted a significant shift in the data market, moving away from the vertically integrated giants of the past towards a more fragmented landscape of specialists. He noted that "specialists outcompete the giants at a lot of steps," including sourcing talent, building environments, designing rewards, and running evaluations. This fragmentation, he believes, is permanent, not transitional.

A key point of his presentation was the concept of "type one" versus "type two" data. Type one data is a pure capture of real workflows, such as GitHub commits or session replays, with minimal reward shaping. Type two data, conversely, involves hiring experts to create contrived examples in arbitrary settings. Cai emphasized that while type two data is useful for initial model training, type one data is crucial for achieving higher levels of performance and expertise, as its realism is inherited from the work itself.

A stark warning was issued regarding the industry's practices: "The dirty secret of the industry though is that everybody sells type two and bills it as type one." This practice, he explained, leads to a situation where quality does not scale linearly with quantity, prompting many labs to diversify their vendor base to mitigate risks.

Verifiability as a Bottleneck and the Problem with Benchmarks

Cai introduced "Verifier's Law," a concept posited by researcher Jason Wei, stating that the ease of training a model is proportional to how verifiable the task is. He broke down verifiability into three axes: asymmetry (decomposing a task into checkable steps), veracity (consensus on what constitutes correctness), and proliferation (frequency of real-world examples). He used coding as an example, citing GitHub's unit tests and commit messages as solving all three axes simultaneously, which contributed to AI's early success in this domain.

Conversely, domains like biology, security, finance, healthcare, and law score lower on these axes, making data acquisition and verification more challenging. Cai also critiqued the prevalent use of "contrived benchmarks," which he described as a "recipe" for creating fake evaluations. This involves hiring experts to generate tasks, solve them with AI, cherry-pick failures, and then sell the data to "hill climb" that same benchmark. He called this "Goodhart's Law with a profit motive," where the measure becomes the target itself, leading to a "fog of war" where no one truly knows which data genuinely improves models.

The Future of Data: Enterprise and Customization

Cai predicted that the next wave of AI application maturity would move into these more complex, less verifiable domains. He also discussed the evolving role of models, noting that as open-source models improve (like GLM 5.2 surpassing GPT in some rubrics), application layer companies can decouple from foundation model labs. This decoupling, he argued, prevents durable lock-in and shifts the focus to efficiency and modality differences.

The most significant trend, however, is the pivot of successful data companies towards enterprise solutions. Cai stated, "Data businesses do not stay data businesses because the durable value accrues to the services and app layer of actual work." This means building specialized infrastructure, including "Antikythera mechanisms" to translate messy business contexts into actionable data, managing RL data sets across base model migrations, and serving route small models efficiently. He concluded with two key takeaways for the AI community: researchers should avoid outsourcing their definition of realism to vendors, and builders should focus on pipelines into real-world work, coupled with the infrastructure for continuous retraining.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.