# Data Quality Is the Compute Multiplier Says Ari Morcos _Ari Morcos explains why high quality data acts as a compute multiplier for AI models, dramatically lowering training and inference costs._ **Published:** 2026-07-31 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/data-quality-is-the-compute-multiplier-says-ari-morcos --- Ari Morcos, founder and chief executive officer of DatologyAI, argues that the AI industry is looking at compute scaling incorrectly. As major research labs spend billions on massive compute clusters, Morcos demonstrates that dataset quality acts as a direct multiplier on compute efficiency. Swapping raw compute for curated data along the scaling curve allows organizations to train superior models at a lower overall cost. AI Compute ScalingDriver industry focuses on massive compute clusters, spending billions on hardwareFrom the article 3 mentionsAri Morcos, founder and chief executive officer of DatologyAI, argues that the AI industry is looking at compute scaling incorrectly.challengesAri MorcosCorefounder of DatologyAI, challenges current compute scaling assumptionsFrom the article 6 mentionsAri Morcos is the founder of DatologyAI, a company building automated data curation technology for foundation models.advocatesData QualityContextacts as a direct multiplier on compute efficiency for AI modelsFrom the article 9+ mentionsTechniques like automated quality classifiers, semantic deduplication, and synthetic data generation each play distinct roles.Data Refinery FrameworkContextdataset preparation framed like oil refining, not an endless firehoseFrom the articleRather than viewing data collection as an endless firehose, Morcos frames dataset preparation as an oil refinery.Lower Training CostsEffectcurated data reduces need for raw compute, lowering overall expensesFrom the article 2 mentionsSwapping raw compute for curated data along the scaling curve allows organizations to train superior models at a lower overall cost.Signal Per TokenEffectimproves inference gains, making models more efficient at runtimeFrom the articleThe scarce resource in frontier AI development is no longer raw token volume, but signal per token.leads toSuperior AI ModelsOutcomehigh-quality data leads to better model performance and generalizationFrom the article 7 mentionsSwapping raw compute for curated data along the scaling curve allows organizations to train superior models at a lower overall cost.shown byReal-World ProofOutcomeThomson Reuters and Arcee demonstrate benefits of data curation ## Who Is Ari Morcos Ari Morcos is the founder of DatologyAI, a company building automated data curation technology for foundation models. Before founding DatologyAI, Morcos was a research scientist at Meta AI (FAIR), where he focused on neural network representations, deep learning efficiency, and model generalization. His technical work centers on solving the data bottleneck facing modern artificial intelligence. ## The Data Refinery Framework Rather than viewing data collection as an endless firehose, Morcos frames dataset preparation as an oil refinery. Raw web scrapes contain noise, duplication, and low quality text that degrade neural network training. A modern data pipeline must clean, curate, create, and compose dataset mixtures. Techniques like automated quality classifiers, semantic deduplication, and synthetic data generation each play distinct roles. Furthermore, Morcos emphasizes that the sequencing of these steps across training stages matters just as much as any single filtering algorithm. ## Signal Per Token and Inference Gains The scarce resource in frontier AI development is no longer raw token volume, but signal per token. Finding data that is optimal for a specific target task provides disproportionate efficiency gains. Curated datasets allow smaller models to outperform far larger architectures trained on unrefined web data. This data advantage extends directly to model deployment. Models trained on dense, high signal data reach benchmark target accuracy with fewer overall parameters. That smaller parameter footprint delivers permanent inference efficiency and lower operational costs in production. ## Real World Proof Points: Thomson Reuters and Arcee The economic benefits of structured data curation are already visible across production deployments. Legal and business media group [Thomson Reuters (NYSE:TRI)](https://www.google.com/finance/quote/TRI:NYSE) applied targeted data curation in mid-training to maximize performance on proprietary legal datasets. Meanwhile, model builder Arcee trained its Trinity model on public data alone, reaching frontier capabilities without proprietary access. StartupHub.ai data tracks market traction across key players in these technical sectors. Thomson Reuters holds a StartupHub score of 38/100. Arcee records a score of 44/100. Tracked competitors in legal intelligence and data processing include Harvey AI with a score of 70/100, alexi at 48/100, and Bloomberg L.P. at 82/100. ## Manufacturing Quality Data Over Buying Compute The core business thesis for data curation rests on basic economics. Manufacturing high quality synthetic and filtered data costs significantly less than purchasing additional GPU clusters. As training runs scale past hundred-million-dollar budgets, data selection and curation quiet determine which teams remain competitive. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.