# Vals: The AI Scorekeeper We Need _Vals is building the essential trust layer for AI, evaluating models on real-world tasks, not just academic benchmarks._ **Updated:** 2026-08-22 **Published:** 2026-08-13 **Source:** https://www.startuphub.ai/ai-news/investors-news/2026/vals-the-ai-scorekeeper-we-need --- The relentless march of artificial intelligence means models are not just getting smarter, they're increasingly expected to *do* things. But how do we know if they're actually good at it? For years, the industry has leaned on academic benchmarks. These offered a common language, but frontier models are now acing them. Public datasets become saturated, leak into training data, or are explicitly optimized against. A model can top a leaderboard and still falter on the messy, multi-step tasks that matter in the real world. This is where [Vals](https://www.a16z.news/p/investing-in-vals) enters the picture, aiming to build a crucial trust layer between AI models and their users. AI models advanceDriver models increasingly expected to perform complex real-world tasksFrom the article 9+ mentionsVals actively retires saturated evaluations and introduces new tasks as the AI frontier advances.leads toAcademic benchmarks failDrivermodels ace public datasets, which are saturated or leak into training dataFrom the articleFor years, the industry has leaned on academic benchmarks.causesNeed real-world testsDrivermodels top leaderboards but falter on messy, multi-step tasks that matterFrom the articleInstead of contrived exams, it tests models on the work people actually need done.solved byVals: AI ScorekeeperCorebuilding the essential trust layer for AI, evaluating models on real-world tasksFrom the article 9 mentionsAs AI adoption accelerates, the need for such an independent scorekeeper becomes not just beneficial, but imperative.Tests real workflowsContextcollaborates with domain experts to translate real-world workflows into rigorous benchmarksBuilds AI trustEffectFrom the article 2 mentionsThis is where Vals enters the picture, aiming to build a crucial trust layer between AI models and their users.withAutomated gradingContextFrom the article 2 mentionsThey then develop automated grading systems capable of evaluating the final output to an expert standard. ## Beyond Academic Exams [Vals](https://www.a16z.news/p/investing-in-vals) takes a different tack. Instead of contrived exams, it tests models on the work people actually need done. The team collaborates with domain experts to translate real-world workflows into rigorous benchmarks. They then develop automated grading systems capable of evaluating the final output to an expert standard. Think of it: not just asking if a legal AI can pass the bar exam, but if it can perform actual legal research. In finance, it's about analyzing complex documents, not just answering trivia. For coders, it's about building a working application. This approach is vital as AI moves from simple Q&A to complex operations: analyzing financial reports, resolving legal cases, writing and debugging software, and navigating intricate enterprise workflows. The stakes are rising. A poor model choice can cost not just token expenses but also significant time and damage customer satisfaction. Vals' infrastructure allows these evaluations to run quickly, reproducibly, and securely at scale. Private test sets are protected from contamination, and results can be delivered within hours of model access. ## The Perishable Benchmark The team at Vals understands that benchmarks have a shelf life. A good benchmark is designed to become obsolete as models improve. When systems consistently score near-perfectly, it signals the benchmark's job is done and a harder one is needed. Vals actively retires saturated evaluations and introduces new tasks as the AI frontier advances. For example, the CorpFin benchmark was recently replaced with a new Excel Modelling Benchmark in May, after CorpFin ceased to offer sufficient differentiation between models. This adaptive nature is essential for the industry to keep pace and make informed AI adoption decisions. StartupHub.ai data shows that while the broader AI evaluation space is seeing activity, with competitors like TradingView scoring 75/100 and Bridgewise at 61/100, Vals' focus on real-world task simulation positions it uniquely. Our data gives Investing a StartupHub score of 20/100, highlighting the gap Vals aims to fill. ## An Independent Scorekeeper for AI Every major market eventually requires an independent arbiter. Credit markets have Moody's and S&P. Public markets rely on auditors. Product manufacturers use UL for safety certification. These institutions emerge because when sellers possess more information and have incentives to present themselves favorably, trusted third-party measurement is essential for market function. AI, poised to become one of the largest markets ever, is at this critical juncture. Vendors grading their own homework is insufficient; buyers need trustworthy assessments. Rayan Krishnan and Langston Nashold, Vals' co-founders, bring a rare combination of technical depth, skepticism, and the patience required to continuously evolve testing. Their background in computer science at Stanford and prior collaboration on real-world measurement problems before starting Vals makes them uniquely suited to build this adaptive evaluation layer. As AI adoption accelerates, the need for such an independent scorekeeper becomes not just beneficial, but imperative. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.