Vals: The AI Scorekeeper We Need

Vals is building the essential trust layer for AI, evaluating models on real-world tasks, not just academic benchmarks.

8 min read
a16z logo with text 'Investing in Vals'
a16z Blog

Visual TL;DR. AI models advance leads to Academic benchmarks fail. Academic benchmarks fail causes Need real-world tests. Need real-world tests solved by Vals: AI Scorekeeper. Academic benchmarks fail solved by Vals: AI Scorekeeper. Vals: AI Scorekeeper by using Tests real workflows. Tests real workflows with Automated grading. Vals: AI Scorekeeper results in Builds AI trust.

  1. AI models advance: models increasingly expected to perform complex real-world tasks
  2. Academic benchmarks fail: models ace public datasets, which are saturated or leak into training data
  3. Need real-world tests: models top leaderboards but falter on messy, multi-step tasks that matter
  4. Vals: AI Scorekeeper: building the essential trust layer for AI, evaluating models on real-world tasks
  5. Tests real workflows: collaborates with domain experts to translate real-world workflows into rigorous benchmarks
  6. Automated grading: develops automated grading systems evaluating final output to an expert standard
  7. Builds AI trust: creates a crucial trust layer between AI models and their users
Visual TL;DR
Visual TL;DR, startuphub.ai AI models advance leads to Academic benchmarks fail. Academic benchmarks fail solved by Vals: AI Scorekeeper. Vals: AI Scorekeeper results in Builds AI trust leads to solved by results in AI models advance Academic benchmarks fail Vals: AI Scorekeeper Builds AI trust From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI models advance leads to Academic benchmarks fail. Academic benchmarks fail solved by Vals: AI Scorekeeper. Vals: AI Scorekeeper results in Builds AI trust leads to solved by results in AI models advance Academicbenchmarks fail Vals: AIScorekeeper Builds AI trust From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI models advance leads to Academic benchmarks fail. Academic benchmarks fail solved by Vals: AI Scorekeeper. Vals: AI Scorekeeper results in Builds AI trust leads to solved by results in AI models advance models increasingly expected to performcomplex real-world tasks Academic benchmarks fail models ace public datasets, which aresaturated or leak into training data Vals: AI Scorekeeper building the essential trust layer for AI,evaluating models on real-world tasks Builds AI trust creates a crucial trust layer between AImodels and their users From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI models advance leads to Academic benchmarks fail. Academic benchmarks fail solved by Vals: AI Scorekeeper. Vals: AI Scorekeeper results in Builds AI trust leads to solved by results in AI models advance models increasinglyexpected to performcomplex real-world… Academicbenchmarks fail models ace publicdatasets, which aresaturated or leak… Vals: AIScorekeeper building theessential trustlayer for AI,… Builds AI trust creates a crucialtrust layer betweenAI models and their… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI models advance leads to Academic benchmarks fail. Academic benchmarks fail causes Need real-world tests. Need real-world tests solved by Vals: AI Scorekeeper. Academic benchmarks fail solved by Vals: AI Scorekeeper. Vals: AI Scorekeeper by using Tests real workflows. Tests real workflows with Automated grading. Vals: AI Scorekeeper results in Builds AI trust leads to causes solved by solved by by using with results in AI models advance models increasingly expected to performcomplex real-world tasks Academic benchmarks fail models ace public datasets, which aresaturated or leak into training data Need real-world tests models top leaderboards but falter onmessy, multi-step tasks that matter Vals: AI Scorekeeper building the essential trust layer for AI,evaluating models on real-world tasks Tests real workflows collaborates with domain experts totranslate real-world workflows intorigorous benchmarks Automated grading develops automated grading systemsevaluating final output to an expertstandard Builds AI trust creates a crucial trust layer between AImodels and their users From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI models advance leads to Academic benchmarks fail. Academic benchmarks fail causes Need real-world tests. Need real-world tests solved by Vals: AI Scorekeeper. Academic benchmarks fail solved by Vals: AI Scorekeeper. Vals: AI Scorekeeper by using Tests real workflows. Tests real workflows with Automated grading. Vals: AI Scorekeeper results in Builds AI trust leads to causes solved by solved by by using with results in AI models advance models increasinglyexpected to performcomplex real-world… Academicbenchmarks fail models ace publicdatasets, which aresaturated or leak… Need real-worldtests models topleaderboards butfalter on messy,… Vals: AIScorekeeper building theessential trustlayer for AI,… Tests realworkflows collaborates withdomain experts totranslate… Automated grading develops automatedgrading systemsevaluating final… Builds AI trust creates a crucialtrust layer betweenAI models and their… From startuphub.ai · The publishers behind this format

The relentless march of artificial intelligence means models are not just getting smarter, they're increasingly expected to do things. But how do we know if they're actually good at it? For years, the industry has leaned on academic benchmarks. These offered a common language, but frontier models are now acing them. Public datasets become saturated, leak into training data, or are explicitly optimized against. A model can top a leaderboard and still falter on the messy, multi-step tasks that matter in the real world. This is where Vals enters the picture, aiming to build a crucial trust layer between AI models and their users.

Beyond Academic Exams

Vals takes a different tack. Instead of contrived exams, it tests models on the work people actually need done. The team collaborates with domain experts to translate real-world workflows into rigorous benchmarks. They then develop automated grading systems capable of evaluating the final output to an expert standard. Think of it: not just asking if a legal AI can pass the bar exam, but if it can perform actual legal research. In finance, it's about analyzing complex documents, not just answering trivia. For coders, it's about building a working application.

This approach is vital as AI moves from simple Q&A to complex operations: analyzing financial reports, resolving legal cases, writing and debugging software, and navigating intricate enterprise workflows. The stakes are rising. A poor model choice can cost not just token expenses but also significant time and damage customer satisfaction. Vals' infrastructure allows these evaluations to run quickly, reproducibly, and securely at scale. Private test sets are protected from contamination, and results can be delivered within hours of model access.

The Perishable Benchmark

The team at Vals understands that benchmarks have a shelf life. A good benchmark is designed to become obsolete as models improve. When systems consistently score near-perfectly, it signals the benchmark's job is done and a harder one is needed. Vals actively retires saturated evaluations and introduces new tasks as the AI frontier advances. For example, the CorpFin benchmark was recently replaced with a new Excel Modelling Benchmark in May, after CorpFin ceased to offer sufficient differentiation between models. This adaptive nature is essential for the industry to keep pace and make informed AI adoption decisions.

StartupHub.ai data shows that while the broader AI evaluation space is seeing activity, with competitors like TradingView scoring 75/100 and Bridgewise at 61/100, Vals' focus on real-world task simulation positions it uniquely. Our data gives Investing a StartupHub score of 20/100, highlighting the gap Vals aims to fill.

An Independent Scorekeeper for AI

Every major market eventually requires an independent arbiter. Credit markets have Moody's and S&P. Public markets rely on auditors. Product manufacturers use UL for safety certification. These institutions emerge because when sellers possess more information and have incentives to present themselves favorably, trusted third-party measurement is essential for market function. AI, poised to become one of the largest markets ever, is at this critical juncture. Vendors grading their own homework is insufficient; buyers need trustworthy assessments.

Rayan Krishnan and Langston Nashold, Vals' co-founders, bring a rare combination of technical depth, skepticism, and the patience required to continuously evolve testing. Their background in computer science at Stanford and prior collaboration on real-world measurement problems before starting Vals makes them uniquely suited to build this adaptive evaluation layer. As AI adoption accelerates, the need for such an independent scorekeeper becomes not just beneficial, but imperative.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.