Vals: The AI Scorekeeper We Need

Vals is building the essential trust layer for AI, evaluating models on real-world tasks, not just academic benchmarks.

a16z logo with text 'Investing in Vals'
a16z Blog
Visual TL;DR
AI models advanceDriver
models increasingly expected to perform complex real-world tasks
From the article 9+ mentionsVals actively retires saturated evaluations and introduces new tasks as the AI frontier advances.
Academic benchmarks failDriver
models ace public datasets, which are saturated or leak into training data
From the articleFor years, the industry has leaned on academic benchmarks.
Need real-world testsDriver
models top leaderboards but falter on messy, multi-step tasks that matter
From the articleInstead of contrived exams, it tests models on the work people actually need done.
Vals: AI ScorekeeperCore
building the essential trust layer for AI, evaluating models on real-world tasks
From the article 9 mentionsAs AI adoption accelerates, the need for such an independent scorekeeper becomes not just beneficial, but imperative.
Tests real workflowsContext
collaborates with domain experts to translate real-world workflows into rigorous benchmarks
Builds AI trustEffect
From the article 2 mentionsThis is where Vals enters the picture, aiming to build a crucial trust layer between AI models and their users.
Automated gradingContext
From the article 2 mentionsThey then develop automated grading systems capable of evaluating the final output to an expert standard.
Contents(3)

The relentless march of artificial intelligence means models are not just getting smarter, they're increasingly expected to do things. But how do we know if they're actually good at it? For years, the industry has leaned on academic benchmarks. These offered a common language, but frontier models are now acing them. Public datasets become saturated, leak into training data, or are explicitly optimized against. A model can top a leaderboard and still falter on the messy, multi-step tasks that matter in the real world. This is where Vals enters the picture, aiming to build a crucial trust layer between AI models and their users.

Beyond Academic Exams

Vals takes a different tack. Instead of contrived exams, it tests models on the work people actually need done. The team collaborates with domain experts to translate real-world workflows into rigorous benchmarks. They then develop automated grading systems capable of evaluating the final output to an expert standard. Think of it: not just asking if a legal AI can pass the bar exam, but if it can perform actual legal research. In finance, it's about analyzing complex documents, not just answering trivia. For coders, it's about building a working application.

This approach is vital as AI moves from simple Q&A to complex operations: analyzing financial reports, resolving legal cases, writing and debugging software, and navigating intricate enterprise workflows. The stakes are rising. A poor model choice can cost not just token expenses but also significant time and damage customer satisfaction. Vals' infrastructure allows these evaluations to run quickly, reproducibly, and securely at scale. Private test sets are protected from contamination, and results can be delivered within hours of model access.

The Perishable Benchmark

The team at Vals understands that benchmarks have a shelf life. A good benchmark is designed to become obsolete as models improve. When systems consistently score near-perfectly, it signals the benchmark's job is done and a harder one is needed. Vals actively retires saturated evaluations and introduces new tasks as the AI frontier advances. For example, the CorpFin benchmark was recently replaced with a new Excel Modelling Benchmark in May, after CorpFin ceased to offer sufficient differentiation between models. This adaptive nature is essential for the industry to keep pace and make informed AI adoption decisions.

StartupHub.ai data shows that while the broader AI evaluation space is seeing activity, with competitors like TradingView scoring 75/100 and Bridgewise at 61/100, Vals' focus on real-world task simulation positions it uniquely. Our data gives Investing a StartupHub score of 20/100, highlighting the gap Vals aims to fill.

An Independent Scorekeeper for AI

Every major market eventually requires an independent arbiter. Credit markets have Moody's and S&P. Public markets rely on auditors. Product manufacturers use UL for safety certification. These institutions emerge because when sellers possess more information and have incentives to present themselves favorably, trusted third-party measurement is essential for market function. AI, poised to become one of the largest markets ever, is at this critical juncture. Vendors grading their own homework is insufficient; buyers need trustworthy assessments.

Rayan Krishnan and Langston Nashold, Vals' co-founders, bring a rare combination of technical depth, skepticism, and the patience required to continuously evolve testing. Their background in computer science at Stanford and prior collaboration on real-world measurement problems before starting Vals makes them uniquely suited to build this adaptive evaluation layer. As AI adoption accelerates, the need for such an independent scorekeeper becomes not just beneficial, but imperative.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.