Vals: The AI Scorekeeper We Need
Vals is building the essential trust layer for AI, evaluating models on real-world tasks, not just academic benchmarks.

Visual TL;DR
models increasingly expected to perform complex real-world tasks
From the article 9+ mentionsVals actively retires saturated evaluations and introduces new tasks as the AI frontier advances.
models ace public datasets, which are saturated or leak into training data
From the articleFor years, the industry has leaned on academic benchmarks.
models top leaderboards but falter on messy, multi-step tasks that matter
From the articleInstead of contrived exams, it tests models on the work people actually need done.
building the essential trust layer for AI, evaluating models on real-world tasks
From the article 9 mentionsAs AI adoption accelerates, the need for such an independent scorekeeper becomes not just beneficial, but imperative.
collaborates with domain experts to translate real-world workflows into rigorous benchmarks
From the article 2 mentionsThis is where Vals enters the picture, aiming to build a crucial trust layer between AI models and their users.
From the article 2 mentionsThey then develop automated grading systems capable of evaluating the final output to an expert standard.
Contents(3)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.