AI Benchmarks Face 'Statistical Precipice'
Pierluca D'Oro of Programma Labs critiques AI benchmarking, highlighting issues with 'replay agents' and deterministic environments, and advocating for more robust, diverse, and accurately measured evaluations.

Visual TL;DR
benchmarks like OSWorld or MobileWorld are easily gamed by replay agents
From the article 6 mentionsD’Oro specifically called out the 'pass@k' metric, stating that on deterministic environments, it effectively becomes a measure of the replay agent's success, rewarding search budget over genuine capability.
From the articlePierluca D'Oro, founder of Programma Labs, delivered a critical assessment of current AI agent evaluation methodologies, highlighting significant flaws that can lead to misleading results.
current evaluation methods lead to misleading results for AI agents
From the article 9+ mentionsThis accuracy is vital for making sound deployment decisions, as flawed confidence intervals can lead to costly mistakes.
simple scripts mimic successful action sequences on deterministic benchmarks
From the article 5 mentionsThese issues point to two core problems in AI agent benchmarks: the environments themselves and the evaluation metrics.
replay agents achieve same or better success rates than frontier models
From the articleHe demonstrated how such agents can outperform sophisticated AI models on deterministic benchmarks, a phenomenon he argues is unacceptable.
advocating for more diverse, verifiable, and accurately measured AI evaluations
proposed framework for addressing benchmark flaws and improving rigor
From the article 2 mentionsD’Oro proposed a set of 'PRISM' principles for designing more robust and trustworthy environments:
Contents(4)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.