Rayan Garg on Why Long Horizon AI Agents Need Better Verifiers
Rayan Garg from Theta Software explains why long horizon AI agent benchmarks need accurate environment design and final-state verifiers.

Visual TL;DR
researcher at Theta Software focusing on dependable evaluations for autonomous agents
From the article 5 mentionsIn a technical talk, Rayan Garg from Theta Software broke down why existing benchmarks for long horizon work fall short and how the industry must rethink environment design and task verification.
AI agents handling tasks stretching over hours or days, hard to measure performance
From the article 5 mentionsIndustry benchmarks often measure long horizon performance using a simple time threshold.
current industry benchmarks use simple time thresholds like 16-hour completion
From the article 4 mentionsHis work addresses the growing gap between headline benchmark metrics and real-world agent performance.
small errors early in long tasks can lead to large failures later
difficulty in tracking all possible states an agent can enter during long tasks
From the article 4 mentionsReal difficulty stems from state complexity and high-stakes choices, not just elapsed time.
time threshold approach produces unreliable data for agent performance evaluation
From the article 2 mentionsGarg argues that this approach produces noisy, unreliable data.
From the article 2 mentionsIn a technical talk, Rayan Garg from Theta Software broke down why existing benchmarks for long horizon work fall short and how the industry must rethink environment design and task verification.
need accurate environment design and final-state verifiers for long horizon tasks
From the article 2 mentionsRayan Garg is a researcher and engineer at Theta Software, where his focus centers on creating dependable evaluations, environments, and verifiers for autonomous agents.
Contents(5)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.