Rayan Garg on Why Long Horizon AI Agents Need Better Verifiers

Rayan Garg from Theta Software explains why long horizon AI agent benchmarks need accurate environment design and final-state verifiers.

8 min read
Rayan Garg speaking about long horizon AI agent environments and verifier design.
Rayan Garg presents on rethinking environments and verifiers for long horizon AI work.· AI Engineer

Visual TL;DR. Rayan Garg studies Long Horizon AI. Long Horizon AI hindered by Flawed Benchmarks. Flawed Benchmarks leads to Noisy Data. Flawed Benchmarks requires Rethink Verification. Rethink Verification means Better Verifiers. Long Horizon AI causes Cascading Errors. Long Horizon AI involves State Space Complexity. Cascading Errors highlights need Rethink Verification. State Space Complexity highlights need Rethink Verification.

  1. Rayan Garg: researcher at Theta Software focusing on dependable evaluations for autonomous agents
  2. Long Horizon AI: AI agents handling tasks stretching over hours or days, hard to measure performance
  3. Flawed Benchmarks: current industry benchmarks use simple time thresholds like 16-hour completion
  4. Noisy Data: time threshold approach produces unreliable data for agent performance evaluation
  5. Rethink Verification: industry must rethink environment design and task verification for agents
  6. Better Verifiers: need accurate environment design and final-state verifiers for long horizon tasks
  7. Cascading Errors: small errors early in long tasks can lead to large failures later
  8. State Space Complexity: difficulty in tracking all possible states an agent can enter during long tasks
Visual TL;DR
Visual TL;DR, startuphub.ai Rayan Garg studies Long Horizon AI. Long Horizon AI hindered by Flawed Benchmarks. Flawed Benchmarks requires Rethink Verification. Rethink Verification means Better Verifiers studies hindered by requires means Rayan Garg Long Horizon AI Flawed Benchmarks Rethink Verification Better Verifiers From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Rayan Garg studies Long Horizon AI. Long Horizon AI hindered by Flawed Benchmarks. Flawed Benchmarks requires Rethink Verification. Rethink Verification means Better Verifiers studies hindered by requires means Rayan Garg Long Horizon AI Flawed Benchmarks RethinkVerification Better Verifiers From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Rayan Garg studies Long Horizon AI. Long Horizon AI hindered by Flawed Benchmarks. Flawed Benchmarks requires Rethink Verification. Rethink Verification means Better Verifiers studies hindered by requires means Rayan Garg researcher at Theta Software focusing ondependable evaluations for autonomousagents Long Horizon AI AI agents handling tasks stretching overhours or days, hard to measure performance Flawed Benchmarks current industry benchmarks use simpletime thresholds like 16-hour completion Rethink Verification industry must rethink environment designand task verification for agents Better Verifiers need accurate environment design andfinal-state verifiers for long horizontasks From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Rayan Garg studies Long Horizon AI. Long Horizon AI hindered by Flawed Benchmarks. Flawed Benchmarks requires Rethink Verification. Rethink Verification means Better Verifiers studies hindered by requires means Rayan Garg researcher at ThetaSoftware focusingon dependable… Long Horizon AI AI agents handlingtasks stretchingover hours or days,… Flawed Benchmarks current industrybenchmarks usesimple time… RethinkVerification industry mustrethink environmentdesign and task… Better Verifiers need accurateenvironment designand final-state… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Rayan Garg studies Long Horizon AI. Long Horizon AI hindered by Flawed Benchmarks. Flawed Benchmarks leads to Noisy Data. Flawed Benchmarks requires Rethink Verification. Rethink Verification means Better Verifiers. Long Horizon AI causes Cascading Errors. Long Horizon AI involves State Space Complexity. Cascading Errors highlights need Rethink Verification. State Space Complexity highlights need Rethink Verification studies hindered by leads to requires means causes involves highlights need highlights need Rayan Garg researcher at Theta Software focusing ondependable evaluations for autonomousagents Long Horizon AI AI agents handling tasks stretching overhours or days, hard to measure performance Flawed Benchmarks current industry benchmarks use simpletime thresholds like 16-hour completion Noisy Data time threshold approach producesunreliable data for agent performanceevaluation Rethink Verification industry must rethink environment designand task verification for agents Better Verifiers need accurate environment design andfinal-state verifiers for long horizontasks Cascading Errors small errors early in long tasks can leadto large failures later State Space Complexity difficulty in tracking all possible statesan agent can enter during long tasks From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Rayan Garg studies Long Horizon AI. Long Horizon AI hindered by Flawed Benchmarks. Flawed Benchmarks leads to Noisy Data. Flawed Benchmarks requires Rethink Verification. Rethink Verification means Better Verifiers. Long Horizon AI causes Cascading Errors. Long Horizon AI involves State Space Complexity. Cascading Errors highlights need Rethink Verification. State Space Complexity highlights need Rethink Verification studies hindered by leads to requires means causes involves highlights need highlights need Rayan Garg researcher at ThetaSoftware focusingon dependable… Long Horizon AI AI agents handlingtasks stretchingover hours or days,… Flawed Benchmarks current industrybenchmarks usesimple time… Noisy Data time thresholdapproach producesunreliable data for… RethinkVerification industry mustrethink environmentdesign and task… Better Verifiers need accurateenvironment designand final-state… Cascading Errors small errors earlyin long tasks canlead to large… State SpaceComplexity difficulty intracking allpossible states an… From startuphub.ai · The publishers behind this format

As autonomous AI agents attempt to handle tasks that stretch over hours or days, measuring their actual performance has become one of the hardest challenges in machine learning. In a technical talk, Rayan Garg from Theta Software broke down why existing benchmarks for long horizon work fall short and how the industry must rethink environment design and task verification.

Rayan Garg on Why Long Horizon AI Agents Need Better Verifiers - AI Engineer
Rayan Garg on Why Long Horizon AI Agents Need Better Verifiers — from AI Engineer

Who Is Rayan Garg

Rayan Garg is a researcher and engineer at Theta Software, where his focus centers on creating dependable evaluations, environments, and verifiers for autonomous agents. His work addresses the growing gap between headline benchmark metrics and real-world agent performance.

The Flaws in Current Long Horizon Metrics

Industry benchmarks often measure long horizon performance using a simple time threshold. For instance, an evaluation might define a model as capable of long horizon tasks if it successfully completes work at the 16-hour mark. Garg argues that this approach produces noisy, unreliable data.

Wall clock time hides the true difficulty of a task. Human time estimates vary wildly depending on skill level and experience. Furthermore, an artificially stretched task with serial dependencies might take hours to complete without testing complex reasoning. Real difficulty stems from state complexity and high-stakes choices, not just elapsed time.

Cascading Errors and State Space Complexity

True long horizon difficulty occurs when a single mistake early in a process cascades through every subsequent step. A bad initial database query or misconfigured environment setting corrupts all downstream actions, turning a simple task into an unrecoverable failure mode.

Evaluating models under these conditions requires creating realistic, dynamic environments. Standardized evaluation becomes exceptionally difficult as state spaces grow. When state spaces balloon, traditional judge models often struggle to accurately evaluate whether an agent truly succeeded.

Verifying Success from Final State

To avoid relying on subjective judge estimates, Garg highlights the importance of verifying agent success from the actual final state of the environment. Rather than grading step-by-step trace files or trusting a judge model to guess correctness, system designers should evaluate concrete final artifacts.

This approach involves collapsing massive state spaces using sample trajectories and structured tools. Evaluators must also ensure LLM judges do not see information they should not, preventing data leakage and unearned success scores.

Reusing Agents for Automated QA

Another major strategy Garg shares is using specialized agents to sift through artifacts like CI logs, build outputs, and execution traces. By pairing final-state verification with specialized agentic grading, teams can build clear rubrics and quality assurance pipelines.

Ultimately, progress in long horizon autonomous work lives or dies on environment and verifier design. Simply pushing headline benchmark scores without honest evaluation frameworks will fail to produce reliable agents in production.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.