# Agent Evals Lag Behind AI Evolution _Ameya Bhatawdekar explains how AI agent evaluations are lagging behind model advancements, creating a need for new methods._ **Updated:** 2026-08-22 **Published:** 2026-08-20 **Source:** https://www.startuphub.ai/ai-news/ai-research/2026/agent-evals-lag-behind-ai-evolution --- The rapid evolution of AI models has outpaced the development of evaluation methods, creating a bottleneck for agent development. Ameya Bhatawdekar, a speaker from Braintrust, argues that the very architectures built to manage AI capabilities are now hindering progress. He traces this loop across five generations of AI architecture, highlighting how each step change in model capability demanded new evaluation strategies. AI Models Evolve RapidlyDriver rapid evolution of AI models outpaces evaluation method developmentEarly AI Evals SimpleContextearly AI systems needed only a single answer to a single promptEvals Lag BehindDriverAI agent evaluations are lagging behind model advancements, creating a bottleneckRetrieval Chains ComplexContextFrom the articleThe complexity increased with retrieval chains, which introduced new failure points like incorrect parsing or irrelevant context retrieval.Orchestration Graphs Add ErrorsContextorchestration graphs with branch logic and classifier nodes added further error surfacesFrom the articleThe introduction of orchestration graphs, designed to manage more complex workflows, added further surfaces for potential errors.Bhatawdekar's ArgumentCoreAmeya Bhatawdekar explains how current architectures hinder progressFrom the article 6 mentionsThe core of Bhatawdekar's argument lies in the distinction between capability and reliability in AI agent evaluation.requiresNew Evals NeededEffecteach step change in model capability demanded new evaluation strategiesFrom the article 2 mentionsHe breaks down the problem into two key questions that a single eval result can hide. ## The Evolution of AI Architecture and Its Evaluation Challenges Bhatawdekar explains that early AI systems, requiring only a single answer to a single prompt, had straightforward evaluations. The complexity increased with retrieval chains, which introduced new failure points like incorrect parsing or irrelevant context retrieval. The introduction of orchestration graphs, designed to manage more complex workflows, added further surfaces for potential errors. These graphs incorporated branch logic, contracts between nodes, and classifier nodes that could fail silently. This increased complexity meant a greater need for rigorous checking. A significant shift occurred when models became reliable enough to operate autonomously. At this stage, the same input could produce different outputs on each run. This variability made single evaluation results less meaningful. Bhatawdekar emphasizes that this is not just an iteration but a replatforming problem. ## Shifting from Capability to Reliability in Evals The core of Bhatawdekar's argument lies in the distinction between capability and reliability in AI agent evaluation. He breaks down the problem into two key questions that a single eval result can hide. The first is **'pass at k'**, which measures whether the system succeeds at least once within 'k' attempts. This is a measure of capability. The second, stricter question asks how many of those 'k' attempts actually succeed. This metric measures reliability. A system might demonstrate strong capability by succeeding once, but fail to be reliable if it only succeeds sporadically. Bhatawdekar stresses that understanding this difference requires running the distribution of results, not just a single sample. He points out that memory, sandboxes, and skills around the AI loop also need careful evaluation as agents become more sophisticated. ## Evals as the Durable Asset Bhatawdekar posits that evaluations are the most durable asset for AI teams, especially through replatformings. However, he warns that evals become stale and lose their value if they are not continuously fed with data from production. This continuous feedback loop is essential for maintaining relevant and effective evaluation strategies that keep pace with AI advancements. The challenge for many teams is that they accept a flywheel of development and deployment but fail to run the crucial evaluation component of that flywheel effectively. This leads to a situation where the agents are advancing rapidly, but the tools to measure their performance and reliability are not keeping up. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.