AI Agents Stall on Core AI Research

Frontier AI agents can automate AI research engineering but fail to make substantial progress on core research questions, according to new shadow evaluations.

Abstract concept art representing artificial intelligence and research.
Visualizing the complex interplay between AI agents and the research process.
Visual TL;DR
AI Agents StallDriver
frontier AI agents fail to make substantial progress on core research questions
From the article 6 mentionsThe promise of explosive AI progress often hinges on AI agents automating AI research itself.
Existing Evals FlawedDriver
current methods are narrow, exclude open-ended research, or use strained peer review
Shadow EvaluationsCore
novel approach places AI agent at heart of unpublished paper's research question
From the article 2 mentionsTo address this measurement gap, researchers Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, and colleagues introduced shadow evaluations.
Author GradingContext
original paper authors directly assess the AI agent's output for research capability
From the article 2 mentionsCrucially, the original authors of the paper then grade the agent's output.
Measure Research GapEffect
addresses the measurement gap for AI's ability to automate core AI research
Engineering-AI DivideOutcome
agents automate research engineering but lack insight for core research problems
Contents(3)

The promise of explosive AI progress often hinges on AI agents automating AI research itself. However, tangible evidence supporting this vision remains surprisingly scarce. Existing evaluation methods fall short. Some focus on narrow, verifiable tasks, inherently excluding the open-ended nature of true research. Others submit AI-generated papers to blind peer review, a process already strained, unpredictable, and prone to inconsistent quality.

Shadow Evaluations: A New Metric for AI R&D Automation

To address this measurement gap, researchers Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, and colleagues introduced shadow evaluations. This novel approach places an AI agent at the heart of a high-quality, unpublished research paper, tasking it with tackling the central, open-ended research question. Crucially, the original authors of the paper then grade the agent's output. This method provides a direct, author-centric assessment of AI's research capabilities.

Frontier Agents Fall Short on Research Insight

In their study, the researchers applied shadow evaluations to two unpublished NeurIPS 2026 submissions. Frontier agents were given six days and substantial compute resources, costing thousands of dollars. While the agents successfully completed all engineering tasks without human intervention, they failed to make significant progress on the core research questions. The authors unambiguously rejected the outputs, highlighting five persistent failure modes: inadequate judgment of publishable research standards, uncreative responses to research design flaws, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A subsequent robustness check using a different model and scaffolding reproduced these failures, reinforcing the findings.

The Engineering-AI Divide in Research

The results offer a sobering early assessment: today's most advanced AI agents excel at the engineering components of AI research, coding, implementation, and execution. Yet, they demonstrably struggle with the more nuanced, critical aspects of the research lifecycle. This includes formulating novel hypotheses, designing robust experiments, interpreting ambiguous results, and critically evaluating the significance of findings. The current limitations suggest that while AI can automate the 'how' of research, the 'what' and 'why' remain firmly in the human domain for now. This distinction is vital for investors and founders considering the trajectory of AI-driven innovation.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.