AI Agents Stall on Core AI Research
Frontier AI agents can automate AI research engineering but fail to make substantial progress on core research questions, according to new shadow evaluations.

Visual TL;DR
frontier AI agents fail to make substantial progress on core research questions
From the article 6 mentionsThe promise of explosive AI progress often hinges on AI agents automating AI research itself.
current methods are narrow, exclude open-ended research, or use strained peer review
novel approach places AI agent at heart of unpublished paper's research question
From the article 2 mentionsTo address this measurement gap, researchers Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, and colleagues introduced shadow evaluations.
original paper authors directly assess the AI agent's output for research capability
From the article 2 mentionsCrucially, the original authors of the paper then grade the agent's output.
addresses the measurement gap for AI's ability to automate core AI research
agents automate research engineering but lack insight for core research problems
Contents(3)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.