AI Agents Stall on Core AI Research

Frontier AI agents can automate AI research engineering but fail to make substantial progress on core research questions, according to new shadow evaluations.

6 min read
Abstract concept art representing artificial intelligence and research.
Visualizing the complex interplay between AI agents and the research process.

Visual TL;DR. AI Agents Stall due to Existing Evals Flawed. Existing Evals Flawed leads to Shadow Evaluations. Shadow Evaluations involves Author Grading. Shadow Evaluations helps Measure Research Gap. Author Grading reveals Engineering-AI Divide. AI Agents Stall confirms Engineering-AI Divide.

  1. AI Agents Stall: frontier AI agents fail to make substantial progress on core research questions
  2. Existing Evals Flawed: current methods are narrow, exclude open-ended research, or use strained peer review
  3. Shadow Evaluations: novel approach places AI agent at heart of unpublished paper's research question
  4. Author Grading: original paper authors directly assess the AI agent's output for research capability
  5. Measure Research Gap: addresses the measurement gap for AI's ability to automate core AI research
  6. Engineering-AI Divide: agents automate research engineering but lack insight for core research problems
Visual TL;DR
Visual TL;DR, startuphub.ai Shadow Evaluations involves Author Grading. Author Grading reveals Engineering-AI Divide. AI Agents Stall confirms Engineering-AI Divide involves reveals confirms AI Agents Stall Shadow Evaluations Author Grading Engineering-AI Divide From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Shadow Evaluations involves Author Grading. Author Grading reveals Engineering-AI Divide. AI Agents Stall confirms Engineering-AI Divide involves reveals confirms AI Agents Stall ShadowEvaluations Author Grading Engineering-AIDivide From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Shadow Evaluations involves Author Grading. Author Grading reveals Engineering-AI Divide. AI Agents Stall confirms Engineering-AI Divide involves reveals confirms AI Agents Stall frontier AI agents fail to makesubstantial progress on core researchquestions Shadow Evaluations novel approach places AI agent at heart ofunpublished paper's research question Author Grading original paper authors directly assess theAI agent's output for research capability Engineering-AI Divide agents automate research engineering butlack insight for core research problems From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Shadow Evaluations involves Author Grading. Author Grading reveals Engineering-AI Divide. AI Agents Stall confirms Engineering-AI Divide involves reveals confirms AI Agents Stall frontier AI agentsfail to makesubstantial… ShadowEvaluations novel approachplaces AI agent atheart of… Author Grading original paperauthors directlyassess the AI… Engineering-AIDivide agents automateresearchengineering but… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Agents Stall due to Existing Evals Flawed. Existing Evals Flawed leads to Shadow Evaluations. Shadow Evaluations involves Author Grading. Shadow Evaluations helps Measure Research Gap. Author Grading reveals Engineering-AI Divide. AI Agents Stall confirms Engineering-AI Divide due to leads to involves helps reveals confirms AI Agents Stall frontier AI agents fail to makesubstantial progress on core researchquestions Existing Evals Flawed current methods are narrow, excludeopen-ended research, or use strained peerreview Shadow Evaluations novel approach places AI agent at heart ofunpublished paper's research question Author Grading original paper authors directly assess theAI agent's output for research capability Measure Research Gap addresses the measurement gap for AI'sability to automate core AI research Engineering-AI Divide agents automate research engineering butlack insight for core research problems From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Agents Stall due to Existing Evals Flawed. Existing Evals Flawed leads to Shadow Evaluations. Shadow Evaluations involves Author Grading. Shadow Evaluations helps Measure Research Gap. Author Grading reveals Engineering-AI Divide. AI Agents Stall confirms Engineering-AI Divide due to leads to involves helps reveals confirms AI Agents Stall frontier AI agentsfail to makesubstantial… Existing EvalsFlawed current methods arenarrow, excludeopen-ended… ShadowEvaluations novel approachplaces AI agent atheart of… Author Grading original paperauthors directlyassess the AI… Measure ResearchGap addresses themeasurement gap forAI's ability to… Engineering-AIDivide agents automateresearchengineering but… From startuphub.ai · The publishers behind this format

The promise of explosive AI progress often hinges on AI agents automating AI research itself. However, tangible evidence supporting this vision remains surprisingly scarce. Existing evaluation methods fall short. Some focus on narrow, verifiable tasks, inherently excluding the open-ended nature of true research. Others submit AI-generated papers to blind peer review, a process already strained, unpredictable, and prone to inconsistent quality.

Shadow Evaluations: A New Metric for AI R&D Automation

To address this measurement gap, researchers Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, and colleagues introduced shadow evaluations. This novel approach places an AI agent at the heart of a high-quality, unpublished research paper, tasking it with tackling the central, open-ended research question. Crucially, the original authors of the paper then grade the agent's output. This method provides a direct, author-centric assessment of AI's research capabilities.

Frontier Agents Fall Short on Research Insight

In their study, the researchers applied shadow evaluations to two unpublished NeurIPS 2026 submissions. Frontier agents were given six days and substantial compute resources, costing thousands of dollars. While the agents successfully completed all engineering tasks without human intervention, they failed to make significant progress on the core research questions. The authors unambiguously rejected the outputs, highlighting five persistent failure modes: inadequate judgment of publishable research standards, uncreative responses to research design flaws, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A subsequent robustness check using a different model and scaffolding reproduced these failures, reinforcing the findings.

The Engineering-AI Divide in Research

The results offer a sobering early assessment: today's most advanced AI agents excel at the engineering components of AI research, coding, implementation, and execution. Yet, they demonstrably struggle with the more nuanced, critical aspects of the research lifecycle. This includes formulating novel hypotheses, designing robust experiments, interpreting ambiguous results, and critically evaluating the significance of findings. The current limitations suggest that while AI can automate the 'how' of research, the 'what' and 'why' remain firmly in the human domain for now. This distinction is vital for investors and founders considering the trajectory of AI-driven innovation.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.