# AI Agents Stall on Core AI Research _Frontier AI agents can automate AI research engineering but fail to make substantial progress on core research questions, according to new shadow evaluations._ **Published:** 2026-07-30 **Source:** https://www.startuphub.ai/ai-news/ai-research/2026/ai-agents-stall-on-core-ai-research --- The promise of explosive AI progress often hinges on AI agents automating AI research itself. However, tangible evidence supporting this vision remains surprisingly scarce. Existing evaluation methods fall short. Some focus on narrow, verifiable tasks, inherently excluding the open-ended nature of true research. Others submit AI-generated papers to blind peer review, a process already strained, unpredictable, and prone to inconsistent quality. AI Agents StallDriver frontier AI agents fail to make substantial progress on core research questionsFrom the article 6 mentionsThe promise of explosive AI progress often hinges on AI agents automating AI research itself.due toExisting Evals FlawedDrivercurrent methods are narrow, exclude open-ended research, or use strained peer reviewleads toShadow EvaluationsCorenovel approach places AI agent at heart of unpublished paper's research questionFrom the article 2 mentionsTo address this measurement gap, researchers Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, and colleagues introduced shadow evaluations.Author GradingContextoriginal paper authors directly assess the AI agent's output for research capabilityFrom the article 2 mentionsCrucially, the original authors of the paper then grade the agent's output.Measure Research GapEffectaddresses the measurement gap for AI's ability to automate core AI researchrevealsEngineering-AI DivideOutcomeagents automate research engineering but lack insight for core research problems ## Shadow Evaluations: A New Metric for AI R&D Automation To address this measurement gap, researchers Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, and colleagues introduced [shadow evaluations](https://arxiv.org/abs/2607.27191v1). This novel approach places an AI agent at the heart of a high-quality, unpublished research paper, tasking it with tackling the central, open-ended research question. Crucially, the original authors of the paper then grade the agent's output. This method provides a direct, author-centric assessment of AI's research capabilities. ## Frontier Agents Fall Short on Research Insight In their study, the researchers applied [shadow evaluations](/ai-news/claude) to two unpublished NeurIPS 2026 submissions. Frontier agents were given six days and substantial compute resources, costing thousands of dollars. While the agents successfully completed all engineering tasks without human intervention, they failed to make significant progress on the core research questions. The authors unambiguously rejected the outputs, highlighting five persistent failure modes: inadequate judgment of publishable research standards, uncreative responses to research design flaws, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A subsequent robustness check using a different model and scaffolding reproduced these failures, reinforcing the findings. ## The Engineering-AI Divide in Research The results offer a sobering early assessment: today's most advanced AI agents excel at the engineering components of AI research, coding, implementation, and execution. Yet, they demonstrably struggle with the more nuanced, critical aspects of the research lifecycle. This includes formulating novel hypotheses, designing robust experiments, interpreting ambiguous results, and critically evaluating the significance of findings. The current limitations suggest that while AI can automate the 'how' of research, the 'what' and 'why' remain firmly in the human domain for now. This distinction is vital for investors and founders considering the trajectory of AI-driven innovation. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.