Visual TL;DR. AI Agents Stall due to Existing Evals Flawed. Existing Evals Flawed leads to Shadow Evaluations. Shadow Evaluations involves Author Grading. Shadow Evaluations helps Measure Research Gap. Author Grading reveals Engineering-AI Divide. AI Agents Stall confirms Engineering-AI Divide.
- AI Agents Stall: frontier AI agents fail to make substantial progress on core research questions
- Existing Evals Flawed: current methods are narrow, exclude open-ended research, or use strained peer review
- Shadow Evaluations: novel approach places AI agent at heart of unpublished paper's research question
- Author Grading: original paper authors directly assess the AI agent's output for research capability
- Measure Research Gap: addresses the measurement gap for AI's ability to automate core AI research
- Engineering-AI Divide: agents automate research engineering but lack insight for core research problems
Visual TL;DR
