SWE-Marathon: Evaluating AI Coding Agents at Scale
Rishi Desai from Abundant AI introduces SWE-Marathon, a benchmark evaluating AI coding agents on billion-token scale tasks, revealing current limitations and the need for robust verification.

Visual TL;DR
from isolated tasks to full-scale projects
From the article 4 mentionsRishi Desai, an ML engineer at Abundant AI, presents SWE-Marathon, a benchmark designed to evaluate AI coding agents on tasks requiring long-horizon reasoning and coherence.
pioneering benchmarks for AI coding agent evaluation
From the article 3 mentionsHe highlights examples from leading AI labs and companies, such as Anthropic building a C compiler, OpenAI's 'Parameter Golf' experiment, Cloudflare's rapid Next.js rebuild, and Cursor's work on autonomous coding.
evaluates AI on billion-token scale tasks
From the article 9+ mentionsThe SWE-Marathon benchmark aims to measure an agent's ability to perform multi-step tasks over extended periods, simulating real-world software engineering workflows.
tests agent focus and functionality over vast codebases
From the article 5 mentionsRishi Desai, an ML engineer at Abundant AI, presents SWE-Marathon, a benchmark designed to evaluate AI coding agents on tasks requiring long-horizon reasoning and coherence.
AI struggles with complex, extended software engineering projects
From the articleThis research delves into the current limitations of AI in handling complex software engineering projects, moving beyond simple code completion or bug fixing.
crucial for ensuring correctness in complex AI code
advances in autonomous coding and agent capabilities
Contents(6)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.