OpenAI Flags Major Flaws in SWE-Bench Pro
OpenAI's audit reveals approximately 30% of SWE-Bench Pro's coding tasks are flawed, prompting the company to retract its recommendation for the benchmark.

Visual TL;DR
SWE-bench Verified had fundamental design and contamination problems identified earlier
From the article 2 mentionsOpenAI has identified significant issues within SWE-Bench Pro, a prominent benchmark for evaluating AI coding agents, estimating that roughly 30% of its tasks are broken.
comprehensive datapoint analysis pipeline reviewed model attempts and failure traces
From the article 2 mentionsOpenAI's methodology for this SWE-Bench Pro audit involved a comprehensive datapoint analysis pipeline.
From the article 5 mentionsOpenAI has identified significant issues within SWE-Bench Pro, a prominent benchmark for evaluating AI coding agents, estimating that roughly 30% of its tasks are broken.
From the articleEach flagged task then underwent scrutiny through multiple investigator-agent passes and independent review by five experienced software engineers.
From the article 2 mentionsOf the 731 public split tasks, the analysis pipeline flagged 200 (27.4%) as broken, while human annotation identified 249 (34.1%).
company previously encouraged adoption, now withdraws its support for the benchmark
From the articleGiven the widespread SWE-Bench Pro task issues, OpenAI has retracted its earlier recommendation for the benchmark.
From the articleThis detailed audit, published by OpenAI News, underscores the challenges in creating reliable coding evaluations for advanced AI models.
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.
More from Daniel Singer