Data Curve Launches DeepSWE Coding Benchmark
Data Curve introduces DeepSWE, a contamination-resistant coding benchmark designed to better evaluate AI coding agents on realistic, long-horizon software engineering tasks.

Visual TL;DR
From the articleSpeaking at the AI Engineer World's Fair, Shi highlighted the shortcomings of existing benchmarks, such as SweetBench Pro, which he claims suffer from data contamination, brittle verification methods, and a tendency for top models to cluster together, making differentiation difficult.
From the article 9+ mentionsJames Shi, a founding engineer at Data Curve, presented DeepSWE, a new benchmark designed to more accurately measure the capabilities of AI coding agents.
113 original, long-horizon tasks, specifically authored, not scraped
From the article 2 mentionsContamination: Tasks are often mined from public pull requests, meaning solutions, tests, and discussions are readily available to AI agents.
includes TypeScript, expanding utility beyond single-language focus
From the articleThe benchmark supports multiple languages, including TypeScript, JavaScript, Python, Rust, and Go, with plans to add more.
From the article 2 mentionsJames Shi, a founding engineer at Data Curve, presented DeepSWE, a new benchmark designed to more accurately measure the capabilities of AI coding agents.
draws from nearly 100 repositories, median one task per repository
From the article 5 mentionsShi explained that DeepSWE comprises 113 original, long-horizon software engineering tasks, specifically authored rather than scraped from existing projects.
more accurately measures capabilities of AI coding agents on realistic tasks
Contents(6)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.
More from Daniel Singer