#AI Benchmarks
11 articles with this tag

David Brumley on Teaching AI to Find Real Zero Day Vulnerabilities
David Brumley details how reinforcement learning sandboxes and deterministic graders allow AI models to reliably discover real software vulnerabilities.

OpenAI Triples Benchmark Scores With Simple Settings
OpenAI found simple API setting tweaks tripled scores on the ARC-AGI-3 benchmark, proving harness design significantly impacts AI evaluation.

Sean Cai on the State of AI Data Markets
Sean Cai dissects the AI data market, highlighting fragmentation, the importance of 'type one' data, and the pitfalls of current benchmarks.

Benchmarks Fail Modern AI, Says OpenAI Scientist
OpenAI's Noam Brown discusses why traditional benchmarks fail modern AI, emphasizing the need for new evaluation methods that account for computational budgets and model capabilities.
OpenAI Unveils LifeSciBench
OpenAI's LifeSciBench is a new benchmark designed to test AI's real-world applicability in complex life science research, moving beyond basic question answering.

Task Fidelity Scaling Laws: Kobie Crawford on AI Data Quality
Kobie Crawford of Snorkel discusses 'Task Fidelity Scaling Laws,' emphasizing how data quality impacts AI model performance and outlining Snorkel's approach to creating verifiable datasets.

Exa Unveils New Code Search Benchmarks
Exa.ai releases 'WebCode', a new benchmark suite for evaluating search performance in coding agents, addressing limitations in existing tools.

Poetiq's Small Team Sets New Together AI Benchmark Records
Poetiq's small team is setting new Together AI benchmark records by leveraging recursively self-improving meta-systems to optimize existing LLMs.

Claude Sonnet 4.6 Ups the AI Ante
Anthropic's Claude Sonnet 4.6 launches with major upgrades in coding, reasoning, and computer use, plus a 1M token context window.

Step 3.5 Flash: AI's New Efficiency Standard
Step 3.5 Flash AI model revolutionizes AI efficiency with a 196B parameter foundation and 11B active parameters, offering competitive performance with lower latency.

Claude Opus 4.6: Smarter, Faster, and Longer Context
Anthropic's Claude Opus 4.6 launches with a 1M token context window, enhanced coding, and state-of-the-art benchmark performance.