#AI Benchmarks

11 articles with this tag

David Brumley on Teaching AI to Find Real Zero Day Vulnerabilities
Cybersecurity

David Brumley on Teaching AI to Find Real Zero Day Vulnerabilities

David Brumley details how reinforcement learning sandboxes and deterministic graders allow AI models to reliably discover real software vulnerabilities.

11 days ago
OpenAI Triples Benchmark Scores With Simple Settings
Artificial Intelligence

OpenAI Triples Benchmark Scores With Simple Settings

OpenAI found simple API setting tweaks tripled scores on the ARC-AGI-3 benchmark, proving harness design significantly impacts AI evaluation.

13 days ago
Sean Cai on the State of AI Data Markets
Artificial Intelligence

Sean Cai on the State of AI Data Markets

Sean Cai dissects the AI data market, highlighting fragmentation, the importance of 'type one' data, and the pitfalls of current benchmarks.

17 days ago
Benchmarks Fail Modern AI, Says OpenAI Scientist
AI Research

Benchmarks Fail Modern AI, Says OpenAI Scientist

OpenAI's Noam Brown discusses why traditional benchmarks fail modern AI, emphasizing the need for new evaluation methods that account for computational budgets and model capabilities.

about 2 months ago
OpenAI Unveils LifeSciBench
Artificial Intelligence

OpenAI Unveils LifeSciBench

OpenAI's LifeSciBench is a new benchmark designed to test AI's real-world applicability in complex life science research, moving beyond basic question answering.

about 2 months ago
Task Fidelity Scaling Laws: Kobie Crawford on AI Data Quality
AI Research

Task Fidelity Scaling Laws: Kobie Crawford on AI Data Quality

Kobie Crawford of Snorkel discusses 'Task Fidelity Scaling Laws,' emphasizing how data quality impacts AI model performance and outlining Snorkel's approach to creating verifiable datasets.

2 months ago
Exa Unveils New Code Search Benchmarks
Artificial Intelligence

Exa Unveils New Code Search Benchmarks

Exa.ai releases 'WebCode', a new benchmark suite for evaluating search performance in coding agents, addressing limitations in existing tools.

5 months ago
Poetiq's Small Team Sets New Together AI Benchmark Records
Technology

Poetiq's Small Team Sets New Together AI Benchmark Records

Poetiq's small team is setting new Together AI benchmark records by leveraging recursively self-improving meta-systems to optimize existing LLMs.

6 months ago
Claude Sonnet 4.6 Ups the AI Ante
Artificial Intelligence

Claude Sonnet 4.6 Ups the AI Ante

Anthropic's Claude Sonnet 4.6 launches with major upgrades in coding, reasoning, and computer use, plus a 1M token context window.

6 months ago
Step 3.5 Flash: AI's New Efficiency Standard
AI Research

Step 3.5 Flash: AI's New Efficiency Standard

Step 3.5 Flash AI model revolutionizes AI efficiency with a 196B parameter foundation and 11B active parameters, offering competitive performance with lower latency.

6 months ago
Claude Opus 4.6: Smarter, Faster, and Longer Context
Artificial Intelligence

Claude Opus 4.6: Smarter, Faster, and Longer Context

Anthropic's Claude Opus 4.6 launches with a 1M token context window, enhanced coding, and state-of-the-art benchmark performance.

6 months ago