#Benchmarks

12 articles with this tag

OpenAI Flags Major Flaws in SWE-Bench Pro
Artificial Intelligence

OpenAI Flags Major Flaws in SWE-Bench Pro

OpenAI's audit reveals approximately 30% of SWE-Bench Pro's coding tasks are flawed, prompting the company to retract its recommendation for the benchmark.

24 days ago
SWE-Marathon: Evaluating AI Coding Agents at Scale
AI Research

SWE-Marathon: Evaluating AI Coding Agents at Scale

Rishi Desai from Abundant AI introduces SWE-Marathon, a benchmark evaluating AI coding agents on billion-token scale tasks, revealing current limitations and the need for robust verification.

28 days ago
OpenAI Unveils Genebench-Pro Benchmark
Artificial Intelligence

OpenAI Unveils Genebench-Pro Benchmark

OpenAI's new Genebench-Pro benchmark rigorously tests AI models on 10 complex, real-world genomics case studies.

about 1 month ago
Personalized AI Agents Now Have a Benchmark
AI Research

Personalized AI Agents Now Have a Benchmark

A new iOSWorld benchmark reveals AI agents' struggles with personalized, multi-app tasks, highlighting the need for richer context and advanced reasoning capabilities.

about 2 months ago
Evaluating Coding Agents: Lessons from SWE-rebench
AI Research

Evaluating Coding Agents: Lessons from SWE-rebench

Ibragim Badertdinov from Nebius shares key lessons from evaluating coding agents using the SWE-rebench benchmark, highlighting the importance of real-world tasks, reliable verification, and cost-effectiveness.

2 months ago
Active Exploration Unlocks Spatial AI
AI Research

Active Exploration Unlocks Spatial AI

New benchmark ESI-BENCH reveals active exploration is key to embodied spatial intelligence, exposing AI's 'action blindness' and metacognitive gaps.

3 months ago
Unlocking AI Agents with Gym-Anything
AI Research

Unlocking AI Agents with Gym-Anything

Gym-Anything enables scalable creation of complex AI agent environments, leading to the vast CUA-World benchmark and more efficient VLM agents.

4 months ago
François Chollet on ARC-AGI-3: The Future of AI Reasoning
AI Research

François Chollet on ARC-AGI-3: The Future of AI Reasoning

François Chollet discusses ARC-AGI-3, a new benchmark for AI reasoning, highlighting current AI's limitations and the path toward general intelligence.

4 months ago
AI Coding Benchmark Scores Skewed by Infrastructure
Artificial Intelligence

AI Coding Benchmark Scores Skewed by Infrastructure

Infrastructure configuration, not just AI model prowess, can significantly skew benchmark results, complicating deployment decisions.

4 months ago
LMArena Series A lands $150M to standardize AI evaluation
Funding Round

LMArena Series A lands $150M to standardize AI evaluation

7 months ago
Anthropic Wins TTFT, But OpenAI Dominates LLM Benchmarks
Market Research

Anthropic Wins TTFT, But OpenAI Dominates LLM Benchmarks

8 months ago
NeuroDiscoveryBench Sets New Standard for Neuroscience AI Benchmarks
AI Research

NeuroDiscoveryBench Sets New Standard for Neuroscience AI Benchmarks

8 months ago