#AI Benchmarking

13 articles with this tag

Kimi K3 Challenges Claude Fable 5 on Code Quality, Slashes Cost
Technology

Kimi K3 Challenges Claude Fable 5 on Code Quality, Slashes Cost

Kimi K3 challenges Claude Fable 5 on coding benchmarks, offering similar quality at a third of the cost and the benefits of an open-weight model.

11 days ago
Alejandro Vidal on Rethinking LLM Evaluation with Psychometrics
AI Research

Alejandro Vidal on Rethinking LLM Evaluation with Psychometrics

Alejandro Vidal of Mindmakers advocates for integrating psychometrics into LLM evaluation to move beyond simplistic accuracy scores and gain deeper insights into model intelligence and benchmark quality.

23 days ago
Benchmarking AI Agents: Snorkel AI's Vincent Chen Explains
AI Research

Benchmarking AI Agents: Snorkel AI's Vincent Chen Explains

Vincent Chen from Snorkel AI explores the art and science of benchmarking AI agents, detailing the complexities and methodologies involved in evaluation.

2 months ago
AI Analysts Lag on Real-World Reasoning
AI Research

AI Analysts Lag on Real-World Reasoning

New Hedge-Bench 1.0 benchmark reveals frontier AI models score under 16% on real-world financial reasoning tasks, exposing a critical gap in expert-level judgment.

2 months ago
Google DeepMind Tackles AI Evaluation Challenges
AI Research

Google DeepMind Tackles AI Evaluation Challenges

Google DeepMind's Nicholas Kang and Michael Aaron discuss the challenges in current AI evaluation and Kaggle's innovative solutions like Hackathons, Agent Exams, and Game Arena.

2 months ago
LinkedIn Tries Real-World AI Benchmarking
tech

LinkedIn Tries Real-World AI Benchmarking

LinkedIn's new Crosscheck platform aims to provide real-world AI model performance insights tailored to professional roles and tasks.

2 months ago
AI's Discovery-to-Application Bottleneck
AI Research

AI's Discovery-to-Application Bottleneck

A new Minecraft benchmark, SciCrafter, reveals frontier AI models plateau at 26% success on causal discovery, highlighting a shift in bottlenecks from problem-solving to problem-raising.

3 months ago
Microsoft's AsgardBench Tests AI's Planning Skills
AI Research

Microsoft's AsgardBench Tests AI's Planning Skills

Microsoft's AsgardBench benchmark tests AI agents' ability to adapt plans using real-time visual feedback, revealing current limitations in perception and state tracking.

4 months ago
Anthropic's Claude 4.6 Found to 'Crack' Benchmarks
AI Research

Anthropic's Claude 4.6 Found to 'Crack' Benchmarks

Anthropic's latest research reveals that Claude Opus 4.6 can detect and exploit "contamination" in AI benchmarks, raising concerns about evaluation integrity.

5 months ago
Engineering AI Prompts: Google's Framework for Benchmarking and Automation
AI Video

Engineering AI Prompts: Google's Framework for Benchmarking and Automation

9 months ago
Qwen-Image-Edit Challenges Image Generation Landscape
AI Video

Qwen-Image-Edit Challenges Image Generation Landscape

10 months ago
Press Release

VERSES® Digital Brain Beats Google’s Top AI At “Gameworld 10k” Atari Challenge

about 1 year ago
Funding Round

LM Arena Secures $100 Million Seed Funding

about 1 year ago