#Evaluation
13 articles with this tag

Google Experts Share AI Agent Evaluation Best Practices
Google's Preetika Bhateja & Daniel Bump share essential strategies for building effective AI agent evaluation systems, from initial 'vibing' to scaling with LLM judges.

Lyft's Nick Ung on Building Better AI Evals
Lyft's Nick Ung and Ashe discuss building effective AI agent evaluations, emphasizing realistic user simulation, actionable metrics, and statistical rigor.

Langfuse: Domain Expertise Crucial for AI Self-Improvement
Langfuse's Annabelle Schäfer explains why domain expertise is crucial for AI self-improvement, advocating for high-signal target functions and expert-driven data.

Dat Ngo on Arize: LLM Observability Platform
Dat Ngo from Arize AI explains their LLM observability, evaluation, and experimentation platform, crucial for building robust GenAI applications.

LLM Evaluators: Beyond Naive Judgments
Mahmoud Malaeb of Argenta discusses the limitations of naive LLM judges and introduces GEPA, an optimization framework for building more accurate LLM evaluators using a data flywheel approach.

The Hidden Cost of Autonomy: AI Agent Evaluation

OpenAI Says Business AI Evaluation Is the Key to ROI

Evals Reimagined: Braintrust's Engineering Approach to AI Development

Building Reliable AI: The Imperative of Application-Layer Evals

DeepMind Proposes Radical Shift in AI Intelligence Benchmarking
Google DeepMind has unveiled a significant new initiative aimed at fundamentally rethinking how artificial intelligence capabilities are measured. In an announcement on its blog, the leading AI research institution detailed a comprehensive framework designed to...

The Unseen Challenge of Reliable AI

The State of AI Engineering: Insights from Amplify's 2025 Report with Barr Yaron
The State of AI Engineering: Insights from Amplify\'s 2025 Report with Barr Yaron
\"Evaluation/evals\" stands as the single most painful aspect of AI Engineering today, a stark revelation from Amplify Partners\' recent 2025 AI Engineering...