The frontier of artificial intelligence demands evaluation metrics that transcend academic benchmarks, a critical pivot highlighted by Tejal Patwardhan and Henry Scott-Green of OpenAI. Their presentation at OpenAI DevDay [2025] unveiled a comprehensive framework for measuring real-world AI progress, emphasizing the necessity of robust evaluation for both cutting-edge model development and practical application building. This shift marks a recognition that traditional tests, while foundational, no longer adequately capture the nuanced capabilities required for AI to perform effectively in economically valuable tasks.
Tejal Patwardhan, a researcher on OpenAI's Reinforcement Learning team leading frontier evals, initiated the discussion by underscoring the intrinsic value of evaluation. For OpenAI, where "training runs are very expensive in terms of compute and researcher time," evaluations provide the essential signal to "measure progress" and "steer model training towards good outcomes." Historically, academic benchmarks like the SAT or the AIME high school math competition served to push reasoning capabilities. However, these quickly reached a ceiling; as Tejal noted, "these evals can only measure so much." Models could achieve near-perfect scores on such tests yet remain incapable of performing real-world work. This stark disconnect necessitated a new approach, moving beyond theoretical aptitude to practical utility.
OpenAI's answer to this challenge is GDPval, a novel evaluation designed to measure model performance on economically viable, real-world tasks. The name itself reflects its core purpose, focusing on tasks relevant to Gross Domestic Product (GDP). This framework represents a fundamental re-calibration of how AI is assessed, moving from the simulated environment of academic testing to the complex, multimodal demands of professional work.
To construct GDPval, OpenAI collaborated with experts averaging 14 years of experience across diverse fields. The tasks included in GDPval are notably long-horizon, often requiring days or even weeks to complete, and are inherently multimodal, integrating computer usage, image analysis, and various tools. Examples shared included a real estate agent designing a sales brochure, a manufacturing engineer creating a 3D model of a cable reel stand, and a film editor crafting a high-energy intro reel with video and audio. These tasks were meticulously selected from the top nine sectors contributing to US GDP, and within each sector, the top five knowledge work occupations were identified based on their contribution to overall wages, as per federal and labor statistics. This rigorous process yielded over a thousand real-world tasks, ensuring a broad and relevant measure of AI's practical capabilities.
The evaluation methodology for GDPval employs pairwise expert grading. Human experts, blinded to the source, compare a model's output against a human expert's deliverable for a given task, selecting the preferred outcome. This process generates an overall "win rate," reflecting the percentage of times the model's output is preferred or deemed equally good. Tejal presented compelling data showing that while GPT-4o scored less than 20% win rate against human professionals, subsequent models like GPT-5 high are nearing a 40% win rate. This trajectory suggests that AI models are rapidly approaching parity with human experts on these complex, economically valuable tasks.
