OpenAI's AI Scorecard

OpenAI proposes a new 'Useful Intelligence per Dollar' metric to measure AI's true business value, focusing on work accomplished, cost, dependability, and scalability.

Diagram illustrating OpenAI's AI work accomplished scorecard with four key questions.
OpenAI's proposed AI scorecard aims to quantify the value of AI beyond basic metrics.· OpenAI News
Visual TL;DR
CFOs ask: AI value?Driver
traditional software metrics like user adoption no longer suffice for artificial intelligence
From the articleCFOs are asking a fundamental question: how to maximize value from AI investments.
OpenAI's AI ScorecardCore
new approach measuring actual work AI accomplishes, not just tokens generated
From the article 2 mentionsThis is the core of OpenAI's proposed AI work accomplished scorecard.
Useful Intelligence / $Context
ultimate metric combining work done, cost, dependability, and scalability of AI
From the article 3 mentionsThe ultimate metric, dubbed 'Useful Intelligence per Dollar', answers four critical questions: Is AI performing valuable work?
Useful work done?Context
did AI resolve customer issues, ship code, or review contracts effectively
From the articleThe ultimate metric, dubbed 'Useful Intelligence per Dollar', answers four critical questions: Is AI performing valuable work?
Cost per task?Context
calculating the true cost of a successful AI task, beyond just token usage
From the article 3 mentionsAs usage increases, the cost per successful task should ideally decrease, or the value generated should outpace costs.
AI dependability?Context
how often AI gets the work right and its outputs can be relied upon
From the articleDependability is crucial for deeper AI integration.
Scalability value?Context
does each AI dollar buy more work as usage and investment grows
From the article 5 mentionsThe final measure examines AI's economic scalability.
Maximize AI ROIOutcome
helps businesses make informed decisions to maximize value from AI investments
From the articleCFOs are asking a fundamental question: how to maximize value from AI investments.
Contents(5)

CFOs are asking a fundamental question: how to maximize value from AI investments. Traditional software metrics like user adoption no longer suffice for artificial intelligence. A new approach is needed, one that measures the actual work AI accomplishes. This is the core of OpenAI's proposed AI work accomplished scorecard.

Companies working on this

StartupHub profiles of the companies this article names, with funding and a one-liner from our database.

OpenAI
$852.0B
OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity.
Kimi
$2.6B
AI assistant by Moonshot AI capable of processing 2 million Chinese characters in a single prompt.
Perplexity AI
$3K
AI-powered search engine that delivers real-time, cited answers to complex questions.
Cohere
$5K
Cohere offers a secure, all-in-one enterprise AI platform with cutting-edge multilingual models, advanced retrieval, and an AI workspace, enabling businesses to build high-impact applications grounded in proprietary data.

The ultimate metric, dubbed 'Useful Intelligence per Dollar', answers four critical questions: Is AI performing valuable work? What is the true cost of a successful AI task? Can AI outputs be relied upon? Does AI's value increase with usage?

1. How much useful work gets done?

This first pillar focuses on tangible outcomes. Did AI help resolve customer issues, ship code, or review contracts? The value of AI lies not in tokens, but in transforming those tokens into actionable work. For instance, AI can automate preparatory tasks for finance teams, freeing them for higher-level analysis.

2. What does a successful task actually cost?

Calculating the cost of an AI task requires looking beyond per-token pricing. It includes compute, employee time, human review, and retries. A cheaper model might incur higher total costs if it requires more iterations or corrections. OpenAI's OpenAI GPT-5.6 family, with tiers like Sol, Terra, and Luna, aims to offer optimized cost-performance options.

Frontier models can provide better value by delivering accurate results in a single pass, reducing overall expenses.

3. How often does AI get the work right?

Dependability is crucial for deeper AI integration. As AI moves from drafting to taking action, its accuracy and consistency become paramount. Tracking outcomes like 'ready to use', 'needs correction', or 'needs escalation' provides a clearer picture than raw model accuracy.

Defining clear boundaries for AI access and actions is essential for safety, security, and trust.

4. Does each AI dollar buy more work as usage grows?

The final measure examines AI's economic scalability. As usage increases, the cost per successful task should ideally decrease, or the value generated should outpace costs. This efficiency is driven by improvements in compute, algorithms, and model architecture.

This virtuous cycle, where infrastructure improvements fuel research, leading to better models and products, ultimately benefits customers through enhanced capabilities and reduced costs. This is the essence of Useful Intelligence per Dollar in practice.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer