Ironclad Rethinks AI Token Spend ROI

Ironclad replaces token leaderboard culture with trusted throughput, measuring merged PRs and review quality to raise ROI on AI token spend.

6 min read
Engineer reviewing AI-generated pull requests and CI metrics on screens
Ironclad frames dashboards as smoke detectors and tracks trusted throughput across review and CI.· AI Engineer
Visual TL;DR
Token leaderboard cultureDriver
Amazon and Meta dashboards where engineers competed to top token usage rankings
Trusted throughput proxyContext
Merged PRs and review quality replace raw token counts as ROI measure
From the articleIronclad will not chase more tokens and is measuring AI token spend against a new proxy it calls trusted throughput, according to AI Engineer.
Legal AI trust ladderCore
From the article 2 mentionsIronclad builds AI for legal contracting, where lawyers test output on known contracts before they trust it on redlines and anomalies.
Cost tracking as smoke detectorEffect
Per-team cost tracking identifies adoption gaps and bursts without competitive pressure
From the articleHis fix is to keep per-team and per-person cost tracking but to treat it as a smoke detector.
Controlled AI spendOutcome
Ironclad measures token spend against trusted throughput raising ROI
From the article 2 mentionsThe goal is not to minimize spend but to raise return per token.
Token leaderboard cultureDriver
Amazon and Meta dashboards where engineers competed to top token usage rankings
Mingsheng Hong frameworkCore
VP of AI Engineering presented trusted throughput model at AI Engineer conference
From the articleMingsheng Hong, VP of Engineering for AI at Ironclad, laid out the framework at AI Engineer after the contract lifecycle company pushed past broad adoption over the last two quarters.
Burst detection triggersEffect
Sudden usage spikes trigger cost investigation not punishment
From the articleLow-use pockets signal an adoption gap to investigate, bursts trigger a check on legitimacy, and teams are compared only in context because platform and UI teams extract value differently.
Trusted throughput proxyContext
Merged PRs and review quality replace raw token counts as ROI measure
From the articleIronclad will not chase more tokens and is measuring AI token spend against a new proxy it calls trusted throughput, according to AI Engineer.
Legal AI trust ladderCore
From the article 2 mentionsIronclad builds AI for legal contracting, where lawyers test output on known contracts before they trust it on redlines and anomalies.
Cost tracking as smoke detectorEffect
Per-team cost tracking identifies adoption gaps and bursts without competitive pressure
From the articleHis fix is to keep per-team and per-person cost tracking but to treat it as a smoke detector.
Controlled AI spendOutcome
Ironclad measures token spend against trusted throughput raising ROI
From the article 2 mentionsThe goal is not to minimize spend but to raise return per token.
Contents(4)

Ironclad will not chase more tokens and is measuring AI token spend against a new proxy it calls trusted throughput, according to AI Engineer.

Ironclad Rethinks AI Token Spend ROI - AI Engineer
Ironclad Rethinks AI Token Spend ROI, from AI Engineer

Mingsheng Hong, VP of Engineering for AI at Ironclad, laid out the framework at AI Engineer after the contract lifecycle company pushed past broad adoption over the last two quarters.

Ironclad builds AI for legal contracting, where lawyers test output on known contracts before they trust it on redlines and anomalies. That same trust ladder now shapes how it judges engineering output.

Why a dashboard is not a leaderboard

Hong cited a voluntary dashboard at Amazon (NASDAQ:AMZN) where engineers competed to top a token usage board, a pattern he said also appeared at Meta Platforms (NASDAQ:META), plus a sensational claim of a company burning $500 million on cloud in a month.

His fix is to keep per-team and per-person cost tracking but to treat it as a smoke detector. Low-use pockets signal an adoption gap to investigate, bursts trigger a check on legitimacy, and teams are compared only in context because platform and UI teams extract value differently.

What trusted throughput actually measures

Ironclad evolved from lines of code to open pull requests to merged pull requests, then added a complexity weight using a prompt that asks one or two models to score each merged PR by t-shirt size.

The qualitative layer has three buckets. Objective gates include test coverage, security checks and canarying, subjective review covers clarity, maintainability and architecture fit, and customer validation looks at rollbacks, fires and tickets for bugs or friction.

Where the bottleneck moved

Generation is now abundant, so pressure has shifted to review and continuous integration. Hong warned against the workaround of shipping large PRs to avoid hour-long CI runs, which thins reviewer attention and hurts quality.

Ironclad puts AI review first as a filter for style and missing tests before human reviewers apply judgment on design and security. For CI, it tracks time from ready-to-submit to submitted and retry counts for flaky tests, and staffs platform work to fix the infrastructure instead of having engineers or agents babysit merges.

How Ironclad controls cost without austerity

The goal is not to minimize spend but to raise return per token. The pragmatic framework sets budgets, quotas and anomaly alerts alongside regular human review, with a learning loop that feeds findings back into institutional best practices.

Team-level practices include capping agentic loops that auto-fix tests, structuring prompts with stable system prefixes first to benefit from prompt caching, and pruning or compacting context in long chats, including with tools like Claude Code and Codex that now do it automatically. StartupHub.ai data shows ROI raised $10M in a Series A in 2026, a useful anchor as private legal-tech peers try to prove that token efficiency scales to real throughput.

Build versus buy is explicit. Ironclad buys non-differentiating IDE and CI infrastructure and builds its internal playbook of prompts for bug fixes, UI features and refactors, while still evaluating whether to wrap Claude Code in a cloud builder agent or buy a vendor alternative.

Token budgets are now an engineering management problem, not a vendor billing detail. Ironclad is betting the teams that instrument review and CI will ship more trusted code per dollar than those that just cap prompts.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.