Inside the Uber Software Factory Scale Play
Uber runs 70% of pull requests via agents and cut cost per session 52% by capping context, tuning cache TTL and routing subagents to cheaper models.
8 min read

Visual TL;DR
Engineers built over 3,600 agent skills for engineers across the company
From the article 9+ mentionsThe Uber Software Factory now handles more than 70% of pull requests via agents, according to Uber Engineering.
Cost per 1,000 requests fell 34% from peak during optimization window
From the article 6 mentionsCost per 1,000 requests fell almost 34% from its peak.
From the articleThe Uber Software Factory now handles more than 70% of pull requests via agents, according to Uber Engineering.
Total AI spend stabilized since April despite 9.4x growth in requests
From the article 2 mentionsThe middle three are where Uber spends its effort.
Trimming context length cut cost per session 52% from June peak
From the articleUber's lesson is portable: benchmark real work, pick Pareto winners, cap context, fix TTL to human idle time, and hide tool schemas until needed.
Engineers built over 3,600 agent skills for engineers across the company
From the article 9+ mentionsThe Uber Software Factory now handles more than 70% of pull requests via agents, according to Uber Engineering.
From the articleThe Uber Software Factory now handles more than 70% of pull requests via agents, according to Uber Engineering.
Holding one model constant from February to July isolated real optimization gains
From the article 3 mentionsUber builds benchmarks from real work, runs every model behind one harness, and moves to whatever is cheapest at equal quality.
Trimming context length cut cost per session 52% from June peak
From the articleUber's lesson is portable: benchmark real work, pick Pareto winners, cap context, fix TTL to human idle time, and hide tool schemas until needed.
From the article 3 mentionsWeekly active users across agentic offerings grew 7x from February to mid-August.
Aggressive cache time-to-live settings reused tokens across requests
Cheaper models handle subagent tasks while flagship models lead reasoning
From the article 5 mentionsUber now routes 1,000+ MCP servers through a gateway and exposes them as CLI commands that resolve at call time, plus tool search that loads only needed tools.
Cost per 1,000 requests fell 34% from peak during optimization window
From the article 6 mentionsCost per 1,000 requests fell almost 34% from its peak.
Total AI spend stabilized since April despite 9.4x growth in requests
From the article 2 mentionsThe middle three are where Uber spends its effort.
Contents(3)
© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.