Ollama open models cost collapse is real

Ollama CEO says AT&T moved 40% of tokens to open models as Chinese flash models and coding agents drive 150x growth.

S
StartupHub.ai Staff
4 min read
Ollama CEO Jeffrey Morgan discusses open models collapsing AI costs on YC Lightcone
Ollama says open models now drive 150x token growth led by coding agents· YC
Contents(11)

Ollama open models cost is collapsing because enterprises are routing the bulk of tokens to cheaper open weights. YC host sat with Jeffrey Morgan, co-founder and CEO of Ollama, which he said serves 9 million developers, 178,000 GitHub stars and 85% of the Fortune 500.

StartupHub data

Companies working on this

Profiles of the companies named in this story, with funding and a one-liner from our database.

DeepSeek
$45.0B
Sakana AI
$2.6B
Japan-focused efficient AI models using bio-inspired evolution to build capable, compact foundation models.
AT&T
A global telecommunications conglomerate leveraging AI for operational efficiency and cost reduction.

Morgan's view is blunt. Cost solves the short term problem. Customization is the north star.

The full discussion can be found on YC's YouTube channel.

Open Models Are Collapsing The Cost Of AI - YC
Open Models Are Collapsing The Cost Of AI, from YC

How the Ollama open models cost collapse actually works

Think partner and associates. The frontier closed model routes and coordinates, the army of cheap open models does the line work.

Morgan said AT&T already shifted 40% of token consumption to open models, mostly US and European weights, with Chinese models under evaluation. Workflows are coding agents first, then co-work agents like OpenClaw and Hermes that automate long horizon tasks for finance, support, marketing and sales.

The token math explains it. Morgan showed per-developer weekly tokens jumping, with a second inflection in April from OpenClaw, context windows growing from 128K to over 1 million, and total cloud tokens up 150x since the start of the year. Flash class models like DeepSeek Flash are built for volume, good enough for 80% of tasks, fast and ultra cheap, so teams stop rationing tokens and start chaining models.

Why this matters and what is not fixed

For builders, abundance creates a new scarcity above the token. Knowledge, coordination of sub-agents, and execution in sandboxes are now the hard problems. Morgan sees 80 to 90% of enterprise tokens going through open models, but only 10 to 20% of budget, which funds wider access instead of less.

Security is the gating factor. Morgan pointed to the GLM 4.5 class capabilities and said safety tooling that closed providers bundle is missing for open weights. Enterprises in the US and Germany will run Chinese origin models if they run in a secure environment they control, but they still screen weights like any other open source dependency. That is solvable, not solved, and startups in governance will matter more than model origin debates.

Is Ollama open models cost legit?

Yes. Morgan ties enterprise shifts directly to price per token and price per task, with AT&T as a named example. The driver is consistent across US and German consumption on Ollama Cloud.

Is Ollama open models cost safe?

Cheaper does not mean automatically safe. Morgan said open models lack the safety stack closed labs ship by default, so buyers must add their own screening, hosting controls and audit. Customers who solve that are comfortable hosting Chinese weights in the US or Europe.

Is Ollama open models cost worth it?

For high volume grunt work it is. Flash models deliver lower latency and far lower cost per task and unlock chaining and routing strategies. Frontier models remain reserved for the hardest tasks.

Is Ollama open models cost good?

It is good enough to power coding agents and co-work agents today. Morgan noted benchmarks like Qwen 3 38B matching Opus 4.6 on coding, runnable on a mid-tier Mac Studio or GB300-class desktop hardware.

Is Ollama open models cost a scam?

No. The cost advantage comes from open weights, competitive inference providers and hardware running 20B to 128B models locally. There is no hidden fee, but you trade managed safety for control.

Is Ollama open models cost real?

Yes, and measurable in token flows. Morgan cited 150x cloud growth this year and a 5x per-user jump around the OpenClaw breakout. Local versus cloud splits prove it is workload dependent, not hype.

Ollama open models cost pros and cons

Pros are lower cost, local control, hybrid cloud and local routing, and better fit for security testing where closed models refuse. Cons are fragmented tooling, day zero launch complexity across engines and harnesses, and the need to build your own safety and memory layers.

Ollama open models cost verdict

Use open models for the 80% of work where speed and volume matter and pay the frontier price for the 20% that needs peak reasoning. The winning architecture Morgan describes is a router, often via Sakana AI style orchestration, that mixes both.

Open weights did not just get cheap enough. They got useful enough that not using them is the expensive choice.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
S

Written by

StartupHub.ai Staff

Editorial team

The staff writers of StartupHub.ai, ranging from investment analysts to avid AI tool users, early adopters and critical enthusiasts. Backgrounds span engineering, business and the arts. We hold every piece to rigorous standards of research and review.