# Ollama open models cost collapse is real _Ollama CEO says AT&T moved 40% of tokens to open models as Chinese flash models and coding agents drive 150x growth._ **Published:** 2026-09-04 **Source:** https://www.startuphub.ai/ai-news/technology/2026/ollama-open-models-cost-collapse-is-real --- Ollama open models cost is collapsing because enterprises are routing the bulk of tokens to cheaper open weights. [YC](https://www.youtube.com/watch?v=rY0wnfFHYbs) host sat with Jeffrey Morgan, co-founder and CEO of Ollama, which he said serves 9 million developers, 178,000 GitHub stars and 85% of the Fortune 500. Morgan's view is blunt. Cost solves the short term problem. Customization is the north star. The full discussion can be found on **YC**'s YouTube channel. ![](https://img.youtube.com/vi/rY0wnfFHYbs/maxresdefault.jpg) Open Models Are Collapsing The Cost Of AI, from YC ## How the Ollama open models cost collapse actually works Think partner and associates. The frontier closed model routes and coordinates, the army of cheap open models does the line work. Morgan said AT&T already shifted 40% of token consumption to open models, mostly US and European weights, with Chinese models under evaluation. Workflows are coding agents first, then co-work agents like [OpenClaw](/startups/openclaw) and Hermes that automate long horizon tasks for finance, support, marketing and sales. The token math explains it. Morgan showed per-developer weekly tokens jumping, with a second inflection in April from [OpenClaw](https://www.startuphub.ai/ai-news/artificial-intelligence/2026/openclaw-s-viral-launch-lessons-for-ai-maintainers), context windows growing from 128K to over 1 million, and total cloud tokens up 150x since the start of the year. Flash class models like DeepSeek Flash are built for volume, good enough for 80% of tasks, fast and ultra cheap, so teams stop rationing tokens and start chaining models. ## Why this matters and what is not fixed For builders, abundance creates a new scarcity above the token. Knowledge, coordination of sub-agents, and execution in sandboxes are now the hard problems. Morgan sees 80 to 90% of enterprise tokens going through open models, but only 10 to 20% of budget, which funds wider access instead of less. Security is the gating factor. Morgan pointed to the GLM 4.5 class capabilities and said safety tooling that closed providers bundle is missing for open weights. Enterprises in the US and Germany will run Chinese origin models if they run in a secure environment they control, but they still screen weights like any other open source dependency. That is solvable, not solved, and startups in governance will matter more than model origin debates. ## Is Ollama open models cost legit? Yes. Morgan ties enterprise shifts directly to price per token and price per task, with AT&T as a named example. The driver is consistent across US and German consumption on [Ollama](https://www.startuphub.ai/startups/ollama-inc) Cloud. ## Is Ollama open models cost safe? Cheaper does not mean automatically safe. Morgan said open models lack the safety stack closed labs ship by default, so buyers must add their own screening, hosting controls and audit. Customers who solve that are comfortable hosting Chinese weights in the US or Europe. ## Is Ollama open models cost worth it? For high volume grunt work it is. Flash models deliver lower latency and far lower cost per task and unlock chaining and routing strategies. Frontier models remain reserved for the hardest tasks. ## Is Ollama open models cost good? It is good enough to power coding agents and co-work agents today. Morgan noted benchmarks like Qwen 3 38B matching Opus 4.6 on coding, runnable on a mid-tier [Mac Studio](https://www.startuphub.ai/hardware/chips/apple-launches-mac-studio-with-m5-ultra-ai-powerhouse) or GB300-class desktop hardware. ## Is Ollama open models cost a scam? No. The cost advantage comes from open weights, competitive inference providers and hardware running 20B to 128B models locally. There is no hidden fee, but you trade managed safety for control. ## Is Ollama open models cost real? Yes, and measurable in token flows. Morgan cited 150x cloud growth this year and a 5x per-user jump around the OpenClaw breakout. Local versus cloud splits prove it is workload dependent, not hype. ## Ollama open models cost pros and cons Pros are lower cost, local control, hybrid cloud and local routing, and better fit for security testing where closed models refuse. Cons are fragmented tooling, day zero launch complexity across engines and harnesses, and the need to build your own safety and memory layers. ## Ollama open models cost verdict Use open models for the 80% of work where speed and volume matter and pay the frontier price for the 20% that needs peak reasoning. The winning architecture Morgan describes is a router, often via [Sakana AI](https://www.startuphub.ai/ai-news/technology/2026/sakana-ai-s-fugu-orchestrates-frontier-models) style orchestration, that mixes both. Open weights did not just get cheap enough. They got useful enough that not using them is the expensive choice. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.