The Daily Read·Y Combinator · Sep 7, 2026

Y Combinator · The YC Podcast

Open Models Change The Economics of AI

Ollama CEO Jeffrey Morgan explains why open models are winning the deployment war, how coding agents broke the frontier-pricing model, and why the bundling cycle is splitting AI infrastructure in two.

thumbnail
57 min
Runtime
67K
Views
9M
Ollama devs
Sep 4
Published

Tap a timestamp pill below to jump the video to that moment.

AI costs are trapping companies in someone else’s ecosystem

Ollama reaches 9 million developers and runs inside 85% of Fortune 500 companies, which means Jeffrey Morgan, its co-founder and CEO, sees the token-level economics of AI deployment more clearly than almost anyone. What he sees right now is a growing revolt against frontier pricing—not a fringe revolt, but a structural one.

The immediate driver is cost. When a product routes every user action through GPT-4 or Claude, margins compress fast. Morgan is blunt about it: cost is the single biggest pain point that open models are positioned to solve. Enterprises are running the numbers, and the numbers are bad enough that switching is worth the friction.

cost is by far the largest pain point that open models can jump in and solve…
cost is by far the largest pain point that open models can jump in and solve. But you know, every business has a vision of getting better control over AI and customizing it for their business. And that’s really their north star. You know, cost is something they can solve in the short term, but it’s the control and customization that keeps them here long term.

But cost is the entry drug, not the full story. Behind the headline number is a structural concern that executives are naming more openly: vendor lock-in. The frontier labs are building walled gardens—their managed agents, their memory layers, their context windows. Every token routed through them makes escape harder. Companies are noticing, and open models are starting to look less like a compromise and more like the sensible default.

Coding agents broke the pricing model, and open infrastructure filled the gap

The demand surge for open models didn’t build slowly. It ignited around 2026, and the fuse was coding agents. When an agent calls a model hundreds of times per user session, cost doesn’t scale linearly—it multiplies. Companies building on frontier APIs discovered that what worked for occasional prompts became economically unsustainable the moment usage became agentic.

driven by coding agents and also AI assistants more co-work cases…
driven by coding agents and also AI assistants more co-work cases like openclaw and Hermes >> and because you sit in the token flow of like so many tokens you have really good data on what models people are actually using and how it’s changing. What are the trends that you’re seeing? Yeah, you know, the biggest shift he sees is toward open models.

The capability gap that once justified frontier pricing is closing fast. Models like DeepSeek arrived with competitive performance at a fraction of the cost, and the Ollama Cloud data tells the story: Morgan describes looking at the usage graph and seeing the Ollama story appear to begin in February 2026 and “just explode thereafter.” Open models in 2024–2025 were mostly fine-tuned custom builds; by 2026, off-the-shelf open models had become the default for a growing share of production workloads.

surge in demand for open models right whereas in 2024 2025 they were mostly custom fine-tunes…
surge in demand for open models right whereas open models I’d say in 2024 2025 from the large models being served were mostly being served as custom models so you’d take an off-the-shelf model like deepseek or kimmy and you’d fine-tune it for your use case um like for example you know cursor had famously done this to build their cursor model.

Morgan frames the longer dynamic as an inevitable unbundling cycle—the same pattern that played out in enterprise software for decades. The frontier labs want total ownership: agents, memory, context, billing, everything in one stack. But the developers building on top of that stack have a different instinct. They want portability, transparency, and the right to swap components. The tension is structural, and it keeps resolving in favour of open tooling.

frontier model labs want you totally in their walled garden…
are just bundling or unbundling like the frontier model frontier model labs want you to be totally in their walled garden of manage agents and their context their memory layer. And then meanwhile little tech and all the founders out there and all the open source developers uh don’t want to be caged in that walled garden.

Practically, Ollama’s value proposition is the same as Docker’s was for containers: make the local setup so straightforward that it stops being a barrier. Pull a model, run it, switch it out. That framing lowered the threshold enough that individual developers could experiment without cloud spend, and enterprises could justify running sensitive workloads on-prem without building custom infrastructure from scratch.

The quick version

  • Cost is the entry point for open models, but control and customisation are what keep companies there long-term.
  • Coding agents are the single biggest driver of adoption—per-session token consumption made frontier pricing untenable at scale.
  • The bundling/unbundling cycle is active: labs want walled gardens, but open tooling keeps winning developer trust through portability.
  • Open model capabilities have converged enough that the “capability premium” justification for frontier pricing is weakening fast.
“Cost is by far the largest pain point that open models can jump in and solve.”— Jeffrey Morgan, Ollama / Y Combinator

The era of AI-as-utility is arriving on its own infrastructure. Ollama’s trajectory suggests open models won’t just compete with frontier labs—they’ll redefine what “deploying AI” means for the next decade.