The Math That Changed Everything
Let me walk you through what happened, because the numbers here are genuinely worth your time to understand. In January 2025, DeepSeek released R1, and what made it notable wasn’t the benchmark scores—it was the training cost. According to the DeepSeek R1 Technical Report, the model cost approximately $5.6 million to train. OpenAI’s GPT-4, by contrast, reportedly consumed hundreds of millions in compute resources. That’s not a marginal difference. That’s an order of magnitude shift in what “frontier-grade capability” actually costs to produce.

Now, if you run an AI startup and you’ve been planning your 2025 budget around inference costs that looked reasonable in November 2024, this number probably made your stomach drop. Because it wasn’t just an academic curiosity. It was proof that the unit economics everyone had built their business plans around were suddenly negotiable.
The market knew it immediately. On January 27, 2025, Nvidia lost $593 billion in a single trading day, the largest single-day market cap destruction in U.S. stock market history. That’s not volatility or sentiment. That’s professional investors repricing an entire ecosystem based on the realization that the compute requirements for competitive AI might be materially lower than the industry had collectively priced in.
The Inference Cost Collapse Nobody Fully Anticipated
Here’s where this gets interesting for anyone actually running an AI product: the efficiency gains weren’t confined to training. The a16z State of AI 2025 report tracked something that should have been front-page news in every VC office. Inference costs for frontier models dropped roughly 90 percent year-over-year between 2024 and 2025. Not 9 percent. Ninety.
If you had a startup spending $500,000 per month on inference, which is not an unusual number for a moderately scaled B2B AI application, you were suddenly looking at a world where that same workload could run for $50,000. Or, if you were willing to embrace the open-weight models that DeepSeek’s release legitimized, potentially far less.
That’s not a tactical optimization. That’s a complete reassessment of what margin structure your business can sustain. It’s the difference between venture math that requires you to sell at a 3x markup and venture math that lets you compete on unit economics alone.
The Provider Switching Cascade
The behavioral response was predictable once you understood the incentives. Lightspeed Venture Partners surveyed AI startups in late 2024 and again in Q4 2025. The shift was striking: over 60 percent of the respondents had either switched their primary model provider or were actively evaluating the switch. That’s not gradual portfolio diversification. That’s a wholesale vendor renegotiation driven by raw economics.
This matters because it broke something that had looked stable six months earlier: the assumption that switching costs or lock-in would keep startups tethered to their initial API provider choice. Turns out, when your inference bill is 80 to 90 percent of your cost of goods sold, and you can cut it by an order of magnitude by switching, the switching cost is essentially zero. You move.
OpenAI understood this, which is why they launched o3 in April 2025 priced at $10 per million input tokens. A meaningful price cut from their previous tiers, yes, but also a defensive move. Because comparable open-weight models running on self-hosted infrastructure were already trading at under $0.50 per million tokens. That’s a 20x cost differential. Even accounting for infrastructure overhead and operational complexity, the math pointed decisively toward decentralized, open-weight models for price-sensitive applications.
What This Means for the Business Model Layer
The real consequence here isn’t that inference got cheaper. It’s that the entire SaaS moat structure for AI applications became fragile. If your value proposition was “we wrapped a good model in a nice UI and business logic,” you were suddenly competing against thousands of other teams with identical model access and radically lower infrastructure costs. The barrier to entry didn’t just lower. It collapsed.
This forced a genuine reckoning in the startup ecosystem. Companies that had been profitable on the back of 20x to 50x gross margins on inference suddenly needed to either build genuine differentiation upstream, in data, fine-tuning, domain expertise, or workflow integration, or accept that they were commodity businesses in a commodity market. A lot of startups discovered they were the latter.
The venture community’s response was telling. The startups that had room to survive this transition were the ones solving actual problems with AI, not the ones using AI to solve problems that already had solutions. Those sound like the same thing, but they’re not. One is a business. The other is a feature.
Where This Leaves You
If you’re building something in this space, the question isn’t whether inference costs will stay low. They will. Competition is commoditizing that layer. The question is what you’re building on top of that commodity. Are you providing domain expertise that nobody else has? Are you capturing data that creates compounding advantage? Are you solving a workflow that was genuinely broken before, or just repackaging an existing solution?
The founders who understood this transition early, who saw DeepSeek’s efficiency numbers and immediately started recalculating their unit economics and business models, those teams are going to have an enormous advantage over the next 18 months. Not because they have faster models or cheaper compute, but because they’re organizing their businesses around the reality of what the market actually values now, rather than the fantasy of what it valued in 2024.
Have you seen this transition play out in your own portfolio or product decisions? The pattern is clear at the aggregate level, but I’m curious about the details of how individual teams navigated this. Reach out if you want to compare notes on how your organization recalculated its own assumptions.