The Cost of Intelligence Is Falling Toward Zero: What That Means for the Future of AI
- David Borish

- 3 days ago
- 5 min read

Tech giants have committed hundreds of billions of dollars to the computing infrastructure behind the AI boom. The intelligence that infrastructure produces is getting cheaper anyway, and the pace of that decline picked up sharply in July.
DeepSeek released V4 Flash on July 31, a coding-focused model that performs close to Claude Opus 4.8 on tests of complex coding and autonomous software tasks, according to reporting from Axios. On Arena.ai's crowdsourced leaderboard for front-end coding, V4 Flash debuted ahead of Opus 4.8 while charging a fraction of the price. DeepSeek bills about 28 cents per million output tokens for output that costs $25 per million on Opus 4.8, a discount of roughly 99%.
That single data point would be notable on its own. It is also a repeat performance. DeepSeek is the same Chinese startup that triggered a market selloff in January 2025 by showing it could train a world-class model with far fewer resources than its US rivals. V4 Flash extends that pattern from training cost to inference cost, and it arrived alongside a broader repricing across the industry, with Moonshot's Kimi K3 adding further pressure on Western labs to defend their pricing in July.
A Full Month of Cuts
OpenAI cut the price of GPT-5.6 Luna, its fastest and cheapest model, by roughly 80% on July 30, just three weeks after the model launched. Luna now costs 20 cents per million input tokens and $1.20 per million output tokens, down from a prior output price of $6. OpenAI also cut GPT-5.6 Terra, the mid-range version, by 20%, bringing it to $2 per million input tokens and $12 per million output tokens. GPT-5.6 Sol, the top-tier version, kept its price unchanged. OpenAI attributed the cuts to efficiency improvements in how it serves the models, and noted that cheaper Chinese open-weight models had increased pressure to justify higher costs.
Google released three new Gemini models built around efficiency in the same window. SpaceXAI shipped Grok 4.5, pricing it at the same rate OpenAI charged for Luna before this week's cut. Meta, which had built its AI strategy around open-weight releases, reversed course with Muse Spark 1.1, a closed-source model priced aggressively for developers. Anthropic has held a different line, keeping its top-tier Claude models at premium pricing on the bet that developers will pay extra for safety and precision.
Zack Kass, OpenAI's former head of go-to-market and now a global AI adviser, has a name for the dynamic driving all of this. He calls it diminishing model returns. Once models cluster close enough in capability, he argues, the next incremental upgrade stops mattering to most buyers, and price becomes the deciding factor. Vinesh Sukumar, Qualcomm's vice president of AI product management, told Axios this could create a lucrative market for what he calls intelligent routers, systems that automatically send each task to whichever model offers the best mix of capability, speed, and price. If that market matures, it further erodes any single lab's ability to command a premium.
Enterprises Are Already Routing Around the Premium
Coinbase gave the clearest public example of what that erosion looks like in practice. CEO Brian Armstrong wrote on X in late June that the company had cut its internal AI spend nearly in half while token usage kept growing. The company did it by changing what loads by default. Coinbase's internal LLM gateway now defaults engineers to open-weight models, specifically GLM 5.2 from Zhipu AI and Kimi 2.7 from Moonshot AI, while still letting engineers escalate to costlier frontier models for tasks that need them. Armstrong said 91% of Coinbase's engineers had never hit their previous usage caps, which suggested the company was paying for capacity it didn't need rather than capability it did.
Armstrong pointed to two other levers alongside the default switch. An automated routing layer preprocesses each prompt and sends it to the most cost-effective model capable of handling the task, factoring in cache status and per-token pricing. And a caching overhaul pushed Coinbase's cache hit rate from 5% to 60%, a twelve-fold improvement that reduces the number of times any model needs to run at all. Armstrong framed the changes as infrastructure for sustainable growth rather than a cost-cutting exercise, writing that the goal was not to suppress usage but to keep costs under control as usage scales.
Coinbase is not alone. Reporting from The Decoder and Business Insider names Snowflake and the AI startup Lindy among other companies that have made similar moves toward Chinese open-weight models for at least part of their workloads.
Microsoft's Hedge
Microsoft is running a parallel experiment inside its own product line. The company is evaluating a self-hosted, fine-tuned version of DeepSeek V4 as a lower-cost model option for Copilot Cowork, its agentic tool built into Microsoft 365, according to Axios. Copilot EVP Charles Lamanna told the outlet that flat-rate pricing does not hold up against heavy users who run hundreds of agentic tasks a week, since each task can burn through tokens quickly. Microsoft is shifting Copilot Cowork to usage-based pricing as a result, and is weighing whether a fine-tuned DeepSeek model, hosted entirely on Azure with added safeguards, could sit alongside its existing OpenAI and Anthropic options as a cheaper tier.
Microsoft has framed any DeepSeek integration as optional and has emphasized that customer data would stay inside its own cloud infrastructure. The company has also disclosed that it is developing its own lower-cost model, internally called Cowork 1, as a further hedge. A final decision on the DeepSeek option is expected within weeks of Axios's reporting.
The pattern across Coinbase and Microsoft is the same one Sukumar described: the model doing the work matters less than the system deciding which model gets the work.
What Would Break the Trend
Axios frames the comparison to commodities like electricity or gasoline. Few people know which power plant supplied their home or which refinery produced the gasoline in their tank, because the product is functionally identical no matter the source. As the performance gap between top-tier language models narrows, AI applications increasingly stop depending on a single provider, and buyers gain the leverage to shop on cost the same way.
That logic holds only as long as the performance gap stays narrow. If a frontier lab pulls meaningfully ahead on a task that actually matters to a customer, the calculus changes and a premium becomes justifiable again. OpenAI is betting on the opposite outcome, that cheaper AI expands demand faster than it compresses margins. CEO Sam Altman said on the Invest Like the Best podcast that the company expects usage volume high enough that it will not need to run as a high-margin business to afford continued model training.
Anthropic's decision to hold premium pricing while competitors cut theirs is its own kind of bet, that enterprises handling sensitive or high-stakes work will keep paying for a smaller gap in reliability. Both bets will be tested by how enterprises actually route their workloads over the next few quarters, and Coinbase's public disclosure of its own routing math gives other companies a concrete playbook to compare their own spending against.
David Borish is the author of the forthcoming book The Tony Hawk Paradox and writes on frontier AI research, enterprise deployment, and technology policy at davidborish.com.
