The surge in artificial intelligence has rested on a foundational belief: larger models yield greater power, and thus the strongest models prevail. However, the industry is on the brink of a crucial realization if this belief begins to falter.
The rising costs associated with these expansive models are prompting users to reconsider smaller, more affordable options. This emerging trend of cost-sensitive model selection is unprecedented and could have profound implications for the sector.
A notable forecast from Brian Armstrong, co-founder of Coinbase, suggests that this shift will involve a dramatic transition to low-cost models for the majority of tasks. He emphasizes on platform X, “The need for intelligence is virtually limitless, yet 80% of workloads will transition to models that are 99% less expensive within a year or so. Only 20% will still rely on cutting-edge models when high intelligence is crucial.”
If Armstrong’s projections materialize, the ramifications for the AI landscape could be monumental. Historically, AI firms have vied for superiority based on the quality of their offerings—often resorting to the most sophisticated models available. Should these tasks prove manageable with less expensive alternatives without compromising quality, it could lead to a significant reconfiguration of AI's economic landscape, potentially impacting major players like OpenAI and Anthropic, especially as they approach their IPOs.
This substantial shift hinges on a pivotal question: Are organizations prepared to adopt smaller models?
Preliminary evaluations indicate that, when configured appropriately, cost-effective models can replace larger ones without diminishing quality. For instance, legal AI company Harvey recently demonstrated a threefold reduction in inference costs while maintaining quality standards. This test was conducted in collaboration with Fireworks AI, utilizing Claude Opus and Fireworks’ GLM 5.1, with Opus taking on the most demanding tasks. This approach resulted in reduced server time and overall expenses.
“Quality is paramount, especially in legal matters,” stated Harvey co-founder Gabe Pereyra in an interview with TechCrunch. He noted that the definition of quality is evolving from simply employing the most powerful model available to identifying the most efficient model that produces the correct results.
This shift in perspective often frames the discussion around established labs versus Chinese or open-weight models, but fails to capture the core issue. The real distinction lies not in the nature of the models—proprietary or open—but rather in the size, i.e., large models versus smaller ones. Transitioning from GPT-5.5 to DeepSeek’s V4 Flash can save costs, but equally effective results can be achieved by opting for GPT-5.4-mini.
Currently, a price competition is underway between the internal inference of major labs and independently operated open-weight models. Ultimately, the question of which small model prevails becomes secondary in the broader context of small versus large.
While it may seem intuitive not to over-utilize computational resources, this principle runs counter to the industry’s predominant scaling-first mindset. Driven by past lessons, labs have aggressively pursued the development of the most computationally demanding models, pushing the boundaries of AI capabilities. With generous investor subsidies in play, clients had little incentive to consider anything less than the most advanced solutions.
However, with token prices on the rise and the pace of subsidies diminishing, users are now confronted with cost pressures for the first time. It's uncertain whether this new financial strain will lead enterprise users to embrace smaller models; they might instead choose to minimize usage, reduce context, or abandon less viable projects altogether.
If it becomes apparent that a majority of deployments can successfully operate on smaller models, this could substantially stifle the growing demand for inference and provoke critical inquiries into the value of developing high-end models.


