Morgan Linton, the chief technology officer of Bold Metrics, an AI startup based in Lake Tahoe, strategically guides his engineering team of 16 on which AI models to deploy, doing so twice each week. Just 50 minutes before his team's daily standup, he discussed his plans to recommend the use of Claude Fable at a low capacity for one group, while directing another to take advantage of GPT-5.5 at a high setting. A third team is achieving "perfect results" with Cursor powered by Composer 2.5.
By specifying which models to use, Linton avoids implementing strict token limits. "My team is using the most effective tools, but they're doing so more efficiently," he explained.
The first half of 2026 saw AI firms embracing a trend known as "tokenmaxxing," where there was a push for maximum AI utilization. However, faced with significant AI expense invoices, corporations such as Uber and Microsoft have begun to reassess this approach.
Individuals from all areas of tech, including founders, software engineers, and UX designers, are quickly adopting a money-saving strategy: model switching. They assign their most complex tasks to premium frontier models while delegating simpler, repetitive tasks to older, cost-effective models.
As companies tighten their AI budgets and implement usage restrictions, this thoughtful approach could prove beneficial, ensuring better value for investment.
Moving Away from Tokenmaxxing
While it is evident that newer models have their merits, including enhanced capabilities that minimize retries and the need for oversight, not all tasks justify their expense. Brian Armstrong, CEO of Coinbase, voiced this perspective in a post on June 7, stating that "80% of workloads will operate on 99% lower-cost models within the next 12-18 months," while the remaining 20% would still require the latest technology where "IQ maxxing" is crucial.
Chris Maconi, cofounder of AI startup Hechura, has always been skeptical of tokenmaxxing. His company operates with a "human-in-the-loop" philosophy, avoiding automated bots for coding tasks. This attitude influences their model selection as well.
Recalling the OpenClaw AI agent's rise and rapid consumption of tokens, Maconi noted his initial use of low-cost Gemini models before transitioning to Anthropic's Haiku.
"I'm open to experimenting with less expensive models to see if they meet our intelligence needs," he remarked.
Innovative Token Usage
Tanvi Pisal, a 29-year-old user-experience designer at a major tech firm, admits she learned the importance of using models wisely the hard way. Utilizing Figma, ChatGPT, and Claude to develop product requirement documents, Pisal originally relied on Claude for brainstorming, only to find that she was wasting tokens without completing the task.
"Now, I generate designs in Figma first and then feed those images into Claude, directing it to maintain the same UI while developing the functionality and flow," she shared. "This design-first approach definitely helps me conserve tokens."
Pisal also prefers to brainstorm ideas using ChatGPT—available at no cost through her company’s enterprise plan—before refining them with Claude.
Alejandra Thomas, a software engineer and tech content creator in New York City, adopts a hands-on approach to evaluating new models. "For routine tasks, I always opt for lighter models or none at all rather than going for the priciest option simply because it's available.”
The CEO of AI sales firm Scoot, Ed Stevens, describes a consistent practice of selecting and testing a model for a few months to assess its effectiveness before considering alternatives. If a more cost-efficient or advanced model emerges, they'll make the switch.
This careful management of token expenditure reflects a scarcity mindset, as noted by Dan Ariely, a behavioral economics researcher and Duke University professor. He compares token budgets to earlier cellphone plans that limited usage, prompting individuals to maximize their minutes regardless of necessity.
"Tokens create a feeling of scarcity, leading users to limit their consumption. This triggers a psychology of waste if they fail to meet their token quota," he explained. Consequently, users tend to switch to alternative models after hitting their token ceiling to avoid additional costs.
Streamlining Model Selection
For those who find constant model switching overwhelming, help is at hand. The growing field of model routing startups provides solutions to automate these decisions, assigning tasks to specific models—sometimes including open-source options—based on their complexity. Companies like OpenRouter are gaining traction and funding in this domain.
David Gilmore, who heads Rayline, noted that their software evaluates whether requests could be assigned to more affordable open-source alternatives. Many clients initially succumb to "FOMO moments," only to be jolted back to reality when they review inflated API bills.
According to Ramp's lead economist, Ara Kharazian, the adoption of model routing technology is increasing, with usage rising from approximately 1% last year to 5% this year.
The San Francisco-based investment firm BlockSpaceForce has been utilizing solutions like OpenRouter, Fireworks, and Together AI. Managing partner Spencer Yang emphasizes the importance of consulting a less expensive model first to determine if a more advanced one is truly necessary for any particular task.
"Models are becoming increasingly adept at evaluating their own complexity," Yang observed.
Still, some businesses persist in defaulting to the latest and most expensive models, with Maconi attributing this tendency to a lack of effort. "People often prefer to ride the hype train rather than invest the time to understand which models excel at which types of tasks."


