
Companies building products and services around artificial intelligence face significant challenges in establishing sustainable pricing models, largely stemming from the unpredictable nature of token consumption in large language models and agentic AI systems.
Tokens are the mathematical units that LLMs break prompts into for processing and generate responses from. While the cost per individual token has declined significantly in recent years, overall token consumption has surged dramatically as businesses and consumers scale their AI usage. Goldman Sachs forecasts that external token consumption will increase 24 times between 2026 and 2030, reaching 120 quadrillion tokens monthly as companies shift toward AI agents. This creates a fundamental problem: the relationship between input and output remains non-deterministic, making it difficult to predict costs. Subtle variations in prompts can yield different results, and multi-agent systems compound this unpredictability further.
Major technology companies themselves have encountered token cost overruns. Microsoft reportedly restricted engineers’ access to certain third-party coding tools, while Uber exhausted its annual AI coding token budget in just a few months earlier this year. Smaller firms have attempted workarounds by using flat-fee personal accounts, though industry observers expect major AI vendors will eventually clamp down on such practices once shareholder pressure for profitability intensifies.
Experts suggest several mitigation strategies. Companies should carefully select which AI models to use, craft more precise prompts to reduce unnecessary token consumption, and consider implementing guardrails and testing protocols. However, these measures become increasingly complex when AI systems are embedded into products serving thousands of users, where costs can balloon unexpectedly.
With no consensus approach yet established, companies are exploring various pricing structures with their customers, including across-the-board price increases, results-based pricing, and bundled incident charges. The challenge remains acute because underlying token costs from major LLM providers continue evolving, making long-term pricing agreements difficult to sustain.
Article Attribution | Read More at Article Source
Article summary produced by Claude AI