Tokenomics: Why Making AI Pay Is Tricky


If you’ve ever used the free version of ChatGPT or any rival AI, you’ve felt the value it offers.


Large‑language models (LLMs) are engineered by cloud giants like Microsoft, Google and Anthropic, with hundreds of billions of dollars invested into their development. The result? A free baseline that costs the company a fraction of the building‑out price.


To recoup those expenses, the same companies now sell “pro” plans with extra coding features, faster response limits and dedicated support. But even beyond that, a new wave of third‑party firms is building AI‑powered agents—tuned to specific tasks and deployed in products—from the ground up.


Setting a price for these services is surprisingly hard, according to experts. “Trying to tie someone into a cost model for the next 12, 24 or 36 months doesn’t make any sense—we simply don’t know how many tokens they’ll burn,” says Simon Gooch (Saviynt).


Tokens are the computational units that an LLM uses to process prompts and produce answers. A single request can vary wildly: the same question may produce different outputs, and different models diverge in cost per token. When an AI agent chains dozens of sub‑tasks together—like scheduling, coding or customer support—the token count can grow astronomically and unpredictably.


Goldman Sachs forecasts that token consumption will jump 24‑fold between 2026 and 2030, reaching 120 quadrillion tokens a month. Yet most organisations struggle to track how many tokens are being used until a bill arrives.


Some firms respond by clamping down. Microsoft reportedly curtailed engineering teams’ use of third‑party coding tools after seeing runaway token usage. Uber, on the other hand, burned through its entire AI token budget in just a few months.


The volatility extends to pricing strategies. A vendor might change token cost every couple of months, confusing customers that depend on a stable budget.


Smaller companies have tried flat‑fee or personal‑account models to stay under the radar, but big providers warn that this practice will be discouraged once profit pressure mounts.


From a product standpoint, companies must also educate themselves on bespoke model selection and prompt engineering. As Rob Steele (CFO, iplicit) notes, “without clear instructions, the AI is like a teenager tossed into grocery shopping—unpredictable and potentially costly.”


If an AI‑driven feature scales to thousands of users, token costs can skyrocket—each request may incur multiple token exchanges for coding, testing or security checks. That scale almost guarantees cost unpredictability, even when the benefit to end‑users is higher.


Ultimately, many firms have no clear path to pass AI costs onto customers. Bill Peterson (Sumo Logic) describes the situation as “fun conversations” while exploring pricing models: flat rate, results‑based or bundled incidents. But these ideas could be upended by shifting token costs from the underlying LLM providers.


In a market where AI delivers outsized value, the industry must still grapple with a fundamental puzzle: how to translate elastic, token‑driven usage into sustainable, predictable revenue streams.