OpenAI and Anthropic made enterprises pay per word for work that didn’t need doing.
It is true that companies were right to try. Generative AI does real work in well-defined domains — code completion, translation, the compression of genuinely large document sets into manageable summaries. The impulse behind “tokenmaxxing,” the maximization of AI-generated output across every business process, was not irrational. The trouble is that the pricing model was never aligned with that impulse. It was aligned with the vendor’s revenue.
Here it is worth being precise about what a “token” actually is, because the pricing architecture is the extraction mechanism. A token is roughly three-quarters of an English word — a sub-word unit the model processes and the customer pays for. When a company routes its customer-service emails, code reviews, and internal memos through a model billed at tokens, the cost scales linearly with output length and has no relationship to whether the output was useful. A ten-paragraph summary that nobody reads costs ten times a one-paragraph summary that solves the problem. The pricing model is indifferent to value. It is a metered connection to a machine that gets paid more the more it talks.
Vincent Gusdorf, who heads AI analytics at Moody’s — not an activist, not a researcher with a book to sell, but a credit analyst whose job is assessing whether investments will pay off — put the finding in plain language: “It’s very easy to create something you don’t need with AI.” The Moody’s report recommends disciplined deployment over indiscriminate volume. That this recommendation is necessary at all tells you something about how the tools were sold. Companies were encouraged to route maximum throughput through the models because the vendors’ revenue is a direct function of token volume. The recommendation to use AI wisely is a recommendation to stop paying the vendors’ preferred pricing model.
What companies initially measured as “productivity gains” was, in most cases, a volume metric — and that metric came from the vendor’s own dashboard. The API console showed tokens consumed, requests served, output delivered. Enterprise buyers measured what the console exposed, not what their workflows needed, because the vendor provided the measurement system and the vendor’s measurement system measured what the vendor sold. More summaries generated. More code suggestions produced. More drafts delivered. The dashboards went up. What was not measured — in most enterprise deployments, what was not yet capable of being measured — was whether any of that output substituted for labour that needed doing, or whether it merely produced documents that were reviewed, corrected, and sometimes discarded. The early metrics rewarded AI activity. The later bills charged for AI activity. The gap between the two is the cost of building a measurement system around what the vendor sells rather than what the customer needs. This phantom interval has a name: the bezzle — the moment before the bill arrives. As earlier MSI coverage noted, major companies are already resuming hiring, having discovered that the AI tools they expected to replace workers still need workers to check their work.
Doctorow has described the current AI investment cycle in terms that fit this moment. AI is “a bubble and it will burst,” he has written, but unlike cryptocurrency, which will leave behind “nothing… but shitty monkey JPEGs and even worse Austrian economics,” the AI buildout will leave durable infrastructure — data centres, GPU clusters, open-source models, a trained workforce of data labellers. The economic tell, he argues, is that the early internet “grew more profitable every day, which workers and young people had to force on their bosses — and AI is a technology that grows less profitable every day, and bosses have to force it on workers.” The tokenmaxxing backlash is the corporate-level version of that inversion. Companies did not discover that AI was unprofitable to them. They discovered that the per-unit pricing made the vendor profitable at the buyer’s expense.
The pricing architecture — per-token, per-word, volume-scaled — is not an accident of early-stage market design. It is the business model. An AI company that charges a flat fee for a useful tool has an incentive to make the tool more useful. An AI company that charges per token has an incentive to make the models produce more tokens. The first incentive aligns vendor revenue with customer value. The second incentive aligns vendor revenue with customer waste. OpenAI and Anthropic chose the second model. The corporate shift to cheaper alternatives is the market beginning to correct for that misalignment. Companies are not abandoning AI. They are refusing to pay per word for content that didn’t need to exist.
What Moody’s is describing is not a technology correction. It is a billing correction — the moment when the companies paying per word discovered that most of the words did not need to exist. The vendors built a meter and sold it as a productivity tool. The only thing it measured was their revenue.