A single image at the center of the AI bubble fight
Of all the charts circulating in the debate over whether artificial intelligence is overheating, one has become something of a Rorschach test: a line showing weekly token consumption on OpenRouter climbing from roughly 0.5 trillion in January 2025 to 126.2 trillion — an increase of more than 25,000 percent. Supporters of the boom thesis point to the near-vertical slope as proof that demand is real and accelerating. Skeptics see a vanity metric dressed up as an industry indicator.
Both readings deserve scrutiny. According to analysis published by Matthias Bastian and highlighted on LinkedIn, the OpenRouter curve is better understood as a story about measurement than as a story about mass adoption or economic value.
What OpenRouter measures — and what it doesn't
OpenRouter is a routing layer: developers use it to connect their applications to a catalog of AI models from multiple providers, often switching between them based on price, speed, or capability. That makes its usage data genuinely interesting, because it captures production traffic rather than benchmark scores. A token is the fundamental unit of AI computation — roughly the equivalent of a gallon of fuel for a car, or a unit of work billed by a utility.
The problem is that tokens are not a proxy for users, sessions, or revenue. A platform can process dramatically more tokens while serving roughly the same number of people, simply because each request now costs far more computation. That gap between volume and value sits at the crux of the bubble argument.
Why token counts inflate so quickly
Reasoning models think out loud
Modern reasoning models generate large volumes of internal "thinking" tokens before they produce a final answer. Those tokens are invisible to the end user and often add nothing to the perceived quality of the output, yet they are counted in full. A model that deliberates at length can multiply token consumption several times over compared with a simpler model answering the same prompt.
Agentic systems burn through tokens
Unoptimized agentic AI systems amplify the effect further. An agent that plans, calls tools, checks results, and retries can loop through many model calls for a single task. A marginal increase in real-world usage — a few more agents running, a slightly longer workflow — can translate into an enormous spike in tokens consumed. In that sense, the OpenRouter chart partly reflects how inefficient some deployments still are.

The leaderboard trap: GPT 5.6 Luna's dominance
Model rankings on the platform invite a similar misreading. OpenAI's GPT 5.6 Luna recently dominated token consumption on OpenRouter, which is easy to interpret as a surge in popularity. It may not be. A model can top a token chart simply by generating more tokens per prompt, which makes it look busier than it is. Consumption rankings measure verbosity and workload as much as adoption.
Follow the revenue instead
Where token volume is ambiguous, spending is harder to fake. On the revenue side, OpenAI's Astra leads, suggesting customers are willing to pay for capability at the top of the market. Meanwhile, Chinese models including Kimi, GLM, and DeepSeek are expanding quickly: monthly spending on those models rose tenfold in 2026, albeit from a much smaller base. Growth rates from a low starting point tend to look spectacular, and that context matters when the numbers are used as evidence of a broader shift.
How to read a token chart without being misled
Token metrics are useful, but only when paired with the right questions:
- Volume versus value. Ask what revenue or completed tasks sit behind the tokens, not just how many were processed.
- Tokens per request. Rising consumption can come from longer reasoning rather than more requests.
- Model mix. A single verbose model entering the top spot can bend the curve without changing underlying demand.
- Agent overhead. Retries, tool calls, and planning loops inflate counts and can mask inefficiency.
- Base effects. Tenfold growth sounds extraordinary until you note where it started.
The bubble question remains open
None of this means the surge is meaningless. A jump from 0.5 trillion to 126.2 trillion weekly tokens shows that a great deal of AI work is being done, and that developers are routing real traffic through a shared platform. It shows inference spending is scaling. But an impressive line on a graph cannot distinguish between efficient growth and expensive sprawl, between millions of users and a handful of token-hungry agents.
The most honest conclusion is that the chart describes the cost of computation, not the state of the market. The AI bubble debate will be settled by revenue, retention, and returns — the measures that token counts are meant to approximate but cannot replace.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com








