A new benchmark asks a harder AI question: not can agents help, but when are they worth it?
As AI systems improve, one of the most important practical questions is becoming less about whether they can complete a task at all and more about whether they can do it at a lower total cost than people. METR, a research organization focused on evaluating advanced AI systems, has proposed a new metric to address that question directly. It calls the measure the “expenditure horizon,” and its premise is straightforward: compare how much an AI system and a human each need to spend to achieve the same level of improvement, then identify the point where their costs are equal.
The idea matters because many common AI benchmarks flatten performance into a simple success-or-failure result. Those tests can show whether a model solved a problem, but they often say little about the economic tradeoffs involved in getting there. For companies, labs, and policymakers deciding where AI automation is actually viable, that omission is becoming harder to ignore.
How the expenditure horizon works
According to the source material, METR’s framework puts human labor costs, the cost of running AI systems, and the compute used in experiments into a single comparison. The result is a more continuous measure of value rather than a binary score. If the budget required by the AI is lower than the budget required by a person for the same improvement, the AI sits on the favorable side of the expenditure horizon. If the AI requires more spending, the human remains the cheaper option.
That framing is useful because it matches how many organizations actually make decisions. They are rarely asking whether AI can produce any output at all. They are asking whether using it reduces cost, time, or both once oversight, experimentation, retries, and infrastructure are included. An agent that looks impressive in a demonstration may still fail that test if it burns too much compute or needs too many supporting runs.
METR argues that the method has two main advantages over more familiar evaluations. First, it produces a fine-grained economic value rather than a pass-fail label. Second, it converts several distinct inputs into one currency, making it easier to compare human and machine effort inside the same frame.
Testing the idea on NanoGPT speedruns
To trial the metric, METR used the NanoGPT speedrun, a public community project where contributors compete to train a language model as quickly as possible on standardized hardware by improving the training approach rather than changing the task itself. The project offers an unusually clean setting for comparison because the goal is stable while performance improvements can be tracked over time.

The source says the speedrun recorded 82 documented improvement steps from May 2024 onward, driving cumulative training time down from roughly 45 minutes to less than two minutes. That historical record gave METR a way to estimate how much human effort had gone into each marginal gain and then compare that with the cost of asking AI systems to produce similar progress.
To estimate human effort, METR interviewed two of the project’s most active contributors and also used an AI model, Opus-4.6, to estimate the work behind each improvement. Both approaches converged on a similar rough figure: about 16 hours of work for each one-percent speedup. The source further characterizes the human side as spending about $2,500 for each one-percent speedup.
Why the early result is sobering for AI agents
The early finding described in the source is not a triumphalist one. On this benchmark, the expenditure horizon appears underwhelming for today’s agents. That does not mean AI systems are useless in optimization-heavy work. It does suggest that once full costs are counted, they may still struggle to outperform human contributors on more difficult or higher-budget improvement tasks.
That pattern lines up with a broader intuition in AI evaluation: agents often look strongest on small, cheap, well-scoped problems, but their advantage weakens as tasks become more complex and expensive. The source text says METR has seen something similar in earlier tests. AI can beat humans on simpler, lower-cost tasks, but as budgets rise and problems get harder, people often become the more economical option.
This is a meaningful corrective to hype-heavy narratives around autonomous AI engineering. If an agent can produce a useful tweak quickly, it may be cost-effective in a narrow operational band. But if it requires extensive experimentation, repeated prompting, large inference bills, or expensive support compute, the business case can erode fast. The expenditure horizon is an attempt to quantify exactly where that erosion begins.

What the metric clarifies, and what it misses
The value of METR’s proposal is that it shifts the conversation from raw capability to cost-adjusted capability. That is a more demanding standard and, for real-world deployment, probably the right one. Enterprises do not adopt tools because they are novel. They adopt them because they improve throughput, lower costs, expand capacity, or some combination of the three. A benchmark that cannot speak to those outcomes has limited practical value.
At the same time, the source text notes that the metric has blind spots. Any framework that compresses human labor, experimental compute, and AI operation into a single dollar comparison will involve assumptions. Human work quality can vary. Hourly rates can vary. Some AI-generated ideas may be easy to discard, while others can unlock outsized gains. The choice of benchmark task also matters. NanoGPT speedrunning is public, measurable, and technically relevant, but it is still only one domain.
There is another important caveat in the source: newer generations of models could change the picture quickly. A disappointing expenditure horizon today does not settle the question for next year. If model quality improves, tool use becomes more reliable, or inference costs fall, the crossover point between human and AI economics could move substantially.
Why this matters beyond one benchmark
Even with those caveats, METR’s proposal arrives at a useful moment. AI discussions are crowded with claims about productivity, acceleration, and self-improving systems. Much of that debate remains abstract because the field lacks shared ways to compare human and machine effort on equal footing. The expenditure horizon is an attempt to build one.
Its biggest contribution may be cultural as much as technical. It pushes evaluation away from theatrical demos and toward budget reality. That is the right pressure to apply in a market where companies increasingly need to justify AI spending with measurable returns. If the metric spreads, it could help separate workflows where agents are genuinely economical from those where they merely appear efficient until all costs are counted.
For now, METR’s early results suggest a restrained conclusion. AI agents may already be useful in some narrow optimization tasks, but their cost advantage should not be assumed. In the places that matter most to buyers and builders, the question is no longer just what the model can do. It is what the model can do cheaper than a person.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com








