Editorial illustration for New AI Cost Metric Finds Human Labor Still Cheaper by USD 250,000
AI Agents Cost $250K More Than Humans, Study Finds
New AI Cost Metric Finds Human Labor Still Cheaper by USD 250,000
METR ran an experiment on the NanoGPT speedrun, a benchmark where researchers try to train a small language model as fast as possible, and found that AI agents need roughly USD 250,000 more in spending than a human would to hit the same performance gains. That gap is the whole point of a new measurement tool the research organization is rolling out, called the "expenditure horizon." The problem it's trying to solve sounds simple but has resisted easy math for years: how do you compare a researcher's salary, a cluster of GPUs, and an API bill for a model like GPT-4 or Claude on the same scale? METR's answer converts all three into dollars, then plots the point where an AI agent's cost curve crosses a human's for the same task.
Below that crossover, the agent is the cheaper option. Above it, paying a person still wins. The framework grew out of METR's earlier work tracking how AI agents handle tasks of increasing difficulty and cost, where models tend to win early and lose ground as budgets climb.
Whether that pattern holds for the latest generation of models is the open question.
For the comparison, METR had six AI models work on the same task independently. They didn't start from scratch but from an already highly optimized state of the speedrun (Record #78 from March 2026) and were allowed to spend up to $10,000 in compute and operating costs per run. The result: estimated expenditure horizons between $0 and $3,300.
Why this matters
The $250,000 gap matters less as a final verdict than as a baseline we now have numbers for. METR's expenditure horizon gives researchers a concrete way to track whether autonomous optimization is actually catching up, rather than relying on vibes about how capable the latest model "feels." Right now the gap is enormous: low four-figure horizons against a quarter-million dollars of human effort on NanoGPT speedrun progress. That's a useful gut check against hype cycles that assume self-improving AI is already here or imminent.
But the metric has blind spots METR itself flags, and the newest model generation could move the numbers fast. For founders betting product roadmaps on AI agents replacing engineering labor, and for researchers studying recursive self-improvement, this is the kind of measurement worth watching quarter over quarter rather than trusting once. A single snapshot showing humans still winning by six figures shouldn't calm anyone down permanently. It should just tell us where the line was on the day it was measured, and how fast the next model closes it.
Common Questions Answered
What is the expenditure horizon metric that METR introduced?
The expenditure horizon is a new measurement tool developed by METR to compare the cost efficiency of AI agents versus human researchers on the same tasks. It provides a concrete way to track whether autonomous optimization is catching up to human performance by quantifying the financial investment required for AI to achieve equivalent results.
How much more expensive is it for AI agents to complete the NanoGPT speedrun compared to humans?
According to METR's experiment, AI agents need roughly USD 250,000 more in spending than a human would to hit the same performance gains on the NanoGPT speedrun benchmark. This significant gap represents the difference between low four-figure expenditure horizons for AI and a quarter-million dollars of human effort required for equivalent progress.
What were the results when METR tested six AI models on the NanoGPT speedrun task?
METR had six AI models work independently on the NanoGPT speedrun starting from an already highly optimized state (Record #78 from March 2026) with a budget of up to $10,000 in compute and operating costs per run. The results showed estimated expenditure horizons ranging between $0 and $3,300 across the different models tested.
Why is the expenditure horizon metric important for tracking AI progress?
The expenditure horizon provides researchers with a concrete, quantifiable way to measure whether autonomous AI optimization is actually catching up to human capabilities, rather than relying on subjective assessments about how capable a model "feels." This metric helps counteract hype cycles by establishing a clear baseline for comparing AI and human performance efficiency on research tasks.