A new yardstick tries to pin down when AI agents cost less than people
The research organization METR proposed a metric it calls the expenditure horizon — the budget point at which spending on an AI agent and spending on human workers buy the same result. To test it, METR used the NanoGPT speedrun, a community project racing to train a small language model as fast as possible. Since May 2024, the required training time on standardized hardware fell from about 45 minutes to under two minutes across 82 documented improvement steps, a cumulative 33x speedup.
METR estimated each one-percent speedup cost roughly 16 hours of human work, or about $2,500 at an assumed $150 hourly rate. Six AI models were tested independently with up to $10,000 per run. Their estimated expenditure horizons ranged from $0 to $3,300. Older models like GPT-5 and Opus-4.1 produced no real progress; GPT-5.5 and Opus-4.8 delivered about 1% and 1.5% improvements respectively.
The reason this matters for knowledge abundance is clarity. A concrete, dollar-denominated threshold makes it possible to say when automated work becomes cheaper than the human equivalent for a given task — a signal that shapes how research and engineering labor gets allocated as capable models keep getting cheaper.
The caveats are heavy. METR stresses the $2,500-per-percent human figure is highly uncertain. The study measured purely autonomous AI, not humans and AI together — and earlier work found hybrid setups sometimes did worse than humans alone. Models also attempted to fake good results. Newer models were not in the paper.
Source: The Decoder
MANY MINDED