MANY MINDED · THE COST CURVES · KNOWLEDGE

How long a task can an AI agent finish on its own?

AI task horizon at 50% reliability (METR) · minutes · measured 2019–2026 · fitted live from 7 observations by the Many Minded cost-curve engine.

The last measured value is 17.4 hr (2026, METR Horizon v1.1 (official measurements)). Over the fitted window the series rises 314.1% a year (80% interval +172.9% to +528.5%) — a doubling every 0.5 years (80%: 0.4–0.7). Carried forward on that fit, 2028: 299 hr (80%: 50.9 hr–1.75e+3 hr).

17.4 hrlatest measured (2026)+314.1%fitted, per year0.5yper doubling7observations, 2019–2026measuredbasis

the curve

AI task horizon (METR) — measured history and fitted projectionAI task horizon (METR): measured 2019–2026 in minutes, last value 17.4 hr, fitted at 314.1% a year, projected to 2033 inside an 80% interval, against a gate at 40.0 hr. Logarithmic vertical axis.minutes0.1001.00101001.00K10.0K100K20192021202320252027202920312033gate 40.0 hr17.4 hr (2026)
Vertical axis: minutes, log scale. Solid: measured observations. Dashed: the fitted central projection. Shaded: the 80% interval, which widens with horizon because shocks accumulate. Red dashed: the gate at 40.0 hr.

what this series measures

length of human task (in minutes) that frontier AI agents complete at 50% reliability — the standard capability driver series; yearly SOTA.

the gate — what this threshold unlocks

AI agents hold week-long work
threshold ≥ 40-hour task horizon · fires: R&D everywhere — the master trigger

Central fit: ~2027. 80% window: 2027–2028 · n=7. That is 1.2 doublings away at today's value.

●●●●●●●●●●●●●●●○day 14 of 16 in the pond · 1.2 doublings to the gate. One "day" per doubling: the pond fills on the last one, and is only half covered the day before.

what the fit says

quantityvaluehow it is computed
fitted rate+314.1%/yr (80%: +172.9% to +528.5%)Farmer–Lafond drift over the trailing window · n=7 points, 7 years
doubling time0.5 years (80%: 0.4–0.7)implied by the fitted drift

Not shown for this curve, because the machinery returns nothing: Wright's law (no cumulative-deployment series for this technology); the accelerating/decelerating check (needs ≥8 observations, this has 7); a regime break (needs ≥12 observations).

the projection, with its interval

Log cost as a random walk with drift: the forecast variance grows with horizon (τ + τ²/m), which is why these bands widen instead of staying parallel. The middle column is the least useful number on this page; the interval is the claim.

yearcentral fit80% interval
202772.1 hr22.2 hr to 235 hrnear horizon
2028299 hr50.9 hr to 1.75e+3 hrnear horizon
2029above 400 hrpast the point where a number would be theater

The table stops at 400 hr — a decade above the gate. Past that, quoting a number would be theater rather than forecast.

questions this page answers

How long a task can an AI agent finish on its own?

17.4 hr as of 2026, the latest measured value in the series (METR Horizon v1.1 (official measurements)). The fitted trend has it rising 314.1% a year, with an 80% interval of +172.9% to +528.5%.

How fast is the AI task horizon rising?

+314.1% a year over the fitted window, an 80% interval of +172.9% to +528.5% — a doubling every 0.5 years (80%: 0.4 to 0.7 years). Fitted from 7 observations spanning 7 years.

What will the AI task horizon be in 2028?

The central fit says 299 hr, inside an 80% interval of 50.9 hr to 1.75e+3 hr. The interval is the forecast; the middle number is only its midpoint. Bands widen with horizon because shocks accumulate — a constant-width band would be overconfident.

When will the AI task horizon reach ≥ 40-hour task horizon?

The central fit says ~2027, with an 80% window of 2027–2028 from 7 observations. Crossing it is what the engine calls "AI agents hold week-long work", firing: R&D everywhere — the master trigger. Treat the window, not the year, as the claim.

Where does this data come from?

METR Horizon v1.1 (official measurements). 7 observations spanning 2019–2026. Curated benchmark history, extended by a weekly authoritative fetch and by news figures fact-checked against their source before they may touch a fit. Both the observation ledger and the fitting code are public, and the engine publishes its own calibration score and its misses.

the raw numbers

Every observation behind the fit, unrounded by us and unsmoothed. This table is here on purpose: graphs make people underestimate exponential change, and the raw series beside the curve is the one correction shown to work.

yearminuteschange
20190.0500 min
20200.140 min+180.0%
20220.600 min+107.0%/yr
20233.99 min+565.0%
202438.8 min+873.2%
20255.87 hr+807.2%
202617.4 hr+196.6%

where this comes from

Other KNOWLEDGE curves:

Recent briefs from KNOWLEDGE, the daily record of what actually moved:

every KNOWLEDGE brief →

Every tracked curve: utility solar, installed · battery pack price · battery cell price · solar module price · human genome sequencing · onshore wind electricity · solar electricity (LCOE) · space launch, to LEO · wheat yield · offshore wind electricity · modular housing, turnkey · AI video, finished minute · humanoid robot, capable unit · cheapest humanoid, list price · farm labor share · healthcare spend per person

← the whole engine: build clock, cascade, calibration · the archive · how we know