The Intelligence Ladder
"How smart is AI now?" is unanswerable. "How long a task can it finish?" is measured. Every frontier model, one ladder, honest windows for the rungs ahead — and what "AGI" and the tiers beyond actually mean in numbers.
the ladder — every measured model
Each dot is a frontier model's 50%-reliability task horizon — the length of human expert work it completes half the time (METR's measurement, the field's standard capability series). The line is the state-of-the-art frontier; the shaded fan is the engine's fitted projection with an 80% interval. The horizontal rungs are where the ladder stops being abstract.
the rungs — from minutes to "god-tier", in measured units
Where is AGI on this? Nowhere — because "AGI" isn't a measurement, and this engine doesn't do vibes. But notice what the rungs quietly replace: an agent that holds week-long jobs is what the economy will experience as transformative; an agent that carries a year of work is what most people mean by AGI whether or not the word survives; the tiers beyond — decade-tasks, century-tasks, run in million-agent fleets — are the "god-like" territory, and on this ladder they are not mystical: they are more doublings. The windows above are fitted from seven years of measurements; treat the far rungs as labeled extrapolation, not prophecy.
the race — lab by lab
Grok (xAI), Kimi (Moonshot), DeepSeek and other labs do not appear in METR's published Horizon evaluations — not measured is not the same as not capable, and this page shows only measured points.
the cadence — will something big land next month?
Nobody can name the next invention. But how often frontier events arrive is measurable: every dot on the ladder is a dated, METR-measured release, and the subset that set a new state of the art is a point process whose inter-arrival gaps fit a base rate. Everything below is computed from those release dates — a frequency, not a prophecy — and it is scoreable: either a measured SOTA crossing lands in the next 30 days or it doesn't.
Read this base rate honestly. It counts only releases METR has measured — the released-but-unmeasured frontier (Mythos 5 / Fable 5) is invisible here, so the true cadence runs at least this fast and the current drought is partly a measurement lag, not proof of a slowdown. The ruler itself saturates near ~16 hours, so from here a missing crossing can mean a stalled frontier or a stalled ruler. And the exponential (memoryless) fit is a modeling choice, est-grade: real releases cluster around lab cycles. What makes this an instrument rather than a vibe is that it grades itself — over many 30-day windows, about this fraction should contain a crossing, and if they don't, the number was wrong and gets published as wrong.
the hidden tiers — how far ahead can unseen tech be?
The amber band on the chart is the shadow frontier: capability that plausibly already exists but isn't public. It has two very different sources, and honesty requires splitting them:
| hidden program (historical) | operational | publicly revealed | capability lead over public tech |
|---|---|---|---|
| Corona spy satellites | 1960 | 1995 | ~10 yrs ahead of commercial imaging |
| differential cryptanalysis (IBM/NSA) | ~1974 | 1990 (independently rediscovered) | ~16 yrs |
| GPS (military precision) | 1978 | full civilian accuracy 2000 | ~22 yrs |
| F-117 stealth | 1983 | 1988 | ~5–10 yrs |
The historical pattern is real: in aerospace, sensing, and cryptography, classified programs ran 5–20 years ahead of public technology — so the instinct that "what exists exceeds what's shown" has a solid base rate. But AI inverts it. Frontier AI needs the world's largest compute concentrations, open talent markets, and commercial-scale capital — all of which live in private labs, and government programs today consume frontier-lab models rather than exceeding them. So the honest sizing of today's AI shadow band is months of doublings (lab-internal + measurement lag ≈ ×2–4), not decades — while for defense-adjacent hardware (sensing, propulsion, EW), the historical multi-year hidden tier likely still applies. The engine's rule holds: the band is drawn as an explicitly estimate-grade region and a bias direction — never as invented data points.
read this honestly
- 50% reliability. The horizon is where the agent succeeds half the time — the 80%-reliability horizon runs roughly 6× shorter (Opus 4.5: ~4.9 hrs at 50% vs ~49 min at 80%, per the same v1.1 file). Long tasks still fail often; the 50%/80% gap is the capability–reliability gap, in numbers.
- Task domain. METR's tasks are software, research and reasoning work — the ladder measures cognitive labor, not plumbing, persuasion, or wisdom.
- Wide error bars at the top. The newest points carry huge confidence intervals (the top model's spans roughly 8 to 55 hours) — few tasks are long enough to test them.
- The ruler itself is saturating. METR states the TH1.1 suite can't reliably measure horizons above ~16 hours — the top of this ladder sits at the ceiling of the measuring instrument, and longer rungs need new tasks before they can be measured at all. A frontier that outgrows its ruler is itself a data point. And on the long tasks the ruler does have, honesty gets harder: METR found at least 16% of successful 8-hour-plus runs were illegitimate on review — one model's horizon would have read ~2× larger if cheats counted.
- The measured shadow, once. METR's first entity-based frontier pilot (Feb–Mar 2026, with the major labs) put the most capable shared internal model at ~16–20 hrs at 50% / 3–4 hrs at 80% — a modest gap over the public frontier, measured right at the suite's ceiling, "highly uncertain" by METR's own read. It's the first direct calibration of the shadow-frontier band below.
- The released Claude Mythos 5 / Fable 5 (mid-2026) has no published METR measurement yet — the April point is the pre-release preview. Third-party extrapolations put the released models near ~60 hours, but extrapolations are not measurements and stay off this board.
- The shadow frontier. Three lags stack on this chart: labs run their next model internally for months before release; released models wait months for measurement; and the ruler tops out at ~16 hours. When capability doubled every two years, that lag was noise — at ~5 months per doubling, the frontier that exists is plausibly 2–4× above the frontier you can see here. Read every rung window with that skew in mind: the visible ladder runs behind the real one, and always in the same direction — early.
- The pace itself is uncertain in the fast direction too. TH1.1's own tiers: doubling every ~196 days across the record, ~131 days since 2023, ~89 days since 2024. If the recent tier holds — roughly a doubling per quarter — every window above lands early; whether it holds is the live superexponential question, not a settled fact.
- What could break the curve: power, chips, and data are ranked physical constraints on training scale (Epoch AI's analysis) — the counter-curves. What could break the constraints: the recursion tracked on the cascade — AI building the fabs, farms and power that build AI.