MANY MINDED · manyminded.com

The Intelligence Ladder

"How smart is AI now?" is unanswerable. "How long a task can it finish?" is measured. Every frontier model, one ladder, honest windows for the rungs ahead — and what "AGI" and the tiers beyond actually mean in numbers.

loading…

the ladder — every measured model

Each dot is a frontier model's 50%-reliability task horizon — the length of human expert work it completes half the time (METR's measurement, the field's standard capability series). The line is the state-of-the-art frontier; the shaded fan is the engine's fitted projection with an 80% interval. The horizontal rungs are where the ladder stops being abstract.

the rungs — from minutes to "god-tier", in measured units

Where is AGI on this? Nowhere — because "AGI" isn't a measurement, and this engine doesn't do vibes. But notice what the rungs quietly replace: an agent that holds week-long jobs is what the economy will experience as transformative; an agent that carries a year of work is what most people mean by AGI whether or not the word survives; the tiers beyond — decade-tasks, century-tasks, run in million-agent fleets — are the "god-like" territory, and on this ladder they are not mystical: they are more doublings. The windows above are fitted from seven years of measurements; treat the far rungs as labeled extrapolation, not prophecy.

the race — lab by lab

Grok (xAI), Kimi (Moonshot), DeepSeek and other labs do not appear in METR's published Horizon evaluations — not measured is not the same as not capable, and this page shows only measured points.

the cadence — will something big land next month?

Nobody can name the next invention. But how often frontier events arrive is measurable: every dot on the ladder is a dated, METR-measured release, and the subset that set a new state of the art is a point process whose inter-arrival gaps fit a base rate. Everything below is computed from those release dates — a frequency, not a prophecy — and it is scoreable: either a measured SOTA crossing lands in the next 30 days or it doesn't.

computing…

Read this base rate honestly. It counts only releases METR has measured — the released-but-unmeasured frontier (Mythos 5 / Fable 5) is invisible here, so the true cadence runs at least this fast and the current drought is partly a measurement lag, not proof of a slowdown. The ruler itself saturates near ~16 hours, so from here a missing crossing can mean a stalled frontier or a stalled ruler. And the exponential (memoryless) fit is a modeling choice, est-grade: real releases cluster around lab cycles. What makes this an instrument rather than a vibe is that it grades itself — over many 30-day windows, about this fraction should contain a crossing, and if they don't, the number was wrong and gets published as wrong.

the hidden tiers — how far ahead can unseen tech be?

The amber band on the chart is the shadow frontier: capability that plausibly already exists but isn't public. It has two very different sources, and honesty requires splitting them:

hidden program (historical)operationalpublicly revealedcapability lead over public tech
Corona spy satellites19601995~10 yrs ahead of commercial imaging
differential cryptanalysis (IBM/NSA)~19741990 (independently rediscovered)~16 yrs
GPS (military precision)1978full civilian accuracy 2000~22 yrs
F-117 stealth19831988~5–10 yrs

The historical pattern is real: in aerospace, sensing, and cryptography, classified programs ran 5–20 years ahead of public technology — so the instinct that "what exists exceeds what's shown" has a solid base rate. But AI inverts it. Frontier AI needs the world's largest compute concentrations, open talent markets, and commercial-scale capital — all of which live in private labs, and government programs today consume frontier-lab models rather than exceeding them. So the honest sizing of today's AI shadow band is months of doublings (lab-internal + measurement lag ≈ ×2–4), not decades — while for defense-adjacent hardware (sensing, propulsion, EW), the historical multi-year hidden tier likely still applies. The engine's rule holds: the band is drawn as an explicitly estimate-grade region and a bias direction — never as invented data points.

read this honestly

← the cost curves · the dashboards