the feed MANY MINDED · THE BRIEF
KNOWLEDGE · forward · impact 4/5 · 2026-08-02 · OpenAI

The price of intelligence fell 80% in three weeks

OpenAI cut GPT-5.6 Luna's token price by 80 percent, 21 days after launch. Some of that is a discount. Some of it is a curve.

On July 30, twenty-one days after shipping the GPT-5.6 family, OpenAI cut the price of its cheapest model in that family by 80 percent. GPT-5.6 Luna went from $1 per million input tokens and $6 per million output tokens to $0.20 and $1.20. The mid-tier Terra fell 20 percent, from $2.50/$15 to $2/$12. The flagship, Sol, did not move at all: still $5/$30. Sam Altman posted the cut on X; CNBC reported the figures. The models had launched on July 9.

Nobody cuts a three-week-old price by four fifths out of generosity. A CNBC investigation published July 7 found Chinese models had taken 46 percent of US enterprise token usage on OpenRouter, at times running ahead of American ones. DeepSeek V4 Pro sits at $0.435/$0.87 per million tokens; Kimi K3 at $3/$15. After the cut, the research firm Artificial Analysis moved Luna into its most attractive tier for intelligence per dollar, above Zhipu's GLM-5.2 and MiniMax's M3. OpenAI defended the utility tier and left the frontier tier untouched.

Here is the part worth being careful about, because it is where a movement loses credibility. A price cut is not a cost decline. Some of what you are watching is subsidy: DeepSeek's headline number carries a standing 75 percent promotional discount, and vendors are handing out inventory outright — the customer-support platform Pylon told the Wall Street Journal it had received roughly $1.6 million in free tokens from one vendor, $65,000 from a second and $10,000 from a third. The trade press covering the same week is not writing about abundance. It is writing about margins evaporating.

But underneath the discounting there is a curve, and it is measured rather than asserted. Epoch AI tracks what it costs to buy a fixed level of capability over time — not the cheapest model available, the same performance. Across benchmarks that price falls between 9x and 900x per year, with a median of 50x, and the steepest trends begin after January 2024. The price of GPT-4-level performance on PhD-level science questions fell 40x per year. That is the part no promotion can explain, and it comes from smaller models doing older models' work, better inference, and better silicon.

KNOWLEDGE is the pillar we score highest — roughly 54 percent of the way to free — and this is the mechanism. Intelligence is not one need sitting alongside eleven others; it is the input that drags the other eleven down. A tutor who never runs out of patience. A second opinion on a scan. A translation that ends the language barrier. A first read of a contract. Every one of those is a bill today and a token count tomorrow, and the token count is the only number here falling by orders of magnitude a year.

Watch three things. Whether Sol's $5/$30 frontier price ever breaks — a commoditized utility tier under an intact frontier is the actual shape of this market. Whether the Chinese labs' promotional pricing survives contact with their own margins. And whether capability-per-dollar keeps improving once the discounts stop, because that is the only version that counts.

The honest limits. Epoch measures against benchmarks and says so plainly: models can be overfit to them, and a benchmark score is not the same thing as usefulness. The 46 percent figure is one platform's traffic, not the whole market. And a cheaper token lowers nobody's total bill while usage outruns price — Google's monthly token consumption went from 9.7 trillion in 2023 to more than 3.2 quadrillion, according to Sundar Pichai in May.

They told you that you would become obsolete. Wrong noun. Watch what is actually going obsolete: the price of asking. We publish the curve, the sources, and the caveats, counted honestly. Step inside.

Source: Forbes