Skip to content
All writing
AI StrategyInfrastructureEfficiency

The Model Is a Menu Now

OpenAI didn't ship a model this month. It shipped a menu — three tiers of the same release on one cost curve. That is the efficiency thesis becoming the org chart of the industry.

Jake Chen··5 min read

Personal perspectives only — does not represent the views of my employer.

Look closely at how the frontier labs shipped this summer and you'll notice something the benchmark coverage mostly missed.

They stopped selling a model. They started selling a menu.

GPT-5.6 didn't arrive as one thing. It arrived as three — Luna, Terra, Sol — a cheap tier, a middle tier, and a flagship, all the same release, priced along a curve from a couple of dollars per million tokens to an order of magnitude more. Terra reportedly matches last year's flagship at roughly half the cost. Gemini has its Flash and Pro lanes. Grok undercuts the field on price. And sitting underneath all of it, the open-weight models are now close enough to the frontier to be the floor.

The release stopped being a model. It became a price-performance curve you order off of.

This is the thesis arriving

I've spent this series arguing one thing from a few directions: the next phase of AI is decided less by how smart the model gets and more by the economics of delivering intelligence. TurboQuant attacked memory cost. ChatJimmy attacked hardware cost. Sora showed what happens when you ignore delivery cost entirely. And convergence showed that once the models are all roughly as smart, intelligence stops being the thing you compete on.

The menu is what happens next. When intelligence converges, the labs can't keep selling it as if smarter is the only axis — so they price along the axis that's left. Same brand, same week, sells you a tier that's cheap and light and a tier that's expensive and brilliant, and asks you to decide which job is which. The pricing page finally admits what the efficiency thesis has been saying all along: the model is a commodity with grades, and the grade you need depends on the work.

Interactive

Order off the menu

One release, three tiers on a cost curve. Pick a job and see which tier the work actually calls for — and what defaulting to the flagship would cost.

Luna

~$2.25/1M

Right tool

Terra

~$5.60/1M

Sol

~$11.25/1M

Rough monthly cost for this workload

Luna — the right tier$900
Sol for everything — the lazy default$4.5k

Same answer, 5.0× the bill if you send this to the flagship out of habit. The skill isn’t picking the smartest model — it’s routing each job to the cheapest tier that still clears the bar.

Illustrative tiers modeled on GPT-5.6’s Luna / Terra / Sol split. The point is the shape, not the exact cents.

Second-order effect one: architecture becomes routing

If the release is a menu, then the core skill is no longer picking a model. It's routing.

The naive move is to choose the flagship and send everything to it, because it's the best and thinking about tiers is annoying. That's also how you light money on fire. Most of what a real system does — tagging, extraction, summarizing, first drafts — is easy work that the cheap tier clears without breaking a sweat. The hard reasoning that actually needs the flagship is a thin slice at the top. Pay flagship rates for the whole pile and you're often paying several times over for an answer the cheap tier would have gotten right.

So the competent architecture stops being "which model do we use" and becomes "which tier does each request deserve, decided per call." That's a router. It's classifiers, fallbacks, confidence thresholds, and a cost budget — a little control plane in front of the menu. The teams that win the next phase will treat model selection as a runtime decision, not a procurement one. FinOps for tokens becomes a real discipline, the way cloud cost management became one once compute got metered.

Second-order effect two: pricing power compresses

Here's the part that should worry the labs more than any single benchmark loss.

When you're the only one with the smart model, you can anchor the price wherever you like. When you yourself also sell a tier at a quarter of that price — and a competitor sells one cheaper still — and an open-weight model sits underneath as the free floor, the anchor slips. Every buyer can now see the whole curve, including your own cheap end of it. The flagship still commands a premium for the hardest work, but the enormous middle of the market gets to ask, every single call, "does this actually need the expensive tier?" More and more often, the honest answer is no.

That's margin migration in slow motion. The revenue doesn't vanish, but it slides toward whoever owns the best price-performance curve and the router that exploits it — not automatically toward whoever holds the single smartest checkpoint.

Second-order effect three: the menu's shape is the product

Which reframes what the labs are even competing on.

For years the competition was a point: the top of the leaderboard. The menu turns it into a shape. What matters now is the whole curve — how cheap the cheap tier is, how good the middle tier is at the price, how cleanly the tiers hand off to each other, how easy it is to route between them. A lab with a mediocre flagship but a devastating cheap tier can win more of the actual market than a lab that owns the crown and nothing underneath it.

The moat was never really the smartest model. This series has said that in five different ways. The menu just makes it literal: the defensible thing is the delivery — the curve, the router, the integration, the economics — not the single impressive number at the top.

The series in one line

When the models converged, the interesting question stopped being which one is smartest. The menu is the answer the market gave: none of them, all the time. You buy a curve now, and the skill is knowing where on it each piece of work belongs.

The first era of AI asked how smart the model could be. This one asks a quieter question, and it's the one the whole efficiency stack has been circling: what is this particular job actually worth paying for?

Order accordingly.

All essays
RSS