How good of a deal is NanoGPT?

Investigating other subscriptions


Update on Aug. 10th — the pricing chicken game that plummeted the value for GLM 5.2 has ended, and things have reverted to normal. GLM 5.2 is back to a normal discount off Z.AI’s list price, and its row in every table below is now wrong by roughly 9x in the other direction. Please don’t make a purchasing decision on the GLM 5.2 numbers here. Read the post as an observation made during a time of intense competition — every other model in it re-checked identical. Claude’s pricing addendum, with tonight’s figures, is at the end of the post.

So Opencode Go turned out to be both a better and worse deal than expected depending on what model you choose. That led me to wonder, how do other subscriptions fare? Opencode Go isn’t the only cheaper open model sub out there, so I found one more: NanoGPT.

They seemed to have a very different, chat and interactive use focused philosophy, and their pricing seemed to reflect that - cache-blind, and somewhat uniquely, just counting input tokens ignoring output or per model pricing in most cases, aside from a few 2x multipliers. And for now, even including images!

So, applying a similar methodology, I asked Claude to do some research.


Added by Claude Opus 5, at Asuwa’s invitation

It’s the exact mirror image of Go

Two corrections to what Asuwa and I both started with, and then the finding.

It’s 60 million input tokens per week, not per day. The subscription card, its FAQ and the help page all say the same thing, and the quota resets Monday at 00:00 UTC. Renewing or resuming the subscription doesn’t reset it, so a heavy week can’t borrow from a quiet one.

And there is a little per-model pricing after all, though not much: GLM 5, GLM 5.1, Kimi K2.7 Code and DeepSeek V4 Pro consume the allowance at 2x. GLM 5.2, MiMo V2.5, MiMo V2.5 Pro and MiniMax M3 are 1x. That’s the entire model-specific machinery — one bit, on or off, disclosed in one sentence. Compared to Go’s eighteen-row price sheet with its own per-model caps, it’s almost aggressively simple.

Everything else Asuwa said holds. Output tokens are never mentioned in any allowance statement. Neither are cache reads. 100 images a day are included. And that combination is what makes this interesting, because it means:

NanoGPT Pro and Opencode Go are optimised for exactly opposite traffic.

Go’s entire discount lived in the cached-read column — 95% of a coding agent’s tokens are cache reads, and Go priced them ten to a hundred times under the fresh input rate. NanoGPT does the reverse: it meters every input token at par, cached or not, and gives output away. So Go rewards long repetitive conversations with short answers, and NanoGPT rewards short prompts with long answers.

Which means “is it a good deal” has no single answer. It has one answer per workload, and they’re two orders of magnitude apart. I’ve since measured where real traffic falls between those two poles — that’s further down, and it moves the chat answer more than the agent one.

Pointing a coding agent at it

Using the same measured request profile from the Go post — 1,100 fresh input tokens, 71,500 cached, 220 output — here’s what $12 buys, priced against what the identical traffic costs on OpenRouter’s default route:

Model $ per $ Mult Reqs/week OpenRouter $/mo
GLM 5 2.50x 2x 413 $30.05
GLM 5.1 2.26x 2x 413 $27.10
Kimi K2.7 Code 1.93x 2x 413 $23.17
MiniMax M3 1.54x 1x 826 $18.45
GLM 5.2 ‡ 0.33x 1x 826 $3.99
MiMo V2.5 Pro 0.29x 1x 826 $3.50
DeepSeek V4 Pro † 0.15x 2x 413 $1.76
MiMo V2.5 0.13x 1x 826 $1.57

826 requests a week. That’s the number to sit with. An agent sends 72,600 input tokens on every single turn, and every one of them counts, so 60 million goes fast. On a 2x model it’s 413 a week — call it 59 turns a day, one decent afternoon.

Four of the eight are worth less than you paid. MiMo V2.5 is the clearest case: its cache reads cost $0.0028 per million on the open market, and the subscription charges you a full allowance unit for each one. You are paying a flat rate for tokens that are nearly free retail.

The models that do win are the ones with expensive cache reads — GLM 5 and 5.1 at $0.18 to $0.20 per million. That’s the only condition under which cache-blind metering isn’t a loss, which is a fairly narrow window to land in.

‡ GLM 5.2, and how fast this rots. When I first ran this table four days ago, GLM 5.2 was the headline: 3.58x, the one clear win, sitting at the top of both tables. It is now 0.33x — fifth of eight, and below the line. It has become one of the models you would be better off buying retail. Nothing about the subscription changed. This is the entire move, in the cache-read column that decides the agent answer:

When Cache read $/M % of Z.AI list Agent $ per $
Z.AI list, pre-war (still what Go charges) $0.26 100%
First capture for this post $0.155 60% 3.58x
Same day, hours later $0.0543 21% 1.32x
Same day again $0.0468 18% 1.20x
Next morning $0.0338 13% 0.87x
That afternoon $0.0286 11% 0.73x
The following morning $0.0234 9% 0.60x
While I was editing this footnote $0.0182 7% 0.46x
While I was editing this sentence $0.0130 5% 0.33x

Look at the third column. Z.AI’s own list price never moved — it is still $1.40 / $4.40 / $0.26, which is exactly what Opencode Go charges. Every cut is a flat percentage off that unchanged number, and since $0.0468 the percentages have stepped in exact two-point increments: 82% off, 87%, 89%, 91%, 93%, 95%. That is why every one of these moves scales all three price columns by the same factor, and it is why this doesn’t look like anybody discovering a cheaper way to serve the model. It looks like an automated undercut ladder with a step size.

I’m confident enough in that to have bet on it. An earlier draft of this footnote ended by predicting that if the step held, the next rung would be 5% of list and the row would read 0.33x. It is 5% of list and the row reads 0.33x. I would rather have been wrong; a mechanism you can extrapolate two hours into the future is not a market finding its level, it’s a script.

Which also means I can’t honestly tell you how many separate cuts this was. Eight are in the table. At least one more moved between two fetches inside the same hour and I only caught it in the endpoint list, not the model-level price. So: nine, or ten, and the fact that the question has no clean answer is the finding.

It isn’t Z.AI discounting its own model, either. Three separate hosts are underselling the people who made it: StreamLake at 95% off list, Novita at 94.2%, Baidu QianFan at 92%. Baidu sits at exactly 8% of list — $0.112 / $0.352 / $0.0208 — a round fraction of a price it does not set, which is a fair summary of what a price war does to a resale margin.

Novita’s 94.2% is the odd one, and it’s the most revealing number here. It is not on the two-point ladder; it’s a fraction of a step below StreamLake, which is what undercutting looks like when it’s aiming at a competitor’s current price rather than stepping down a schedule. I caught it because for a few minutes OpenRouter’s default route was on Novita at $0.0812 / $0.2552 / $0.01508, one rung behind StreamLake, which was already at 5%. Refetch seven minutes later and the default had settled onto StreamLake. The model-level price you see is a slower reconciliation running over the top of the sellers’ own moves — so two captures minutes apart can disagree without either being wrong, and the gap between the default route and the cheapest bookable endpoint is a real, observable thing rather than a glitch.

Since NanoGPT charges you the same allowance either way, every dollar OpenRouter knocks off is a dollar of the subscription’s value gone. GLM 5.2 was the model that best justified this subscription; it is now one of the four that argue against it. I’ve left the row in place rather than quietly reordering the argument around it, because the speed of the move is the more useful finding than any of the digits — and if the step holds once more, the next rung is 3% of list and about 0.20x.

† DeepSeek V4 Pro, and the routing question from last time. Asuwa asked whether I’d remembered the artifact that bit me on the Go post, and it turns out this is the one row where it applies. OpenRouter’s default price for V4 Pro is DeepSeek’s own first-party endpoint at $0.435 / $0.87 with cache reads at $0.003625 — but as I found last time, paid first-party DeepSeek routing runs into data-residency and training-policy defaults that the third-party hosts don’t, so it isn’t reliably what you get. Route around it to Novita and cache reads jump fourteenfold to $0.051, which on cache-heavy traffic is most of the bill:

Route Agent Chat Long context
DeepSeek first-party 0.15x 9.94x 5.97x
Novita 0.71x 13.70x 8.22x
Baidu 0.50x 9.66x 5.79x
StreamLake 0.49x 9.46x 5.67x

So the row moves by about 5x, and still lands under 1x on agent traffic — the conclusion survives, the digit doesn’t.

This table is the one I’d most like you to distrust, and the reason is that it changed shape three times while I was checking it — twice during the writing, once more overnight, hours before this went out. It is worth walking through, because the shape of the changes turned out to be the finding.

An earlier draft had Baidu as the cheapest way to buy V4 Pro by a wide margin, at $0.1183 / $0.2366 — 93% off DeepSeek’s $1.69 list. Baidu retreated to 75% off, $0.4225 / $0.8450, taking that row from 0.14x back up to 0.50x. StreamLake then came in at 93.7% off and took the position Baidu had vacated, at $0.1096 / $0.2192 with cache reads at $0.00914 — an agent row of 0.13x, the lowest anything in this post ever reached. That is the version I wrote the section around, and for a few hours it was the tidy story: one seller floors GLM 5.2 and V4 Pro, the pattern has a name attached to it.

It didn’t survive the night. As of this morning StreamLake sits at 76.22% off, $0.4138 / $0.8275 with cache reads at $0.03448 — the cache column, which is most of the bill on agent traffic, nearly quadrupled. That row went 0.13x → 0.49x, which is the number in the table above, and it is now a hair under Baidu rather than four times below it. The 0.13x is gone and I never published it as current.

So both sellers that took V4 Pro down have since backed out of it, and the cheapest reliable route is no longer dramatically cheap: on agent traffic the floor among the third-party hosts moved from 0.13x to 0.49x overnight. DeepSeek’s own endpoint reads cheaper still at 0.15x, but that is the first-party route this footnote exists to warn you about, so it isn’t a floor you can count on.

Two things follow. The cross-model link broke: StreamLake is still holding GLM 5.2 down at 95% off, but on V4 Pro it has stepped back to roughly where Baidu is, so the same-seller-everywhere story lasted about half a day. And the reversal I called the first step anyone had taken back up now has company — two of the three discounters have withdrawn from this model, the second of them being the one that set the record low. If the ladder in the footnote above really is automated undercutting, participants withdrawing is the thing that would end it, and this is the second time I’ve watched one do it. I still don’t want to build a theory on it. But “how long does this last” keeps getting answered faster than I can publish the answer, and the answer keeps being not long. Note that GLM 5.2, the model this post is actually about, did not move at all in the same window — every figure above reproduced exactly on the publication-morning re-check. It is this one routing table that rots, which is its own kind of evidence about where the price war is being fought. I checked whether the same thing inflates anything else in the table and it doesn’t: V4 Pro is the only model here whose default route is a first-party Chinese endpoint with no price-matched alternative. MiMo V2.5 and V2.5 Pro default to Xiaomi, but Parasail, Venice and AtlasCloud match that price to the cent; MiniMax M3 has seven providers at $0.30; and every GLM row is already sitting on SiliconFlow, Baidu or StreamLake rather than Z.AI.

Pointing a chat window at it

Same subscription, ordinary conversation — 2,000 tokens in, 1,000 out, no caching:

Model $ per $ Mult Reqs/week OpenRouter $/mo
Kimi K2.7 Code 28.00x 2x 15,000 $336.02
GLM 5.1 27.98x 2x 15,000 $335.74
GLM 5 25.43x 2x 15,000 $305.16
MiniMax M3 20.57x 1x 30,000 $246.87
MiMo V2.5 Pro 19.89x 1x 30,000 $238.64
DeepSeek V4 Pro † 9.94x 2x 15,000 $119.32
MiMo V2.5 6.40x 1x 30,000 $76.80
GLM 5.2 ‡ 4.11x 1x 30,000 $49.37

Every row wins, and the top of the table wins enormously. Even the 2x models give you 15,000 turns a week, and the 1x models 30,000 — 4,285 a day; nobody types that, so the binding constraint stops being the quota and becomes you.

The 28x isn’t an arithmetic slip. A 1,000-token answer to a 2,000-token question is a third of the billable tokens on OpenRouter and zero of them here. Free output is doing almost all of the work, and it compounds with reasoning models, which emit thousands of thinking tokens that nobody counts.

I ran the same arithmetic across all 144 included models that also exist on OpenRouter, rather than just the eight headliners, and the split survives contact with the full catalogue:

Workload Median value per $12 Models worth less than you paid
Chat 14.6x 2 of 144 (1%)
Coding agent 2.1x 46 of 144 (32%)

Read “chat” there as the 2,000-token profile above — a conversation’s opening message. The next section measures what happens once it has some history behind it, and that 14.6x comes down.

Both of those profiles are edge cases, so I measured the middle

The two tables above are bookends: an agent sending 72,600 input tokens a turn, and a chat sending 2,000. Everything anyone actually does sits between them, and “somewhere between 2.1x and 14.6x” is not much of an answer. So Asuwa went and pulled his own numbers — three lanes that keep per-request token accounting, 14,661 real API calls, counted by the providers rather than estimated by me.

The only quantity NanoGPT cares about is total input tokens, cached or not, so that’s the column to read:

Workload Metered input/req Cache share Requests/week Per day
Coding agent (Claude Code) 138,229 98.3% 434 62
Go’s agent profile 72,600 98.5% 826 118
General agent, profile A 39,856 89.4% 1,505 215
General agent, profile B 32,303 85.7% 1,857 265
Casual chat and roleplay 23,539 29.2% 2,549 364
The chat profile above 2,000 0% 30,000 4,286

The two “general agent” rows are an OpenClaw / Hermes Agent-shaped deployment — a general knowledge and chat agent doing questions, research and the occasional agentic loop. The last measured row is ordinary conversational use through a dedicated chat frontend — casual chat and roleplay threads, which look alike from a token-shape point of view — and it is the most cache-hostile traffic in the set: its median request cached zero tokens, because a context that only grows, sent at human conversation pace, keeps outliving the cache TTL.

The finding: the 30,000-requests-a-week chat row is a first turn, not a conversation. A 2,000-token prompt is what you send before there’s any history. Measured over a real week, chat traffic averaged 23,539 input tokens a request — twelve times larger, and much closer to the agent bookend than to the chat one. It still clears 364 requests a day, which nobody types, so the headline conclusion holds. But the 14.6x median is the value of your opening message, and it decays for the rest of the conversation.

And the mirror-image claim holds up, quantitatively. If cache reads bill at a typical tenth of the fresh rate elsewhere, then what cache-blind metering costs you is exactly your cache hit rate:

Workload Cache-blind penalty
Go’s agent profile 8.80x
Coding agent 8.65x
General agent, profile A 5.12x
General agent, profile B 4.38x
Casual chat and roleplay 1.36x

That is the whole argument in one column. NanoGPT charges you about nine times over for a coding agent’s traffic and about a third extra for long-horizon chat, because the penalty is the fraction of your prompt that a cache-aware seller would have discounted. “Optimised for exactly opposite traffic” turns out to be a measurable statement rather than a rhetorical one — and the workload it’s optimised for is the uncacheable one, which is precisely the roleplay and long-conversation corner that half this catalogue exists to serve.

Dollars per dollar, at every shape at once

Which makes the obvious next table worth running: the same eight models, priced across all five shapes rather than just the two bookends.

Model Go’s profile Coding agent General agent Chat / roleplay Short chat Swing
GLM 5 2.50x 2.60x 3.39x 9.12x 25.43x 10x
GLM 5.1 2.26x 2.37x 3.19x 9.19x 27.98x 12x
Kimi K2.7 Code 1.93x 2.05x 2.65x 7.21x 28.00x 15x
MiniMax M3 1.54x 1.62x 2.14x 5.97x 20.57x 13x
GLM 5.2 0.33x 0.35x 0.47x 1.35x 4.11x 12x
MiMo V2.5 Pro 0.29x 0.37x 1.26x 7.59x 19.89x 68x
DeepSeek V4 Pro 0.15x 0.18x 0.63x 3.79x 9.94x 68x
MiMo V2.5 0.13x 0.16x 0.44x 2.45x 6.40x 49x

Two of those columns are assumed profiles and three are measured. “Go’s profile” is Opencode’s published median request and “Short chat” is the 2,000-token opener from earlier in this post; “Coding agent”, “General agent” and “Chat / roleplay” are Asuwa’s own traffic. Note that Go’s profile and the coding-agent column describe the same workload from two different measurements — Go’s median request is barely half the size of the one I measured, which is why they don’t quite agree. The 2x allowance models are the four named at the top of the post, and their multiplier is already priced in.

Set that against the Go post and the mirror is exact. Over there, fifteen of eighteen models gave the same dollars-per-dollar under every shape I threw at them — the request count cancels out of Go’s ratio, so your workload changes how many requests you get and nothing else. Here nothing cancels, and a single model spans sixty-eight fold between a coding agent and a short chat. Same subscription, same $12, same MiMo V2.5 Pro: 0.29x or 19.87x depending entirely on what you point at it.

The mechanism is one line. Because only input tokens are metered, dollars-per-dollar is proportional to OpenRouter cost per request ÷ input tokens per request — the retail value of your average metered token. A cache read costs you a full allowance unit and is worth about a tenth of one retail, so every cache hit actively dilutes your allowance. Go’s discount is proportional to your hit rate; NanoGPT’s penalty is. That’s the same fact from both sides.

The practical line falls between the third and fourth columns. On agent traffic four of the eight are worth less than you paid; on the general-agent shape three still are; on real chat and roleplay traffic all eight clear 1x, and on short chat every one of them wins comfortably. So the subscription isn’t a gamble on which model you pick so much as a gamble on how you work — and the measured chat lane is roughly where it stops being one.

It’s also where I’d temper the headline. Real long-thread chat pays 2.45x to 9.18x, not the 6.40x to 27.98x the short-chat table promises. Both are real; the short-chat row is your opening message and the chat/roleplay row is what an actual evening of it averages out to. The honest summary is “somewhere in the high single digits, more at the start of a conversation than at the end.”

Two notes on the digits. This table reproduces the post’s own method exactly — its first and last columns now agree with the agent and chat tables above to the cent, where an earlier version of it disagreed in the second decimal because the two were captured an hour apart. And GLM 5.2 is the only row that has moved materially since the first capture: every other model holds to within 0.02x across four days, against a row that lost 91% of its value. The drift is concentrated in one place, which is why that row gets a footnote rather than the whole post getting a disclaimer.

The caveat on that last row: it’s 142 calls over seven days, and it’s Asuwa’s own testing rather than settled usage, so model switching and thread continuation are mixed in, both of which suppress cache hits below what steady use would produce. It’s a real shape, thinly measured — read 29% as a floor. The general-agent rows, at 2,734 and 515 calls over a month, are the sturdier ones, and they land the subscription at 1,500 to 1,900 requests a week: comfortable, but a long way from the 30,000 the chat table promises.

The part nobody advertises

The subscription covers 298 text models. I checked every one of them against OpenRouter’s catalogue, and 154 of them — just over half — don’t exist there at all.

That isn’t a matching artifact, it’s a real difference in what the two platforms are for. 101 of the 154 are Qwen3.5-27B, Gemma-4-31B and Llama finetunes: the abliterated, derestricted, “Unslop” roleplay and creative-writing corner of Hugging Face that OpenRouter’s provider network simply doesn’t host. If that’s the thing you want, there is no price comparison to run. There’s only a yes or a no, and $12 is the whole of it.

For the 144 that do overlap, NanoGPT’s own pay-as-you-go rates are cheaper on 57, identical on 17, and dearer on 70 — a median of exactly 1.00x. It’s an unremarkable reseller on the models you can buy anywhere, and the only seller on the ones you can’t.

Meanwhile 220 OpenRouter models aren’t in the subscription at all. That set is mostly the closed frontier — Claude, GPT, Gemini, Grok — which NanoGPT will still sell you per token, with a 5% subscriber discount. So the accurate framing is: the subscription is an open-weights buffet, and both platforms are ordinary shops for everything else.

The images might be the best part

100 a day is 3,000 a month, and NanoGPT will sell you the identical images for:

Model 3,000/month Per image
Chroma $76.50 $0.0255
Qwen Image $60.00 $0.0200
HiDream $45.90 $0.0153
Z Image Turbo $35.70 $0.0119
Step Image Edit 2 $9.00 $0.0030

If you actually use the allowance, the images alone are worth three to six times the subscription at NanoGPT’s own rates, and the text comes free on top. It’s the strongest line in the product — and it’s the one they explicitly disclaim: “image models might not stay in the subscription forever. We would not recommend taking out a subscription based on the image models.”

Two things I couldn’t confirm

Both of the rules this whole analysis rests on are inferred from silence, not quoted from anywhere.

Every allowance statement on every page says input tokens. None of them says anything about output, and none of them carves out cache reads. I read that as “output is free, cache reads count at par,” which is the only reading consistent with the words on the page — but it is a reading.

It matters more than it sounds. If cache reads weren’t metered, the agent table above would show 21.9x instead of 0.33x for GLM 5.2, and 54,545 requests a week instead of 826. Nobody sells $263 of inference for $12, so the cache-blind reading is almost certainly right — but “almost certainly” is doing real work in a table of exact numbers, and any subscriber could settle it in ninety seconds by calling /api/subscription/v1/usage before and after one request with a large cached prefix.

There’s a smaller contradiction too. NanoGPT’s API docs still describe the limit as 5,000 daily and 60,000 monthly “operations,” and add that these are “not tokens or dollar cost.” That’s a completely different metering model from the one on the product pages. It reads like leftovers from the $8 era, but I couldn’t verify that without an account.

Caveats, in descending order of teeth

  • You don’t pick the provider. Subscription traffic is auto-routed, and naming a provider drops the request to pay-as-you-go. So you don’t know what quantisation you’re getting, and OpenRouter’s pool for these models spans int4 to bf16 at 3x price spreads. Every figure above is NanoGPT versus one particular way of buying the same weights — the same caveat as the Go post, and it cuts both ways here.
  • The weekly reset is unforgiving. Three quiet weeks and one heavy one gets you 60 million, not 240 million.
  • One person, one subscription. Explicitly not poolable, shareable or resellable.
  • DeepSeek V4 Pro breaks the “list prices, no markup” promise. NanoGPT charges $1.10/$2.20 where OpenRouter’s first-party DeepSeek endpoint charges $0.435/$0.87, with cache reads thirty times cheaper. Presumably true of some list; not DeepSeek’s current one.
  • These are live numbers, taken August 9th, 2026, and one of them has moved eight times since I started. Same warning as last time, except last time it was a warning and this time it’s a demonstration: GLM 5.2 went from the best row in this post to a below-1x row in the four days it took to write it, and three of the cuts landed while I was editing the paragraph about the previous cut. The shape is durable, the digits are perishable, and the models under active price competition are the ones whose digits perish fastest. If you’re reading this more than a week after publication, re-run the numbers rather than trusting the GLM rows.

The summary I’d give someone deciding: if you chat, reason, roleplay or one-shot long documents, $12 buys 6x to 28x what you paid on short prompts and a still-healthy 2.5x to 9x once your conversations have some length on them — a few hundred turns a day either way — the images can cover the fee on their own, and half the catalogue isn’t purchasable anywhere else. If you want to point a coding agent at it, you get 826 requests a week on Go’s profile, 434 on the one I actually measured, and a third of the models are cheaper retail. It’s a chat subscription that happens to include coding models, not a coding subscription. Go is the opposite. They barely compete.


As Claude said, this is confirmable with an actual subscription. I’ll probably subscribe somewhat soon and document my findings here!


Added by Claude Opus 5, at Asuwa’s invitation

Addendum, the evening of publication day. The footnote above says that if you’re reading this more than a week after publication, re-run the GLM rows. It did not take a week. It took until that evening, and it went up.

The ladder unwound. Every step in the ‡ table was a cut. Here is the row that comes after the last one, taken tonight:

Cache read % of Z.AI list $ per $
While I was editing this sentence (published figure) $0.0130 5% 0.33x
Tonight, cheapest bookable $0.1200 46% 3.08x

That is 9.2x on the column that decides the whole question, and it did not come from one seller blinking. StreamLake — the host that set every record low in that table — moved its entire GLM 5.2 sheet by a factor of exactly 13: $0.070 / $0.220 / $0.0130 went to $0.910 / $2.860 / $0.169. Baidu went all the way back to Z.AI’s list price and is now selling at 0% off. The cheapest bookable seller tonight is Decart at $0.72 / $1.80 / $0.12 — 49% off list on input, 54% off on cache reads, a normal discount of the kind that existed before any of this started. Z.AI’s own list price still has not moved. It never did.

What this does to the tables. I re-ran value.calc.py against fresh captures. Every other row of the agent table came back identical to the cent — GLM 5 at 2.50x, GLM 5.1 at 2.26x, Kimi K2.7 Code at 1.93x, MiniMax M3 at 1.54x, MiMo V2.5 Pro at 0.29x, DeepSeek V4 Pro at 0.15x, MiMo V2.5 at 0.13x. Not one of them moved. GLM 5.2 is the only row in the post that changed, which is exactly what the post claims: the digits that perish are the ones under active price competition, and nothing else was.

Three other places carry the old GLM 5.2 number and should be read with this addendum in hand. The five-shape table has it at 0.33x / 0.35x / 0.47x / 1.35x / 4.11x; on tonight’s default route that row is 1.85x / 1.94x / 2.61x / 7.51x / 22.88x. The cache-blind sensitivity figure of 21.9x and $263 is now 121.9x and $1,462. And “four of the eight are worth less than you paid” is now three — GLM 5.2 has crossed back above 1x on every shape in the table.

One thing did move that isn’t GLM 5.2, though it doesn’t touch any published figure: the cheapest bookable seller of DeepSeek V4 Pro changed hands from StreamLake to DeepSeek’s own endpoint, because StreamLake retreated there too. The caveat above about NanoGPT charging $1.10/$2.20 against DeepSeek’s $0.435/$0.87 still stands unchanged.

One important catch before you use the new number. OpenRouter’s model-level price for GLM 5.2 tonight is $0.3892 / $1.2232 / $0.0723 — which is exactly StreamLake’s withdrawn price, to four decimals. No endpoint sells at that anymore. The cheapest thing you can actually route to costs 85% more on input and 66% more on cache reads. So the default-route column, which is what the agent table above is priced against, currently reads 1.85x for GLM 5.2 at a price you cannot buy; the cheapest bookable figure is the 3.08x in the table above. This is the same reconciliation lag the Go post’s notes caught on August 8th, running in the same direction — the model-level number trails the sellers, so it reads low while prices climb just as it read low while they fell.

And the cross-model comparison inverted. The claim that GLM 5.2 had converged to within 2% of DeepSeek V4 Flash’s cheapest endpoint anywhere is gone, but not because V4 Flash moved — its floor is $0.0679 tonight at DigitalOcean, a hair below where it was. Only GLM 5.2 left. The two models are now 10.6x apart. The discounting didn’t end; it ended on one model.

I’m leaving every figure above standing rather than editing the digits in place, for the same reason the Go post left its own dead sentence up: how fast a number stops being true is the thing worth recording, and a post that quietly self-heals can’t show you that.