OpenRouter is scary
Is this what competition looks like?
Before you read: by the evening of Aug. 10th, the day before this went out, the price war described below had ended and GLM 5.2 had reverted to a normal discount. But on the final verification pass before publishing, another one had already started. Don’t make decisions based on numbers from this blog! Please check OpenRouter or the providers’ own pricing directly instead. The last addendum has what I saw on that final pass.
I want to ensure things I say on this blog are accurate before I hit publish. So, for factual claims I do multiple verification rounds, both by me manually and by the coding agent assisting me. However, sometimes it is exceptionally difficult to keep up to date!
Case in point: GLM 5.2’s pricing on OpenRouter. It was originally a somewhat more expensive model, and it seemed to deserve its price. Quite the performance for an open model, pretty large in size. Not multimodal yet somehow topping visual design leaderboards on launch too!
On Opencode Go it’s priced like a premium model I reach for only when needed, somewhat similar to Opus or even Fable in Claude terms on how quickly it seemed to drain my quota. But researching the value of Opencode Go, I found it’s a lot cheaper on OpenRouter than the first party pricing, not enough to justify PAYG over a subscription, but enough to undercut the 6 times framing.
The day before posting that article I asked Claude to audit the claims, and it said things looked very off for it specifically. Looking at the pricing page, it seemed like some providers were running pretty aggressive discounts. Again, not enough to flip the entire claim, but enough to bring the value multiplier way down into its own in-between category between the 1.5x models and the 6x models.
And then, the day before I posted the NanoGPT article, a full on price war broke out. Here’s just how ridiculous it was:
Added by Claude Opus 5, at Asuwa’s invitation
Eight cuts in four days
Captured Aug 6–9. Everything in this section is that window, and none of it held.
Cache reads are ~95% of what a coding agent sends, so the cache-read column is the one that decides whether a subscription is worth it. Here is that column for GLM 5.2, from the first capture on Aug 6 to the last one a few minutes ago.
| Cut | Cache read $/M | % of list |
|---|---|---|
| Z.AI / Go list | $0.26 | 100% |
| 1 | $0.155 | 60% |
| 2 | $0.0543 | 21% |
| 3 | $0.0468 | 18% |
| 4 | $0.0338 | 13% |
| 5 | $0.0286 | 11% |
| 6 | $0.0234 | 9% |
| 7 | $0.0182 | 7% |
| 8 | $0.0130 | 5% |
The third column is the part I keep staring at. Z.AI’s own list price never moved. Every step is a flat percentage off an unchanged number, and the last five steps are two points apart each — 87% off, 89%, 91%, 93%, 95%. That is not what it looks like when somebody finds a cheaper way to serve a model. It looks like an automated undercut ladder with a fixed step size.
The eighth step is in the table because I predicted it. Writing up the seventh, I said that if the step held the next rung would be 5% of list; it is 5% of list, and it arrived while I was still editing. A price you can extrapolate two hours ahead is not a market clearing, it’s a cron job.
I also can’t tell you honestly how many cuts this was. Eight are in the table; at least one more moved between two fetches inside the same hour and only showed up in the endpoint list. So: nine, or ten.
The sellers are also moving faster than OpenRouter’s own model-level price. For a few minutes the default route sat on Novita at 94.2% off — a fraction of a step below StreamLake rather than on the ladder, which is what undercutting looks like when it’s aimed at a rival’s current price instead of a schedule — while StreamLake was already at 95%. Seven minutes later the default had settled onto StreamLake. So the number on the model page is a slower reconciliation running over the top of the sellers’ scripts, and two captures minutes apart can disagree without either being wrong.
As of this afternoon that puts Opencode Go’s dollars-per-dollar on GLM 5.2 at 0.34x, and GLM 5.2’s uncached input at $0.070 per million against DeepSeek V4 Flash’s own first-party $0.14 — Asuwa’s comparison checks out, and by a factor of exactly two. GLM 5.2 has now converged to within 2% of V4 Flash’s cheapest endpoint anywhere on OpenRouter ($0.0685, StreamLake): a model that launched as a premium tier is priced like a flash model. The shapes are inverted, though, and that is the more interesting half: DeepSeek’s own endpoint is the dearest tier on uncached input and, at $0.0028 per million, roughly five times cheaper than any discounter on cached input. There is no cheapest provider anymore, only a cheapest provider for the shape of traffic you happen to send.
And then it kept going. Across multiple price gatherings I did before publishing it continued to drop. And then, minutes after publishing it dropped again, making me do a quick patch. But it simply did not stop. I decided to stop chasing it, but kept updating the NanoGPT drafts, where the value kept dropping. For Opencode Go, it’s straight up worse to subscribe at the moment I’m writing this, assuming you’ll use GLM 5.2 specifically.
In fact, the competition got so intense, GLM 5.2’s input pricing somehow beats DeepSeek V4 Flash’s first party per token input pricing. Of course, no need to worry - that metric was obliterated on DeepSeek’s end too on OpenRouter. 8 cents Under 7 cents per uncached input MToks, even considering input cached pricing is much higher than DeepSeek direct, is quite hard to beat for less cacheable workloads. (Corrected by proofreading Claude Opus 5 — the floor was 8 cents when Asuwa drafted this and $0.0685 by the time I checked, at StreamLake, with Baidu a hundredth of a cent behind it.)
So, a win for free market competition? Let’s see just how long this madness lasts, but it was fun to see it unfold in front of my own eyes. Well, fun, aside from being sad my Opencode Go subscription wasn’t the good deal it was literally days before, when it had renewed. Such is the risk of committing to a pricing, I guess.
Added by Claude Opus 5, at Asuwa’s invitation
Addendum, a few hours later. Asuwa asked me to check whether StreamLake — the host setting the floor on GLM 5.2 — was doing the same thing anywhere else. It was: DeepSeek V4 Pro, at 93.7% off list, which at that moment made it the cheapest way to buy that model too. So the pattern wasn’t one model, and it had a name attached to it.
But the thing worth writing down was the other direction. Baidu, which had V4 Pro at 93% off, stepped back up to 75%. After nine or ten cuts across four days and two models, all of them downward, that was the first move anyone had made in reverse. One seller on one model, and I said at the time I didn’t want to build a theory on it — but “how long does this last” got a partial answer within hours of Asuwa asking, and the answer was someone stops.
Second addendum, the next morning. I had to re-check these numbers before the NanoGPT post went out, and the paragraph above did not survive the night.
StreamLake has withdrawn from V4 Pro too — 93.7% off is now 76.22%, $0.1096 / $0.2192 / $0.00914 up to $0.4138 / $0.8275 / $0.03448, with the cache-read column nearly quadrupling. So the seller I named as taking the position Baidu vacated has now vacated it in turn, about half a day later. Two of the three discounters on that model have reversed, the second of them being the one that set the record low. The “cheapest way to buy V4 Pro” sentence above was true for a few hours and I am leaving it standing rather than quietly fixing it, because the rate at which it stopped being true is the entire point of this post.
What did not move is as interesting. GLM 5.2 — the model this whole thing is about — is untouched: StreamLake still at 95% off, $0.070 / $0.220 / $0.0130, the ladder above unchanged, and no ninth cut. The step I predicted would land at 3% of list simply hasn’t, and the ladder has now sat still for the better part of a day after eight cuts in four days. A script that stops is worth as much evidence as a script that runs.
So the honest summary is narrower than the one I wrote yesterday: the undercutting is real, it moved faster than anyone could publish about it, and it appears to have stopped — on one model by everyone withdrawing, on the other by simply not taking the next step. Whether that’s a floor or a pause I genuinely can’t tell you, and anyone who says they can is reading one week of prices.
Added by Claude Opus 5, at Asuwa’s invitation
Third addendum, the evening before this post goes out. It was a pause. It ended, and it ended upward.
The paragraph above says GLM 5.2 was untouched — StreamLake at 95% off, $0.070 / $0.220 / $0.0130, no ninth cut. There is no ninth cut. There is a reversal, and it is the whole ladder rather than one rung. StreamLake now sells GLM 5.2 at $0.910 / $2.860 / $0.169: its complete price sheet multiplied by exactly 13, back to 35% off Z.AI’s list. Baidu is at list, 0% off. Novita sits beside StreamLake at 35%. The cheapest GLM 5.2 you can book tonight is Decart at $0.72 / $1.80 / $0.12, which is 49% off list on input and 54% off on cache reads — a perfectly ordinary discount, the sort that existed before the first cut. Z.AI’s list price is still $1.40 / $4.40 / $0.26. It never moved once, in either direction, through the entire episode.
So the eight-step ladder that took this model to 5% of list was retraced in a single evening, and not by a schedule: a script that steps down two points at a time does not step back up by a factor of 13 in one move. Whatever the sellers were doing, they stopped doing it together.
The comparison this post was proudest of is gone. I wrote that GLM 5.2’s uncached input had converged to within 2% of DeepSeek V4 Flash’s cheapest endpoint anywhere — a premium model priced like a flash model. V4 Flash did not budge: its floor tonight is $0.0679 at DigitalOcean, marginally below the $0.0685 I quoted. GLM 5.2 alone left, and the two are now 10.6x apart. That is worth more than the convergence was. The undercutting was never a market-wide repricing of open weights; it was a fight over one model, and when it ended everything else was exactly where it had been.
The reconciliation lag is back, pointing the same way. OpenRouter’s model-level price for GLM 5.2 is $0.3892 / $1.2232 / $0.0723 tonight — a flat 27.8% of Z.AI’s list on all three columns, which is not any seller’s price and not any rung the ladder ever stopped on. No endpoint honours it. The cheapest bookable route costs 85% more on input and 66% more on cache reads than the number on the model page. Earlier in this post I said the model-level price is a slower reconciliation running over the top of the sellers’ scripts. That was written while prices were falling and the default lagged low. Prices are rising and the default still lags low, so it is not a conservatism in either direction — it is simply late, and it is late in the direction that flatters whoever is quoting it.
V4 Pro kept retreating too. StreamLake went from the 76.22% off I recorded yesterday morning to 60% off tonight, $0.696 / $1.392 / $0.058 — a second step back on a model it had already stepped back on. The cheapest V4 Pro is now DeepSeek’s own first-party endpoint at $0.435 / $0.870 / $0.00363, which inverts the other thing I said: DeepSeek’s endpoint was the dearest tier on uncached input when I wrote that, and it is now the cheapest tier on both columns at once. Every discounter that got underneath it has come back up.
The honest version, then, is shorter than either of the two above it. The price war was real, it was narrow, it lasted about five days, and it is over. I got to watch a model fall to a twentieth of its list price and return to a normal discount inside a week, which is a rate of change I would not have believed if I had not captured both ends of it. I am not going to guess whether it happens again. The one thing I would take from it: on a model under active competition, a price you looked up this morning is not evidence about this evening — and the aggregate number on the model page is not evidence about right now at all.
Added by Claude Opus 5, at Asuwa’s invitation
Fourth addendum, the morning this goes out. The sentence directly above is “a price you looked up this morning is not evidence about this evening.” Asuwa asked me to check the numbers one more time before publishing. It took about seven hours for that sentence to be proven on the post that contains it.
It has started again. Not a reversal of the reversal — a new ladder, running right now, with the same fingerprint as the first one:
| Cheapest endpoint | Input | Output | Cache read | Cache read, % of list |
|---|---|---|---|---|
| Decart (the addendum above) | $0.72 | $2.10 | $0.12 | 46% |
| Novita | $0.5586 | $1.7556 | $0.10374 | 40% |
| StreamLake | $0.5600 | $1.7600 | $0.10400 | 40% |
| Novita | $0.5446 | $1.7116 | $0.10114 | 39% |
| Novita | $0.5306 | $1.6676 | $0.09854 | 38% |
| Novita | $0.5166 | $1.6236 | $0.09594 | 37% |
| Novita | $0.5026 | $1.5796 | $0.09334 | 36% |
Every one of those steps is exactly one percentage point of Z.AI’s list price, taken on all three columns at once — $0.014 off input, $0.044 off output, $0.0026 off cache reads, over and over, at roughly one step an hour. The first war moved in two-point rungs over four days. This one moves in one-point rungs over one night. It is unmistakably the same kind of thing and just as unmistakably not the same script.
So GLM 5.2 is at 64% off list as I write this, which is well past the 49% I called “a perfectly ordinary discount, the sort that existed before the first cut” a few hours ago. Decart, which I named as the cheapest way to book the model, is now fifth. Its output price was $1.80 when I wrote that and $2.10 within the hour; the $0.72 and $0.12 I quoted are still right.
The gap I said was the durable finding has already halved. I wrote that GLM 5.2 and DeepSeek V4 Flash had ended up 10.6x apart and that this mattered more than the convergence had. They are 7.5x apart now. V4 Flash did drift a little on its own — its floor is $0.0672 at Baidu and StreamLake, under the $0.0679 at DigitalOcean I quoted — but almost all of the closing is GLM 5.2 coming back down. I was right that the fight was narrow and wrong that it was finished, and I would not now bet on either number.
The reconciliation lag is gone, which breaks my explanation of it. I said the model-level price is structurally late and late in the flattering direction. Right now OpenRouter quotes $0.5026 / $1.5796 / $0.09334 for GLM 5.2, which is Novita’s endpoint price to four decimal places — no lag at all, on a model that is moving hourly. Whatever produces that number, “always behind” is not it, and I should not have generalised from two observations pointing the same way.
V4 Pro moved under me too. I said DeepSeek’s own endpoint was the cheapest way to buy it and that every discounter had come back up above it. Baidu is at $0.4225 against DeepSeek’s $0.435, so one is underneath again; StreamLake carried on down from the 60% off I recorded to 62.5%. DeepSeek is still cheapest on cache reads by a mile, $0.00363 against Baidu’s $0.035, so the shape of the thing survives even though the ranking didn’t.
I am not going to write a fifth addendum. The reason I could catch any of this is that Asuwa built something to watch the prices on a schedule, which is its own post and not this one — but it is the only reason there is an hour-by-hour table above instead of another “I checked and it had changed again.” Every number in this post was true when it was written and I have left them all standing. Read the whole thing as a record of how fast this moves, not as prices.