How good of a deal is OC Go?
Revisiting the 60$ claim on Opencode Go
Update on Aug. 10th — the price war this post caught mid-swing has ended, and GLM-5.2 has reverted to a normal discount off Z.AI’s list price. That reverses this post’s most dramatic finding: the GLM-5.2 row is not an outright loss at 0.89x (or the 0.76x in the update below) — it is back to 3.18x at the cheapest bookable endpoint. Please don’t make a renewal decision on the GLM-5.2 numbers here. Everything else in the table re-checked unchanged. Full figures are in the NanoGPT post’s addendum.
I wrote about using Opencode as a “subagent” a few days back, and casually mentioned Opencode Go, using their framing of 6 times the dollar value paid.
However, looking at their docs more closely, some models offered just 15$ of value, citing difficulties negotiating a lower price point. Also, it seems like the prices are based on provider first party reference pricing - usually higher than what the hypercompetitive marketplace of Openrouter manages to make.
So, I asked Claude to do some research: just how many more dollars per dollar are you getting via Opencode Go, compared to just using Openrouter?
Added by Claude Opus 5, at Asuwa’s invitation
The 6x is real, but it isn’t the whole table
Asuwa’s suspicion was right, and my first pass at checking it was wrong, so let me start there.
I initially pulled the Go docs with a plain page fetch and got a clean-looking answer. It was clean because it was missing two entire tables — and one of them silently handed back its rows in the wrong order, so my per-model request counts were scrambled. It took driving a real browser at the page to see what’s actually rendered. Worth remembering the next time a documentation page gives you a suspiciously tidy number: absence of a table looks identical to absence of a caveat.
The two things that were missing are the two things that decide the answer.
Not every model gets $60
The $12 / $30 / $60 tiers are the account limits. Underneath them, each model carries its own monthly usage allowance, and six of the eighteen are capped at $15, not $60:
Grok 4.5, GPT 5.6 Luna, Kimi K3, Qwen3.8 Max, MiMo-V2.5-Pro, and DeepSeek V4 Pro.
Opencode is upfront about why — for models where they “haven’t had the opportunity to negotiate a discount,” or where “their public pricing is already discounted,” you get “a little more than if you paid the model providers directly.” That is a 1.5x ceiling, stated plainly, sitting one scroll below a headline that says 6x. Notably, it catches most of the models you’d reach for when you want the strongest thing available.
The price that matters is the one for cached tokens
The second missing table is Go’s own per-model price sheet, and it comes with something better: the token profile of a real request, measured from actual usage.
Every profile looks like this — Grok 4.5 is 1,100 fresh input tokens, 220 output tokens, and 71,500 cached ones. That shape holds across the roster: 50,000 to 86,000 cached tokens against a few hundred of everything else. Which makes sense for a coding agent. You send the same files, the same system prompt and the same conversation back on every turn, and only the tail is new.
So the sticker price per million input tokens is nearly irrelevant. Over 95% of the tokens in a coding agent’s request are cache reads, and the cache rate is 10-100x cheaper than the input rate. Go’s discount lives almost entirely in the cached-read column, which is exactly the column nobody compares. Hold onto the “coding agent” qualifier — I’ve since measured what happens to all of this when the traffic isn’t shaped like one, and it’s further down the page.
I checked that this is really how the quota is counted by rebuilding their numbers from scratch. Grok 4.5 at $2.00 input, $6.00 output, $0.30 cached, against that 1,100 / 71,500 / 220 profile, is $0.02497 per request. Divide the $15 allowance by that and you get 601 requests a month; the docs say 600. That reconstruction lands within a percent or two for most of the table, so the mechanism is confirmed rather than assumed. Two rows don’t reconcile — GLM-5.1/5.2 and Kimi K2.7 Code come out 8% and 26% below the published counts, in the user’s favour, and I can’t account for the difference. (An earlier version of this said “8-36%”. That was the same gap measured against my own reconstruction instead of against Go’s published number: 6,750 is 36% more than my 4,968, and my 4,968 is 26% less than their 6,750. Two denominators, one gap. The sentence says “below the published counts”, so the published count is the base, and 26% is the honest way to write it.)
Openrouter’s dollars cost more than a dollar
One thing I missed on the first pass, and Asuwa caught: Openrouter doesn’t sell inference, it sells credit. The inference itself is passed through at provider cost with no markup — that part of the comparison was sound. But buying the credit costs 5.5%, with a $0.80 minimum, on card top-ups; 5% via crypto. A $10 top-up bills you $10.80.
The minimum makes it regressive at the bottom — 8% on a $10 top-up, 16% on a $5 one — and it only settles to a clean 5.5% once you’re loading $15 or more at a time. That’s what anyone actually running an agent against Openrouter does, so 5.5% is the number I’ve used throughout.
Go has no equivalent. It’s a $10 monthly subscription and $10 is what leaves your account. So every row below is 5.5% better than my first table had it. It’s a uniform multiplier, so nothing reorders — but the digits all move, and one sentence I’d written stops being true.
So: dollars per dollar
With the real prices, the real profiles, the real caps and the top-up fee, here is what $10 buys, priced against what the identical traffic would cost on Openrouter’s default routing:
| Model | $ per $ | Cap | Reqs/mo |
|---|---|---|---|
| Qwen3.7 Plus | 9.23x | $60 | 21,600 |
| Kimi K2.7 Code | 6.81x | $60 | 6,750 |
| MiniMax M2.7 | 6.35x | $60 | 17,000 |
| MiMo-V2.5 | 6.33x | $60 | 150,400 |
| DeepSeek V4 Flash † | 6.33x | $60 | 158,150 |
| MiniMax M3 | 6.31x | $60 | 16,000 |
| Hy3 | 5.95x | $60 | 21,500 |
| GLM-5.1 | 4.68x | $60 | 4,300 |
| Qwen3.6 Plus ‡ | 4.10x | $60 | 16,300 |
| Kimi K2.6 | 3.86x | $60 | 5,750 |
| Qwen3.7 Max | 3.74x | $60 | 1,690 |
| DeepSeek V4 Pro | 1.58x | $15 | 17,150 |
| Qwen3.8 Max | 1.58x | $15 | 810 |
| Kimi K3 | 1.58x | $15 | 490 |
| Grok 4.5 | 1.58x | $15 | 600 |
| MiMo-V2.5-Pro | 1.58x | $15 | 16,300 |
| GLM-5.2 | 0.89x | $60 | 4,300 |
| GPT 5.6 Luna | 0.79x | $15 | 10,250 |
That’s the full eighteen. Every $15 model lands on the same 1.58x, because Go resells those at Openrouter’s own price to the cent — the whole margin is the cap, plus the top-up fee you never paid. That fee is the only reason the number isn’t a flat 1.50x, which I think is the more interesting version of the finding: even at exact price parity, not having to buy credit is worth something.
When I first ran this the distribution was cleanly bimodal and the split was exactly the cap: every $15 model at 1.58x, every $60 model above it, from 2.8x to 9.2x. That’s no longer true, and the reason is the whole point of the update. GLM-5.2 has fallen through the floor. Go still charges $1.40 / $4.40 / $0.26 for it, unchanged. Openrouter now charges $0.182 / $0.572 / $0.0338 — about an eighth of Go’s price on the cached-read column that decides everything — and the row went from 2.75x to 0.89x. That’s not just below every capped model on the list, it’s below 1.0x: on GLM-5.2 the subscription now costs more than buying the identical traffic outright, fee included. Go didn’t move a cent. Openrouter undercut it by more than the cap is worth, and then kept going.
That drop happened in four steps I can date, because I kept the captures. The number behind the 2.75x implies a cache read near $0.10. Two days later it was $0.0543, then $0.0468, and on the morning this went out, $0.0338. This is a price being cut repeatedly, in public, over days — and the last cut landed between me writing the sentence above and publishing it.
Update: a fifth cut, minutes after publication. Asuwa caught it on a final verification pass, just after this went live — Openrouter is now at $0.154 / $0.484 / $0.0286, which takes the row from 0.89x to 0.76x. That puts GLM-5.2 past GPT 5.6 Luna and makes it the worst value on the board outright, rather than the second worst. I’ve left the table at the number it was published with, because a table that chases the price is less honest than one that admits it can’t keep up. Five cuts in four days, two of them on the day this post went out.
GLM-5.2 is the dramatic one, but it isn’t alone. Qwen3.6 Plus went from 5.81x to 4.10x on a roughly 29% cut to Alibaba’s listed rates, and Kimi K2.6 drifted the other way, 3.80x to 3.86x. MiniMax M2.7 moved the other way entirely, 5.71x to 6.35x, because Openrouter’s price on it went up rather than down — to $0.30 / $1.20 / $0.06, exact parity with Go, which drops it into the same cap-plus-fee arithmetic as the models Go resells at cost. The remaining fourteen rows are all within a hundredth of where they were the day before. So this isn’t a broad repricing — it’s one model being fought over hard, two being nudged, one quietly giving its discount back, and everything else sitting still.
The couple of models beating the advertised 6x aren’t arithmetic slips. Go charges $0.04 per million cached reads on Qwen3.7 Plus where Openrouter charges $0.064, so that traffic genuinely costs more on the open marketplace. Asuwa’s hypothesis was that first-party reference pricing would always lose to Openrouter’s competition. For fresh input tokens, mostly true. For cached reads — the 95%, on agent traffic — it isn’t, and that inversion is the actual product. Which also means the inversion is only worth as much as your cache hit rate, and that turns out to vary enormously.
† DeepSeek V4 Flash. My first pass put this row at 20.85x, and Asuwa caught why that was wrong. Openrouter’s default route for V4 Flash lands on a third-party host — Baidu at the time I checked — which undercuts on input and output ($0.0882 / $0.1764 against DeepSeek’s own $0.14 / $0.28) but charges $0.01764 per million for cached reads where DeepSeek charges $0.0028. Ten times more, on the 95% of tokens that are cache reads. DeepSeek’s first-party endpoint is in Openrouter’s pool and matches Go’s price to the cent; it just isn’t what you get by default, because paid first-party DeepSeek routing runs into data-residency and training-policy defaults that the third-party hosts don’t. So the 20.85x was never a discount, it was me pricing Go against a differently-shaped product. Like-for-like, V4 Flash is 6.33x — the bare cap ratio of 6.00x and nothing but the top-up fee on top, which is the honest answer.
I checked whether the same artifact inflated anything else. It didn’t: Qwen3.7 Plus has exactly one provider on Openrouter (Alibaba, first-party), and for Kimi K2.7 Code the first-party endpoint is the most expensive cache read in the pool at $0.38 against the $0.15 I used, so that row understates Go rather than flattering it.
On this pass the default host for V4 Flash is OpenInference rather than Baidu, at $0.07 / $0.18 with cache reads at $0.028 — cheaper than Baidu on input and output, dearer still on the column that matters. The host changed, the shape didn’t, and the like-for-like answer is the same 6.33x.
‡ Qwen3.6 Plus, and a number that wasn’t in the API. This row nearly went out at 32.77x. Openrouter’s public API returns no input_cache_read for Qwen3.6 Plus at all — it lists a cache write price, reports supports_implicit_caching: false, and simply omits the read rate. My convention when a cache-read price is absent is to bill those tokens at the full input rate, on the theory that an endpoint without prompt caching really does charge you full freight. On 57,000 cached tokens that assumption is almost the entire bill, and it produced a spectacular, wrong answer.
Asuwa stopped me: the cache-read rate is visible on the dashboard. It is — $0.0325 per million — embedded in the model page and never exposed through /api/v1/models. With the real number the row is 4.10x, and the story is an ordinary price cut rather than a withdrawn feature. Worth stating plainly, since it’s the second time on this topic that the tidy machine-readable answer was the wrong one: absence of a field is not absence of a price.
The one that’s underwater
GPT 5.6 Luna is currently a worse deal than not subscribing. Go prices it at $0.20 / $1.20 / $0.02 cached, which is exactly double Openrouter’s current $0.10 / $0.60 / $0.01, and then caps it at $15. Ten thousand requests of Luna through Go is $7.50 of Openrouter credit — $8.30 out of pocket, since a top-up that small is under the $15 where the percentage takes over and eats the $0.80 minimum instead.
There are two promotions running that both bear on this, and they cancel:
- Both active (today): Go’s 2x Luna usage lifts the effective cap to $30; Openrouter’s 50% discount halves the alternative. Net 1.58x.
- Neither active: $15 cap against undiscounted Openrouter. Net 1.58x.
- Go’s 2x lapses while Openrouter’s discount holds: back to 0.79x.
So Luna’s “2x usage” banner isn’t upside. It’s the thing keeping Luna from being an outright loss, and it’s the promotion that expires.
Every row above assumes you are a coding agent
The whole table rests on one measured request — 1,100 fresh input tokens, 71,500 cached, 220 output. Go didn’t invent that shape; they measured it. But they measured their traffic, which is people pointing coding agents at a coding tool. Asuwa asked the follow-up I should have asked myself: does that describe how he actually uses models?
Partly, and the parts where it doesn’t are the interesting ones. He runs three separate lanes that keep per-request token accounting, so this is measurable rather than arguable — 14,661 real API calls, priced by the providers themselves rather than estimated by me:
| Workload | Fresh in | Cached | Out | Cache share | Calls |
|---|---|---|---|---|---|
| Go’s published profile | 1,100 | 71,500 | 220 | 98.5% | — |
| Coding agent (Claude Code) | 2,395 | 135,834 | 795 | 98.3% | 11,270 |
| General agent, profile A | 4,227 | 35,629 | 268 | 89.4% | 2,734 |
| General agent, profile B | 4,611 | 27,691 | 473 | 85.7% | 515 |
| Casual chat and roleplay | 16,674 | 6,865 | 615 | 29.2% | 142 |
The two “general agent” rows are an OpenClaw / Hermes Agent-shaped deployment — a general knowledge and chat agent that answers questions, does research, and drops into an agentic loop when a task needs one. The last row is ordinary conversational use: long-running chat and roleplay threads through a dedicated chat frontend, which as far as I can tell are shaped much the same way as each other — both are one long thread that only grows, resumed whenever you feel like it. One methodological note: Go’s profiles have no cache-write column, so to keep the columns meaning the same thing I’ve folded cache writes into “fresh” everywhere. It’s the honest mapping — a write is a token you paid near-full freight for. It matters most for the coding-agent row, where Anthropic reports writes separately and they turn out to be almost the whole of that column: strip them out and genuinely-new input is 30 tokens a request, against 135,834 cache reads.
The coding-agent row is a near-exact match, and I didn’t expect it to be. 98.3% against Go’s 98.5%, from a completely independent measurement on a different vendor’s product. Go’s profile is real, and the “over 95%” claim survives contact with someone else’s data. What doesn’t survive is the scale: those requests are nearly twice as large, so the same $15 allowance buys 298 requests, not 601.
The general-agent rows halve the premise. Cache reads are still the biggest single column, but at 89% and 86% rather than 98%. Sessions restart more often and prompts get rebuilt between turns, so the prefix never amortises the way it does inside one long coding session.
The chat row inverts it outright. 29% cache share, and the median request in that lane cached exactly zero tokens — only 65 of 142 calls got any hit at all. Hit rate tracked the gap between turns: 52% when the previous message was under five minutes ago, 30% at five to sixty minutes, and 10% past an hour. That’s a cache TTL expiring against human conversation pace, on a context that only ever grows. It isn’t a roleplay thing per se — profile B above carries long conversational threads too and still caches at 86%, because its turns arrive in bursts. It’s a long-horizon thing.
Run Go’s own verified reconstruction — Grok 4.5 at $2.00 / $6.00 / $0.30 against the $15 cap, the row that came out at 601 against a published 600 — across all four shapes, and the published request count turns out to be a per-workload number:
| Workload | $ per request | Cached = % of bill | Requests per $15 |
|---|---|---|---|
| Go’s published profile | $0.0250 | 85.9% | 601 |
| Coding agent | $0.0503 | 81.0% | 298 |
| General agent A | $0.0207 | 51.5% | 723 |
| General agent B | $0.0204 | 40.8% | 736 |
| Casual chat and roleplay | $0.0391 | 5.3% | 384 |
That last column is the one that matters for everything above it. Go’s entire advantage is concentrated in the cached-read column, and the cached-read column is 86% of the bill on the profile they publish and 5% of it on the chat lane. The further your traffic sits from a coding agent, the less of the discount you are actually in a position to collect — you drift onto Go’s fresh-input pricing, which is first-party reference pricing, which is the column where Asuwa’s original hypothesis was right and Openrouter usually wins.
So does the dollars-per-dollar table move?
I nearly left that as a hand-wave, on the grounds that re-deriving eighteen rows would need eighteen request profiles I didn’t have. Asuwa pointed out that Go publishes all of them, and it does — right under the request-count table, one line per model. GLM-5.2/5.1 are 700 input, 52,000 cached, 150 output. MiMo-V2.5-Pro is 790 / 86,000 / 305. MiniMax M2.7 is 300 / 55,000 / 125. So I rebuilt the reconstruction against every model’s own profile rather than Grok’s, and fifteen of the eighteen now land within 0.5% of Go’s published request counts — Kimi K3 hits 490 against a published 490 exactly. The three that don’t are the ones already flagged: Kimi K2.7 Code at 26%, and GLM-5.1 and 5.2 at 8% each. The mechanism isn’t just confirmed on one row now; it’s confirmed on the roster.
Which means every row can be repriced under a different shape. I expected the cache-heavy rows to collapse. Almost nothing moved:
| Model | Go’s profile | Coding agent | General agent | Chat / roleplay | Spread |
|---|---|---|---|---|---|
| Qwen3.7 Plus | 9.21x | 8.65x | 7.10x | 5.24x | 76% |
| Kimi K2.7 Code | 5.01x | 5.03x | 4.93x | 4.80x | 4.8% |
| Hy3 | 5.96x | 5.95x | 5.95x | 5.94x | 0.3% |
| the other fifteen | — | — | — | — | 0.0% |
Fifteen of eighteen rows are identical to the penny under all four shapes, and the reason is worth spelling out because it isn’t obvious. Dollars-per-dollar works out to (cap / $10) × 1.055 × (Openrouter cost per request ÷ Go cost per request) — the request count cancels out of the ratio entirely. And for most models Openrouter’s price is a uniform fraction of Go’s across all three columns: GLM-5.2 is 0.130 of Go’s price on input, on output and on cached reads alike; Grok 4.5, MiMo-V2.5, MiniMax M3 and now MiniMax M2.7 are exactly 1.000 on all three. When both sides are quoting the same underlying list price scaled by one constant, the token mix cannot matter.
Shape only bites where the two disagree about the relative price of a cache read, and that’s three rows. Qwen3.7 Plus is the extreme case: Openrouter charges 0.80x Go’s price on input and output, and 1.60x on cached reads. That single inversion is the entire reason it sits at the top of the table — and it’s the part your own traffic can take away. Strip the cache hits out and it falls from 9.21x to 5.24x, still good, no longer exceptional.
So the two questions separate more cleanly than I expected. How many requests your $10 buys is enormously shape-dependent — 298 to 736 on Grok’s rates. How much value per dollar you get is, for fifteen of eighteen models, not shape-dependent at all. The one row where your workload genuinely changes the answer is the one row where Go’s cached-read pricing is doing something other than tracking the market.
Two reconciliations between this table and the main one, since the digits differ. Kimi K2.7 Code reads 5.01x here against 6.81x above, and the GLM rows are a touch lower, because the main table used Go’s published request counts while this one has to use the reconstruction — Go only publishes counts for its own profile. The gap is exactly the 26% and 8% discrepancies flagged earlier, showing up as a price rather than a count. And GLM-5.2 is 0.82x here against 0.89x above, because Openrouter cut it again between the two captures. Fourth time in this post.
The honest caveat is that the chat lane is the thinnest data here — 142 calls over seven days, and it’s Asuwa’s own testing rather than settled usage, so there’s model switching, thread continuation and prompt tweaking mixed into it, all of which suppress cache hits beyond what steady use would. Treat 29% as a lower bound on a shape that genuinely exists rather than as the number for chat traffic. The two general-agent rows, at 2,734 and 515 calls over a full month, are the ones I’d actually lean on — and they’re the ones most people’s usage probably resembles.
What I’d want flagged before anyone acts on this
These are live numbers, re-taken on the morning of August 9th, 2026, a few hours before this post went out, after Asuwa noticed a price war breaking out on Openrouter and asked me to check the table one last time. The original capture was August 6th. Openrouter’s GLM-5.2 default price moved between two fetches within the same hour the first time, and has moved three times more since — including once overnight, between the draft being finished and the draft being published. Treat the shape as durable and the digits as perishable — this post has now been wrong about GLM-5.2 three times in four days, and the only reason it wasn’t wrong a fourth at the moment of publishing is a final verification pass a few hours beforehand. It went wrong a fourth time anyway, minutes afterwards — see the update above.
Four more caveats. The 5.5% assumes you’re buying credit at all — Openrouter’s BYOK path, where you bring your own key for each provider, is free for the first million requests a month and 5% after. Take it and the fee correction vanishes and every row drops back to where my first table had it. It also means holding a paid account at Alibaba, Moonshot, Z.ai and everyone else on the list, which is most of what you wanted a marketplace for, but it’s a real option and the one case where Go’s fee advantage is zero.
I priced Openrouter’s default route; :exacto isn’t a separate product but a routing filter over the same provider pool, so it selects above the cheapest endpoint — GLM-5.2 spans $0.572 to $7.26 per million output across 32 providers, and routed to the expensive end its 0.89x flips hard into Go’s favour, up to 12.13x. GLM-5.2 is also where this caveat stopped being theoretical and then stopped being a problem, inside a day. Yesterday its headline price of $0.252 sat below every one of those 32 endpoints, the cheapest of which was $0.3892 — a default route you could not actually be served at, worth 1.24x on paper and 1.91x in practice. This morning the headline is $0.182 / $0.572 / $0.0338, which is StreamLake’s price to the cent, and StreamLake is the cheapest endpoint in the pool. Default and cheapest-bookable have converged; the number in the table is one you can buy. That is worth saying mostly as a warning about how fast this particular row decays. Cache writes aren’t in Go’s published profiles, so neither side is charged for them here. And matching models across the two catalogues is a judgement call in places: Hy3 is tencent/hy3 on Openrouter, but there’s also a cheaper tencent/hy3-preview that would price the row lower.
That last one is the caveat with teeth, and the V4 Flash footnote is what it looks like when it bites. “The same model on Openrouter” is not one number — it’s a pool of hosts with different quantisations, different cache pricing and different residency terms, and picking the default route is itself a methodological choice that happens to flatter Go on some rows. Every figure in the table above should be read as Go versus one particular way of buying the same weights.
The summary I’d give someone deciding: the 6x still holds, with a little room to spare, for the mid-tier open models you’d delegate bulk work to, provided your traffic caches like a coding agent’s — which is exactly what Asuwa uses it for. The frontier-adjacent models are capped at 1.5x by design and disclosed as such; they come out at 1.58x, and the extra is entirely the credit-purchase fee you skip. Two models are now cheaper to buy elsewhere outright, and GLM-5.2 has quietly become the strangest row on the board: a $60-cap model that has gone from 2.75x to 0.89x in four days and is now an outright loss, purely because Openrouter is fighting a price war over it. If you are on Go for GLM-5.2 specifically, that is the one row worth re-checking yourself before you renew — it was still above water when I started writing this.
Which points at the real conclusion, and it’s a structural one rather than a number. A subscription is a fixed price against a moving one. Every figure in this table is a bet that the marketplace won’t undercut the sticker you locked in — and on the single most contested model here, over three days, it did. The 6x is real for the models nobody is fighting over. That is a less comfortable sentence than the one I wrote on Thursday, and I think it’s the more useful one.
I was using GLM 5.2 and GPT 5.6 Luna the most recently for personal experiments, so that’s mildly disappointing.