Kimi K3 Pricing Breaks the Cheap-Chinese-Model Pattern at $3/$15 Per Million Tokens
- Moonshot AI released Kimi K3 on July 16, 2026, a 2.8 trillion parameter open-weight reasoning model with a 1 million token context window, according to TechCrunch and Tom's Hardware.
- Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens, with a discounted $0.30 per million on cache hits, the same tier as Anthropic's Claude Sonnet 5 and well above Moonshot's earlier Kimi K2.6 at $0.95/$4.
- Kimi K3 reached number one on Arena.ai's Frontend Code Arena with a 76 percent pairwise win rate, beating Anthropic's Claude Fable 5, per reporting from Simon Willison and Tom's Hardware in mid-July 2026.
- Full open weights for Kimi K3 are scheduled for release by July 27, 2026, letting anyone self-host the model instead of relying only on Moonshot's hosted API.
Chinese AI labs have spent 2026 competing almost entirely on price, undercutting US frontier models by 60-90% on a per-token basis. Moonshot AI's Kimi K3, released July 16, 2026, breaks that pattern. At $3 per million input tokens and $15 per million output tokens, Kimi K3 pricing lands in the same range as Anthropic's Claude Sonnet 5, not underneath it, even though it beats several frontier rivals on coding benchmarks.
Kimi K3 Pricing: $3 In, $15 Out, Not the Usual Discount
Moonshot's own Kimi K2.6, released earlier this year, priced at $0.95 per million input tokens and $4 per million output tokens. Kimi K3 more than triples that on both ends, landing at $3/$15, with a steeper discount of $0.30 per million on cache-hit input tokens. That puts it roughly in line with Claude Sonnet 5's published rate of $2 per million input and $10 per million output, and makes Kimi K3 the most expensive model any Chinese lab has released to date, according to Tom's Hardware. For a category that built its reputation on undercutting US labs, that's a deliberate repositioning: Moonshot is betting the model is good enough to charge flagship rates rather than compete purely on cost.
What 2.8 Trillion Parameters Actually Buys
Kimi K3 is a mixture-of-experts model with 2.8 trillion total parameters, a 1 million token context window, and an always-on "thinking mode" baked into the default behavior rather than offered as a toggle. That scale makes it the largest open-weight model released to date, per coverage from Cryptobriefing and Tom's Hardware. Independent benchmark trackers place it fourth among all frontier models overall, trailing only Claude Fable 5 and GPT-5.6 Sol, and edging past Claude Opus 4.8 on several evaluations. Whether that benchmark position holds up on messy, real-world tasks rather than curated test suites is the usual caveat with any week-old release.
Beating Claude Fable 5 on Code, Not on Price
The clearest signal behind Kimi K3's pricing confidence is its coding performance. The model reached number one on Arena.ai's Frontend Code Arena with a 76% pairwise win rate, beating Anthropic's Claude Fable 5 head-to-head, a result flagged by independent developer Simon Willison and covered by Tom's Hardware in mid-July 2026. Winning a benchmark against a top-tier proprietary model is a strong opening claim for any lab, but it's a different claim than "cheapest option," which is the lane Chinese open-weight models have occupied almost without exception since early 2025. Kimi K3 is explicitly asking developers to pay for quality rather than for a discount.
Open Weights Are Coming July 27
Kimi K3 is live now via Moonshot's own apps and API, Kimi.com, the Kimi Work desktop client, Kimi Code, and aggregators like OpenRouter. But the full open weights are not out yet, they're scheduled to release by July 27, 2026. Once that happens, anyone with the hardware can self-host Kimi K3 rather than paying Moonshot's hosted rate at all, which is a meaningful difference from closed models like GPT-5.6 or Claude Sonnet 5 where the API price is the only price. That eleven-day gap between hosted launch and open weights is also a window in which Moonshot can observe real usage and adjust pricing before the self-hosting option removes their leverage entirely.
Why the Pricing Choice Matters Beyond Moonshot
Kimi K3 lands in a market where model pricing changes on a timescale of weeks, and where the "cheapest" and "best" model for a given task can flip within days of a new release. A model that intentionally prices at flagship levels rather than discount levels is a useful signal that per-token cost and model quality are decoupling somewhat, at least at the top of the market, which makes blind provider loyalty a worse bet than actually comparing options. That's the same logic behind bring-your-own-key tools like ByteChat, which let you hold API keys for Kimi, Claude, GPT, and other providers side by side and see what each one actually costs for the work you're doing, instead of assuming a Chinese-labeled model automatically means a cheaper bill.
Frequently asked questions
How much does Kimi K3 cost?
Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens, with a discounted $0.30 per million on cache-hit input tokens, according to Moonshot AI's published API pricing as of July 2026.
Is Kimi K3 cheaper than Claude Sonnet 5?
No, they're close. Claude Sonnet 5 is priced at $2 per million input tokens and $10 per million output tokens, slightly below Kimi K3's $3/$15, putting the two models in the same general pricing tier rather than Kimi undercutting Claude as earlier Chinese models typically did.
When do Kimi K3's open weights release?
Moonshot AI has said full open weights for Kimi K3 will be available by July 27, 2026, roughly eleven days after the model's July 16 hosted launch on Kimi's apps and API.
A Chinese lab pricing at flagship rates instead of discount rates is the more interesting story here than the benchmark win.