Pricing Comparison: Kimi K3 vs the frontier: what the open-source flagship actually costs
Kimi K3 was launched on July 16, 2026 as the largest open-weight model shipped to date, and its rate card sits right in GPT-5.6 Terra territory. The number that matters isn't on the rate card at all: how many tokens it burns to actually finish a task.

Kimi K3 vs the frontier at a glance
Question | Answer |
|---|---|
Rate card | $3.00 input / $15.00 output per 1M tokens, in GPT-5.6 Terra territory |
Cache pricing | $0.30 cached input, a 90% discount, DeepSeek-style |
Context window | 1M tokens, matching Claude Fable 5 and Gemini 3.1 Pro |
What's new | The largest open-weight model released to date: 2.8T-parameter MoE, weights public |
The catch | Token-inefficient. Per task, independent benchmarking puts it meaningfully above its rate card |
Takeaway | K3 prices the open frontier at closed-frontier rates. The story is margins, and it's bigger than the rate card |
Model overview
Kimi K3 is Moonshot AI's flagship, released July 16, 2026: a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per query, with a 1M-token context window and reasoning always on. It benchmarks in genuine frontier territory (third on GDPval-AA v2 behind Claude Fable 5 Max and GPT-5.6 Sol Max, second on AA-Briefcase), and the weights are public. That last fact is the entire reason this page exists.

The comparison set: OpenAI's gpt-5.6 family, Anthropic's Fable 5 and Opus 4.8, xAI's Grok 4.5, Google's Gemini 3.1 Pro, and DeepSeek V4, the other open-weight contender.
Key features
Capability | Kimi K3 | Closest comparisons |
|---|---|---|
Context window | 1M tokens | Fable 5 (1M), Gemini 3.1 Pro (1M), DeepSeek V4 (1M), Grok 4.5 (500k) |
Reasoning | Always on, reasoning_effort currently max only | Configurable on gpt-5.6 and Grok 4.5 |
Caching | Automatic, $0.30 per 1M cached input | DeepSeek V4 ($0.0028 cache hit), Anthropic (0.1x reads) |
Open weights | Yes, public | DeepSeek V4 yes, everything else no |
Tool calling | Tool choice constraints, dynamically loaded tools | Standard across the frontier |
Pricing snapshot
Prices verified July 19, 2026. Providers reprice often, see the update timeline below.
Model | Input / 1M | Output / 1M | Open weights |
|---|---|---|---|
Claude Fable 5 | $10.00 | $50.00 | No |
gpt-5.6-sol | $5.00 | $30.00 | No |
Kimi K3 | $3.00 ($0.30 cached) | $15.00 | Yes |
gpt-5.6-terra | $2.50 | $15.00 | No |
Gemini 3.1 Pro | $2.00 | $12.00 | No |
Claude Sonnet 5 | $2.00 (intro, $3.00 from Sep 1) | $10.00 (intro, $15.00 from Sep 1) | No |
Grok 4.5 | $2.00 | $6.00 | No |
gpt-5.6-luna | $1.00 | $6.00 | No |
DeepSeek V4 Pro | $0.435 | $0.87 | Yes |
Read the table twice. On the first pass, K3 sits comfortably mid-frontier: Terra-class rates for Sol-adjacent benchmarks looks like a bargain, but when you dig deeper it gets more complicated.
Here's the Intelligence vs. Cost per Intelligence Index Task, which puts Kimi at a very favourable and attractive quadrant

(See more on ArtificalAnalysis.ai)
What this costs at a real workload
The standardized scenario from the AI Pricing Index: 30M input + 30M output tokens per month.
Model | Monthly cost at rate card |
|---|---|
gpt-5.6-sol | $1,050 |
Kimi K3 | $540 |
gpt-5.6-terra | $525 |
Gemini 3.1 Pro | $420 |
Grok 4.5 | $240 |
DeepSeek V4 Pro | $39.15 |
Tokenizer caveat for the Claude rows above: Fable 5, Sonnet 5, and Opus 4.7+ use a new tokenizer that produces roughly 30% more tokens for the same text. Comparing per-token rates across providers understates Anthropic's effective cost per request by about that margin.
Now the catch. That table prices tokens, and tasks don't consume fixed token counts. K3 reasons on every request and reasons verbosely: per Artificial Analysis, completing comparable tasks runs 50-70% more expensive on K3 than on GPT-5.6, despite the near-identical rate card, because K3 spends far more tokens getting there. Apply that to the scenario and K3's effective cost lands roughly in the $810-920 range for the same work Terra does for $525. Grok 4.5's low rate card compounds the same way in the other direction, since it's notably token-efficient per task.
This is the general lesson hiding in the K3 launch, and it goes beyond one model: cost per token and token efficiency (how much intelligence each token carries) multiply into cost per task, and only cost per task is real. Rate cards are the price of the ingredient, not the meal. It's the same lesson as Anthropic's tokenizer change, which raised effective costs roughly 30% without touching the rate card. The invoice is decided by things the pricing page doesn't show.
Timeline of past updates
Moonshot AI, most recent first:
July 16, 2026: Kimi K3 released at $3.00 / $15.00, cached input $0.30. Weights public by July 27
November 2025 (verify): Kimi K2 Thinking released, extending K2 to reasoning workloads
July 2025: Kimi K2 released open-weight at aggressive sub-$1 input pricing, establishing Moonshot as the open-frontier price setter alongside DeepSeek
Trajectory: upward. K3 prices roughly 3-5x above K2's launch rates while moving from challenger to frontier class.
For the other providers' timelines, see their comparison pages or the AI Pricing Index.
The AI Pricing Index
Moonshot AI enters the index provisionally: 3 documented price events in the trailing 12 months (K2 launch pricing, K2 Thinking, K3 repricing), flagship direction up, reprice risk for builders medium and rising with each release. The index counts list-price events per provider per year; methodology and all providers at the AI Pricing Index.
One index observation worth making explicit: K3's launch is itself a pricing event for everyone else. Open-weight models at the frontier compress what closed labs can charge for equivalent capability, which historically shows up in the index as price cuts and cheaper mid-tier launches within 1-2 quarters. Watch the September Sonnet 5 increase for whether it survives contact with this launch.
What this means for your own pricing
The margin argument went around within hours of the launch, and it goes like this: a model layer dominated by 2-3 closed labs at very high inference margins pulls value out of every other layer, while every competitive force at the model layer (open weights at the frontier, vertically integrated entrants) pushes margin toward infrastructure and software instead. An open model needs the same compute as a closed one of similar size, so what open weights remove is margin rather than cost. Whether K3 specifically triggers that compression is an open question (its token inefficiency blunts the threat, the closed labs' product harnesses carry real weight, and frontier demand may even rise as cheaper intelligence makes more tasks economical), but the direction of pressure is not ambiguous.
If you build on these models you should consider these limitations:
Cost per task is your COGS, and it isn't on any pricing page. Budgeting on rate cards misprices token-hungry models against efficient ones. Your metering has to capture tokens per completed task per model, or your margins are estimates.
The orchestration should be your default architecture. The strongest model plans the task, cheaper or open models execute the bulk tokens. That's the model mix problem multiplied: rating, margin, and credit conversion per model per request, on one invoice.
Repricing velocity just went up again. A new frontier entrant with public weights means more price events per year across every provider you use. Whatever you charge customers, the conversion layer between provider prices and your prices gets re-derived more often, which is a billing-infrastructure property, covered under AI token pricing.
Intelligence per dollar is becoming the axis of competition at the model layer. For everyone building on top, the equivalent discipline is margin per task per model, and you can only manage what your billing ledger can compute. Solvimon for AI is built for exactly that arithmetic.
Related
AI Pricing Index. All providers, all timelines.
OpenAI vs DeepSeek. The other open-weight contender.
OpenAI vs Anthropic. The closed frontier under pressure.
Token economics. How token pricing plays out downstream.
Ready to Solve Monetization?
Solvimon monetizes small and large companies alike to drive more revenue through effective pricing and billing.
Why Solvimon
AI monetization that drives innovation
The Solvimon platform is extremely flexible allowing us to bill the most tailored enterprise deals automatically.
Ciaran O'Kane
Head of Finance
Solvimon is not only building the most flexible billing platform in the space but also a truly global platform.
Juan Pablo Ortega
CEO
I was skeptical if there was any solution out there that could relieve the team from an eternity of manual billing. Solvimon impressed me with their flexibility and user-friendliness.
János Mátyásfalvi
CFO
Working with Solvimon is a different experience than working with other vendors. Not only because of the product they offer, but also because of their very senior team that knows what they are talking about.
Steven Burgemeister
Product Lead, Billing



