Claude vs. Kimi K3: should you switch?

Moonshot's 2.8-trillion-parameter Kimi K3 is the largest open-weight model ever released. It tops Claude on a frontend-coding leaderboard, undercuts Opus on price, and slots into Claude Code with three environment variables. A comparison of capability, cost, and the switching path — and the case for staying.

TL;DR

K3 is the credible answer to a high Claude bill. It matches or beats Claude on several coding benchmarks at ~40% below Opus blended cost, and the weights are open — nobody can deprecate it out from under you. On the hardest long-horizon work, Opus 5 is still the safer default; Fable 5 sits above both. Keep Opus or Fable where a failed run is expensive. Give the routine work to K3 if it wins your eval.
2.8TK3 parameters — largest open-weight model ever
$2.90K3 input per 1M tokens, vs $5.00 for Claude Opus 5
1Mtoken context window — both models
32.9K3 output tokens/sec, vs 53.7 for Opus 5

1The contenders

Claude is Anthropic's closed-weight family: Haiku 4.5 for cheap fast tasks, Sonnet 5 as the volume workhorse, Opus 5 for hard agentic coding, and Fable 5 at the top of the range. The relevant tiers here are Opus 5 and Fable 5; that's what K3 has to beat. All current tiers run a 1M-token context window with 128K max output. The moat isn't just the models — it's the harness around them: Claude Code, adaptive thinking with effort control, prompt caching, server-side tools, and a first-party API that the whole toolchain is tuned against.

Kimi K3 is Moonshot AI's flagship, released July 16, 2026, with the full 2.8-trillion-parameter weights published on July 27. It's a mixture-of-experts model (KDA plus attention residuals for computational efficiency), natively multimodal, with a 1M-token context window — up from 256K on its 1.1T-parameter predecessor K2.7. It's positioned squarely at agentic coding: navigating large repositories, using tools, debugging, and iterating against screenshots, logs, and test output.

2What the benchmarks say

K3 placed in the top three across six coding benchmarks at launch, and took the #1 spot on Arena.AI's Frontend Code Arena — ahead of both Claude Fable 5 and GPT-5.6 Sol. On the independent Artificial Analysis index, though, Claude Opus 5 still edges it out on raw intelligence and generates substantially faster.

MeasureKimi K3ComparisonRead
Frontend Code Arena (Arena.AI)1,679 — #1ahead of Claude Fable 5, GPT-5.6 SolK3's headline win; visual/frontend work is its strongest suit
Terminal-Bench 2.188.3−0.5 vs GPT-5.6 Soleffectively tied at the top
SWE Marathon42.0 — leadsleads all modelslong-horizon software tasks
ProgramBench (raw pass)77.8 — leadsleads all modelsvendor-reported
DeepSWE / FrontierSWE67.5 / 81.2top-3vendor-reported
AA Intelligence Index57Opus 5 (high effort): 59independent; Claude ahead
Output speed32.9 tok/sOpus 5: 53.7 tok/sClaude ~60% faster in generation
Time to first token4.2sOpus 5 (high): 12.8sOpus figure includes thinking time
Benchmark caveats Several of the coding scores above are Moonshot's own reported numbers, and launch-week benchmarks are marketing as much as measurement. The independent signal (Artificial Analysis, early hands-on comparisons) is more modest: K3 is genuinely in the frontier pack, slightly behind Opus 5 on intelligence, clearly behind on throughput, clearly ahead on price. For your workload, the only benchmark that matters is running both on a week of your real tasks.

The long view: capability over time

Zoom out and the question looks different. On SWE-bench Verified — the share of real GitHub issues a model can resolve end-to-end — frontier models went from a third of issues in mid-2024 to over 90% today. The line to watch is the aqua one: open-weight models trailed the frontier by 20–30 points for two years, and K3 is the first to close the gap almost entirely.

Real-world coding ability, mid-2024 → mid-2026
SWE-bench Verified — % of 500 real GitHub issues resolved, by model release date
AnthropicOpenAIOpen-weight
0%25%50%75%100%Jul '24Jan '25Jul '25Jan '26Jul '26Claude 3.5 Sonnet 33.4%o1 48.9%DeepSeek V3 42.0%Kimi K2 65.8%GPT-5.6 Sol 96.2%Claude Opus 5 96.0%Kimi K3 93.4%
Data tablechart data
ModelLabReleasedSWE-bench Verified
Claude 3.5 SonnetAnthropicJun 202433.4%
Claude 3.5 Sonnet v2AnthropicOct 202449.0%
DeepSeek V3Open-weightDec 202442.0%
o1OpenAIDec 202448.9%
DeepSeek R1Open-weightJan 202549.2%
Claude 3.7 SonnetAnthropicFeb 202562.3%
o3OpenAIApr 202569.1%
Claude Sonnet 4AnthropicMay 202572.7%
Kimi K2Open-weightJul 202565.8%
GPT-5OpenAIAug 202574.9%
Claude Sonnet 4.5AnthropicSep 202577.2%
Claude Opus 4.5AnthropicNov 202580.9%
Claude Opus 4.7Anthropicearly 202687.6%
GPT-5.3OpenAI202685.0%
Kimi K3Open-weightJul 202693.4%
Claude Opus 5Anthropic202696.0%
GPT-5.6 SolOpenAI202696.2%
Best published agentic-harness score at or near each model's release; 2026 release dates approximate. The benchmark is nearing saturation above 90%, which is why newer suites (Terminal-Bench, SWE Marathon) exist at all.

3What it costs

ModelInput $/1MOutput $/1MContextNotes
Kimi K3$2.90$15.001Mvia OpenRouter; caching cuts effective cost 60–80%
Claude Haiku 4.5$1.00$5.00200Kcheap tier, not a K3 peer
Claude Sonnet 5$3.00$15.001M$2.00 / $10.00 intro through Aug 31, 2026
Claude Opus 5$5.00$25.001Mthe model most switchers are leaving
Claude Fable 5$10.00$50.001Mtop tier

Artificial Analysis puts the blended real-world cost at $2.31 per 1M tokens for K3 versus $3.85 for Opus 5 at high effort — roughly a 40% saving against the model most people are thinking of leaving. Fable 5 doubles Opus again. The question isn't whether K3 is cheaper per token (it is); it's how many Opus-or-Fable-grade tasks you'd feed to a model that fails them. Price the retries, not the tokens.

The long view: price over time

Capability climbed while price collapsed. GPT-4 launched at $30 per million input tokens in March 2023; today's flagships sit at $5, and yesterday's frontier is nearly free — Epoch AI estimates GPT-4-level performance has fallen from ~$20 to ~$0.40 per million tokens. Note where K3 sits: at $2.90 it is expensive for an open-weight model — DeepSeek and Kimi's own K2 launched at a tenth of that — because it's priced as what it is, a frontier model that happens to publish its weights.

Flagship input price at launch, 2023 → 2026
$ per 1M input tokens, log scale, by model release date
AnthropicOpenAIOpen-weight
$0.25$1$5$302023202420252026GPT-4 $30Claude 3 Opus $15GPT-4o $5Claude Opus 4 $15DeepSeek V3 $0.27Kimi K2 $0.6GPT-5 $1.25Claude Opus 5 $5Kimi K3 $2.90
Data tablechart data
ModelLabReleasedInput $/1M at launch
GPT-4OpenAIMar 2023$30
Claude 2AnthropicJul 2023$11.02
GPT-4 TurboOpenAINov 2023$10
Claude 3 OpusAnthropicMar 2024$15
GPT-4oOpenAIMay 2024$5
DeepSeek V3Open-weightDec 2024$0.27
o1OpenAIDec 2024$15
Claude Opus 4AnthropicMay 2025$15
Kimi K2Open-weightJul 2025$0.6
GPT-5OpenAIAug 2025$1.25
Claude Opus 4.5AnthropicNov 2025$5
Kimi K3Open-weightJul 2026$2.90
Claude Opus 5Anthropic2026$5
First-party list price at launch, input tokens, before caching discounts. Log scale — each gridline is a different order of magnitude.

4Where each one wins

DimensionClaudeKimi K3
Hardest agentic coding, repo-scale debuggingStronger safer default when a failed run is expensiveClose top-3 on most suites
Frontend & visual iterationStrong#1 inspects screenshots and iterates against what it sees
Generation speed53.7 tok/s32.9 tok/s
Latency to first token12.8s at high effort4.2s
Price vs. Opus~40% cheaper blended
Open weights, self-hosting, fine-tuningNoYes 2.8T weights published
Ecosystem & harness maturityFirst-party Claude Code, thinking/effort controls, server toolsCompatible rides Anthropic-compatible endpoints
Native multimodal / videoImagesBuilt-in from the ground up
The practical part

5The switching path

The reason this question is even live: Claude Code doesn't care whose model answers it. It reads a base URL, a token, and a model name from the environment, and Moonshot ships an Anthropic-compatible endpoint built specifically to be a drop-in target. Switching is three variables, not a new tool:

shell
# Point Claude Code at Kimi K3 — Moonshot first-party or OpenRouter,
# one API key from either console is enough.
export ANTHROPIC_BASE_URL="<anthropic-compatible endpoint>"
export ANTHROPIC_AUTH_TOKEN="<your Moonshot or OpenRouter key>"
export ANTHROPIC_MODEL="kimi-k3"
claude   # same harness, different brain

Get the endpoint URL and a key from Moonshot's console (platform.kimi.ai) or from OpenRouter. Unset the three variables — or keep them in a separate shell profile — and you're back on Claude. Switching cost is near zero in both directions.

On OpenRouter, pick the routing mode deliberately: Exacto routes for tool-calling accuracy, which is what an agentic harness needs; Nitro optimizes speed, Balanced price. Flaky tool calls that look like model weaknesses are often just provider variance.

6What you give up

And what you gain beyond price: weights you control. No deprecation schedule, no retention policy you didn't write, fine-tuning on your own data, and the option — if you ever have the hardware — of running the thing yourself.

7What happens next

Everything above is a snapshot: K3's weights shipped a week ago, and Anthropic's intro pricing ends August 31. Switching today is a bet on where both curves go. Six forecasts, each with the signal that confirms or kills it.

PredictionConfidenceWatch forWhat it means for you
K3 serving prices fall below list as independent hosts come online, the way DeepSeek and Qwen prices did within weeks of their weights dropping. K3 is cheap to serve for its size: ~104B active parameters, weights shipped natively in MXFP4.HighA second and third provider on OpenRouter under $3/$15 by mid-August. As of July 30 there is exactly one listing, at effectively Moonshot's list price.The price case gets stronger if you wait; don't lock in anything annual this month.
The "Kimi K3 License" becomes the real fight. The weights are downloadable but the license is proprietary, not Apache. Whether it permits commercial re-hosting decides whether the row above happens at full speed.Med-highIndependent hosts appearing (the license permits it) or still absent by mid-August (it doesn't). A definitional "is this even open?" dispute either way.Read the license before building on self-hosting or fine-tuning rights; "open weights" may be narrower than it sounds.
Routing beats picking a side. Fireworks ran 1,030 agentic tasks through a K3-vs-frontier router and it beat either model alone, sending 72–96% of tasks to K3. The Exacto/Nitro modes in section 5 are an early version of the same idea.HighRouter layers becoming the default way teams consume models rather than a power-user trick.The hybrid verdict below is the end state, not a compromise.
Washington restricts government use, not yours. A draft executive order on Chinese open models was reported in July and remains unsigned; ~200 startups, then Nvidia, Microsoft, and Meta, pushed back publicly. A blanket ban on a model already downloaded worldwide is hollow; procurement rules and hosting-liability conditions are the live options.MediumA Commerce or OMB action scoped to government procurement or federal contractors, likely before September.If your work touches government or regulated clients, consume K3 through a US host — that's the compliance hedge.
Claude pricing keeps falling. Opus 5 launched at half of Fable's price with selectable effort levels — a margin response to exactly this competition, and it won't be the last cut.Med-highAn Opus price cut, a cheaper fast-Opus tier, or intro rates spreading across the range after August 31.The savings from switching shrink even if you never switch. Re-run the math September 1.
Derivatives bring scrutiny. Community quantizations of K3 hit Hugging Face within days of the weights, uncensored fine-tunes follow every major open release, and open weights keep no capability gate a fine-tune can't remove. A security incident traced to a K3 derivative would put employer model policies in motion quickly.MediumA documented intrusion or worm attributed to a K3-based agent within a quarter.If your org might restrict Chinese-origin models later, keep the Claude path warm — switching back takes minutes.
The thread: open weights turn "which model is better" into "who captures inference margin." Moonshot gave up pricing power the day the weights shipped; it bought ubiquity. Every prediction above is that trade playing out — hosts competing the margin away, licenses clawing some back, incumbents cutting prices to defend theirs, and regulators deciding whose infrastructure it runs on.

Forecasts as of July 30, 2026; confidence labels are judgment, not measurement. The macro side of this — what commodity inference does to an AI buildout financed on off-balance-sheet debt — is tracked weekly in the newsletter.

8The verdict

This is a portfolio decision, not a clean switch.

Sources

  • Epoch AI — LLM inference prices have fallen rapidly but unequally
  • TokenMix — AI API pricing history, GPT-4 to 2026
  • CodeAnt — SWE-bench leaderboard 2026
  • MorphLLM — Claude benchmark scores & SWE-bench Verified history
  • Claude pricing and model specs from Anthropic's published API pricing as of July 2026.