Kimi K3 (Moonshot AI): The Open Model That Made Me Rethink “Open-Source AI”
Quick summary: Kimi K3 is Moonshot AI’s new 2.8-trillion-parameter model, released July 16, 2026, and it’s now officially the largest open-weight AI model ever shipped. It trails Claude Fable 5 and GPT-5.6 Sol on overall benchmarks but beats Opus 4.8 and GPT-5.5 on several coding and agentic tasks – at roughly half the per-task cost of Opus. Full open weights land July 27, 2026, under a modified MIT license.
Last updated: July 23, 2026. Open weights are scheduled for July 27 – check back after that date for confirmation the release happened on schedule.
I’ve been skeptical of “open-source AI” headlines for a while now, mostly because so many of them turn out to be a smaller model dressed up with an impressive parameter count and not much else behind it. Kimi K3 is the first one this year that made me actually stop and re-read the benchmark tables twice. (If you’re tracking the broader AI news cycle alongside this one, our AI News hub has the rest of what shipped this month.)
Moonshot AI – the Beijing-based startup behind the Kimi chatbot – dropped K3 on July 16, 2026, timed just ahead of the World Artificial Intelligence Conference in Shanghai. It’s not a quiet incremental update. It’s a 2.8-trillion-parameter model, about 75% larger than DeepSeek’s V4 Pro, and Moonshot is calling it the first “open 3T-class” system in the world, according to VentureBeat’s launch coverage. Full weights are due July 27, so as of writing, you can use it through the API and Moonshot’s own site, but you can’t download it yet.
Here’s what I actually found once I dug past the headline number.
What Kimi K3 Actually Is?
K3 runs on a sparse mixture-of-experts architecture – 896 experts total, with only 16 activated per token, which works out to roughly 1.8% of the whole model firing at once, per Tom’s Hardware’s technical breakdown. That’s the trick that makes a 2.8T-parameter model even remotely practical to serve. It also carries a full 1-million-token context window and native multimodal vision, so it’s not just a text model wearing a big number.
Two architectural details are worth knowing if you’re technical at all: Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), both aimed at squeezing more reasoning quality out of less compute. Moonshot says this combination lets K3 use about 21% fewer output tokens than its predecessor, K2.6, on equivalent tasks – which matters more than it sounds, because output tokens are where API costs actually pile up.
Where It Actually Ranks Against Claude and GPT?
This is the part most coverage gets lazy about, so I pulled the actual numbers instead of repeating “rivals top U.S. models.”

On Artificial Analysis’s composite leaderboard, K3 scored an Elo of 1,547 – a 732-point jump over K2.6, and second only to Claude Fable 5, per MLQ News’s report on the release. On GDPval-AA v2, a benchmark built around real-world tasks across 44 occupations, K3 landed third overall (1,687), behind Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), but ahead of Claude Opus 4.8 (1,600), as VentureBeat detailed in its coverage. On Arena’s Frontend Code benchmark specifically, K3 actually took first place at 1,679 points, ahead of Fable 5, in blind developer testing.
So the honest verdict: K3 doesn’t beat the best closed models overall, but it beats Opus 4.8 and GPT-5.5 on a meaningful chunk of coding and agentic work, and it outright wins on frontend code generation. For an open-weight model, that’s not a rounding error – that’s a real result. (If you’re more focused on prompting than benchmarks, our Best ChatGPT Prompts guide covers the prompt-structure side of getting good output out of any of these models, K3 included.)
What It Costs (And the Catch Nobody’s Talking About)
API pricing is $3 per million input tokens and $15 per million output tokens – matching Claude Sonnet 5, and about half the per-task cost of Opus 4.8, according to Moonshot’s own comparison. There’s also a cache-hit discount: $0.30 per million tokens on cached input versus $3 on a cache miss, which rewards repeated-context workloads like long coding sessions, per OpenRouter’s pricing page for the model.
The catch: K3 currently only offers one reasoning effort level, “max,” and independent testers have flagged heavy reasoning-token consumption as a result. One tester reported 13,241 reasoning tokens just to generate a simple SVG – about $0.25 for a task that should cost a fraction of that on a model with adjustable effort, as noted in MLQ News’s breakdown of independent testing. If you’re planning to run K3 at scale, budget for that, not just the headline per-token price.
Why This Matters Beyond the Benchmarks?
The bigger story here isn’t really about K3’s Elo score – it’s about what it signals. Moonshot’s Kimi chatbot has become one of China’s most-used consumer AI products, crossing $200 million in annualized recurring revenue back in April. The company is reportedly in funding talks targeting a $30 billion valuation, up from roughly $3.77 billion raised to date, backed by Alibaba, Tencent, and several major Chinese VC firms, according to MLQ News’s reporting on the company’s funding history.
That context matters because K3 landing this close to Claude Fable 5 and GPT-5.6 Sol – while being fully open-weight – is exactly the kind of result that’s been fueling the broader “is the US AI lead shrinking” conversation all month. Whether you buy that framing or not, K3 is a real data point in it, not just a press release. (For how this shift is showing up on the consumer side, see our breakdown of Google AI Mode, which covers a similar competitive squeeze from Google’s angle.)
FAQ: Kimi K3
Is Kimi K3 better than Claude or GPT?
Not overall - it trails Claude Fable 5 and GPT-5.6 Sol on composite benchmarks. But it beats Claude Opus 4.8 and GPT-5.5 on several coding and agentic tasks, and it currently leads the Frontend Code Arena benchmark outright.
How much does Kimi K3 cost to use?
$3 per million input tokens and $15 per million output tokens on standard pricing, with a cache-hit rate of $0.30 per million tokens for repeated context - roughly half of what Opus 4.8 costs per task.
Is Kimi K3 really open-source?
It's open-weight under a modified MIT license, meaning the trained model weights will be publicly downloadable (from July 27), which is a meaningfully more open release than a closed API-only model - though "open-weight" isn't identical to fully open training data and code.
Where This Leaves You
If you’re building something that leans heavily on coding or agentic workflows and you’ve been priced out of Opus-tier models, K3 is genuinely worth testing once the open weights land on July 27 – the frontend-code result alone makes it worth a look. If you just want the single best model regardless of cost, Fable 5 and GPT-5.6 Sol are still ahead on paper.
I’ll be running K3 through some real coding tasks once the weights are out on July 27 and updating this article with hands-on results – drop a comment if there’s a specific workload you want me to test it against.



