News

Kimi K3 Release: Benchmarks, Pricing and Open Weights Date

Published: July 17, 2026 · Updated: August 11, 2026

Moonshot AI has released Kimi K3, and the company is calling it the largest open-weight AI model built so far. The Beijing-based startup, backed by Alibaba, rolled the model out on July 16, 2026, just ahead of the World Artificial Intelligence Conference in Shanghai — timing that looks anything but accidental.

Kimi K3 carries 2.8 trillion total parameters, roughly 75% larger than DeepSeek’s V4-Pro and well ahead of Zhipu AI’s GLM-5.2. That single number is why so many headlines are calling this a watershed moment for open-source AI. It also marks a comeback of sorts for Moonshot, whose position had slipped over the past year and a half after DeepSeek’s rapid rise reshaped the Chinese AI market.

What Happened

Moonshot describes Kimi K3 as its most capable model yet, built for long-horizon coding, knowledge work, and complex reasoning. It’s a Mixture-of-Experts model that activates only 16 of 896 experts per request, which keeps inference relatively efficient despite the enormous total parameter count.

The model went live on kimi.com, the Kimi app, Kimi Code, and via API on launch day. It’s also reachable through OpenRouter, which is useful if you’d rather not set up a direct Moonshot account just to try it out.

Moonshot says this is the ninth time in the past year that a Kimi release has pushed open-source model scale to a new record, a run that started with Kimi K2 at roughly 1 trillion parameters. K3 nearly triples that number in one jump.

Key Details and Context

Kimi K3 supports a 1-million-token context window and understands text, images, and some video natively, rather than bolting vision on as a separate module. That’s a big step up from the Kimi K2 family, which topped out at a 256,000-token window.

The model is built on two techniques Moonshot developed and had already published as open research: Kimi Delta Attention, a hybrid linear-attention mechanism, and Attention Residuals, described as a drop-in replacement for standard residual connections that improves scaling. On the API side, Kimi K3 is compatible with the OpenAI SDK, which lowers the integration barrier for developers already building on GPT-style tooling.

On pricing, K3 costs $3 per million input tokens and $15 per million output tokens, with a discounted $0.30 per million rate for cached input. Those rates are flat across the model’s full 1-million-token context window — there’s no premium for longer prompts, unlike some competing APIs. That pricing puts K3 roughly in line with Anthropic’s Sonnet-tier models — noticeably more expensive than Moonshot’s own earlier Kimi K2.6 model, but still well under what the very top US frontier models charge.

kimi 3

Kimi K3 Benchmarks: How It Compares

Independent evaluator Artificial Analysis ranks Kimi K3 fourth among the models it has tested, with an Intelligence Index score of 57. That places it behind Claude Fable 5 and GPT-5.6 Sol, but ahead of Claude Opus 4.8, GPT-5.5, Claude Sonnet 5, and GLM-5.2. Moonshot’s own launch materials describe similar positioning, claiming K3 outperforms Opus 4.8 and GPT-5.5 on several tasks while trailing Fable 5 and GPT-5.6 Sol.

It’s worth treating the newest numbers with some caution. Runtime testing so far shows around 62 output tokens per second and just under 2 seconds to first token on the provider configurations tested solid figures, but not yet verified across every deployment or every benchmark suite different outlets are citing. Independent, large-scale testing of a 2.8-trillion-parameter model takes time, and most of what’s public right now is only a few hours to a day old.

Is Kimi K3 Actually Open Source Yet?

Here’s the detail a lot of the coverage blurs together: Kimi K3 is usable right now through Moonshot’s hosted app and API, but the full open weights the actual downloadable model files aren’t out yet. Moonshot has committed to releasing them by July 27, 2026. Until that date arrives, “open-weight” is a promise with a deadline attached, not something you can actually download today.

That distinction matters for the “largest open-source model” framing running through most headlines. The claim is Moonshot’s own, and it will only be fully checkable once independent labs can pull the weights and run their own tests.

How to Try Kimi K3 Right Now

You don’t have to wait for the weights to test Kimi K3 yourself. It’s accessible today through a few different routes:

Subscription plans on Kimi start at roughly ¥199, or about $28, with launch top-up bonuses running through mid-August 2026.

Why This Matters

Kimi K3 lands in one of the most crowded stretches the open-model space has seen, arriving within days of other major releases from rival labs. Its scale and pricing show Chinese AI developers are now willing to spend more to compete head-on with the biggest US systems, rather than only undercutting them on price, which was the strategy behind most earlier open-model releases.

For everyday users and developers in Pakistan, the practical takeaway is simpler than the geopolitics: a genuinely frontier-class model is now reachable through a normal API key and a relatively modest budget, without needing a US-based subscription or payment method.

What Happens Next

The date to watch is July 27, 2026, when Moonshot has said the full model weights will be published. That’s the point where “largest open-source model” becomes something independent labs can actually verify, instead of relying on Moonshot’s own self-reported benchmark tables.

It’s also worth watching how quickly inference providers can scale a model this size to competitive speeds. A 2.8-trillion-parameter model is a genuinely heavy thing to serve at scale, and whether third-party providers beyond Moonshot itself can offer it at solid tokens-per-second economics is still an open question.

Final Takeaway

Kimi K3 is a real leap for Moonshot AI bigger, more capable, and more expensive than anything the company has shipped before. It’s live and usable today, but the full open-weight release that justifies the “largest open-source model” headline is still roughly two weeks away. That gap is worth keeping in mind before putting too much weight on the record-breaking claims currently making the rounds.

```