Moonshot AI launches Kimi K3 — 2.8T open frontier MoE with 1M context

The open-weight ceiling just jumped a class. Moonshot AI’s Kimi K3 is live as a 2.8-trillion-parameter Mixture-of-Experts model with about 104 billion parameters active per token, native vision, and a 1-million-token context window — and Moonshot is promising the full weights later this month. Source: kimi.com/blog/kimi-k3 and arXiv abs/2607.24653.
That is not another mid-size open MoE with a louder blog. It is the first open model Moonshot is willing to call 3T-class, aimed at long-horizon coding, knowledge work, and agent loops that used to live behind closed APIs.
Architecture numbers that matter
Per Moonshot’s tech blog and the Kimi K3 technical report abstract on arXiv, the stack is built around three bets:
- Scale: 2.8T total MoE parameters, ~104B activated
- Context: native 1M-token window, with vision in the same backbone
- Sparsity: Stable LatentMoE activating 16 of 896 routed experts per token
Moonshot says Kimi Delta Attention (KDA) and Attention Residuals improve information flow across sequence length and depth, and that the combined recipe delivers roughly 2.5× overall scaling efficiency versus Kimi K2. Treat the efficiency multiple as a vendor scaling-law claim until independent recreations exist. The parameter counts and expert routing ratios are the hard product facts.
The lab’s own evaluation framing is blunt for a launch post: K3 reaches frontier-level scores across coding, agentic, knowledge, reasoning, and vision suites, but still trails the strongest proprietary systems in their comparison set — Claude Fable 5 and GPT-5.6 Sol — while beating the other open and proprietary models they tested. That is a competitive map, not a trophy tour.
What is live now vs weights later
K3 is already usable on Kimi.com, Kimi Work, Kimi Code, and the Kimi API (kimi-k3). At launch, thinking runs at max effort by default; low- and high-effort modes are promised later. API pricing on the blog: $0.30 per million tokens for cache-hit input, $3.00 for cache-miss input, and $15.00 for output.
The open drop is staged. Moonshot says it is aligning with inference partners and open-source maintainers first, with full model weights due by July 27, 2026, and a fuller technical report alongside that release. Hugging Face is named as the destination (moonshotai/Kimi-K3 in the arXiv abstract). Until the checkpoint lands, “open frontier” means API-and-product open, not downloadable weights on your cluster.
Kimi Delta Attention claims to test
KDA is the architecture story builders should stress-test, not just cite. Moonshot positions it as hybrid linear attention paired with periodic global layers, designed to make million-token contexts and deep stacks practical. Attention Residuals add selective retrieval across depth instead of stacking residuals blindly.
What to verify once weights land: whether 1M context stays useful past retrieval demos; whether KDA serving stacks match API latency; and whether the open license and hardware floor make self-hosting realistic. Moonshot itself recommends large multi-accelerator deployments for local serving. Open weights that only run in a few labs are still a research event.
How K3 changes the open-frontier race
For years, open models scaled reasoning tricks faster than raw pretrain size, and the gap to the best closed systems widened on the foundation axis. K3 is Moonshot’s answer: push the open foundation into multi-trillion MoE territory and keep agentic RL and 1M context in the same release.
Builders get a hosted frontier option today and a promised weight drop before month-end. Closed labs get a public reminder that “open” no longer means permanently stuck near 1T. July 27 is the test: clean checkpoint and recipes move the open ceiling; a slip leaves a strong API with a delayed research gift.



