Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Moonshot AI launches Kimi K3 — 2.8T open frontier MoE with 1M context

Listen to this article

The open-weight ceiling just jumped a class. Moonshot AI’s Kimi K3 is live as a 2.8-trillion-parameter Mixture-of-Experts model with about 104 billion parameters active per token, native vision, and a 1-million-token context window — and Moonshot is promising the full weights later this month. Source: kimi.com/blog/kimi-k3 and arXiv abs/2607.24653.

That is not another mid-size open MoE with a louder blog. It is the first open model Moonshot is willing to call 3T-class, aimed at long-horizon coding, knowledge work, and agent loops that used to live behind closed APIs.

Architecture numbers that matter

Per Moonshot’s tech blog and the Kimi K3 technical report abstract on arXiv, the stack is built around three bets:

  • Scale: 2.8T total MoE parameters, ~104B activated
  • Context: native 1M-token window, with vision in the same backbone
  • Sparsity: Stable LatentMoE activating 16 of 896 routed experts per token

Moonshot says Kimi Delta Attention (KDA) and Attention Residuals improve information flow across sequence length and depth, and that the combined recipe delivers roughly 2.5× overall scaling efficiency versus Kimi K2. Treat the efficiency multiple as a vendor scaling-law claim until independent recreations exist. The parameter counts and expert routing ratios are the hard product facts.

The lab’s own evaluation framing is blunt for a launch post: K3 reaches frontier-level scores across coding, agentic, knowledge, reasoning, and vision suites, but still trails the strongest proprietary systems in their comparison set — Claude Fable 5 and GPT-5.6 Sol — while beating the other open and proprietary models they tested. That is a competitive map, not a trophy tour.

What is live now vs weights later

K3 is already usable on Kimi.com, Kimi Work, Kimi Code, and the Kimi API (kimi-k3). At launch, thinking runs at max effort by default; low- and high-effort modes are promised later. API pricing on the blog: $0.30 per million tokens for cache-hit input, $3.00 for cache-miss input, and $15.00 for output.

The open drop is staged. Moonshot says it is aligning with inference partners and open-source maintainers first, with full model weights due by July 27, 2026, and a fuller technical report alongside that release. Hugging Face is named as the destination (moonshotai/Kimi-K3 in the arXiv abstract). Until the checkpoint lands, “open frontier” means API-and-product open, not downloadable weights on your cluster.

Kimi Delta Attention claims to test

KDA is the architecture story builders should stress-test, not just cite. Moonshot positions it as hybrid linear attention paired with periodic global layers, designed to make million-token contexts and deep stacks practical. Attention Residuals add selective retrieval across depth instead of stacking residuals blindly.

What to verify once weights land: whether 1M context stays useful past retrieval demos; whether KDA serving stacks match API latency; and whether the open license and hardware floor make self-hosting realistic. Moonshot itself recommends large multi-accelerator deployments for local serving. Open weights that only run in a few labs are still a research event.

How K3 changes the open-frontier race

For years, open models scaled reasoning tricks faster than raw pretrain size, and the gap to the best closed systems widened on the foundation axis. K3 is Moonshot’s answer: push the open foundation into multi-trillion MoE territory and keep agentic RL and 1M context in the same release.

Builders get a hosted frontier option today and a promised weight drop before month-end. Closed labs get a public reminder that “open” no longer means permanently stuck near 1T. July 27 is the test: clean checkpoint and recipes move the open ceiling; a slip leaves a strong API with a delayed research gift.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile