Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Thinking Machines Lab releases open-weights Inkling — 975B MoE enters the arena

Listen to this article

Thinking Machines Lab just put a frontier-scale multimodal MoE on Hugging Face and said the quiet part out loud: it is not the strongest model on the board. The bet is customization, not crown.

In a July 15 announcement, the lab released the Inkling model — a Mixture-of-Experts transformer with 975B total parameters, 41B active, context up to 1M tokens, pretrained on 45 trillion tokens of text, images, audio, and video. Full weights are public. Fine-tuning is live on Tinker. A lighter Inkling-Small preview (276B total / 12B active) ships beside it, with full Small weights promised after testing.

What Inkling ships in the open weights

Inkling reasons natively over text, images, and audio. Controllable thinking effort lets developers trade tokens for score instead of locking one operating point. The training brief is generalist — agentic coding, tool use, instruction following, factuality, vision, audio — not a single-benchmark specialist.

Architecture details matter for anyone who will host it. MoE layers follow a DeepSeek-V3-like recipe: 256 routed experts plus 2 shared, with 6 routed experts active per token. Attention interleaves sliding-window and global layers at 5:1 and uses relative positional embeddings instead of RoPE. Audio enters as dMel spectrograms; vision uses 40×40 patch embeddings in an encoder-free stack. Training ran on NVIDIA GB300 NVL72 systems with a Muon/Adam hybrid optimizer and asynchronous RL past 30 million rollouts.

Safety and calibration get real space. Inkling posts strong FORTRESS adversarial refusal among open-weights peers without tanking benign look-alikes, and Cognition’s censorship eval reportedly shows low compliance with propaganda-style refusals. Forecasting benches put it near closed models on calibration.

Active params vs total MoE size

975B total / 41B active is the efficiency claim in one line. You eat MoE storage and routing complexity; you run closer to a ~40B dense forward pass when the router behaves. Controllable effort is the second lever: on Terminal Bench 2.1, HLE, and IFBench sweeps, the lab says Inkling hits Nemotron 3 Ultra’s Terminal Bench score at roughly a third of the tokens.

Inkling-Small pushes harder: 12B active, similar recipe, early numbers that match or beat full Inkling on several reasoning and vision rows while trailing on factuality and some agentic suites. For latency-sensitive coding, grading, or synthetic-data jobs, Small is the obvious fork once weights land.

Thinking Machines is blunt: “Inkling is not the strongest overall model available today, open or closed.” Against Kimi K2.6, GLM 5.2, DeepSeek V4 Pro, and closed Sol / Fable / Gemini lines, it sits in the competitive middle — Design Arena Agentic Web Dev near the open-weights pack, SWE-Bench Verified at 77.6% in their bash-only harness, not at the absolute tip.

Tinker fine-tuning path on Hugging Face

The product play is weights plus a first-party fine-tune path. Inkling is on Tinker today at 64K and 256K context options, with a temporary 50% discount. An Inkling Playground in the Tinker console lets teams chat before they burn a run. Cookbook recipes now cover Inkling natively, including audio. The lab even demoed Inkling writing and running its own Tinker fine-tune job.

Full checkpoints sit on Hugging Face in original and NVFP4 forms for Blackwell inference. Serving partners named include Together AI, Fireworks, Modal, Databricks, and Baseten, with open inference and RL work across SGLang, vLLM, llama.cpp, and Hugging Face transformers.

How it positions against other open frontier MoEs

Chinese open MoEs have been setting the pace on raw leaderboard spikes. NVIDIA’s Nemotron line owns much of the efficient open narrative. Inkling’s counter is multimodal from scratch, effort control, Tinker as a first-party fine-tune surface, and a lab willing to say “customize us” louder than “beat everyone.”

Open weights at this scale change procurement math for teams that refuse API-only dependency. They do not automatically change the quality crown. If Inkling becomes the default base people actually fine-tune — because Tinker is easier than wrestling someone else’s MoE — Thinking Machines wins without topping every spider chart.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile