← /pulse/thinking-machines-inkling-975b-open-weights-release
INFOMODEL RELEASE·2026-07-18

Thinking Machines releases Inkling: 975B/41B-active MoE, Apache 2.0

▼ WHAT HAPPENED

Thinking Machines Lab released Inkling on July 15, 2026 — its first model, open-weight under Apache 2.0 with weights on Hugging Face at launch. It is a Mixture-of-Experts transformer: 975B total parameters, 41B active (6 of 256 routed experts plus 2 shared), context window up to 1M tokens, and multimodal input (text, image, audio in; text out), pretrained on 45T tokens. Two checkpoints ship: BF16, requiring at least 2TB of aggregated VRAM (8x NVIDIA B300 or 16x H200), and NVFP4, cutting the floor to 600GB (W4A4 on 4x B300, or W4A16 on 8x H200). Officially supported inference stacks at launch: SGLang, vLLM, TokenSpeed, Unsloth, and Hugging Face. A smaller sibling, Inkling-Small (276B total, 12B active), was previewed alongside it, with full weights promised once testing completes.

▼ OPERATOR ANGLE

Even the NVFP4 checkpoint needs 600GB of VRAM, so Inkling is multi-GPU server territory (4x B300 / 8x H200), not a home-rig model — vLLM and SGLang support is official at launch. The watch item for local operators is Inkling-Small: 276B total with 12B active, weights promised after testing, which puts a quantized build within reach of large unified-memory or multi-GPU workstation setups.

▼ ENTITIES REFERENCED

[pulse item] · runlocalai.co/pulse/thinking-machines-inkling-975b-open-weights-release