Kimi K3: 2.8T-param MoE announced, open weights due July 27
▼ WHAT HAPPENED
Moonshot AI announced Kimi K3 on 2026-07-16: a 2.8-trillion-parameter MoE with 896 experts, 16 active per token (~1.8% of the pool), a 1M-token context window, native vision, and always-on reasoning. The architecture pairs Kimi Delta Attention (hybrid linear attention) with "Attention Residuals"; Moonshot claims ~2.5x scaling efficiency over K2, and quantization-aware training uses MXFP4 weights and MXFP8 activations from the SFT stage. At launch it leads Arena.ai's Frontend Code arena ahead of Claude Fable 5, and Moonshot's self-reported benchmarks put it above Opus 4.8 and GPT-5.5 while trailing Fable 5 and GPT-5.6 Sol. API pricing is $3/$15 per million input/output tokens. Full open weights are promised by 2026-07-27, reportedly under a Modified MIT license; until they land, every number is Moonshot-reported.
▼ OPERATOR ANGLE
At 4-bit, 2.8T parameters is roughly 1.4 TB of weights, so K3 stays an API model for anyone without a multi-node vLLM cluster even after the July 27 drop. The practical upside is the MXFP4 quantization-aware training: weights trained for 4-bit from the SFT stage should hold quality better in the aggressive community quants that follow.