RUNLOCALAIv38
->Will it run?Best GPUCompareTroubleshootStartLearnPulseModelsHardwareToolsBench
Run check
RUNLOCALAI

Independently operated catalog for local-AI hardware and software. Hand-written verdicts. Source-cited claims. Reproducible commands when we have them.

OP·Eruo Fredoline
DIR
  • Models
  • Hardware
  • Tools
  • Benchmarks
TOOLS
  • Will it run?
  • Compare hardware
  • Cost vs cloud
  • Choose my GPU
  • Prompting kits
  • Quick answers
REF
  • All buyer guides
  • Learn local AI
  • Methodology
  • Glossary
  • Errors KB
  • Trust
  • Suggest a feature
EDITOR
  • About
  • Author
  • How we make money
  • Editorial policy
  • Contact
LEGAL
  • Privacy
  • Terms
  • Sitemap
MAIL · MONTHLY DIGEST
Get monthly local AI changes
Monthly recap. No spam.
DISCLOSURE

Some links on this site are affiliate links (Amazon Associates and other first-class retailers). When you buy through them, we earn a small commission at no extra cost to you. Affiliate links do not influence our verdicts — there are cards we rate highly that we don't have affiliate relationships with, and cards that sell well that we refuse to recommend. Read more →

© 2026 runlocalai.coIndependently operated
RUNLOCALAI · v38
← /pulse/kimi-k3-2-8t-moe-announced-open-weights-july-27
INFOMODEL RELEASE·2026-07-18

Kimi K3: 2.8T-param MoE announced, open weights due July 27

▼ WHAT HAPPENED

Moonshot AI announced Kimi K3 on 2026-07-16: a 2.8-trillion-parameter MoE with 896 experts, 16 active per token (~1.8% of the pool), a 1M-token context window, native vision, and always-on reasoning. The architecture pairs Kimi Delta Attention (hybrid linear attention) with "Attention Residuals"; Moonshot claims ~2.5x scaling efficiency over K2, and quantization-aware training uses MXFP4 weights and MXFP8 activations from the SFT stage. At launch it leads Arena.ai's Frontend Code arena ahead of Claude Fable 5, and Moonshot's self-reported benchmarks put it above Opus 4.8 and GPT-5.5 while trailing Fable 5 and GPT-5.6 Sol. API pricing is $3/$15 per million input/output tokens. Full open weights are promised by 2026-07-27, reportedly under a Modified MIT license; until they land, every number is Moonshot-reported.

▼ OPERATOR ANGLE

At 4-bit, 2.8T parameters is roughly 1.4 TB of weights, so K3 stays an API model for anyone without a multi-node vLLM cluster even after the July 27 drop. The practical upside is the MXFP4 quantization-aware training: weights trained for 4-bit from the SFT stage should hold quality better in the aggressive community quants that follow.
SOURCE: https://simonwillison.net/2026/Jul/16/kimi-k3/[BLOG]

▼ ENTITIES REFERENCED

TOOLvLLMTOOLllama.cpp
[pulse item] · runlocalai.co/pulse/kimi-k3-2-8t-moe-announced-open-weights-july-27