Qwen 3.6 35B-A3B (MTP)
Qwen 3.6 35B-A3B with Multi-Token Prediction (MTP). The "A3B" suffix means ~3B activated parameters per token via Mixture-of-Experts — inference cost stays mid-tier while total parameter count climbs to 35B. MTP enables the model to predict multiple tokens per forward pass, materially speeding up generation throughput on supported runtimes (vLLM 0.20+, llama.cpp recent builds). Currently trending #1 on HuggingFace via unsloth's GGUF quantizations.
Our verdict
The first MoE-with-MTP combination to hit mainstream attention. Activated-param count (~3B) means a 12GB GPU can serve this model at usable speed, while the 35B total gives quality that punches above the activated-param tier. The MTP feature requires runtime support — only land it on vLLM 0.20+ or llama.cpp post-b9148. Big caveat: GGUF quants of MoE-with-MTP are still maturing; expect rough edges in the first month. Production-deploy after Aug 2026.
Overview
Qwen 3.6 35B-A3B with Multi-Token Prediction (MTP). The "A3B" suffix means ~3B activated parameters per token via Mixture-of-Experts — inference cost stays mid-tier while total parameter count climbs to 35B. MTP enables the model to predict multiple tokens per forward pass, materially speeding up generation throughput on supported runtimes (vLLM 0.20+, llama.cpp recent builds). Currently trending #1 on HuggingFace via unsloth's GGUF quantizations.
How to run it
Recommended runtime: vLLM 0.20+ (best MTP perf) or llama.cpp post-b9148 (best CPU compatibility). For Ollama users, the unsloth/Qwen3.6-35B-A3B-MTP-GGUF Q4_K_M is the practical default — pull with ollama pull hf.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF:Q4_K_M. VRAM math: model file ~20GB at Q4, KV cache ~3-5GB at 16K context with FP16 — comfortable on a 24GB GPU; tight on 16GB.
Family & lineage
How this model relates to others in its lineage. Family members share architecture and training-data roots; parent / children edges record direct distillation or fine-tune relationships.
Strengths
Weaknesses
Quantization variants
Each quantization trades model quality for file size and VRAM. Q4_K_M is the most popular starting point.
| Quantization | File size | VRAM required |
|---|
Get the model
HuggingFace
Original weights
Source repository — direct quantization required.
Hardware that runs this
Cards with enough VRAM for at least one quantization of Qwen 3.6 35B-A3B (MTP).
Models worth comparing
Same parameter band, plus what's one tier above and below — so you can decide what actually fits your hardware.
Frequently asked
Can I use Qwen 3.6 35B-A3B (MTP) commercially?
What's the context length of Qwen 3.6 35B-A3B (MTP)?
Source: huggingface.co/Qwen/Qwen3.6-35B-A3B-MTP
Reviewed by RunLocalAI Editorial. See our editorial policy for how we research and verify model claims.
Related — keep moving
Verify Qwen 3.6 35B-A3B (MTP) runs on your specific hardware before committing money.