Qwen 3.6 27B (MTP)
Qwen 3.6 27B dense (not MoE) with Multi-Token Prediction. Sits between the 14B and 35B-A3B as a "single dense model with MTP throughput acceleration." Targets workloads where the MoE activated-param dance isn't ideal but you still want MTP's throughput gains. Released alongside the 35B-A3B and trending on HuggingFace via unsloth's MTP GGUF quants.
Our verdict
The "I want MTP but not MoE" pick in the Qwen 3.6 family. Dense 27B at Q4 is roughly 16GB on disk — fits 24GB VRAM with 16K context comfortably. MTP gains 30-50% effective tok/s on supported runtimes (vLLM 0.20+ / llama.cpp post-b9148). If you tested the 32B from the Qwen 3 generation and liked it, this is the natural upgrade.
Overview
Qwen 3.6 27B dense (not MoE) with Multi-Token Prediction. Sits between the 14B and 35B-A3B as a "single dense model with MTP throughput acceleration." Targets workloads where the MoE activated-param dance isn't ideal but you still want MTP's throughput gains. Released alongside the 35B-A3B and trending on HuggingFace via unsloth's MTP GGUF quants.
How to run it
Same runtime story as the 35B-A3B: vLLM 0.20+ or llama.cpp post-b9148 for MTP support. Without MTP support, the model still runs but loses the throughput acceleration. On Ollama, ollama pull hf.co/unsloth/Qwen3.6-27B-MTP-GGUF:Q4_K_M gets you up. VRAM math: ~16GB weights + ~3GB KV at 16K context = 19GB usable footprint.
Family & lineage
How this model relates to others in its lineage. Family members share architecture and training-data roots; parent / children edges record direct distillation or fine-tune relationships.
Strengths
Weaknesses
Quantization variants
Each quantization trades model quality for file size and VRAM. Q4_K_M is the most popular starting point.
| Quantization | File size | VRAM required |
|---|
Get the model
HuggingFace
Original weights
Source repository — direct quantization required.
Hardware that runs this
Cards with enough VRAM for at least one quantization of Qwen 3.6 27B (MTP).
Models worth comparing
Same parameter band, plus what's one tier above and below — so you can decide what actually fits your hardware.
Frequently asked
Can I use Qwen 3.6 27B (MTP) commercially?
What's the context length of Qwen 3.6 27B (MTP)?
Source: huggingface.co/Qwen/Qwen3.6-27B
Reviewed by RunLocalAI Editorial. See our editorial policy for how we research and verify model claims.
Related — keep moving
Verify Qwen 3.6 27B (MTP) runs on your specific hardware before committing money.