← /pulse/qwen3-8-27b-open-weights-apache
INFOMODEL RELEASE·2026-10-07

Qwen3.8-27B: Apache 2.0, vision, 16.5 GB at Q4_K_M

▼ WHAT HAPPENED

Alibaba's Qwen team published Qwen3.8-27B on Hugging Face in August 2026 (the model repo was last updated on August 14). It is a dense 27B model with a vision encoder under the Apache 2.0 license. The model card lists 64 layers and a context length of 262,144 tokens natively, extensible up to 1,000,000. GGUF builds are available: in unsloth/Qwen3.8-27B-GGUF the Q4_K_M file is 16.5 GB, Q5_K_M 19.8 GB, Q6_K 22.0 GB and Q8_0 29.0 GB, with a 0.9 GB vision projector. Ollama carries it as `qwen3.8:27b` and `qwen3.8:27b-mlx`.

▼ OPERATOR ANGLE

This is the size class that fits a single 24 GB card: Q4_K_M at 16.5 GB leaves several gigabytes for context and the vision projector on an RTX 3090, 4090 or 5090. On a 16 GB card it will not fit fully on the GPU at Q4. We have no first-party benchmark for it yet, so treat any tokens-per-second figure you see elsewhere as someone else's measurement.
[pulse item] · runlocalai.co/pulse/qwen3-8-27b-open-weights-apache