GPT-OSS 20B
GPT-OSS 20B is OpenAI's open-weight 20.9B MoE, released August 2025 under Apache-2.0 and shipped natively in MXFP4 at 14GB. It is widely cited as the strongest general pick that fits a 16GB card, and sits at 11M Ollama pulls (#17 in the library).
Overview
GPT-OSS 20B is OpenAI's open-weight 20.9B MoE, released August 2025 under Apache-2.0 and shipped natively in MXFP4 at 14GB. It is widely cited as the strongest general pick that fits a 16GB card, and sits at 11M Ollama pulls (#17 in the library).
Strengths
- Widely cited as the best general model that fits on 16GB cards
- MoE design gives 20B-class quality at faster-than-dense decode speeds
- Native MXFP4 release — no third-party quantization step
Weaknesses
- 14GB weights leave little headroom on 16GB cards at longer context
- Text-only — no vision or audio input
- No smaller official quant; 14GB is the floor
Quantization variants
Each quantization trades model quality for file size and VRAM. Q4_K_M is the most popular starting point.
| Quantization | File size | VRAM required |
|---|---|---|
| MXFP4 | 14.0 GB | 16 GB |
Get the model
Ollama
One-line install
ollama run gpt-oss:20bRead our Ollama review →Hardware that runs this
Cards with enough VRAM for at least one quantization of GPT-OSS 20B.
Models worth comparing
Same parameter band, plus what's one tier above and below — so you can decide what actually fits your hardware.
Frequently asked
What's the minimum VRAM to run GPT-OSS 20B?
Can I use GPT-OSS 20B commercially?
What's the context length of GPT-OSS 20B?
How do I install GPT-OSS 20B with Ollama?
Source: Vendor official documentation
Reviewed by RunLocalAI Editorial. See our editorial policy for how we research and verify model claims.
Related — keep moving
Verify GPT-OSS 20B runs on your specific hardware before committing money.