Fits comfortably

Running DeepSeek R1 Distill Qwen 7B on NVIDIA GeForce RTX 3080 16GB (Mobile)

NVIDIA GeForce RTX 3080 16GB (Mobile) runs DeepSeek R1 Distill Qwen 7B comfortably at Q8_0 with 6 GB of headroom for context.

By Eruo Fredoline·Latest benchmark evidence Jun 2, 2026

Recommended quant

Q8_0
Highest quality that fits

Quick start with Ollama

1. Install
ollama pull deepseek-r1:7b
2. Run
ollama run deepseek-r1:7b

Default quant in Ollama is Q4_K_M. To use a different quant, append it: deepseek-r1:7b-q5_K_M.

Variants and what fits

QuantizationFile sizeVRAM requiredFits on NVIDIA GeForce RTX 3080 16GB (Mobile)?
Q4_K_M4.7 GB6 GB
Yes
Q8_08.1 GB10 GB
Yes

Real benchmarks

ToolQuantContexttok/sVRAM usedDateEvidenceExport
Q4_K_M4,09680.3 tok/sJun 2, 2026Measured here
operator: fred-oline

Frequently asked

Can NVIDIA GeForce RTX 3080 16GB (Mobile) run DeepSeek R1 Distill Qwen 7B?

NVIDIA GeForce RTX 3080 16GB (Mobile) runs DeepSeek R1 Distill Qwen 7B comfortably at Q8_0 with 6 GB of headroom for context.

What quantization should I use?

Q8_0 is the highest-quality variant of DeepSeek R1 Distill Qwen 7B that fits in 16 GB VRAM. Lower-bit quants will be smaller but lose some quality.

How fast will it be?

Measured at 80.3 tok/s on this combination in our testing.

See also: DeepSeek R1 Distill Qwen 7B, NVIDIA GeForce RTX 3080 16GB (Mobile), all benchmarks.

Reviewed by RunLocalAI Editorial. See our editorial policy.