BLK · COMPARE · MODELS

DeepSeek V3 vs Qwen 3 235B-A22B — flagship MoE showdown

Reviewed 2026-05-152 min read
TL;DR

DeepSeek V3 for reasoning + coding heritage. Qwen 3 235B-A22B for instruction-following + faster wall-clock. Both need 192 GB unified memory or multi-GPU SXM.

Option A

DeepSeek V3 (671B MoE)

D

671B params · DeepSeek License · deepseek

64K ctx · ~405.1 GB @ Q4 · Commercial OK
Option B

Qwen 3 235B-A22B

S

235B params · Apache 2.0 · qwen

128K ctx · ~141.9 GB @ Q4 · Commercial OK
WINNER
VERDICT
Qwen 3 235B-A22B wins 6 of 6 dimensions for local AI workloads.
MODEL · A
DeepSeek V3 (671B MoE)
PARAMS: 671BCTX: 64KFAMILY: deepseekLICENSE: commercial OK
MODEL · B★ EDGE
Qwen 3 235B-A22B
PARAMS: 235BCTX: 128KFAMILY: qwenLICENSE: commercial OK
WORKLOAD WINNERS

Who wins each use case

Each row is the dimension-weighted verdict for that use case. Use case weights live in src/lib/model-battle/comparator.ts and are public.

7 workloads
Chat
Daily-driver assistant — multi-turn conversation
Qwen 3 235B-A22B
Qwen 3 235B-A22B wins. Context length (tokens): 131K.
Coding agent
Aider / Cline / Cursor — diff edits + refactors
Qwen 3 235B-A22B
Qwen 3 235B-A22B wins. Context length (tokens): 131K.
Agentic workflows
Long-running tool-using agent loops
Qwen 3 235B-A22B
Qwen 3 235B-A22B wins. Context length (tokens): 131K.
RAG / retrieval
Long-context document QA
Qwen 3 235B-A22B
Qwen 3 235B-A22B wins. Context length (tokens): 131K.
Reasoning / math
Chain-of-thought heavy, output-token-heavy
Qwen 3 235B-A22B
Qwen 3 235B-A22B wins. Context length (tokens): 131K.
Creative writing
Style + tone, long-form generation
Qwen 3 235B-A22B
Qwen 3 235B-A22B wins. Context length (tokens): 131K.
Vision-language
Neither model is multimodal — text-only.
Neither
Both are text-only LLMs. Pick a vision-language model for image input.
SPEC RATIOS
Parameters
Total parameter count (active + inactive for MoE)
671B
235B
DeepSeek+186%
Context length
Max input + output the model can handle
65536tokens
131072tokens
Qwen+100%
VRAM footprint @ Q4
Weights only — add ~20% for KV cache + overhead
405GB
142GB
Qwen+186%
Our rating
RunLocalAI editorial rating (when set)
9.0/100
0.0/100
DeepSeek
FIT MATRIX

What hardware actually runs each model

VRAM math against the canonical hardware ladder. The largest context window that fits with headroom at Q4_K_M appears in each cell.

Hardware tierDeepSeek V3 (671B MoE)Qwen 3 235B-A22B
RTX 3090 (24 GB)
Used $700-1,000 — the local-AI workhorse
OOM
OOM
RTX 4090 (24 GB)
Used $1,400-1,900 — current consumer flagship
OOM
OOM
RTX 5090 (32 GB)
Retail $2,000-2,500 — Blackwell consumer
OOM
OOM
Mac M4 Max (64 GB unified)
$4,000-5,000 — Apple Silicon flagship
OOM
OOM
Dual RTX 3090 (48 GB pooled)
~$1,500-2,000 — workstation budget build
OOM
OOM
H100 (80 GB)
$25K+ — datacenter / cloud rental tier
OOM
OOM
Comfortable — fits with headroom Borderline — tight, may need quant downgrade Doesn't fit — needs bigger card or CPU offload
COST PER MILLION TOKENS

On RTX 4090 @ Q4_K_M — bandwidth-derived estimate

Computed from each option's sustained TDP × predicted tok/s at $0.16/kWh. Cloud baseline: Claude Sonnet 4.6 (input + output).

DeepSeek V3 (671B MoE)
$12.366/M tok
Qwen 3 235B-A22B
$4.331/M tok
Claude Sonnet 4.6 (input + output)
$9.000/M tok

Electricity-only cost — excludes the upfront hardware purchase, cooling, and amortized component depreciation. Hardware ROI math lives at /cost-vs-cloud; this line is for "is the marginal token cheaper than Claude?" not "should I buy this rig instead of paying Anthropic." MODELED ESTIMATE.

Both are flagship open-weight Mixture-of-Experts models targeting frontier capability. DeepSeek V3 is larger overall (671B params total, 37B active) with strong reasoning + coding heritage. Qwen 3 235B-A22B is more compact (235B total, 22B active) with sharper instruction-following.

Realistic local deployment for either is Mac Studio M-Ultra-class (192 GB+ unified memory) or multi-GPU NVLink/SXM rigs. Both ship under permissive licenses. The pick is workload + budget.

The verdict for reasoning workloadsPick → Qwen 3 235B-A22B

decisive edge for Qwen 3 235B-A22B wins 5 of 10 dimensions (1 loss, 4 ties). Verdict reasoning below — no percentage shown on purpose (why).

Qwen 3 235B-A22B is the better fit for reasoning on the dimensions we score, taking 5 of 10 rows. The weighted score (5% vs 40%) reflects use-case priorities: reasoning (40%) outweighs everything else. Both models are worth running — this just tells you which one to reach for first.

DIMENSION MATRIX
DimensionDeepSeek V3 (671B MoE)Qwen 3 235B-A22BEdge
Editorial rating (1-10)
Editor rating — single human assessment across reasoning, fluency, tool-use, instruction-following.
9.0unratedtie
Parameters (B)
671.0B235.0BDeepSeek
Context length (tokens)
66K131KQwen
License (commercial OK?)
✓ DeepSeek License✓ Apache 2.0tie
Decode tok/s on NVIDIA GeForce RTX 4090 (Q4_K_M)
Bandwidth-derived estimate. Smaller models stream faster on the same hardware.
1.4 tok/s3.9 tok/sQwen
Fits comfortably on NVIDIA GeForce RTX 4090?
✕ 543.2 GB short✕ 174.6 GB shortQwen
Cost to run (local, Q4)
Smaller model → less VRAM + less electricity per token. Cross-reference with /cost-vs-cloud for $-anchored math.
405.1 GB at Q4_K_M141.9 GB at Q4_K_MQwen
Community popularity
Editorial popularity score — proxy for runtime support breadth + community recipe availability.
8896tie
Multimodal support
text onlytext onlytie
Released
2024-12-262025-04-29Qwen
DECISION BY HARDWARE TIER

Which model wins on which VRAM tier. Picks update based on which one fits comfortably + which one’s strengths are unlocked by the available headroom.

VRAM tierPickWhy
Under 128 GBQwen 3 235B-A22BQwen's smaller footprint makes it the only realistic local option in this tier. V3 needs offload that tanks throughput.
192 GB Mac Studio M3 UltraQwen 3 235B-A22BQwen 3 235B fits comfortably at Q4 with headroom; V3 tight.
Multi-H100 SXM (640 GB+)DeepSeek V3 (671B MoE)Now V3's full capability is unlocked. Pick it for the reasoning + coding edge.
QUESTIONS OPERATORS ASK

DeepSeek V3 or Qwen 3 235B-A22B for high-end local AI?

DeepSeek V3 for reasoning + coding-heavy workloads (its R1 lineage shows). Qwen 3 235B-A22B for instruction-following + agentic loops where the smaller active-parameter count translates to faster wall-clock. Both fit on a 192 GB Mac Studio M-Ultra at heavy quant; neither is a single-card consumer rig.

What hardware can actually run these?

DeepSeek V3 needs roughly 350-400 GB for FP8 weights; Qwen 3 235B-A22B needs roughly 140 GB. At Q4 both shrink: V3 to ~170 GB, Qwen 3 235B-A22B to ~90 GB. Realistic deployments: Mac Studio M3 Ultra 192 GB unified (Qwen 3 235B comfortably, V3 tight), multi-A100/H100 SXM nodes, or rented cloud.

Are these worth running locally vs the cloud API?

For privacy-sensitive workloads or sustained-load deployments where API cost compounds, yes. For occasional use, the API is cheaper. Run /cost-vs-cloud math with your actual monthly token volume before committing to the hardware.

CUSTOM
Swap either model →
Pick different models + see fit across 8 hardware tiers.
DETAIL
DeepSeek V3 (671B MoE)
Editorial verdict, how to run, hardware guidance.
DETAIL
Qwen 3 235B-A22B
Editorial verdict, how to run, hardware guidance.

Comparison data computed from live catalog rows + the model-battle comparator (src/lib/model-battle/comparator.ts). For arbitrary pairings outside this curated list, use /model-battle to pick any two models + your hardware.