Distributed10 GbEexpert

What runs on Ray Serve multi-node distributed inference (4 nodes × 2× RTX 4090)?

Distributed serving across 4 machines, each with 2× RTX 4090. Ray Serve orchestrates replicas. 192 GB total / ~80 GB per replica. Built for high-concurrency request routing, not single-large-model deployment.

At a glance
Effective VRAM
80 / 192 GB
Not pooled
Speed penalty
~25%
vs ideal single-card
Recommended runtime
ray-serve
request routing
Setup difficulty
expert
~3600W peak
24
Models fit
7
Borderline
8
Not practical
Deployment recipe
Distributed inference homelab

Ray Serve replica orchestration recipe — multi-node aggregate throughput pattern.

Memory budget
Total VRAM
192 GB
Effective for inference
80 GB
42% of total
Not pooled

Multi-node Ray Serve clusters do NOT pool VRAM across machines for a single model. Each node hosts its own replica (or tensor-parallel rank within a tensor-parallel-2 group on dual-4090 nodes). Effective VRAM 'for a single model' is the per-replica capacity (~45 GB), not the cluster total. The 192 GB total is meaningful only for **aggregate throughput** — 4 replicas serving 4× the requests, not 4× the model size. This is the pattern that prosumer multi-machine deployments most often misunderstand. If your goal is 'run a 200B model that doesn't fit on one machine,' Ray Serve is the wrong tool — you want SGLang distributed or Exo-style layer split. Ray Serve's value is replica orchestration, autoscaling, and request routing.

Why total VRAM is not the whole story

Multi-node deployment. Each replica holds a full copy of the model — aggregate throughput scales, but single-model size is capped by per-replica capacity. Effective single-replica VRAM ~80 GB.

See the multi-GPU guide for topology tradeoffs, and the RunLocalAI Will-It-Run Framework for the citable fit-tier method.

Topology

Topology
distributed
Interconnect
ethernet-10g~10 GB/s
Component count
8 units
Components
Recommended runtime
ray-serve
Also: vllm, sglang
Recommended split strategy
request-routing
Also: tensor-parallel
Setup difficulty
expert
~3600W peak

Models that fit comfortably (24)

Effective VRAM utilization ≤ 85% at the smallest production quant. Comfortable headroom for KV cache.

90B·Q4_K_M60 GB·75% of effective VRAM·~25% speed penalty vs ideal
90B·AWQ-INT464 GB·80% of effective VRAM·~25% speed penalty vs ideal
78B·Q4_K_M52 GB·65% of effective VRAM·~25% speed penalty vs ideal
72B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
72B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
72B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
72B·AWQ-INT448 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·AWQ-INT448 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·AWQ-INT448 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·AWQ-INT448 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·Q4_K_M49 GB·61% of effective VRAM·~25% speed penalty vs ideal
70B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·AWQ-INT448 GB·60% of effective VRAM·~25% speed penalty vs ideal
70B·Q4_K_M48 GB·60% of effective VRAM·~25% speed penalty vs ideal
52B·Q4_K_M36 GB·45% of effective VRAM·~25% speed penalty vs ideal
49B·AWQ-INT432 GB·40% of effective VRAM·~25% speed penalty vs ideal
47B·Q4_K_M32 GB·40% of effective VRAM·~25% speed penalty vs ideal
46.7B·Q4_K_M33 GB·41% of effective VRAM·~25% speed penalty vs ideal
40B·Q4_K_M28 GB·35% of effective VRAM·~25% speed penalty vs ideal

Borderline (7)

Fits but with little headroom. KV cache for long context may not fit; verify before deployment.

123B·Q4_K_M88 GB·110% of effective VRAM·~25% speed penalty vs ideal

Effective VRAM utilization >110% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

120B·Q4_K_M84 GB·105% of effective VRAM·~25% speed penalty vs ideal

Effective VRAM utilization >105% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

Llama 4 Scout
Borderline
109B·Q4_K_M80 GB·100% of effective VRAM·~25% speed penalty vs ideal

Effective VRAM utilization >100% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

Sarvam 105B
Borderline
105B·Q4_K_M74 GB·93% of effective VRAM·~25% speed penalty vs ideal

Effective VRAM utilization >93% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

Sarvam 105B FP8
Borderline
105B·Q4_K_M74 GB·93% of effective VRAM·~25% speed penalty vs ideal

Effective VRAM utilization >93% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

Command R+ 104B
Borderline
104B·Q4_K_M70 GB·88% of effective VRAM·~25% speed penalty vs ideal

Combination fits but with little headroom. Verify KV cache budget for your target context window before committing.

104B·AWQ-INT472 GB·90% of effective VRAM·~25% speed penalty vs ideal

Combination fits but with little headroom. Verify KV cache budget for your target context window before committing.

Not practical (8)

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly. Drop to a smaller quant or move to a larger combo.

1600B·Q4_K_M1024 GB·1280% of effective VRAM·~25% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

Step-3
Not practical
1000B·AWQ-INT4640 GB·800% of effective VRAM·~25% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

Kimi K2.6
Not practical
1000B·Q4_K_M700 GB·875% of effective VRAM·~25% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

DeepSeek V4
Not practical
745B·AWQ-INT4480 GB·600% of effective VRAM·~25% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

675B·Q4_K_M448 GB·560% of effective VRAM·~25% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

671B·Q4_K_M420 GB·525% of effective VRAM·~25% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

671B·Q4_K_M420 GB·525% of effective VRAM·~25% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

Llama 4 405B
Not practical
405B·AWQ-INT4280 GB·350% of effective VRAM·~25% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

Benchmark opportunities

estimates, not measurements

Pending benchmark targets for this combo. Once measured, results land in the catalog as benchmarks.

Ray Serve 4-node × 2× 4090 + Qwen 3 32B (concurrency scan)
pending
Estimate: 30-45 tok/s per stream × 4-32 concurrent

Ray Serve replica orchestration. Each replica runs vLLM tensor-parallel-2; 4 replicas = 4 parallel serving paths. Measure aggregate throughput vs concurrency scan.

Going deeper