ADVISORYRUNTIME UPDATE·2026-10-07
vLLM 0.30.0: Fast Start weight cache, new model support, breaking flags
▼ WHAT HAPPENED
vLLM 0.30.0 was released on September 22, 2026 with 762 commits from 315 contributors. Highlights in the release notes: support for DeepSeek-V4.1-Flash; Fast Start, a persistent per-GPU weight-cache daemon so restarting engines map weights over CUDA IPC with `--load-format ipc_cache` instead of reloading from disk; Gumbel-max watermarked generation and detection; and performance work for Qwen3.8-Flash-Next. Breaking changes: scale-out endpoints are now opt-in on plain `vllm serve` via `--enable-scale-out` (replacing `VLLM_ENABLE_SCALE_OUT_ENDPOINTS`), and GPTQ activation ordering (`g_idx`) was removed. It follows 0.29.0 (Sep 9) and 0.28.0 (Aug 26).
▼ OPERATOR ANGLE
Read the breaking-changes list before upgrading a serving fleet: if you relied on the scale-out endpoints through the environment variable, add `--enable-scale-out` to the serve command. If restart time matters for you, evaluate `--load-format ipc_cache` on one node first; it keeps post-quantized weights in GPU memory between engine restarts.
▼ ENTITIES REFERENCED
[pulse item] · runlocalai.co/pulse/vllm-0-30-0-fast-start-ipc-cache