Serving DeepSeek-V4-Flash on 4x DGX Spark: the 3.76.1 recipe
An in-house vLLM image serving DeepSeek-V4-Flash at ~71 tok/s decode, 400K context, thinking enabled by default — distilled from a month of benchmarking, three confirmed bugs, and one configuration audit. Extracted verbatim from production.
- DeepSeek-V4-Flash
- vLLM
- multi-node
- TP4
- DGX Spark
- FP8
- MoE
- thinking mode
- long context
- time~45 min per node, plus 15–20 min first boot
- yields4-node cluster serving DSV4-Flash at ~71 tok/s, 400K context, thinking on by default