← back to the cookbook
August 19, 2026·10 min read

Serving DeepSeek-V4-Flash on 4x DGX Spark: the 3.76.1 recipe

An in-house vLLM image serving DeepSeek-V4-Flash at ~71 tok/s decode, 400K context, thinking enabled by default — distilled from a month of benchmarking, three confirmed bugs, and one configuration audit. Extracted verbatim from production.

  • DeepSeek-V4-Flash
  • vLLM
  • multi-node
  • TP4
  • DGX Spark
  • FP8
  • MoE
  • thinking mode
  • long context
  • time~45 min per node, plus 15–20 min first boot
  • yields4-node cluster serving DSV4-Flash at ~71 tok/s, 400K context, thinking on by default