← back to the cookbook
September 2, 2026·14 min read

Serving DeepSeek-V4-Flash-Vision on 4× DGX Spark: images, DSpark and DCP in one stack

DeepSeek-V4-Flash-Vision-Exp on four GB10 nodes, ported by hand onto the sparkinfer B12X stack before upstream support landed. The interesting part was not the vision tower — it was making DSpark speculative decoding and decode-context-parallel run together, which the stock stack refuses. Decode kept, KV doubled, images cost about a second.

  • DeepSeek-V4-Flash
  • vLLM
  • multi-node
  • TP4
  • DCP
  • DGX Spark
  • FP8
  • MoE
  • MLA
  • speculative decoding
  • thinking mode
  • tool calling
  • long context
  • multimodal
  • time~30 min per node, plus ~10 min per boot
  • yields4-node cluster serving DeepSeek-V4-Flash-Vision-Exp at 66–71 tok/s decode, 393K context, 3.5M-token KV pool, images in and thinking on