Serving DeepSeek-V4-Flash-Vision on 4× DGX Spark: images, DSpark and DCP in one stack
DeepSeek-V4-Flash-Vision-Exp on four GB10 nodes, ported by hand onto the sparkinfer B12X stack before upstream support landed. The interesting part was not the vision tower — it was making DSpark speculative decoding and decode-context-parallel run together, which the stock stack refuses. Decode kept, KV doubled, images cost about a second.
- DeepSeek-V4-Flash
- vLLM
- multi-node
- TP4
- DCP
- DGX Spark
- FP8
- MoE
- MLA
- speculative decoding
- thinking mode
- tool calling
- long context
- multimodal
- time~30 min per node, plus ~10 min per boot
- yields4-node cluster serving DeepSeek-V4-Flash-Vision-Exp at 66–71 tok/s decode, 393K context, 3.5M-token KV pool, images in and thinking on