Hugging Face
The Hugging Face model card for DeepSeek-V4-Flash-Vision-Exp introduces it as the first experimental multimodal model in the DeepSeek-V4 family, building on the V4-Flash architecture with added visual modules and continued training to unlock visual understanding. On text-agent benchmarks (Terminal Bench 2.1 83.9, NL2Re The repository layout section confirms the release ships the tokenizer, prompt encoding reference, and a minimal PyTorch inference implementation covering the vision encoder and aligner, DFlash attention, MoE, Hyper-Connections, and the DSpark forward path, deliberately split between encoding/ and inference/ directorie