Currently listed through these providers:
Model details
DeepSeek V4.1 Flash Flex
DeepSeek V4.1 Flash Flex is a 552-billion-parameter Mixture of Experts model built on a new Causal Encoder Decoder architecture, activating only 8B parameters for input and 16B for output. It is positioned as the smallest and most efficient member of its architecture family, and is described as the first Flash variant to ship with native vision and multimodal understanding, making it suitable for assistants that need to reason over both text and images. Open weights are published on Hugging Face, allowing researchers and developers to inspect, fine-tune, and deploy the model on their own infrastructure rather than depending solely on a hosted endpoint.
The model is designed around practical agent and high-throughput workloads, pairing a one-million-token context window with a heavily compressed KV cache of roughly 890 bytes per token that is claimed to be about four times smaller than its predecessor and orders of magnitude smaller than V1. It supports thinking and non-thinking modes, tool calling, and structured JSON output, and is reported to surpass DeepSeek V4 Pro on performance, cost, speed, and task completion time while reaching a score of 88.1 on CyberGym. These characteristics make V4.1 Flash Flex a strong fit for coding agents, long-context retrieval-augmented generation, batch processing pipelines, and multimodal assistants where context length, throughput, and operating cost dominate the design.
Quick Info
Powered by- Provider
- Neuralwatt
- Model key
- deepseek-v4.1-flash-flex
- Release date
- Sep 10, 2026
- Last updated
- Sep 10, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.0975
- Output token cost
- $0.39
Limits
- Output tokens
- 393,216 tokens
- Context window
- 1,048,560 tokens
Latest news about DeepSeek V4.1 Flash Flex
No articles yet. Fetch the latest news to show it here.