DeepSeek V3.1 is an open-weights large language model positioned for general text tasks, with reasoning and tool-calling capabilities. On Baseten it was served as an FP8 endpoint, and independent benchmarking on Artificial Analysis placed that Baseten deployment at the top of the field for time-to-first-token latency at 0.72 seconds and second for output speed at 187.8 tokens per second across seven providers running a 10k input-token workload, with a blended-price ranking of third. These figures suggest the Baseten offering competed well on responsiveness and throughput, qualities that typically suit chat assistants, code generation, and API-driven agents that benefit from snappy first responses. The same benchmark noted pricing variance up to 14.9x across providers for V3.1, reinforcing how endpoint choice materially shapes total cost.
DeepSeek V3.1 reached a broad cloud footprint in late 2025, with AWS announcing its availability in Amazon Bedrock on September 18, 2025, alongside other managed deployments that signaled enterprise interest in the open-weights release. Community commentary framed V3.1 as a measured step rather than a breakthrough, but it consolidated support for reasoning and tool use in an accessible open model family. Practical fit therefore centered on teams wanting self-hostable or hosted access to a capable text model without proprietary lock-in. On Baseten specifically, the Model API for DeepSeek V3.1 was deprecated at 5pm PT on June 24, 2026 per Baseten's changelog, so developers should plan migrations to alternate providers rather than relying on this endpoint for new workloads.