Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Baseten logo

Model details

DeepSeek V3.1

DeepSeek V3.1 is an open-weights large language model positioned for general text tasks, with reasoning and tool-calling capabilities. On Baseten it was served as an FP8 endpoint, and independent benchmarking on Artificial Analysis placed that Baseten deployment at the top of the field for time-to-first-token latency at 0.72 seconds and second for output speed at 187.8 tokens per second across seven providers running a 10k input-token workload, with a blended-price ranking of third. These figures suggest the Baseten offering competed well on responsiveness and throughput, qualities that typically suit chat assistants, code generation, and API-driven agents that benefit from snappy first responses. The same benchmark noted pricing variance up to 14.9x across providers for V3.1, reinforcing how endpoint choice materially shapes total cost.

DeepSeek V3.1 reached a broad cloud footprint in late 2025, with AWS announcing its availability in Amazon Bedrock on September 18, 2025, alongside other managed deployments that signaled enterprise interest in the open-weights release. Community commentary framed V3.1 as a measured step rather than a breakthrough, but it consolidated support for reasoning and tool use in an accessible open model family. Practical fit therefore centered on teams wanting self-hostable or hosted access to a capable text model without proprietary lock-in. On Baseten specifically, the Model API for DeepSeek V3.1 was deprecated at 5pm PT on June 24, 2026 per Baseten's changelog, so developers should plan migrations to alternate providers rather than relying on this endpoint for new workloads.

Basetendeepseek-ai/DeepSeek-V3.1deepseekdeprecated

Quick Info

Powered by
Provider
Baseten
Model key
deepseek-ai/DeepSeek-V3.1
Release date
Aug 25, 2025
Last updated
Aug 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$1.50

Limits

Output tokens
131,000 tokens
Context window
164,000 tokens

Transparent token rates

Compare DeepSeek V3.1 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V3.1

Baseten

Official sourceRelease Notes

The Baseten changelog contains a decisive first-party entry dated Jun 24, 2026 stating: "Model Deprecation (DeepSeek v3.1, MiniMax m2.5) — The DeepSeek v3.1 and MiniMax M2.5 Model API(s) were deprecated at 5pm PT on June 24th." This is the single most decision-relevant fact for the subject: developers currently pointed The same changelog also documents adjacent Baseten platform changes through mid-July 2026, including new model availability (Inkling on Jul 15), admin visibility for personal API keys, events overlay on metrics/logs graphs, a new Baseten CLI, coding-agent MCP integration, configurable scale-down rate, log downloads, an

Baseten

CoverageBenchmark

Artificial Analysis provides third-party benchmarking of DeepSeek V3.1 (non-reasoning) across 7 API providers, including Baseten (FP8). Baseten ranks #2 on output speed at 187.8 tokens/second (behind Google Vertex at 194.7 t/s) and #1 on time-to-first-token latency at 0.72 seconds, while placing #3 on blended price at The benchmark workload uses 10k input tokens, and pricing spans up to 14.9x across providers for DeepSeek V3.1, illustrating significant variance in production cost. While this Artificial Analysis page gives useful comparative performance data for Baseten's V3.1 FP8 endpoint, the data is third-party and the model itsel

Videos about DeepSeek V3.1

More models around DeepSeek V3.1