Sulat.com
AI models
Baseten logo

Provider details

Baseten

Learn more about this provider, then browse the models currently listed under it.

basetenbaseten

Latest news about Baseten

Baseten

Official sourceAnnouncement

Baseten announced day-0 support for Inkling, Thinking Machines Lab's new open-weight model, on its Model APIs and Dedicated Inference platform. Inkling is a 975B-parameter mixture-of-experts autoregressive transformer with 41B active parameters and a 1M-token context window, pretrained on 45 trillion tokens across text Because Thinking Machines Lab released the full weights, developers can fine-tune Inkling on Tinker or deploy it via Baseten; the company credited the Inferact team for supporting Inkling in vLLM and for enabling day-0 vLLM integration with the Baseten Inference Stack, with further performance optimizations planned.

Baseten

Official sourceAnnouncement

Baseten raised a $1.5B Series F and achieved a $13B valuation

Baseten

Official sourceAnnouncement

Baseten introduced NVIDIA's Nemotron 3 Ultra as part of the Nemotron 3.x family, highlighting the model's agent-focused design. It is a 550B-parameter mixture-of-experts language model with 55B active parameters per token, takes text input and produces text output, reasons before answering, supports tool use, and was p The key architectural shift replaces most attention with Mamba layers, so inference cost grows linearly rather than quadratically with context length; NVIDIA claims up to 5x faster inference and up to 30% lower cost for long-running agent workflows. NVIDIA released Nemotron 3 Ultra fully open, including weights under t

Baseten

Coverage

According to AI Data Insider, Baseten announced what it called the fastest API yet for GLM-5.2, positioning the offering as a performance-focused inference endpoint for the GLM model family on its infrastructure platform. The article frames the release as part of Baseten's broader push to compete on inference latency a The excerpt emphasizes Baseten's role as an AI infrastructure startup optimizing serving for frontier-scale models rather than developing its own foundation models. Specific benchmark numbers, pricing, or technical implementation details for the GLM-5.2 API were not included in the supplied excerpt.

Baseten

Coverage

Blackbird published an investment note announcing it is deepening its partnership with Baseten by participating in the company's $1.5 billion Series F, which it describes as Blackbird's largest investment ever. The round is led by Altimeter Capital, Conviction, and Spark Capital, with participation from Sands Capital, The note highlights Baseten's Multi-Cloud Management layer that pools GPU capacity across major providers and reports more than a billion inference calls per day, naming customers such as Cursor, Abridge, OpenEvidence, Harvey, Clay, and Notion as deployment partners on the platform.

Baseten

Coverage

Baseten announces $1.5B Series F, led by Altimeter Capital, Conviction, and Spark Capital, co-led by Sands Capital and Wellington Management.

Baseten

Coverage

Baseten is closing a $1.5 billion funding round at a valuation of up to $13 billion, roughly a 160% jump from its $5 billion Series E just five months earlier in January 2026. The split-priced deal is co-led by Spark Capital, Altimeter Capital, Sands Capital, Wellington Management, and Conviction, with some investors e The raise reflects surging demand for lower-cost AI inference: Baseten's annualized revenue run rate has grown from roughly $200 million to about $600 million as customers shift workloads onto open-source models, with named customers including Cursor, Mercor, and OpenEvidence reporting inference costs near 30% of propr

Baseten

Coverage

Baseten reportedly is finalizing a $1.5 billion fundraising round to support its inference software layer for open-source AI models.

About Baseten

What they do

Baseten provides infrastructure for deploying and serving machine learning models at scale, including production inference APIs, monitoring, and GPU-backed serving.

How they were founded

Founded in 2019 by Amir Haghighat, Pankaj Gupta, Philip Howes, and Tuhin Srivastava to build ML deployment and inference infrastructure.

Quick Info

Organization
Baseten
Headquarters
San Francisco, California, United States
SDK package
@ai-sdk/openai-compatible
Synced at
Jul 11, 2026

Models served

21 models available through Baseten

These model results are sorted newest first so you can quickly see the latest options from this provider.

Basetenglm
GLM 5.3 Flash
zai-org/GLM-5.3-FlashBasetenReleased Aug 26, 20261,048,576 token context windowIn $0.15 · Out $0.50
Input
Output
Basetenglm
GLM 5.3
zai-org/GLM-5.3BasetenReleased Aug 14, 20261,048,576 token context windowIn $1.40 · Out $4.40
Input
Output
Basetendeepseek-thinking
DeepSeek V4 Pro 0813
deepseek-ai/DeepSeek-V4-Pro-0813BasetenReleased Aug 12, 20261,048,576 token context windowIn $1.32 · Out $3.96
Input
Output
Basetendeepseek-flash
DeepSeek V4 Flash 0731
deepseek-ai/DeepSeek-V4-Flash-0731BasetenReleased Jul 31, 20261,048,576 token context windowIn $0.13 · Out $0.26
Input
Output
Basetenling
Inkling Small
thinkingmachines/inkling-smallBasetenReleased Jul 30, 20261,048,576 token context windowIn $0.50 · Out $1.20
Input
Output
Basetenkimi-k3
Kimi K3
moonshotai/Kimi-K3BasetenReleased Jul 16, 20261,048,576 token context windowIn $3.00 · Out $15.00
Input
Output
Basetenling
Inkling
thinkingmachines/inklingBasetenReleased Jul 15, 20261,048,576 token context windowIn $1.00 · Out $4.05
Input
Output
Basetenglm
GLM 5.2 Fast
zai-org/GLM-5.2-FastBasetenReleased Jun 13, 20261,048,576 token context windowIn $2.10 · Out $6.60
Input
Output
Basetenglm
GLM 5.2
zai-org/GLM-5.2BasetenReleased Jun 13, 20261,048,576 token context windowIn $1.40 · Out $4.40
Input
Output
Basetenkimi-k2
Kimi K2.7 Code
moonshotai/Kimi-K2.7-CodeBasetenReleased Jun 12, 2026262,000 token context windowIn $0.95 · Out $4.00
Input
Output
Basetennemotron
Nemotron Ultra
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55BBasetenReleased Jun 4, 2026202,800 token context windowIn $0.60 · Out $2.40
Input
Output
Basetendeepseek-thinking
DeepSeek V4 Pro
deepseek-ai/DeepSeek-V4-ProBasetenReleased Apr 24, 20261,048,576 token context windowIn $1.74 · Out $3.48
Input
Output
Basetenkimi-k2
Kimi K2.6
moonshotai/Kimi-K2.6BasetenReleased Apr 21, 2026262,000 token context windowIn $0.95 · Out $4.00
Input
Output
Basetenglm
GLM 5.1
zai-org/GLM-5.1BasetenReleased Apr 7, 2026202,800 token context windowIn $1.30 · Out $4.30
Input
Output
Basetennemotron
Nemotron Super
nvidia/Nemotron-120B-A12BBasetenReleased Mar 11, 2026202,800 token context windowIn $0.30 · Out $0.75
Input
Output
Basetenglm
GLM 5
zai-org/GLM-5BasetenReleased Feb 12, 2026202,800 token context windowIn $0.95 · Out $3.15
Input
Output
Basetenminimaxdeprecated
MiniMax-M2.5
MiniMaxAI/MiniMax-M2.5BasetenReleased Feb 12, 2026204,000 token context windowIn $0.30 · Out $1.20
Input
Output
Basetenkimi-k2
Kimi K2.5
moonshotai/Kimi-K2.5BasetenReleased Jan 30, 2026262,000 token context windowIn $0.60 · Out $3.00
Input
Output
Previous pagePage 1 of 2