Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Ember-1

Ember-1 is a specialized reasoning model from Fireworks Research, introduced through a dedicated announcement post that frames the work around reducing the token overhead of "chain of thought" style thinking. The announcement explicitly positions Ember-1 as delivering Kimi K3's quality while using 40% fewer tokens, and the underlying narrative treats over-thinking as the central problem the model is designed to solve, walking readers from the observation to what Fireworks calls a "premium model." This positions Ember-1 for workflows where reasoning quality matters but verbose internal deliberation is a cost or latency concern, making it attractive for production agents and pipelines that need disciplined, efficient reasoning rather than maximal verbosity.

The model's evaluation story is organized around Pareto-frontier thinking rather than single-score leaderboard claims. Fireworks Research highlights a Specialized Intelligence Index on the Bedside Bench suite where Ember-1 is presented as setting a Pareto frontier, then extends the same efficiency-versus-quality framing to additional industry benchmarks, and grounds the claims further in live customer A/B tests plus internal validation in which their own developers did not notice the switch. Together this signals a model intended for applied, latency- and cost-sensitive deployments where maintaining answer quality while shrinking reasoning traces is the primary value proposition, and where real-world A/B evidence is treated as a first-class signal alongside academic benchmarks.

Vercel AI Gatewayfireworks/ember-1

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
fireworks/ember-1
Release date
Sep 23, 2026
Last updated
Sep 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$3.00
Output token cost
$15.00

Limits

Output tokens
1,048,576 tokens
Context window
1,048,576 tokens

Latest news about Ember-1

OpenRouter

Coverage

Fireworks announced Ember-1 on September 23, 2026, as a specialized model post-trained from Moonshot AI's Kimi K3 open weights, training reasoning length as a behavior rather than leaving it as a fixed cost of model quality. The objective was narrower than building a new foundation model: teach K3 to use shorter reasoning traces without abandoning useful analysis. Fireworks reports completing more than 50 training experiments and over 200 evaluations across mathematics, coding, instruction following, conversation, search, tool use, and software engineering, using task and environment feedback for on-policy learning. However, the evidence is primarily vendor-run, and Ember-1 is API-only with no released weights or reproducible training pipeline, leaving outside researchers unable to independently inspect or reproduce the result.

OpenRouter

CoverageRelease Notes

Fireworks AI released Ember-1, a specialized model from Fireworks Research that post-trains Moonshot AI's open-weight Kimi K3 to produce shorter reasoning traces while preserving task accuracy. According to Fireworks' release, Ember-1 delivers K3's quality with roughly 40% fewer tokens, addressing customer complaints that reasoning models like K3 sometimes spend over 90% of tokens on internal reasoning. Ember-1 is deployable only via the Fireworks serverless API as a Research Preview, with no released weights, training code, or exact training algorithm, so self-hosting is not currently possible. Fireworks reports that lowering K3's inference-time reasoning effort lost too much quality, so the team instead trained Ember-1 to cut redundant loops while keeping useful self-reflection, spanning math, coding, instruction following, and conversation workloads.

OpenRouter

Coverage

Fireworks AI released Ember-1 on September 23, 2026, a specialized model built by post-training Moonshot AI's 2.78-trillion-parameter Kimi K3 base weights. The post-training trims verbose reasoning traces, letting Ember-1 match K3 quality while using roughly 40% fewer tokens. List pricing sits at $3.00 input, $0.30 cached input, and $15.00 output per million tokens. Ember-1 targets the "double-billing trap" in multi-turn agent workflows, where reasoning tokens billed as output on turn one are re-sent as input on subsequent turns, compounding costs. The piece frames Ember-1 as Fireworks's playbook for curbing token inflation in coding-agent deployments. Note: the article carries promotional directory CTAs that do not affect the technical claims.

OpenRouter

Coverage

Fireworks announced Ember-1 on September 23, 2026, as a specialized model built by post-training Kimi K3 to produce shorter reasoning traces while preserving task quality. The company reports roughly 40 percent fewer tokens across its evaluations and about 35 percent fewer tokens per task in live customer A/B tests. Hosted on Fireworks Serverless at $3 per million input tokens and $15 per million output, Ember-1 ships as a time-limited research preview with two-week serverless access and no downloadable weights. All published comparisons are Fireworks-run, so developers should measure cost and failure rates against their own agentic workloads before swapping out K3.

OpenRouter

Coverage

Fireworks AI released Ember-1 on September 23, 2026, as the first model from Fireworks Research. Rather than a new base model, Ember-1 is Moonshot AI's open-weight Kimi K3 post-trained on Fireworks' training service to reason in fewer tokens. Fireworks reports it matches K3's quality with 35 to 50 percent shorter reasoning across seven benchmarks. Pricing stays identical to K3 on Fireworks at $3 per million input tokens, $0.30 cached, and $15 per million output, so any bill reduction comes purely from shorter outputs. Ember-1 is a research preview with a two-week serverless access window, no downloadable weights, and Fireworks-run benchmarks without independent replication.

OpenRouter

Coverage

Fireworks AI introduced Ember-1, a model further trained on Moonshot AI's Kimi K3 to address excessive reasoning-token consumption in cutting-edge models. Ember-1 is the first release in Fireworks's initiative to build developer-specialized models, targeting the problem that advanced models devote roughly 90% of output to thinking, inflating agent costs. According to Fireworks AI, Ember-1 retains Kimi K3 performance while cutting output tokens by approximately 35%. On the Bedside Bench medical task benchmark, Ember-1 matched Kimi K3 scores at lower per-task cost. The base Kimi K3, released July 2026, had surpassed GPT-5.6 Sol and Claude Fable 5 on some tests but ran slower and pricier, especially on the AA-Briefcase agent benchmark.

Videos about Ember-1