Fireworks Ember-1: Kimi K3 With ~40% Fewer Tokens — What It Is, Benchmarks, and Who Should Care
Fireworks AI just released Ember-1 (September 28, 2026), a post-trained version of Moonshot AI's open-weight Kimi K3 that, according to Fireworks, delivers Kimi K3's quality with about 40% fewer tokens. If you run AI coding agents or any multi-step agentic workload, this one is worth understanding — because the savings come from the model learning to think more efficiently, not from a cheaper settings toggle.
What is Fireworks Ember-1?
Ember-1 is a specialized model from Fireworks Research, built by post-training Moonshot AI's open-weight Kimi K3. The core idea is simple to state and hard to execute: teach the model to produce shorter reasoning traces while keeping task accuracy the same.
This is explicitly not the same as lowering the reasoning-effort setting at inference time. Fireworks says its customers tried that with K3 and it didn't work — lower effort settings gave up too much quality. So instead of asking the model to think less at runtime, the team trained it to reason more efficiently. Useful self-reflection (revisiting an assumption, reacting to feedback) is kept; redundant reasoning and unproductive loops are cut.
Why do reasoning models "think too much" — and why does it cost you?
Fireworks reports that reasoning models like Kimi K3 sometimes spend more than 90% of generated tokens on internal reasoning. That's the invisible "thinking" text you never see in the final answer — but you still pay for every token of it.
The cost compounds in multi-turn agentic workloads. Each turn replays prior reasoning back to the model, so context grows roughly quadratically with the number of turns. Long traces from early turns get re-read — and re-billed — on every later call. If you're running a coding agent that takes 20+ steps to finish a task, you're paying for the model's thinking many times over.
How did Fireworks build Ember-1?
According to the release details reported by MarkTechPost, the training collection spans mathematics, coding, instruction following, conversation, search, tool use, and software engineering — covering both standalone problems and extended multi-step interactions. Task and environment feedback guided on-policy planning and learning.
The scale is notable: Fireworks says it ran more than 50 training experiments and over 200 evaluations, all on its own Fireworks Serverless Training infrastructure, using its own data and no customer data. The team also developed new training algorithms for this — which it has not published. So the technique is proprietary even though the base model (Kimi K3) is open-weight.
What do the benchmarks actually show?
Fireworks compared Ember-1 against Kimi K3 at three reasoning-effort levels. Important honesty note: these are Fireworks' own published evaluations, not independent reproductions. Cost was computed with public Kimi K3 API pricing.
| Benchmark | Kimi K3 Max | Ember-1 |
|---|---|---|
| Terminal Bench 2.1 | 80.9% | 82.0% |
| SWE-bench Verified | 93.2% | 92.2% |
| SWE-Interact | 21.3% | 20.0% |
| DeepSWE 1.1 | 66.4% | 75.2% |
| τ-2 Bench Airline | 64% | 66% |
Ember-1 beats K3 Max on Terminal Bench 2.1 and DeepSWE 1.1, and trails slightly on SWE-bench Verified and SWE-Interact. Across seven benchmarks plus two customers' production traffic, Fireworks says K3's reasoning was shortened by 35–50% without sacrificing accuracy.
The production evidence is arguably more interesting than the benchmarks: Fireworks ran live A/B tests with two customers on real coding workloads. Output tokens fell from 49.3K to 29.9K per task, reasoning tokens dropped 71.3%, total tokens dropped 39% — while the task score was essentially unchanged (0.753 vs 0.751). One of those customers now runs Ember-1 in production.
There's also a medical angle: on Doximity's Bedside Bench — a physician-validated set of 500 clinical cases — Ember-1 set a new cost-per-task Pareto frontier, according to Fireworks' new Specialized Intelligence Index.
How much money does it really save?
Ember-1 costs the same per token as Kimi K3 on Fireworks: $3.00 input, $0.30 cached input, and $15.00 output per 1M tokens. All savings come from generating fewer tokens. At that output rate, the A/B figures work out to roughly $0.45 vs $0.74 in output cost per task — about 40% cheaper per completed task, with the same quality score.
The savings are workload-dependent. Simple single-turn questions won't benefit much — there's little reasoning to trim. The win is concentrated in agentic, multi-step workloads: coding agents, tool-using assistants, anything where the model thinks, acts, observes, and thinks again.
Who should use Ember-1 — and who shouldn't?
Worth trying if: you run Kimi K3 (or a similar reasoning model) in production on Fireworks and your bills are dominated by reasoning tokens — especially agentic coding or automation workloads.
Skip it if: you need open weights or self-hosting. Ember-1 is API-only via Fireworks serverless, as a Research Preview. Fireworks has not released the weights, training code, or exact training algorithms, so you can't run it yourself today.
Wait and verify if: you want independent confirmation. All performance numbers so far come from Fireworks itself. The production A/B results are encouraging, but independent benchmarks haven't landed yet.
If you just want to experiment with Kimi K3 itself for free before committing to any API, you can try it in your browser — Toolxz AI Chat offers Kimi K3 alongside other models with no login required.
For the bigger picture this week, OpenAI's DevDay lands tomorrow (September 29) with GPT-6 Cyber expected — I covered what we actually know before the event yesterday.
FAQ
Is Ember-1 open source?
No. It's built on the open-weight Kimi K3, but Fireworks has not released Ember-1's weights, training code, or training algorithms. It's available only through the Fireworks serverless API as a Research Preview.
Is Ember-1 just Kimi K3 with lower reasoning effort?
No — that's the key distinction Fireworks makes. Lowering the effort setting at inference time hurt quality for their customers. Ember-1 was post-trained to reason more efficiently, keeping accuracy while using fewer tokens.
How much cheaper is it really?
Per-token pricing is identical to Kimi K3. The ~40% saving comes from generating fewer tokens per task. In Fireworks' production A/B test, total tokens per task dropped 39% at an unchanged quality score.
Can I self-host Ember-1?
Not today. API-only via Fireworks. If you need self-hostable, the base Kimi K3 weights remain the open option.
Founder of Toolxz (toolxz.com) — 45+ free browser-based tools. I write about practical AI tooling and developer workflows.
Top comments (0)