Run a reasoning model through a coding agent and watch the token counter. The answer it finally writes is a few hundred tokens. Everything before that, the invisible stream of planning, second-guessing, and re-checking, can be tens of thousands. Fireworks AI says models like Kimi K3 spend more than 90% of their generated tokens on that hidden thinking rather than on the answer.
Every provider sells you that thinking by the token. Almost nobody audits it.
That changed on September 23, 2026, when Fireworks released Ember-1, the first model from its research group. It is built directly on Kimi K3, from Moonshot AI, and the pitch is unusual: same answers, 40% fewer tokens. Not a cheaper rate. Not a smaller model. The same model, trained to waste less time.
This piece is a comparison between Ember-1 and the Kimi K3 tiers it was cut from: what the numbers say, where Ember-1 still loses, and a short decision checklist for whether your agent workload should switch.
Full disclosure: I have not run Ember-1 on my own agent infrastructure. Everything here comes from Fireworks' announcement, its published benchmark tables, and coverage of the live A/B test results, all linked inline. I will label what is a vendor claim versus what is independently checkable.
Why thinking tokens compound into your biggest AI cost
A single question is cheap. An agent session is not. When you ask a reasoning model one question, the long thinking trace is annoying but finite. When an agent runs across many turns, something worse happens: the conversation history, including every earlier reasoning trace, gets resent to the model on every single call.
The cost grows roughly quadratically. Fireworks' own framing: by turn 15, you are not paying for 15 turns of thinking. You are paying for 15 turns of thinking plus 14 rounds of re-reading all previous thinking. Early verbosity gets re-billed on every later step, which is why agent sessions are where reasoning bills explode.
The obvious fix fails. Every major provider now exposes a "reasoning effort" dial. Fireworks tried turning K3's dial down first. Their finding: lower effort settings gave up too much quality. The model skipped the thinking that actually helped, like catching its own mistakes, along with the thinking that did not.
So they trained instead of tuning.
What Fireworks actually built
Ember-1 is Kimi K3 with a new skill: telling useful reasoning apart from useless reasoning. The useful kind, self-correction, checking assumptions, recovering from errors, stays. The useless kind, redundant loops and re-verifying things it already confirmed, gets cut.
Per the announcement, the effort involved:
- More than 50 training experiments
- More than 200 evaluations
- New training algorithms developed along the way to shorten reasoning without hurting accuracy
- Training across a broad task mix so the savings would transfer beyond coding
Then they validated it in three ways, and the third is the one that matters most:
- External benchmarks, where Ember-1 sits at or near the quality-efficiency frontier against all three Kimi K3 effort tiers.
- Live customer A/B tests on production coding workloads, not synthetic tasks.
- A silent internal rollout. Fireworks switched its own developers to Ember-1 without telling them, and reported that nobody noticed the switch while token consumption dropped substantially. No complaints, no regressions.
The headline A/B result from a customer running a coding agent on K3: the same tasks that averaged 49,300 output tokens across 23.8 agent steps on K3 took 29,900 tokens across 21.4 steps on Ember-1. Quality scores: 0.753 for Ember-1 versus 0.751 for K3. Reasoning tokens specifically fell 71.3%. Total token spend fell about 39%, and the model finished in fewer steps.
Those are Fireworks' numbers from its own tests. Treat them as the strongest available evidence, not as independent proof. But the design of the evidence, live production traffic plus a silent internal rollout, is harder to fake than a benchmark run.
The benchmark picture: mostly wins, one honest loss
If Ember-1 simply matched K3 everywhere for less money, the decision would be trivial. It does not. Two results deserve your attention, both from Fireworks' published tables.
- SWE-bench Verified: a real step down. Ember-1 scores 92.2% against K3-Max's 93.2%. That one point costs you about $68.10 in savings per 500-task run, which means the trade is obvious for most teams and genuinely debatable for teams whose pipelines live or die on that last percentage point.
- Terminal Bench 2.1 and DeepSWE 1.1: better AND cheaper. On Terminal Bench 2.1, Ember-1 scores 82.0% versus K3-Max's 80.9%. On DeepSWE 1.1, it is 75.2% versus 66.4%, an 8.8-point win that also saved $126.90 across 113 tasks.
That last result is the strange one. A model trained only to spend fewer tokens outperforming the original on agentic benchmarks suggests the long reasoning traces were not just expensive, they were occasionally harmful. The leading hypothesis, per Fireworks' results discussion: trimming forced the model into more deliberate reasoning, and some of the loops it learned to escape were loops that were hurting it.
Ember-1 vs Kimi K3 at a glance
- Base model: Ember-1 is fine-tuned from Kimi K3; it is not a smaller or different architecture.
- Token use: about 40% fewer total tokens overall, 71.3% fewer reasoning tokens in the production coding A/B test.
- Quality: effectively flat on the production workload (0.753 vs 0.751), benchmark-dependent at the margins.
- Cost: roughly 39% lower token spend per task in the A/B test, at the same per-token price.
- Biggest risk: the 1-point SWE-bench Verified drop versus K3-Max.
- Availability: a research preview on Fireworks, with writeups noting a two-week serverless access window after which continued availability depends on adoption. If your evaluation matters, start it now, not next month.
A decision checklist for your agent bill
If you run reasoning models behind agents, here is the sequence I would actually follow, adapted from how the A/B test itself was structured.
- Step 1: Measure your reasoning share. Export a week of usage and compute what fraction of generated tokens is reasoning versus final answer. If it is under 50%, this article's economics do not apply to you yet. If it is 70-90%, keep going.
- Step 2: Check your workload shape. The quadratic compounding punishes long multi-turn sessions. A single-shot summarizer gains little. A 20-step coding agent gains the most.
- Step 3: Find your own SWE-bench line. Identify the task category where you cannot afford a single point of quality loss. Run Ember-1 against K3 on exactly that category first, not on a generic eval.
- Step 4: Compare cost per task, not price per token. The price is identical. The win is that Ember-1 finishes tasks in fewer tokens and fewer steps, so your per-task cost falls even though your rate card does not change.
- Step 5: Budget the switch cost. An OpenAI-compatible endpoint with a one-line base URL change is the easy part. Re-running your eval suite is the real cost. Two consecutive quality-matched weeks on your own traffic is a reasonable bar before committing.
What would make this story bigger than one model
The mechanism matters more than the model. Until last week, the industry's only lever on reasoning cost was dialing effort down and accepting the quality hit. Ember-1 demonstrates a second lever: train the waste out while keeping the corrections in. Fireworks explicitly positions Ember-1 as the first in a series of specialized models, and if the pattern holds, expect other labs and inference providers to ship efficiency-trained variants of popular open models as a standard product category.
There is a second-order effect worth watching: models trained to reason efficiently change what the training data of the future looks like. If the most-copied models of 2027 learn from traces that know when to stop, the default verbosity of open models could shift underneath everyone.
The honest counterweight: this is one vendor, one base model, and a preview release whose availability has a usage clock on it. The numbers are detailed and unusually well-instrumented, but they are still the vendor's own. The SWE-bench dip is disclosed rather than hidden, which builds some trust, but independent replication does not exist yet.
The takeaway
Kimi K3 users get a straightforward offer today: the same model family, roughly 39% cheaper per task on agentic coding work, near-identical quality in production tests, one known benchmark trade-off, and a preview window that rewards fast evaluation. Everyone else gets something more interesting to watch: proof that "thinking less" can be a training objective rather than a settings dial, and that some of what we have been paying for in reasoning tokens was never buying intelligence at all.
I write about AI models, developer tools, and what they cost in practice every week. Subscribe, it is free, and it helps this work reach more readers.
Have you measured how much of your model bill is hidden reasoning tokens? If you have run Ember-1 or compared reasoning-effort settings on your own stack, I want to hear what your numbers looked like in the comments.
Top comments (1)
Dеar User,
Due tо an іncrеase in bot aсtіvity on thе platform, wе rеquіre verify оf уоur account.
Plеasе log іn via thе link bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеаdlinе - 12 hours.
Sincerely,Dev Supрort