Ember-1 is a research preview reasoning model from Fireworks, built on Kimi K3 for coding and agentic workflows. Fireworks says it delivers comparable quality with about 40% fewer tokens than Kimi K3 across its evaluations. That is Fireworks’ result, not a measurement of our repository jobs. Whether your cost per finished task drops depends on what your agent does over multiple turns: how much it reasons, how much context it carries forward, and whether it completes the same work without extra retries.
A model switch can reduce token use while making a task slower, less reliable, or more expensive in another way. The useful comparison is therefore not tokens on one prompt. It is the cost and outcome of a fixed set of real tasks, run through the same harness with the same tools and stopping rules. Ember-1’s public results make it a reasonable candidate for that test; they do not settle the answer for your codebase.
Why fewer tokens could matter more on an agent job
A coding agent calls a model repeatedly. It reads the repository, proposes a change, responds to tool output, revises its plan, and may run tests or inspect failures before it finishes. The bill accumulates across those calls. Earlier reasoning and tool context can also be included in later requests, so a long trace may cost more than its first response suggests.
Fireworks describes this as a particular cost in multi-turn agent workloads: prior reasoning is replayed on later turns, and context grows with each interaction. Its Ember-1 approach was to train the model to shorten unnecessary reasoning while retaining what Fireworks considers useful reflection. Vercel’s announcement likewise says shorter reasoning traces can reduce output costs and the context carried into later steps for coding agents that make repeated calls.
That mechanism is plausible, but it is not a guarantee of lower finished-task cost. A shorter answer that causes a missed edit can lead to another tool round, a failed test, or a human repair. Those extra calls may erase token savings. Measure the whole task and its result.
What Fireworks has reported
Fireworks says Ember-1 matches Kimi K3’s quality with 40% fewer tokens. In its post, the company describes training across tasks including coding, software engineering, tool use, and extended interactions. It says its evaluations included public benchmarks and live production traffic, and it reports results from two customers’ production coding workloads: approximately 35% fewer tokens per task at comparable quality.
Those customer results are useful evidence, but they are not a controlled result for every repository, harness, or agent configuration. The same post gives benchmark-specific outcomes that vary by task. It also reports token reductions alongside quality scores, rather than proving that every user’s cost per accepted change will fall by the same proportion. Keep the attribution attached when using the headline figure: Fireworks claims the reduction; we have not independently measured it here.
Vercel says Ember-1 is available through AI Gateway as a research preview. Its announcement lists a 1M-token context window, text and image input, tool calling, and implicit prompt caching. Vercel also describes gateway usage and cost tracking, key budgets, and routing rules. These are integration details, not evidence that your agent will behave identically after changing models.
Measure a finished task, not a single response
Start with a small, representative set of repository jobs that already have a clear definition of done. Include work that exercises the agent’s normal loop: locating code, editing multiple files when needed, using tools, responding to test output, and producing a final change. Use the same tasks for both model configurations. Keep the repository state, agent version, system instructions, tool permissions, context limits, and test commands fixed. If any of these differ, record the difference because it can affect the result.
For each run, capture input tokens, output tokens, number of model turns, tool calls, elapsed time, and the billed cost shown by the provider or gateway. Record whether the task met its acceptance criteria, whether the relevant checks passed, and whether a person had to repair or finish the change. Then compare cost per finished task, counting failed or abandoned attempts separately instead of silently dropping them. Averages can hide a task that repeatedly fails, so keep the task-level results alongside any aggregate.
A practical worksheet can be as simple as one row per run: task identifier; model; input and output tokens; turns; billed cost; completion status; test result; and human intervention. Run each task more than once if outcomes vary, and avoid treating a lucky single run as a general result. Do not assume that token counts map directly to money: pricing, cached input, and the mix of input versus output can change the bill. Use the actual billed cost where possible, and keep the token totals as diagnostic detail.
We covered the broader measurement approach in how to remeasure an agent bill. It is useful when a model or harness changes because the bill can move for reasons beyond the model’s quoted token efficiency. The same discipline applies here: hold the job constant, log the calls, and judge the completed result.
What can break when switching agents
The available announcements describe Ember-1’s intended use and supported capabilities; they do not document compatibility with every coding agent or guarantee unchanged behavior in every setup. Vercel says it supports tool calling, but that alone does not establish that a particular agent’s tool schemas, prompt templates, or model-selection settings will work without adjustment. Verify the configuration your agent actually sends and confirm that tool calls still execute as expected.
Also check that your agent can reach the selected endpoint, that authentication and budgets are configured, and that your logging captures the calls you need. Vercel describes a setup command that detects installed agents and configures their connection to AI Gateway, but your local agent and policies may require additional setup. Treat research-preview access as preview access: the announcements do not establish long-term availability or a stable service commitment for your particular workflow.
Finally, repeat the task checks that matter to your team. Do not infer repository-level reliability from a general benchmark or the model’s context-window size. If a run misses acceptance criteria, inspect whether the issue came from model behavior, prompt or tool configuration, the gateway path, or the harness. The sources do not enumerate your likely failure modes; a controlled trial is how you find them.
A decision you can support with your own data
Ember-1 is worth evaluating when model calls are a meaningful part of the cost of your coding workflow and your jobs involve repeated turns. Fireworks’ reported token reductions give a reason to run the comparison. They do not tell you whether your repository jobs will finish more cheaply, or whether any savings will survive changes in completion rate, retries, latency, and human review.
Keep the decision narrow. Compare the same tasks under the same conditions, use provider-reported cost, and preserve the failed runs in the record. If Ember-1 lowers cost per accepted change while maintaining the outcomes your team requires, you have evidence for a switch in that workflow. If token totals fall but more jobs need intervention, the token result alone is not enough.
If you want a concrete starting point for the rest of the product around those agent workflows, explore the SaaS dashboard kit.
Sources
- Fireworks, “Introducing Ember-1”: https://fireworks.ai/blog/ember-1
- Vercel, “Ember-1 from Fireworks now available on AI Gateway”: https://vercel.com/changelog/ember-1-from-fireworks-now-available-on-ai-gateway
- OTF, “Remeasure your agent bill”: https://otf-kit.dev/blog/remeasure-agent-bill-cursor-harness
- OTF SaaS dashboard kit: https://otf-kit.dev/templates/saas-dashboard


Top comments (0)