Fireworks Research, the model team inside inference startup Fireworks AI, has introduced Ember-1, a specialized model that the company says produces the same answers as Moonshot AI's Kimi K3 while emitting roughly 40% fewer tokens. The model went live on Fireworks' platform on September 23, and over the past few days it has become one of the most discussed releases in the developer community, drawing hundreds of upvotes on Hacker News as coders debated whether reasoning tokens are the next cost frontier.

The pitch comes at a moment when inference bills, not model capability, increasingly decide what teams can afford to ship. For teams watching agentic workloads consume budgets, the latest AI developments have made one thing clear: the industry's most expensive habit right now is not training frontier models, it is running them. Ember-1 is Fireworks' attempt to attack that problem directly, and the company backed it with an unusually detailed set of benchmark and production numbers.

Thinking Models Think Too Much

The core observation behind Ember-1 is that reasoning models like Kimi K3 spend the majority of their generated tokens — sometimes more than 90%, according to Fireworks — on internal reasoning rather than on the answer itself. On a single request, that is expensive. In agentic workloads, it compounds: every turn of an agent conversation replays all prior reasoning back to the model, so context grows roughly quadratically with the number of turns, and long reasoning traces from early turns get re-read and re-billed on every subsequent call.

Fireworks says customers told them they wanted Kimi K3's coding capability at a lower cost, but that simply turning down the model's reasoning effort did not work — the low-effort settings gave up too much quality. The alternative was to train a model that reasons more efficiently while keeping the reasoning that matters.

Training the Waste Out

Building Ember-1 took more than 50 training experiments and over 200 evaluations, run entirely on Fireworks' own Serverless Training platform, according to the company's blog post. The team says it developed new training algorithms along the way to shorten reasoning without losing accuracy.

Notably, the company did not try to eliminate reasoning wholesale. Some of Kimi K3's apparent verbosity is productive self-reflection — revisiting an assumption, responding to feedback, or tracing an outcome back to an earlier decision helps the model recover from mistakes. The training regime was designed to preserve that ability while trimming unproductive loops.

The training data spanned mathematics, coding, instruction following, conversation, search, tool use, and software engineering, covering both standalone problems and extended multi-turn interactions. Fireworks says it used only its own data and no customer data.

Benchmarks: On or Near the Pareto Frontier

To make the cost case, Fireworks computed per-benchmark cost using Kimi K3's public API pricing — $3 per million uncached input tokens, $0.30 per million cached input tokens, and $15 per million output tokens — and compared Ember-1 against K3 at low, high, and maximum reasoning effort. Across every benchmark with more than 50 test samples, the company reports that Ember-1 sits on or near the Pareto frontier, matching K3 at maximum effort for a fraction of the cost and strictly dominating K3 at low effort.

The headline comparisons against K3 at maximum effort, as reported by Fireworks:

| Benchmark | Samples | K3 (max effort) | Ember-1 |
|---|---|---|---|
| Terminal Bench 2.1 | 89 | 80.9% | 82.0% |
| SWE-bench Verified | 500 | 93.2% | 92.2% |
| SWE-Interact | 75 | 21.3% | 20.0% |
| DeepSWE 1.1 | 113 | 66.4% | 75.2% |
| τ²-Bench Airline | 50 | 64% | 66% |

On cost per task, Fireworks reports Ember-1 cut spending by 51.9% on Terminal Bench 2.1, 32.5% on SWE-Interact, 23.7% on DeepSWE 1.1, and 15.5% on SWE-bench Verified relative to K3 at maximum effort — in the case of DeepSWE, a difference of more than $126 per unit of work at the company's calculated rates.

The company also evaluated Ember-1 on Doximity's Bedside Bench, a physician-validated benchmark of 500 clinical cases across 10 specialties that Fireworks runs through its newly launched Specialized Intelligence Index. There, Fireworks says Ember-1 set a new Pareto frontier on cost per task across both open and closed models, including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5. That claim comes from Fireworks' own index and has not been independently verified.

Live Traffic: Fewer Tokens, Same Scores

Benchmarks are one thing; production is another. Fireworks ran live A/B tests with two customers on their production coding workloads and reports that Ember-1 used approximately 35% fewer tokens per task at comparable quality. In one measured comparison, Kimi K3 scored 0.751 while generating 49,300 output tokens per task, against 0.753 for Ember-1 on 29,900 tokens — a 71.3% reduction in reasoning tokens and 39% fewer total tokens. Following the tests, one customer moved Ember-1 into live production with plans to scale it up to replace the base model entirely.

Inside Fireworks itself, the model was quietly swapped into the company's own coding workflows before launch. The result the team says it is proudest of: nobody noticed. Developers carried on their work without detecting the switch while consuming substantially fewer tokens.

Specialized Intelligence as a Strategy

Ember-1 is the first model in a planned series from Fireworks Research, and it arrives alongside the company's Specialized Intelligence Index, a benchmarking effort introduced earlier this month to compare open, closed, and specialized models on expert-created real-world tasks.

The strategic play is easy to read. Rather than train a frontier model from scratch, Fireworks took a strong open-weight model, trained token efficiency into it, and is now selling the same capability at a lower cost per task. As open-weight releases from labs like Moonshot keep closing the capability gap, inference providers are discovering that the durable differentiation may lie in cost per completed task rather than leaderboard position — a shift that favors companies with strong training and serving infrastructure.

The caveats are worth stating plainly: the numbers above come from Fireworks' own blog, its own evaluations, and its own paying customers. Independent replication of the token savings on outside workloads will be the real test. But if the results hold, the most cost-effective way to run a frontier-class open model may no longer be to make it think less — it may be to run a model that learned to think efficiently.
---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →