the short version
- Fireworks says Ember-1 delivers Kimi K3's quality with 40% fewer tokens, and that K3's reasoning could be shortened by 35 to 50% across seven benchmarks and two customers' production traffic without losing accuracy.
- The Ember-1 vs K3 max differences in the post's benchmark table range from -51.9% on Terminal Bench 2.1 to -5.9% on τ-2 Bench Airline, so one headline figure will not predict your bill.
- The post announces no weight release, no licence and no Ember-specific per-token price, and its benchmark cost figures are computed at public Kimi K3 API rates.
- Commenters on Hacker News argue that Sol and GLM 5.3 already beat K3 on price, so Ember-1 is most relevant to teams that have already settled on Kimi K3.
Ember-1 is a Fireworks Research model built on Kimi K3. Fireworks says it delivers K3's quality with 40% fewer tokens, and that K3's reasoning could be shortened by 35 to 50% without sacrificing accuracy. In live A/B tests on two customers' production coding workloads, Fireworks reports approximately 35% fewer tokens per task at comparable quality. It is available as a two-week Research Preview on Fireworks Serverless. The announcement mentions no downloadable weights, no licence and no per-token price for the model itself.
The blog post is dated 9/23/2026 and frames Ember-1 as the first in a series of specialized models from Fireworks Research. Fireworks says the model learned to cut unnecessary reasoning while keeping the thinking that matters. It reports more than 50 training experiments and over 200 evaluations, all run on its Serverless Training product. It says it used its own data and no customer data.
Ember-1 is a hosted preview, not open weights
The release is a hosted serving option that sits alongside base Kimi K3 as a Research Preview on Serverless. Fireworks describes research releases as two-week serverless access to new research models, made permanent based on community demand. The blog's page header also styles the model as Ember 1. The post calls Ember-1 Fireworks' own model, but it does not say whether the weights will be published or what licence governs them.
Fireworks is also launching training support for Ember-1. This lets enterprises build customized, token-efficient models tailored to their needs with their own data. That offer and the serverless endpoint are the only access paths the source describes. Teams that need to self-host have nothing to download. Teams that need a model to stay available beyond a two-week window have no commitment to rely on.
Why shorter reasoning matters in agent loops
Fireworks says reasoning models like Kimi K3 spend the majority of their generated tokens, sometimes more than 90%, on internal reasoning. In multi-turn agent work, every turn replays prior reasoning back to the model. Context therefore grows roughly quadratically with the number of turns, and early traces are re-read and re-billed on every later call. Fireworks says that for agentic coding and other workloads where reasoning tokens account for most of the cost, Ember-1 delivers the same quality at roughly half the token cost.
A table in the customer section of the post compares the two models. Kimi K3 scored 0.751 with 23.8 steps and 49.3K output tokens. Ember-1 scored 0.753 with 21.4 steps and 29.9K output tokens, a 71.3% reasoning token reduction and a 39% total token reduction. The post does not say whether this table covers one customer or both. Fireworks says most downstream product metrics held or improved, including task completion, success scores and failure rates. It also says one customer now runs Ember-1 in live production, with plans to scale it up to replace the base model entirely.
Fireworks says it tried the obvious alternative first. Turning down K3's reasoning effort gave up too much quality, so the model had to be trained to reason more efficiently. The training collection spans mathematics, coding, instruction following, conversation, search, tool use and software engineering, and covers both standalone problems and extended interactions. Fireworks also says the model shows restrained token use on unsuccessful attempts, which reduces prolonged, unproductive reasoning.
Ember-1 benchmarks against K3 at every effort level
Fireworks compared Ember-1 against K3 at low, high and max reasoning effort, where max is K3's default. Costs were computed at public Kimi K3 API pricing: $3 per million uncached input tokens, $0.30 cached and $15 output. The list below gives pass rates, with the post's Ember-1 vs K3 max difference in brackets. The post reports that difference as a percentage and a USD amount without labelling what the percentage measures.
- Terminal Bench 2.1 (N=89): K3 low 76.4%, high 77.6%, max 80.9%, Ember-1 82.0% (-51.9% / -23.1 USD)
- SWE-bench Verified (N=500): K3 low 80.4%, high 86.0%, max 93.2%, Ember-1 92.2% (-15.5% / -68.1 USD)
- SWE-Interact (N=75): K3 low 6.7%, high 13.3%, max 21.3%, Ember-1 20.0% (-32.5% / -60.8 USD)
- DeepSWE 1.1 (N=113): K3 low 55.8%, high 62.8%, max 66.4%, Ember-1 75.2% (-23.7% / -126.9 USD)
- τ-2 Bench Airline (N=50): K3 64% at all three efforts, Ember-1 66% (-5.9% / -0.3 USD)
Ember-1 trails K3 max on SWE-bench Verified and SWE-Interact and leads on the other three. On SWE-bench Verified, the largest sample, the difference is -15.5%. The blog says Ember-1 sits on or near the Pareto frontier on every benchmark with more than 50 test samples, strictly dominating K3-low. By that wording, τ-2 Bench Airline, at exactly 50 samples, is excluded. The text cites seven benchmarks, but only these five appear in the table.
Fireworks also says it analyzed GPT-6 Astra, Claude Opus-5 and GLM 5.3 and found Ember-1 was a leader on the Pareto frontier, averaged across five industry benchmarks. On its own Specialized Intelligence Index, which Fireworks introduced earlier the same week, it says Ember-1 set a cost-per-task Pareto frontier on Doximity's Bedside Bench. That benchmark is physician-validated and spans 500 clinical cases across 10 specialized categories. The comparison includes GPT-5.6 Sol, GPT-6 Astra and Claude Opus 5. The post shows charts but gives no scores in its text.
Hacker News commenters question weights, method and base model
So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?
The source does not contradict that reading. One reply argued that Kimi K3's own licence carries enough restrictions that it is 'weight open' rather than open weight. Another asked whether Cursor's Composer models are like this too. Other commenters asked what capability the token reduction costs that benchmarks would miss, for example a model that was strong at Golang and no longer is. Fireworks reports only benchmark and aggregate results, so the post does not answer this.
One commenter called the method description uninformative, citing phrases such as 'on-policy planning and learning'. The blog says Fireworks developed new training algorithms to shorten reasoning, but it does not name or describe them.
Other commenters questioned whether K3 is the right base at all. One wrote that 'Sol is at 2/10 vs kimi's 3/15' and reported better quality at lower cost from Sol in internal tests. They added that they had heard Kimi sets its pricing across all the neoclouds. Another said GLM 5.3 performs roughly like K3 for less than half the cost. These are commenter reports, not measurements from the source.
One commenter who runs a model-routing benchmark every couple of days reported that it did not pick Ember. In that benchmark, Opus 5.5 won under the planning weights. On code, GPT-6 Sol scored 10/10, the same as Ember, but had a higher quality score and a lower estimated cost. The commenter said Ember has no intelligence index, so its starting score was only 0.73, which held its result to 0.954 against Sol's 0.975.
Should you switch from Kimi K3 to Ember-1
If you already run Kimi K3 at max effort on Fireworks for agentic coding, Ember-1 is a direct A/B candidate. It is offered as a serving option alongside the base model. The customer table (0.753 vs 0.751 score, 39% fewer total tokens) is the closest published match to that workload. Measure tokens and cost per completed task on your own traffic, because the benchmark differences range from -5.9% to -51.9% depending on the task.
Several gaps block a firm decision. The post does not state what Fireworks charges for Ember-1, and its benchmark cost figures assume K3's public pricing. It publishes no latency, time-to-first-token or throughput figures, only a score-versus-duration chart for Bedside Bench with no values in the text. Teams that need self-hosting, a stated licence or availability past the two-week preview cannot build on it yet. If K3 is not already your model, the commenter comparisons to Sol and GLM 5.3 suggest benchmarking those first.
What Fireworks has not said about Ember-1
Fireworks has not given a serverless price for Ember-1 or said whether it will become permanent after the preview window. It has not said whether any weights or licence will follow. Commenters also asked whether the method transfers to smaller models such as Qwen 3.8. One commenter suspected, by their own account on vibes, that K3's size, which they put at 2.8 trillion parameters, is what makes shorter reasoning possible without quality loss. The blog promises more Ember models but names no dates or base models.
Questions this raises
Are Ember-1 weights open or downloadable?
No. The announcement mentions no downloadable weights and no licence. Ember-1 is available only as a two-week Research Preview on Fireworks Serverless, plus a training offer for enterprises.
How does Ember-1 compare to Kimi K3 on benchmarks?
Ember-1 leads K3 at max effort on Terminal Bench 2.1, DeepSWE 1.1 and τ-2 Bench Airline, and trails it on SWE-bench Verified (92.2% vs 93.2%) and SWE-Interact (20.0% vs 21.3%). Fireworks reports lower cost than K3 max on all five.
How much does Ember-1 cost per token?
Fireworks has not published an Ember-specific price. Its cost comparisons use public Kimi K3 API pricing of $3 per million uncached input tokens, $0.30 cached and $15 output.
These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.
