Fireworks AI 發布 Ember-1:後訓練 Kimi K3,Token 用量減少約 40%

2026年9月28日 07:22
站內 AI 整理稿

Fireworks AI has released Ember-1, a specialized model from Fireworks Research built by post-training Moonshot AI’s open-weight Kimi K3.Ember-1 learns to produce shorter reasoning traces while keeping task accuracy.This is different from lowering the reasoning effort setting at inference time.

According to the Fireworks release post, Ember-1 delivers Kimi K3’s quality with about 40% fewer tokens.Is it deployable?Yes, but only through the Fireworks serverless API as a Research Preview.

Fireworks has not released Ember-1’s weights, training code, or exact training algorithms, so self-hosting is not an option today.

The Problem: Reasoning Models Think Too Much Fireworks team reports that reasoning models like Kimi K3 sometimes spend more than 90% of generated tokens on internal reasoning.That cost compounds in multi-turn agentic workloads.Each turn replays prior reasoning back to the model.

Context grows roughly quadratically with the number of turns.Long traces from early turns get re-read, and re-billed, on every later call.Fireworks team explains how customers wanted K3’s coding capability at lower cost.Turning down K3’s reasoning effort did not solve it.

Lower effort settings gave up too much quality.So the team trained the model to reason more efficiently instead.How Fireworks Research Built Ember-1 Not all of K3’s reasoning is waste.Some of it is useful self-reflection, like revisiting an assumption or reacting to feedback.

Ember-1 keep that behavior while cutting redundant reasoning and unproductive loops.The training collection spans mathematics, coding, instruction following, conversation, search, tool use, and software engineering.It covers both standalone problems and extended multi-step interactions.

Task and environment feedback guides on-policy planning and learning.Fireworks team ran more than 50 training experiments and over 200 evaluations.They also developed new training algorithms, which it has not published.All training ran on Fireworks Serverless Training.

Fireworks states it used its own data and no customer data.Benchmark Results Fireworks compared Ember-1 with Kimi K3 at three reasoning effort levels.Cost was computed with public Kimi K3 API pricing.These are Fireworks’ own published evaluations.

BenchmarkNK3 LowK3 HighK3 MaxEmber-1Ember-1 vs K3 Max (cost)Terminal Bench 2.18976.4%77.6%80.9%82.0%-51.9% / -23.1 USDSWE-bench Verified50080.4%86.0%93.2%92.2%-15.5% / -68.1 USDSWE-Interact756.7%13.3%21.3%20.0%-32.5% / -60.8 USDDeepSWE 1.111355.8%62.8%66.4%75.2%-23.7% / -126.

9 USDτ-2 Bench Airline5064%64%64%66%-5.9% / -0.3 USD Ember-1 leads K3 Max on Terminal Bench 2.1 and DeepSWE 1.1.It trails slightly on SWE-bench Verified and SWE-Interact.

Across seven benchmarks and two customers’ production traffic, Fireworks says K3’s reasoning was shortened by 35 to 50% without sacrificing accuracy.On Doximity’s Bedside Bench, a physician-validated set of 500 clinical cases, Ember-1 set a new cost-per-task Pareto frontier.

That result comes from Fireworks’ new Specialized Intelligence Index.Production A/B Test Results Fireworks ran live A/B tests with 2 customers on production coding workloads.Both saw roughly 35% fewer tokens per task at comparable quality.In the published run, output tokens fell from 49.3K to 29.

9K per task.Reasoning tokens dropped 71.3% and total tokens dropped 39%.The task score was essentially unchanged: 0.753 for Ember-1 versus 0.751 for K3.Average steps fell from 23.8 to 21.4.One customer now runs Ember-1 in production.Ember-1 costs the same per token as Kimi K3 on Fireworks: $3.

00 input, $0.30 cached input, and $15.00 output per 1M tokens.The savings come entirely from generating fewer tokens.At that output rate, the A/B figures work out to about $0.74 versus $0.45 in output cost per task (our calculation, output only).Interactive Explainer window.

addEventListener("message",function(e){var d=e.data;if(!d){return;}if(d.type!=="mtp-ember-resize"){return;}var f=document.getElementById("mtp-ember-frame");if(f){f.style.height=d.h+"px";}}); Key Takeaways Ember-1 is Kimi K3 post-trained to reason in fewer tokens, not run at lower effort.

Fireworks reports about 40% fewer tokens with accuracy held across its evaluations.In a production A/B test, output fell from 49.3K to 29.9K tokens per task at a 0.753 vs 0.751 score.It beats K3 Max on Terminal Bench 2.1 (82.0%) and DeepSWE 1.1 (75.2%), trailing on SWE-bench Verified (92.2%).

API-only Research Preview at K3 pricing; weights and training code are not released.Check out the Technical Details.All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!

are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens appeared first on MarkTechPost.

Related

相關文章

雷峰網模型更新

三箭齊發!Sharpa三大新品IROS全球首秀,三位一體再次刷新靈巧操作上限

美國匹茲堡,9月28日——美國匹茲堡智能機器人與系統國際會議(IROS)現場,Sharpa正式發佈三款重磅全自研新品:首個一體式觸覺感知靈巧操作機器人D01、新一代全觸覺超緊湊、輕量化靈巧手W02,以及高保真外骨骼觸感數據手套AE01。三款產品分別面向全身動作執行、手部精細操作與高質量操作數據採集,覆蓋機器人如何運動、如何接觸和操控物體,以及如何藉助人的自然動作完成遙操作與數據獲取三大關鍵環節,為真實環境中的複雜富接觸操作提供了一套完整的硬件閉環。

2 小時前

智譜ZCode落地新輪補償:贈付費用戶8張重置卡,已徹底移除代碼快照上傳鏈路

根據官方公告,補償措施於9月28日10時30分生效,付費用戶及一個月內迴歸的付費用戶將獲贈4張周額度與4張5小時額度重置卡,有效期一個月;同時在9月28日至10月7日期間,平臺面向全體用戶每日發放“1億Token×10萬份”。資本市場上,受前序風波影響,智譜28日港股開盤走低,盤中報620.

2 小時前
量子位模型更新

量子計算走上桌面!“小盒子”跑通端到端,數據全程不出門

酉術量子發布全球首個異構計算硬件底座UnitarySpark,並上線以自然語言驅動的量子科學計算平台UnitaryLab 2.5公測版。該平台旨在降低量子計算的使用門檻,讓開發者能用大白話描述科學問題,由機器自動完成量子與經典算力的協同運算,無需自行撰寫量子程式或處理數據上雲的合規問題。

10 小時前
量子位模型更新

索辰科技加碼世界模型,與戰略投資企業美夢空間聯合發佈具身模型與物理測評標準

索辰科技投資的美夢空間在數貿會上發布具物理感知能力的Physical-WAM世界動作模型及RoboTwin-Phys物理漂移評測基準,旨在解決現行VLA模型缺乏物理理解、實際操作成功率低的問題。Physical-WAM透過物理Token與三大模塊讓機器人預判與修正動作,RoboTwin-Phys則提供可控的物理擾動評測標準,且已開源,希望加速具身智能產業落地。

1 天前