Fireworks AI 發布 Ember-1:後訓練 Kimi K3,Token 用量減少約 40%

2026年9月28日 07:22
站內 AI 整理稿

Fireworks AI has released Ember-1, a specialized model from Fireworks Research built by post-training Moonshot AI’s open-weight Kimi K3.Ember-1 learns to produce shorter reasoning traces while keeping task accuracy.This is different from lowering the reasoning effort setting at inference time.

According to the Fireworks release post, Ember-1 delivers Kimi K3’s quality with about 40% fewer tokens.Is it deployable?Yes, but only through the Fireworks serverless API as a Research Preview.

Fireworks has not released Ember-1’s weights, training code, or exact training algorithms, so self-hosting is not an option today.

The Problem: Reasoning Models Think Too Much Fireworks team reports that reasoning models like Kimi K3 sometimes spend more than 90% of generated tokens on internal reasoning.That cost compounds in multi-turn agentic workloads.Each turn replays prior reasoning back to the model.

Context grows roughly quadratically with the number of turns.Long traces from early turns get re-read, and re-billed, on every later call.Fireworks team explains how customers wanted K3’s coding capability at lower cost.Turning down K3’s reasoning effort did not solve it.

Lower effort settings gave up too much quality.So the team trained the model to reason more efficiently instead.How Fireworks Research Built Ember-1 Not all of K3’s reasoning is waste.Some of it is useful self-reflection, like revisiting an assumption or reacting to feedback.

Ember-1 keep that behavior while cutting redundant reasoning and unproductive loops.The training collection spans mathematics, coding, instruction following, conversation, search, tool use, and software engineering.It covers both standalone problems and extended multi-step interactions.

Task and environment feedback guides on-policy planning and learning.Fireworks team ran more than 50 training experiments and over 200 evaluations.They also developed new training algorithms, which it has not published.All training ran on Fireworks Serverless Training.

Fireworks states it used its own data and no customer data.Benchmark Results Fireworks compared Ember-1 with Kimi K3 at three reasoning effort levels.Cost was computed with public Kimi K3 API pricing.These are Fireworks’ own published evaluations.

BenchmarkNK3 LowK3 HighK3 MaxEmber-1Ember-1 vs K3 Max (cost)Terminal Bench 2.18976.4%77.6%80.9%82.0%-51.9% / -23.1 USDSWE-bench Verified50080.4%86.0%93.2%92.2%-15.5% / -68.1 USDSWE-Interact756.7%13.3%21.3%20.0%-32.5% / -60.8 USDDeepSWE 1.111355.8%62.8%66.4%75.2%-23.7% / -126.

9 USDτ-2 Bench Airline5064%64%64%66%-5.9% / -0.3 USD Ember-1 leads K3 Max on Terminal Bench 2.1 and DeepSWE 1.1.It trails slightly on SWE-bench Verified and SWE-Interact.

Across seven benchmarks and two customers’ production traffic, Fireworks says K3’s reasoning was shortened by 35 to 50% without sacrificing accuracy.On Doximity’s Bedside Bench, a physician-validated set of 500 clinical cases, Ember-1 set a new cost-per-task Pareto frontier.

That result comes from Fireworks’ new Specialized Intelligence Index.Production A/B Test Results Fireworks ran live A/B tests with 2 customers on production coding workloads.Both saw roughly 35% fewer tokens per task at comparable quality.In the published run, output tokens fell from 49.3K to 29.

9K per task.Reasoning tokens dropped 71.3% and total tokens dropped 39%.The task score was essentially unchanged: 0.753 for Ember-1 versus 0.751 for K3.Average steps fell from 23.8 to 21.4.One customer now runs Ember-1 in production.Ember-1 costs the same per token as Kimi K3 on Fireworks: $3.

00 input, $0.30 cached input, and $15.00 output per 1M tokens.The savings come entirely from generating fewer tokens.At that output rate, the A/B figures work out to about $0.74 versus $0.45 in output cost per task (our calculation, output only).Interactive Explainer window.

addEventListener("message",function(e){var d=e.data;if(!d){return;}if(d.type!=="mtp-ember-resize"){return;}var f=document.getElementById("mtp-ember-frame");if(f){f.style.height=d.h+"px";}}); Key Takeaways Ember-1 is Kimi K3 post-trained to reason in fewer tokens, not run at lower effort.

Fireworks reports about 40% fewer tokens with accuracy held across its evaluations.In a production A/B test, output fell from 49.3K to 29.9K tokens per task at a 0.753 vs 0.751 score.It beats K3 Max on Terminal Bench 2.1 (82.0%) and DeepSWE 1.1 (75.2%), trailing on SWE-bench Verified (92.2%).

API-only Research Preview at K3 pricing; weights and training code are not released.Check out the Technical Details.All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!

are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens appeared first on MarkTechPost.

Related

相關文章

千問牽手夸克網盤:授權後翻找整理資料,照片還能秒變海報

AI資訊AI新聞資訊正文千問牽手夸克網盤:授權後翻找整理資料,照片還能秒變海報發佈於AI新聞資訊發佈時間 :2026年9月28號 16:46閱讀 :1分鐘阿里旗下的千問 App 宣佈與夸克網盤打通得更深了。用戶在千問裡授權之後,就能直接在對話中翻查、歸置和讀取網盤裡的文件,並據此生成學習工具、工作文檔以及可互動的網頁。從官方說明看,在千問首頁 @ 一下夸克網盤,就能調起夸克 Agent 的技能來"料理"網盤資料。比如忘了某份文件丟在哪個角落,只要用大白話描述要找的東西,千問就會替你去搜;若想幹更復雜的活,也能在工作助理裡啟用"夸克網盤管理"技能,基於盤內資料接著往下做。照片變素材,資料沉澱成共享知識庫網盤裡那些照片,不再只是存著翻看。舉例說,你在杭州旅遊拍下不少風景,想做成一張紀念海報,只要告訴千問想用哪些照片、偏好什麼風格,它便會找出對應圖片並完成創作。另一方面,無論是平時囤的外部資料,還是對話中新生成的內容,都可以交給千問回傳到夸克網盤,順手生成一條可分享鏈接,一點點攢成隨時能查、實時更新的共享知識庫。像備考衝刺時,讓千問照著網盤裡的課程講義排一份複習計劃、拉一張進度看板,做完再存回網盤、生成共享鏈接發給同伴,整套流程都在一個入口裡轉完。對阿里而言,這相當於把"對話入口"和"個人雲存儲"縫到了一處:千問負責理解與執行,夸克網盤充當記憶與分發中樞,讓原本散落、沉睡的資料真正流動起來、被反覆利用,而不是安靜躺在文件夾裡等著落灰。對既用千問又用夸克的用戶,少切幾次應用、資料多一處去處,體驗上確實更順。相關推薦千問App接入夸克網盤,AI直接翻文件成新內容9月28日,阿里旗下AI助手千問App與夸克網盤深度打通。用戶授權後,可在千問對話中@夸克網盤,直接搜索、讀取並處理網盤內照片、視頻和文檔,無需切換應用。核心亮點是AI直接處理文件:長視頻可自動提煉重點並生成看板,照片可用於A

剛剛

千問牽手夸克網盤:授權後翻找整理資料,照片還能秒變海報

AI資訊AI新閒資訊正文千問牽手夸克網盤:授權後翻找整理資料,照片還能秒變海報發布於AI新閒資訊時間 :Sep 28, 2026閱讀 :1分鐘阿里旗下的千問 App 宣佈與夸克網盤打通得更深了。用戶在千問裡授權之後,就能直接在對話中翻查、歸置和讀取網盤裡的文件,並據此生成學習工具、工作文檔以及可互動的網頁。從官方說明看,在千問首頁 @ 一下夸克網盤,就能調起夸克 Agent 的技能來"料理"網盤資料。比如忘了某份文件丟在哪個角落,只要用大白話描述要找的東西,千問就會替你去搜;若想幹更復雜的活,也能在工作助理裡啟用"夸克網盤管理"技能,基於盤內資料接著往下做。照片變素材,資料沉澱成共享知識庫網盤裡那些照片,不再只是存著翻看。舉例說,你在杭州旅遊拍下不少風景,想做成一張紀念海報,只要告訴千問想用哪些照片、偏好什麼風格,它便會找出對應圖片並完成創作。另一方面,無論是平時囤的外部資料,還是對話中新生成的內容,都可以交給千問回傳到夸克網盤,順手生成一條可分享鏈接,一點點攢成隨時能查、實時更新的共享知識庫。像備考衝刺時,讓千問照著網盤裡的課程講義排一份複習計劃、拉一張進度看板,做完再存回網盤、生成共享鏈接發給同伴,整套流程都在一個入口裡轉完。對阿里而言,這相當於把"對話入口"和"個人雲存儲"縫到了一處:千問負責理解與執行,夸克網盤充當記憶與分發中樞,讓原本散落、沉睡的資料真正流動起來、被反覆利用,而不是安靜躺在文件夾裡等著落灰。對既用千問又用夸克的用戶,少切幾次應用、資料多一處去處,體驗上確實更順。相關推薦千問App接入夸克網盤,AI直接翻文件成新內容9月28日,阿里旗下AI助手千問App與夸克網盤深度打通。用戶授權後,可在千問對話中@夸克網盤,直接搜索、讀取並處理網盤內照片、視頻和文檔,無需切換應用。核心亮點是AI直接處理文件:長視頻可自動提煉重點並生成看板,照片可用於AI海報製作,

剛剛5700
雷峰網模型更新

三箭齊發!Sharpa三大新品IROS全球首秀,三位一體再次刷新靈巧操作上限

美國匹茲堡,9月28日——美國匹茲堡智能機器人與系統國際會議(IROS)現場,Sharpa正式發佈三款重磅全自研新品:首個一體式觸覺感知靈巧操作機器人D01、新一代全觸覺超緊湊、輕量化靈巧手W02,以及高保真外骨骼觸感數據手套AE01。三款產品分別面向全身動作執行、手部精細操作與高質量操作數據採集,覆蓋機器人如何運動、如何接觸和操控物體,以及如何藉助人的自然動作完成遙操作與數據獲取三大關鍵環節,為真實環境中的複雜富接觸操作提供了一套完整的硬件閉環。

2 小時前

智譜ZCode落地新輪補償:贈付費用戶8張重置卡,已徹底移除代碼快照上傳鏈路

根據官方公告,補償措施於9月28日10時30分生效,付費用戶及一個月內迴歸的付費用戶將獲贈4張周額度與4張5小時額度重置卡,有效期一個月;同時在9月28日至10月7日期間,平臺面向全體用戶每日發放“1億Token×10萬份”。資本市場上,受前序風波影響,智譜28日港股開盤走低,盤中報620.

3 小時前
量子位模型更新

量子計算走上桌面!“小盒子”跑通端到端,數據全程不出門

酉術量子發布全球首個異構計算硬件底座UnitarySpark,並上線以自然語言驅動的量子科學計算平台UnitaryLab 2.5公測版。該平台旨在降低量子計算的使用門檻,讓開發者能用大白話描述科學問題,由機器自動完成量子與經典算力的協同運算,無需自行撰寫量子程式或處理數據上雲的合規問題。

11 小時前