LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation

2026年8月19日 13:48
站內 AI 整理稿

Back to Articles LFM2.5 Q40 Checkpoints from Quantization-Aware Distillation Team Article Published August 19, 2026 Upvote - Aditya Tadimeti adityatadimeti Follow LiquidAI Leonie Monigatti iamleonie Follow LiquidAI Today, we release QAD Q40 GGUFs.These are updated 4-bit checkpoints for LFM2.

5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B.They allow developers to run LFM2.

5 models at Q40 memory and speed without the usual quality drop: Trained with Quantization-Aware Distillation (QAD): a high-precision teacher model is distilled into a quantized student model Same memory and speed as native Q40: They keep the low memory footprint and high throughput of Q40 GGUFs Recovery: 97% of their BF16 average accuracy lost to quantization is recovered Benchmark results For all four models, we compare their released GGUFs produced with post-training quantization (PTQ) against the trained QAD Q40 checkpoints on a benchmark suite spanning reasoning, instruction-following, tool use, and agentic capabilities: GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, and BFCLv4.

The BF16 GGUF serves as the in-format ceiling.We also add one scale-appropriate math evaluation: GSM8K for LFM2.5-230M and LFM2.5-350M, and AIME25 for LFM2.5-1.2B-Instruct and LFM2.5-2.6B.We report the mean across five repeats.Across all four models, QAD substantially improves the Q40 checkpoint.

The QAD checkpoints retain 97.1%, 96.5%, 97.4%, and 96.6% of their respective BF16 baseline performance.Speed and size on real edge hardware We measure decode throughput for the four models LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.

6B across four targets: MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5.MacBook Pro and NucBox use GPU inference, while Samsung and Raspberry Pi use Arm CPU inference.BF16 and F16 are shown as full-precision references where profiled.

The 230M and 350M QAD Q40 checkpoints match Q5KM quality within evaluation variance at a 4-33% higher decode throughput.The 1.2B and 2.6B QAD Q40 checkpoints match Q4KM quality at a 3-14% higher throughput.The QAD Q40 checkpoints also match Unsloth's UD-Q4KXL (where applicable, for the 230M and 1.

2B), a strong external post-training quantization checkpoint.How to use QAD GGUFs Use the files with llama.cpp or any runtime that supports GGUF Q40 artifacts.llama-cli -hf LiquidAI/LFM2.5-350M \ --hf-file LFM2.5-350M-QAD-Q40.gguf \ -p "What is C.elegans?

" Get Started with QAD GGUFs The QAD GGUFs are available on Hugging Face today: LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B.We can't wait to see what you build.Citation For citations, please use the following reference or BibTeX: Liquid AI, "LFM2.

5 Q40: Quantization-Aware Distillation for Edge Deployment", Liquid AI Blog, Aug 2026.Or use the BibTeX citation @article{liquidAI2026Q40, author = {Liquid AI}, title = {LFM2.5 Q40: Quantization-Aware Distillation for Edge Deployment}, journal = {Liquid AI Blog}, year = {2026}, note = {www.liquid.

ai/blog/qad}, } Models mentioned in this article 4 More from this author LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge 45 August 12, 2026 Deploy local agents everywhere with LFM2.5-2.

6B 90 August 4, 2026 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.Tap or paste here to upload images Comment · Sign up or log in to comment Upvote - Models mentioned in this article 4

Related

相關文章

WRC 2026|原生全模態世界模型:從模擬世界到交互世界

世界機器人大會期間,智象未來創辦人梅濤於「物理AI引領者論壇」發表演講,提出原生全模態世界模型從「模擬世界」走向「交互世界」的觀點。他強調即使AI模型智商接近140,高IQ不代表全能,需具備在真實物理世界中穩定完成任務的能力,此為Physical AI發展的關鍵。論壇聚焦通用物理智慧的技術演進與產業路徑,匯聚眾多專家參與。

剛剛

阿里巴巴達摩院推出肝癌 AI 模型:可精準識別 1 釐米微小腫瘤

作者:遠洋 責編:遠洋 評論: 感謝網友 HH_KK 的線索投遞!8 月 24 日消息,阿里巴巴達摩院聯合中國醫科大學附屬盛京醫院等機構研發出肝癌診斷 AI 模型 DAMO LiON,可通過 CT 影像識別微小的肝臟癌變病灶。在兩個月的真實世界前瞻臨床試驗中,該 AI 模型發現了 15 例原本被遺漏的惡性腫瘤,絕大部分為 1 釐米左右的病灶,幫助患者得到及時的手術或藥物治療。

剛剛
何夕2077研究與前沿

棋類模型可解釋

在人工智慧研究領域,模型的可解釋性一直是備受關注的課題。近期有觀點指出,棋類模型具備可解釋的特性,這意味著此類模型的決策過程與內部運作機制,能夠被研究者或使用者以相對直觀的方式理解與分析。相較於許多深度學習模型常被視為「黑箱」,棋類模型在處理圍棋、象棋等棋類遊戲時,其每一步的選擇與策略推演,往往能透過棋譜或演算法邏輯加以回溯,從而為AI的透明化提供了一個具體的觀察窗口。

8 小時前

美國專家示警:學生依賴“AI 代寫”會削弱思考能力

作者:清源 責編:清源 評論: 8 月 23 日消息,美國學生使用 AI 完成作業、甚至代寫整篇論文的現象已經十分普遍,也有不少學校允許學生在一定範圍內藉助 AI 工具。據《紐約時報》當地時間 17 日報道,越來越多專家擔心,問題可能不只是學生會不會寫文章,而是長期依賴 AI 可能削弱他們本身的思考能力。

12 小時前