LFM2.5-2.6B:讓本地代理無所不在
重點摘要
LFM2.5-2.6B 專為在裝置端運行高效能代理而設計,支援工具呼叫與多步驟工作流程,同時保持輕量高速,適用於從筆電到手機的日常硬體。開發者因此能隨處部署代理、確保資料隱私,並無需雲端運算成本即可擴展使用。在工具使用、指令遵循及多步驟代理任務上,其效能可與規模大四倍的模型匹敵,堪稱同級最佳代理。
Back to Articles Deploy local agents everywhere with LFM2.5-2.6B Team Article Published August 4, 2026 Upvote 1 Leonie Monigatti iamleonie Follow LiquidAI Sergei Tilga tilgasergey Follow LiquidAI Sinoué GAD GAD-cell Follow LiquidAI Song Duong sduong Follow LiquidAI Tim Seyde tseyde Follow LiquidAI Maxime Labonne mlabonne Follow LiquidAI LFM2.5-2.6B is built to power capable agents entirely on-device. It supports tool calling and multi-step workflows while staying small and fast enough for everyday hardware, from laptops to phones. This enables developers to deploy agents everywhere, keep data private on the device, and scale usage without a cloud inference bill. Best-in-class agent: Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks. Agentic reinforcement learning: Trained inside the most popular agentic harnesses to improve compatibility. Efficient inference: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, in under 2.5 GB of memory. How we built a reliable agentic model for edge devices LFM2.5-2.6B is pre-trained on ~34T tokens, with a mid-training phase that extends the context window to 128K. Post-training then turns the base model into an agent in four stages: Supervised fine-tuning (SFT): two rounds of SFT, weighted heavily toward agentic data like tool use, web search, and harness trajectories. Teacher specialization: train one specialist teacher per domain (math, code, tool use, and more). Multi-domain on-policy distillation (MOPD): distill the specialist teachers into a single student. Agentic Reinforcement Learning (Agentic RL): run multi-turn RL inside real agent harnesses, where the model learns to work across different tools, system prompts, and multi-turn task environments. The Agentic RL pipeline separates model optimization, inference, and environment execution into distinct components. The Training Engine optimizes the model, while the Rollout Engine generates actions using the latest policy. The RL framework orchestrates the training loop by launching rollouts, collecting trajectories and rewards, and updating the model. Actions are executed within a Sandbox Service, where the Blackbox Harness hosts the agent (e.g., OpenClaw or Hermes Agent) and coordinates interactions with the task environment. The Harness Proxy lets us treat agentic harnesses as black boxes with no modification, while transparently capturing the token-level trajectories needed to reconstruct and validate RL training samples. Benchmark results We evaluated LFM2.5-2.6B against models up to ~4x its size on STEM, instruction following, tool use, and agentic tasks. It is the smallest model in the group, yet it competes with and often beats the rest. Benchmark LFM2.5-2.6B (2.6B) gemma-4-E2B-it (5.1B) gemma-4-E4B-it (8B) Qwen3.5-4B (4.7B) Qwen3.5-9B (9.7B) AA Omniscience -29.50 -74.47 -49.03 -54.30 -50.43 AIME25 51.87 26.33 34.27 49.33 56.07 LiveCodeBenchv6 59.41 54.92 63.77 60.85 69.86 IFBench 59.17 34.08 39.24 48.40 56.47 Multi-IF 80.07 69.44 77.35 55.67 62.55 IFStruct 85.49 64.85 76.65 36.25 78.50 BFCLv4 56.88 36.98 46.39 50.56 60.13 ToolSandbox 77.83 52.40 65.00 75.55 76.44 τ³-Bench Banking 5.67 3.35 4.12 5.45 5.15 Claw-Eval average (EN) 62.85 53.14 58.02 62.28 66.53 PinchBench 68.22 44.24 55.09 71.26 71.45 BrowseComp+ (OpenClaw) 26.89 8.31 15.90 24.46 27.23 For your app, the strengths are instruction following and tool use. LFM2.5-2.6B tops every instruction-following benchmark here, and every tool-use benchmark except BFCLv4, where only the 9.7B Qwen edges ahead. On agentic tasks, it beats both Gemma models and stays even with the Qwens. It also leads on knowledge and stays close on math. Coding is the one place the larger models keep a clear lead, so reach for something bigger there. Inference speed on CPU and GPU LFM2.5-2.6B ships with day-one support across the inference ecosystem, including llama.cpp, MLX, vLLM, SGLang, and ONNX. CPU inference. Due to its efficient LFM2 architecture, LFM2.5-2.6B is the fastest model we tested, with decode speeds of 220 tokens/s on an M5 Max and 113 tokens/s on a Ryzen AI Max+ 395. At 30 tokens/s, it allows you to run capable agents even on a phone. GPU inference. LFM2.5-2.6B is the fastest model in its size class, reaching almost 15K output tokens per second at high concurrency, roughly 1.3B tokens per day on a single H100. How to use LFM2.5-2.6B Reach for LFM2.5-2.6B when you need on-device agents for high-volume workloads. Install the latest version of transformers (compatible with transformers>=5.0.0): pip install -U transformers Then load and run the model: from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "LiquidAI/LFM2.5-2.6B" model = AutoModelForCausalLM.from_pretrained( model_id, device_map="auto", dtype="bfloat16", # attn_implementation="flash_attention_2" # uncomment on a compatible GPU ) tokenizer = AutoTokenizer.from_pretrained(model_id) prompt = "What is C. elegans?" input_ids = tokenizer.apply_chat_template( [{"role": "user", "content": prompt}], add_generation_prompt=True, return_tensors="pt", tokenize=True, ).to(model.device) output = model.generate( input_ids, do_sample=True, temperature=0.2, top_k=80, repetition_penalty=1.05, max_new_tokens=512, ) print(tokenizer.decode(output[0], skip_special_tokens=False)) LFM2.5-2.6B demo Check out this browser demo of LFM2.5-2.6B powering a research agent. The agent helps you research specific questions and generates a summary. Get Started Both LFM2.5-2.6B and LFM2.5-2.6B-Base are available on Hugging Face today. With LFM2.5, we're delivering on our vision of AI that runs anywhere. These models are: Download: LFM2.5-2.6B-Base and LFM2.5-2.6B on Hugging Face. Try: run the WebGPU demo in your browser, no setup needed. Use in your harness: follow our guide on how to run a local agent, like OpenClaw, Hermes Agent, and Pi. We can't wait to see what you build. Citation Please cite this article as: Liquid AI, "LFM2.5-2.6B: Deploy Agents Everywhere", Liquid AI Blog, Aug 2026. Or use the BibTeX citation: @article{liquidAI202626B, author = {Liquid AI}, title = {LFM2.5-2.6B: Deploy Agents Everywhere}, journal = {Liquid AI Blog}, year = {2026}, note = {www.liquid.ai/blog/lfm2-5-2-6b}, } Models mentioned in this article 2 Spaces mentioned in this article 1 More from this author LFM2.5-Encoders for Fast Long-Context Inference on CPU 65 July 28, 2026 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here. Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 1 Models mentioned in this article 2 Spaces mentioned in this article 1
Related
相關文章

最大AI黑洞,Anthropic年入1萬億美元后,算力價格狂飆10倍
賬號設置我的關注我的收藏申請的項目退出登錄登錄搜索36氪Auto數字時氪未來消費智能湧現未來城市啟動Power on36氪出海36氪研究院潮生TIDE36氪企服點評36氪財經職場bonus36碳後浪研究所暗湧Waves硬氪氪睿研究院媒體品牌企業號企服點評36Kr研究院36Kr創新諮詢企業服務核心服務城市之窗政府服務創投發佈LP源計劃VClubVClub投資機。

8位工程師花一年、耗資近千萬的工作,這個開源AI系統7天搞定了
智東西(公眾號:zhidxcom) 作者 | 程茜 編輯 | 雲鵬 智東西8月4日消息,今日,面壁智能聯合開源社區OpenBMB開源全球首個支持Stencil(模板計算)自動研究、自動部署的AI優化系統ForgeStencil。 根據官方披露,ForgeStencil一週內完成了超100個真實工業和科學計算軟件的自動優化,這相當於每年投入約4-8個工程師,等效研發投入為300萬~1000萬元。
All in AI全面衝刺,夯實智能體時代AI基建底座,安謀科技三大業務線集體交卷
智東西(公眾號:zhidxcom) 作者 | 雲鵬 編輯 | 漠影 “龍蝦”和“愛馬仕”的火爆拉開了智能體(Agentic AI)時代的序幕,從年初發展至今,AI智能體幾乎已成為貫穿智能終端、應用、模型、算力、具身等領域的唯一主線。 如何築好智能體時代的“AI算力基建”就成為各路科技巨頭和創企關注的焦點。隨著算力範式告別追求單一芯片極致性能的階段,系統整體效率、能效、成本等因素成為企業更關鍵的考量。

GitHub AI Agent 翻車:攻擊者不用黑客技術,只寫一句話就能竊取數據
賬號設置我的關注我的收藏申請的項目退出登錄登錄搜索36氪Auto數字時氪未來消費智能湧現未來城市啟動Power on36氪出海36氪研究院潮生TIDE36氪企服點評36氪財經職場bonus36碳後浪研究所暗湧Waves硬氪氪睿研究院媒體品牌企業號企服點評36Kr研究院36Kr創新諮詢企業服務核心服務城市之窗政府服務創投發佈LP源計劃VClubVClub投資機。
開發者苦 “造輪子” 久矣,HarmonyOS 7 正在抹平系統能力的接入鴻溝
HarmonyOS 7 透過 Skill 與 Agent 等高階抽象元件,讓開發者能直接呼叫語音辨識、意圖識別等系統能力,大幅降低整合門檻。此舉旨在解決開發者長期重複「造輪子」的痛點,並藉由統一的模組化介面,加速應用開發週期,促進生態創新。
