Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device
Yesterday, Liquid AI released LFM2.5-VL-3B.It is a 3.1B-parameter vision-language model built for on-device deployment.The model reads digital screens across mobile, web, and desktop.It grounds objects to coordinates, parses documents and charts, and calls tools from text or image input.
Liquid AI reports an average of 69.4 across 28 vision benchmarks.That matches InternVL-3.5-4B and sits 0.7 points behind Qwen3.5-4B, both 4.7B models.The model is non-reasoning, so it answers directly and keeps latency low.
It fits in roughly 3 GB of memory and decodes 228 tokens/s on an Apple M5 Max.Is it deployable?Yes, the checkpoint ships in four formats: native, GGUF, ONNX, and MLX.Day-one runtimes include llama.cpp, MLX, vLLM, SGLang, and ONNX.It fits in roughly 3 GB of memory.
Which company levels: The LFM Open License v1.0 is Apache-2.0-based with one change: free commercial use ends once a company’s annual revenue reaches $10M USD.So indie developers, startups, and SMBs under that line can ship commercially at no cost.
Enterprises above it must negotiate a commercial license with Liquid AI.Research, education, and non-profit use carry no revenue limit.Industries: Consumer electronics, automotive, industrial and robotics, financial services, healthcare, and e-commerce.Also QA and RPA vendors that automate GUIs.
Applications: On-device screen agents, GUI test automation, PDF-to-structured-text with layout labels, invoice and receipt OCR, near-real-time object detection in vehicles, offline translation of menus and road signs, and multi-image comparison.So, What is new?LFM2.
5-VL-3B extends LFM2-VL-3B along four axes.Screen and UI understanding: The model averages 80.7 on ScreenSpot-v2 across desktop (78.7), mobile (81.2), and web (82.2).Liquid AI reports Gemma-4-E4B at 51.2 and Qwen3.5-4B at 78.5, with the larger InternVL-3.5-4B ahead at 84.1.
Function calling: This is new to the VL line.ToolSandbox moves from 26.4 to 59.5.BFCL v4 moves from 20.5 to 32.5.Tool calls are emitted as Pythonic calls between <|toolcallstart|> and <|toolcallend|> tokens.Grounding: RefCOCO-avg precision@1 rises from 57.1 to 87.
9, a 30-point gain driven by scaled synthetic grounding data.Multi-image input: BLINK improves from 50.2 to 61.5, and MuirBench from 34.9 to 58.3.Architecture and training The language backbone is LFM2.5-2.6B.The vision tower is a SigLIP2 NaFlex shape-optimized 400M encoder.
NaFlex handles native resolution by splitting large images into non-overlapping 512×512 patches plus a resized whole-image thumbnail.Context length is 32,768 tokens, and 16 languages are supported.Pre-training used approximately 34T tokens.
Vocabulary was doubled to 128K by extending the existing tokenizer in place, which improves non-Latin script coverage.Vision pre-training was scaled 4× in tokens with curated and synthetic caption, OCR, grounding, and instruction-following data.
Post-training is SFT with knowledge distillation from a larger teacher and Antidoom training, followed by multi-reward reinforcement learning.The model is non-reasoning.It answers directly, which is the design choice behind its latency profile.
Benchmarks Liquid AI evaluated across 28 vision benchmarks using vLLM 0.26.0 in non-reasoning mode.LFM2.5-VL-3B averages 69.4, matching InternVL-3.5-4B (69.4) and landing 0.7 points behind Qwen3.5-4B (70.1).Both comparison models are 4.7B parameters.Notable individual results: RealWorldQA 73.
1 against InternVL-3.5-4B at 67.7, TextVQA 84.3 against Qwen3.5-4B at 81.2, MMStar 63.3, MathVista-mini 68.5, ChartQA 81.3, DocVQA 91.1, and OCRBench v1 84.2.CountBenchQA regressed to 87.3 from 92.2 in the prior release.On text-only evaluation, IFEval reaches 82.3, up from 72.9.
Gemma-4-E4B still leads there at 87.9.(function(){ var f=document.getElementById('lfmx-frame'); window.addEventListener('message',function(e){ if(e&&e.data&&e.data.lfmxHeight){ f.style.height=e.data.lfmxHeight+'px'; } },false); })(); Key Takeaways LFM2.5-VL-3B hits a 69.
4 average across 28 vision benchmarks, matching 4.7B-class models.ScreenSpot-v2 jumps to 80.7 and RefCOCO-avg to 87.9, from 57.1 in the prior release.Function calling is new to the VL line: ToolSandbox 26.4 → 59.5, BFCL v4 20.5 → 32.5.
Runs in ~3 GB, decoding 228 tok/s on M5 Max and 20 tok/s on a Galaxy S26 Ultra.LFM Open License v1.0 is free commercially only under $10M annual revenue.Check out the Technical Details and Model Weights.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device appeared first on MarkTechPost.
Related
相關文章

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”
阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

月之暗面第一代萬億參數多模態模型 Kimi K2.5 官宣月底結束服役
作者:歸瀧 責編:歸瀧 評論: 8 月 24 日消息,月之暗面 Kimi 官方微博今日宣佈,其第一代萬億參數多模態模型 —— Kimi K2.5 本月底即將結束服役。據此前報道,今年 1 月,月之暗面宣佈推出並開源了其最新的 Kimi K2.

消息稱知名 AI 研究員 Luke Metz 離開 OpenAI,加入 Meta 超級智能實驗室
作者:遠洋 責編:遠洋 評論: 感謝網友 華南吳彥祖 的線索投遞!8 月 24 日消息,據知情人士向 Axios 證實,知名 AI 研究員 Luke Metz 已加入 Meta 的超級智能實驗室(Superintelligence Labs)。

Anthropic 最強大模型 Fable 5 遇冷,企業用戶轉向更便宜 AI 產品
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,據英國《金融時報》報道,Anthropic 的美國客戶正在使用更便宜的替代品來替代其最強大的 AI 工具,這在其預計將實現有史以來規模最大的 IPO 之前,對其高支出的商業模式提出了質疑。

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。

企業AI最後一公里:三路人馬在此交鋒
鄭敏芳 發表於 2026年08月24日 03:09 摘要:尋找自己的位置 2026年世界機器人大會現場,談到這一輪突然走紅的FDE(前線部署工程師),明略科技CEO吳明輝先把時間往回撥了十多年。“12年前我們就在非常認真地研究。”當華爾街見聞·問及FDE與傳統軟件部署有什麼區別時,吳明輝說,兩者都會進入客戶現場,但今天的FDE需要做得更深:一邊把Agent接進真實業務,一邊把現場形成的能力繼續沉澱回後臺。