Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together
Model cards report quality under server-class, full-precision conditions.Those numbers rarely predict how the same model behaves on a phone.This week, Liquid AI released Pipette.
It is an open-source platform for benchmarking foundation models on edge devices, built in partnership with Artificial Analysis as an independent methodology validator.Pipette treats on-device behavior as a property of the deployed system, not the model in isolation.
Its unit of measurement is a full configuration: model + quantization + runtime + device.The launch dataset covers five on-device performance metrics across more than 1,000 model × quantization × runtime × device × context configurations, spanning 30+ models, llama.
cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to 8,192 tokens.Initial verified results come from a MacBook Pro with M5 Max, an iPhone 17 Pro and a Galaxy S26 Ultra.The practical claim is testable: two 350M models at the same quantization on the same phone retain 78.
4% and 33.8% of decode throughput at 4,096 tokens.Is it deployable?Yes, Pipette ships as Apache 2.0 infrastructure (pipette-mgmt, pipette-clients, pipette-scores), a public results dataset, a hosted dashboard, and native iOS and Android benchmark apps.Nothing is waitlisted.
Publication of community-submitted results is still in beta.Which companies: Any team shipping a model onto hardware it does not own.Solo developers and seed-stage startups can use the dashboard and apps without infrastructure.
Mid-market product teams can run the clients across an internal device fleet.Large OEMs, chip vendors and enterprises can operate the whole pipeline behind their own firewall.
Industries: Consumer electronics and smartphone OEMs, automotive, industrial and robotics, healthcare devices, financial services, defense — anywhere latency, privacy or connectivity forces inference onto the device.
Applications: Model and quantization selection before a sprint commits; SoC and hardware procurement validation; regression testing when a runtime, OS or driver updates; context-length capacity planning; independent verification of vendor performance claims.
What Liquid AI shipped Liquid AI released Pipette in partnership with Artificial Analysis, an independent validator that reviewed and verified the methodology.The premise is narrow and useful: on-device behavior is a property of the deployed system, not of the model in isolation.
The launch dataset covers five on-device performance metrics across more than 1,000 model × quantization × runtime × device × context configurations.It spans 30+ models, multiple quantization formats, llama.cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to 8,192 tokens.
Initial published results come from a MacBook Pro with M5 Max, an iPhone 17 Pro and a Galaxy S26 Ultra, with AMD Ryzen AI Max+ 395 and Radeon 8060S results listed as coming soon.In Pipette, the unit of measurement is a deployment configuration: model + quantization + runtime + device.
A benchmark then defines the metric and token shape, producing a latency, throughput or memory result.Quality is tracked separately on IFBench, GPQA Diamond and MATH-500.Those quality scores currently come from llama.
cpp evaluation runs on NVIDIA H100 80GB reference systems, then get matched to on-device runs sharing the same model and quantization — a quality number shown next to phone throughput was not produced on the phone.
Why the deployment context changes the answer Four published comparisons show how far a configuration can move a decision: Context scaling can diverge at identical parameter counts.At Q4KM on Galaxy S26 Ultra, Granite-4.0-H-350M retains 78.
4% of its decode throughput from 256 to 4,096 input tokens, while Granite-4.0-350M retains only 33.8%.Sparse activation buys speed, not memory.At 2,048 input tokens on the same phone, LFM2.5-8B-A1B decodes 2.4x faster than Qwen3.5-4B and 2.6x faster than Ministral-3-3B-Instruct-2512.It activates 1.
5B of 8.5B parameters per token, yet still peaks at 5.29 GiB because all expert weights occupy memory.Speed and quality do not co-locate.On iPhone 17 Pro at Q4KM, MiniCPM5-1B completes a 2,048-in / 256-out workload in 3.47 seconds versus 4.12 seconds for LFM2.5-1.2B-Instruct, a 15.
8% reduction in elapsed time.On the same artifacts, LFM scores 9.0 points higher on MATH-500.Near-identical system profiles can hide task-level reversals.At Q4KM and 2,048 input tokens on M5 Max, Granite-4.1-8B and Ministral-3-8B-Instruct-2512 differ by 2.4% in decode throughput and 1.
2% in peak RAM.Granite leads IFBench by 7.3 points; Ministral leads GPQA Diamond by 14.0 points.How the measurements are produced Performance runs follow a published methodology: fixed token shapes, greedy decoding, a discarded warm-up, five measured repetitions and readiness gating.
Before each timed repetition, a platform-specific check verifies thermal and load conditions; failing runs are not published.Evaluations use a separate protocol with deterministic, model-blind scoring, and pipette-scores never sees generation provenance.
Every submission records benchmark version, token shape, model artifact, quantization, runtime version and settings, and device hardware and OS.Interactive explainer #mtp-pipette-x7k2 { background:#0B0D12 !important; border:1px solid #232833 !important; border-radius:14px !important; padding:0 !
important; margin:22px 0 !important; overflow:hidden !important; } #mtp-pipette-x7k2 iframe { width:100% !important; border:0 !important; display:block !important; background:#0B0D12 !
important; } #mtp-pipette-x7k2 p:empty, #mtp-pipette-x7k2 hr, #mtp-pipette-x7k2 del, #mtp-pipette-x7k2 s { display:none !important; } (function(){ var f=document.getElementById("mtp-pipette-x7k2-frame"); window.addEventListener("message",function(e){ if(e&&e.data&&e.data.pipetteHeight){ f.style.
height=e.data.pipetteHeight+"px"; f.setAttribute("height",e.data.pipetteHeight); } }); })(); Key Takeaways Pipette benchmarks configurations, not models: model + quantization + runtime + device.Apache 2.0 stack, 1,000+ configurations, 30+ models, three verified devices at launch.
Quality evals run on H100 references and are matched to on-device performance, not measured on-device.Identical parameter counts can differ 78.4% vs 33.8% in context-scaling retention.Check out the Technical Details and Leaderboard.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
The post Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together appeared first on MarkTechPost.
Related
相關文章

Ilya新模型要來了?投資人爆料:今年最重要的發佈
機器之心·2026年08月25日 21:18沉默兩年,SSI即將交卷。Ilya Sutskever 的第一款模型,可能真的要來了。昨天,a16z 合夥人 Martin Casado 突然發出一條預告:「剛剛獲得了一個新模型的訪問權限。這將是今年最重要的模型發佈,甚至可以去掉‘之一’。

叛變,OpenAI親兒子竟然選了Kimi K3
AI圈近來最熱的話題,莫過於一場被外界戲稱為「叛變」的產品路線選擇。向來被視為OpenAI陣營重點扶持的「親兒子」級應用,竟然在關鍵的模型選型上,捨棄了來自母公司陣營的技術,轉而投入了國產模型Kimi K3的懷抱。這個消息一出,立刻在開發者社群與AI從業者之間炸開了鍋,許多人開始重新審視目前大模型競爭的版圖。 所謂的「親兒子」,指的是市場上與OpenAI有深厚資本或技術綁定關係的明星產品。

Claude兩個新模型實測流出,空間推理封神,算力直接榨乾了
27分鐘前AI開始替人類調用AI,token用量已是人類5.2倍4小時前讓Claude改報錯,它卻把紅燈換成了黃燈,三星芯片驗證,AI三次闖禍9小時前閱讀更多內容,狠戳這裡選靠譜AI,看真實評測查看AI測評官方交流社區加入諮詢項目審核和入駐聯繫項目推薦訂閱號關注下一篇高盛大佬警告,過度使用AI,人會慢慢喪失推理能力AI能幹活,但會把人變廢?

四位掌權者,四種命運:美國四大AI巨頭全景
更關鍵的是節奏。8 月 13 日,Google 又拋出了 Gemini 3.7 Flash,主打編碼場景與商業自動化,輸入價格低至每百萬 token 0.75 美元、輸出 3.75 美元——僅為前代的一半。這套模型已經入駐 Gemini Spark 訂閱服務、面向 Google AI Pro 與 Ultra 用戶開放,並行接入 Google AI Studio及新上線的 agentic 開發平臺 Antigravity。但皮查伊麵對的並非全是好消息。第一是旗艦模型的“跳票焦慮”。市場長期以 Gemini 3.
Granite 4.2 LLM:建構技術全解析
本文深入探討 IBM 如何打造 Granite 4.2 推理模型系列。Granite 4.2 是我們首個密集、僅解碼器的推理大型語言模型家族,提供 3B、8B 和 30B 三種規模。每個模型皆從零開始,以約 15 兆個 token 進行預訓練,並採用五階段策略將上下文窗口擴展至 512K tokens。隨後,我們在思維鏈、推理與代理軌跡資料上進行監督式微調,最後透過多階段強化學習管線完成後訓練。
微軟 CEO 納德拉示警:企業要掌控 Token 資本,使用 AI 不能押注單一模型
納德拉認為,有別於面向普通消費者的 AI 運營邏輯,企業會持續創造知識,這些知識應留在企業內部,而不應被綁定到某個模型或服務商。IT之家援引博文介紹,這裡的“Token 資本”涵蓋提示詞上下文、員工與模型的交互記錄,以及員工使用模型的方式等,這是企業在使用 AI 時所積累的知識。