Supersonic Labs 推出 Julia 1:一個 1.443 億參數、可在 CPU 上運行的開放決策模型

2026年9月26日 19:50
站內 AI 整理稿

Supersonic Labs, a small AI lab from Brazil, has released Julia 1.It is a compact decision model, not a chatbot.You pass it context, a question, and 2 to 20 candidate answers.It picks one and returns a probability for every option.The model has 144.3M parameters and runs on a plain CPU.

Is it deployable?Yes.The weights are on Hugging Face under Apache 2.0 and run locally with Python 3.11+ on CPU or a BF16-capable GPU.An ONNX build also runs in the browser via WebGPU.A hosted API is announced but not open yet.

What Julia 1 Does Julia 1 handles three decision types through one API: choice: pick one label from 2 to 20 described options (classification, routing).score: return the expected index on an ordered rubric, such as low, medium, high.noul: return the probability that a yes-or-no statement is true.

Results come back in the caller’s option order with full softmax probabilities.Caller IDs such as billing are returned unchanged.The model does not generate text.

Architecture and Training Budget Julia 1 starts from JHU CLSP’s mmBERT-small, a 140M-parameter multilingual ModernBERT encoder trained on 1,800+ languages.Supersonic Labs kept the encoder and tokenizer, added a decision head, and trained on decision-format examples.

The lab states Julia 1 is not a fine-tuned Qwen model.The runtime supports 8,192 combined tokens, but published benchmarks used a 1,024-token limit.Total cloud GPU spend for training and experiments was about R$540 (US$104.08).The FP32 weights occupy 550.5 MiB.

The private training pipeline is not released.Julia 2, with the lab’s own foundation architecture, is in development.Benchmark Results The September 24, 2026 evaluation ran on H200 BF16 with strict encoding.

The comparison baseline is TypeSafe’s Jev, using reference values from the Jev benchmark protocol, not a new Jev run.Typed Decisions: 73.15% (1,463/2,000) vs 72.70% reference.AG News, 4 labels: 94/100 vs 91% reference.DAIR Emotion, 6 labels: 86/100 vs 48% reference.

Banking77, 72 labels: 64/100 vs 87% reference.This is the clear failure.MASSIVE, 18 scenarios: 71.50% macro accuracy across 52 locales; 86.25% pt-PT, 86.75% en-US.The classification pilots use only 100 examples each.A September 25 CPU run reproduced most numbers: 72.

55% on Typed Decisions and 60/100 on Banking77 with 3 abstentions.On-Device Latency The lab published per-device measurements.On an Apple M4, one decision per call took a 33.15 ms median.On a Samsung SM-X510 tablet via ONNX Runtime, the median was 203 ms with 393.1 MB peak RSS.

On an Intel Core i5-1235U, AG News decisions took a 107.83 ms median.Banking77 took 3,713.54 ms because it narrows 72 labels first.On X, @supersonicai claims Julia 1 classifies 5x faster than Jev on an i5 laptop.Treat that carefully.

The Jev pilot measured Jev as a hosted service called from France, so latencies are not like-for-like.Introducing Julia-1:Our first classification model that runs on almost anything.Learn more https://t.co/YJCdEeIBSo pic.twitter.

com/3cizsbG9ZB— Supersonic Labs (@supersonicai) September 26, 2026 Interactive Explainer (function(){var f=document.getElementById('mtp-julia1-frame');window.addEventListener('message',function(e){if(e.source===f.contentWindow&&e.data&&e.data.j1xHeight){f.style.height=e.data.

j1xHeight+'px';}});})(); Julia 1 vs Closest Competitors FeatureJulia 1TypeSafe JevGLiNER2.5 MultiDeveloperSupersonic LabsTypeSafe AIFastinoAccessOpen weightsHosted API, early accessOpen weightsLicenseApache 2.0ProprietaryApache 2.0Parameters144.

3MNot disclosed287MBase encodermmBERT-smallNot disclosedmDeBERTa-v3-baseDecision typesChoice, score, yes/noTyped structured decisionsClassification, NER, relations, recordsOptions per call2 to 20 (Router for more)Up to 255Label list per schemaRuns locally on CPUYesNoYesInput price per 1M tokens$0.

025 (planned API)$0.042Free (self-hosted)AG News pilot94%91%70%DAIR Emotion pilot86%48%44%Banking77 pilot64%87%61% Sources: Julia 1 model card, TypeSafe launch post, GLiNER2.5 Multi card, Jev benchmark pilot.Julia 1 pilots ran separately from the Jev and GLiNER runs.

Limitations Julia 1 compares the answers you supply.It cannot be counted on for missing facts, algebra, or multi-step calculation.The Router can drop the correct label during narrowing.It is not a drop-in Transformers pipeline, and no Hugging Face inference provider serves it.

Supersonic Labs advises evaluating on your own questions and keeping humans in the loop for consequential decisions.Key Takeaways Julia 1 is a 144.3M-parameter, Apache 2.0 decision model that runs on CPU.One API covers choice, ordered score, and yes-or-no decisions over 2 to 20 options.

It beat Jev references on 3 of 4 pilots but trailed badly on 72-label Banking77.Median latency hit 33.15 ms per decision on an Apple M4.Training cost about US$104 in cloud GPUs; a $0.025/MTok API is planned.Check out the Model Weights, ONNX/WebGPU build, and Technical details.

All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU appeared first on MarkTechPost.

Related

相關文章

量子位模型更新

量子計算走上桌面!“小盒子”跑通端到端,數據全程不出門

酉術量子發布全球首個異構計算硬件底座UnitarySpark,並上線以自然語言驅動的量子科學計算平台UnitaryLab 2.5公測版。該平台旨在降低量子計算的使用門檻,讓開發者能用大白話描述科學問題,由機器自動完成量子與經典算力的協同運算,無需自行撰寫量子程式或處理數據上雲的合規問題。

1 小時前
量子位模型更新

索辰科技加碼世界模型,與戰略投資企業美夢空間聯合發佈具身模型與物理測評標準

索辰科技投資的美夢空間在數貿會上發布具物理感知能力的Physical-WAM世界動作模型及RoboTwin-Phys物理漂移評測基準,旨在解決現行VLA模型缺乏物理理解、實際操作成功率低的問題。Physical-WAM透過物理Token與三大模塊讓機器人預判與修正動作,RoboTwin-Phys則提供可控的物理擾動評測標準,且已開源,希望加速具身智能產業落地。

1 天前
MarkTechPost AI模型更新

Liquid AI 推出 LFM2.5-VL-3B-DSpark:推測解碼加速視覺語言模型,最高提升 3.13 倍

Liquid AI 發布 LFM2.5-VL-3B-DSpark,這是針對其 LFM2.5-VL-3B 視覺語言模型所設計的實驗性推測解碼草稿模型。該草稿模型約增加 2.8 億個參數,在不改變模型輸出的前提下加快解碼速度。官方數據顯示,在 Apple 晶片上解碼速度最高提升 3.13 倍,在 NVIDIA H100 上則最高提升 2.66 倍。此模型已可部署,權重以 Safetensors 與 GGUF 格式發布於 Hugging Face,並支援 SGLang、MLX-VLM 與 llama.cpp。Liquid AI 團隊將其標註為實驗性版本,採用 LFM 開放授權 v1.0,僅允許年營收低於 1000 萬美元的公司免費商用。推測解碼的原理是:標準模型每次前向傳播僅生成一個 token,而此技術加入一個小型草擬模型,以加速整個生成流程。

2 天前
IT之家模型更新

德國柏林警方藉助人工智能監控攝像頭打擊犯罪,自動識別暴力和破壞行為

作者:浩渺 責編:浩渺 評論: 9 月 25 日消息,據央視新聞報道,德國柏林警方 24 日在一處犯罪高發區域啟動一項計劃,將藉助人工智能監控攝像頭打擊犯罪。據悉,柏林警方將在這項計劃啟動的前四周內,在科特布斯門地區安裝並調試好人工智能監控攝像頭,隨後正式投入使用。

2 天前