webAI 發表 TwIL-LM:專為本地硬體自動形式化設計的 1.7B 與 3B 形式邏輯模型家族
webAI has released TwIL-LM, a two-model family of formal-logic reasoners at 1.7B and 3B parameters.The 3B member, TwIL-LM3, is a merged fine-tune of SmolLM3-3B; the 1.7B member is a PEFT LoRA adapter for SmolLM2-1.7B-Instruct.
Both target autoformalization: translating English into first-order logic and checking whether a conclusion follows from its premises.Both run locally, with a 1.06 GB quantized build for the 1.7B and a 1.78 GiB Q4KM GGUF for the 3B.
webAI’s announcement frames the release around beating gpt-oss-120b on four of five formal-reasoning lanes.Is it deployable?Partially.Non-commercial use only, as of now.Both checkpoints ship under the webAI Non-Commercial License ver.1.0.
Revenue-generating deployment requires a separate agreement with webAI.Company level: any size.The 3B Q4KM GGUF is 1.78 GiB and runs on CPU or 4 GB of VRAM.The 1.7B Q4KM is 1.06 GB.
Industries: compliance and RegTech, financial services, healthcare and pharma, legal and contract operations, formal-methods research.webAI positions local execution for environments where data cannot leave the device.
Applications: first-order logic (FOL) translation, entailment classification over premise sets, natural language to structured query, Lean formalization drafting and critique, and a verifier layer that checks a larger model’s output.How TwIL-LM3 was built?Four stages sit on top of the base model.
LoRA supervised fine-tuning on a synthetic formal-logic corpus.Checkpoint fusion, averaging intermediate SFT checkpoints in parameter space.WiSE-FT interpolation back toward the pretrained base at λ = 0.25.Then MGPO, an entropy-weighted GRPO stage run against a programmatic verifier.
The published checkpoint is step 2071.That λ is load-bearing: only a quarter of the fine-tuned delta is retained.A sibling arm that skipped the interpolation scored higher in-domain, at macro gate 0.515, but gave back roughly twelve points of held-out capability.webAI did not publish that arm.
(function(){ window.addEventListener("message", function(e){ if(!e.data || typeof e.data.twilHeight !== "number") return; var f = document.getElementById("twil-explainer-frame"); if(f && e.data.twilHeight > 200 && e.data.twilHeight < 6000){ f.style.height = e.data.
twilHeight + "px"; } }, false); })(); Performance webAI's announcement lists 96.4 on rule induction, 87.6 on semantic parsing, 64.6 on Lean formalization, 52.0 on exact-format answering, and 68.7 on entailment labeling.It reports two tracks.On Track A, in-domain formal logic, TwIL-LM3 scores 0.
4488 on the six-lane average and 0.4218 on the macro gate, the metric the training pipeline gates on.It leads every arm up to and including LFM2.5-8B-A1B on all six objective lanes, at 0.4218 against 0.3757 with a third of the parameters.It does not lead the two largest arms.
Qwen3-8B takes the gate 0.5336 to 0.4218, but most of that is loose-match credit; under strict-7 the two sit at 0.2093 and 0.1971.gpt-oss-120b takes the six-lane average 0.5192 to 0.4488.Efficiency is where the model card is unambiguous.
TwIL-LM3 produces the shortest generations of any arm, 482 tokens on Track B, and consequently the most answers per second at 32.9 against the 120B's 4.2.https://www.webai.
com/blog/webai-releases-twil-lm-a-family-of-formal-logic-models-that-outreason-a-120b-model-and-run-on-an-iphone Held-out transfer TwIL-LM3 improves in-domain by +26% relative, macro gate 0.336 to 0.422, while also gaining +0.022 on the held-out core average.
The model card calls it the only arm in the project that gains on both tracks.LogicBench moves to 0.7167 from 0.6467.GSM8K slips slightly to 0.8733 from 0.8833, and IFEval regresses to 0.6433 from 0.6767.The 1.7B is a different trade.Its macro-primary score is 0.361 against 0.
185 for the unadapted base.Out-of-distribution results are mixed: LogicBench BQA improves to 0.590 from 0.563, while GSM8K falls to 0.380 from 0.413 and ARC-C chain-of-thought falls to 0.463 from 0.587.Key Takeaways TwIL-LM3 (3B) and TwIL-LM (1.
7B) target formal logic, both under a non-commercial license.Shipping TwIL-LM3 trails gpt-oss-120b on the six-lane average, 0.4488 to 0.5192.Its real edge is efficiency: 32.9 answers/sec from 482-token generations.WiSE-FT at λ = 0.25 is why in-domain gains do not collapse held-out performance.
Check out the Model weights and Technical details.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware appeared first on MarkTechPost.
Related
相關文章

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”
阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

月之暗面第一代萬億參數多模態模型 Kimi K2.5 官宣月底結束服役
作者:歸瀧 責編:歸瀧 評論: 8 月 24 日消息,月之暗面 Kimi 官方微博今日宣佈,其第一代萬億參數多模態模型 —— Kimi K2.5 本月底即將結束服役。據此前報道,今年 1 月,月之暗面宣佈推出並開源了其最新的 Kimi K2.

消息稱知名 AI 研究員 Luke Metz 離開 OpenAI,加入 Meta 超級智能實驗室
作者:遠洋 責編:遠洋 評論: 感謝網友 華南吳彥祖 的線索投遞!8 月 24 日消息,據知情人士向 Axios 證實,知名 AI 研究員 Luke Metz 已加入 Meta 的超級智能實驗室(Superintelligence Labs)。

Anthropic 最強大模型 Fable 5 遇冷,企業用戶轉向更便宜 AI 產品
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,據英國《金融時報》報道,Anthropic 的美國客戶正在使用更便宜的替代品來替代其最強大的 AI 工具,這在其預計將實現有史以來規模最大的 IPO 之前,對其高支出的商業模式提出了質疑。

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。

企業AI最後一公里:三路人馬在此交鋒
鄭敏芳 發表於 2026年08月24日 03:09 摘要:尋找自己的位置 2026年世界機器人大會現場,談到這一輪突然走紅的FDE(前線部署工程師),明略科技CEO吳明輝先把時間往回撥了十多年。“12年前我們就在非常認真地研究。”當華爾街見聞·問及FDE與傳統軟件部署有什麼區別時,吳明輝說,兩者都會進入客戶現場,但今天的FDE需要做得更深:一邊把Agent接進真實業務,一邊把現場形成的能力繼續沉澱回後臺。