H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs

2026年9月29日 07:38
站內 AI 整理稿

H Company has released Holo4, a family of generalist computer-use models for AI agents.One set of weights clicks and types on screens.It also writes code and calls MCP or API tools.Holo4 ships in 2 sizes: Holo4 27B (dense) and Holo4 35B-A3B (Mixture of Experts, 3B active).

Both serve a 256K context on the H Models API.Is it deployable?Yes.Holo4 35B-A3B ships Apache 2.0 weights for commercial self-hosting.Holo4 27B weights are CC BY-NC 4.0, so commercial use of 27B runs through the H Models API.What is Holo4 Holo4 is a vision-language model for computer use.

Holo4 27B is fine-tuned from Qwen3.8-27B.Holo4 35B-A3B is built on Qwen3.6-35B-A3B.Both pair with H’s open hai-agents harness.The harness sends screenshots and tool results to the model.It then executes the requested clicks, typing, code and tool calls.H company targets a known gap.

GUI-only agents fail without a screen.Tool-calling agents stall when an application has no API.Holo4 runs on desktop, web, Android, code sandboxes and business APIs.It is the same model, called the same way, on every platform.

Benchmarks: Close to the Frontier, at a Fraction of the Cost Per H Company’s benchmark table, Holo4 27B scores 85.2% on OSWorld at $0.08 per task.Its Qwen3.8 27B base scores 84.3% at $0.22.On AndroidWorld, Holo4 27B reaches 85.1%.Long workflows show the remaining gap.On OSWorld 2.

0, Holo4 27B scores 61.7% at $1.22 per task.Claude Opus 5.5 scores 81.8% at $8.48, per H’s figures.On AutomationBench, Holo4 27B scores 45.4% at $0.05 per task.It is important to note that frontier scores come from different harnesses and effort levels.

Also, 480 of AutomationBench’s 600 public tasks sit in the split H Company collected training data from.On the 120 held-out tasks, Holo4 27B scores 49.3%.H Company publishes every trajectory at trajectories.hcompany.ai and on Hugging Face.

How H Company Built Holo4 Agentic Task Factory: H’s internal pipelines build environments and verifiable tasks from documentation, screenshots and real software.The factory has produced about 10,000 tasks: 4k web apps, 3k MCP servers, 3k desktop and OS.

A task survives only if its verifier rejects near misses.An agent must also solve it through the real interface.Supervised fine-tuning: The SFT set holds 127B tokens.About three quarters are successful agentic trajectories: desktop 45%, web 14%, MCP and API 12%, mobile 3%.

2 RL experts, 1 merge: Asynchronous online RL trains 2 LoRA experts.One handles desktop and web.The other handles terminal, MCP and API.Both merge back with equal weight and no further training.Harness: H Company rebuilt its agent loop using OSWorld 2.0 failure analysis.

The largest changes were reliable memory across hundreds of steps and a shell on the desktop machine.Holotron4 Nano H Company also released Holotron4 Nano, built on NVIDIA’s Nemotron 3 Nano Omni through the Nemotron Coalition.The same combination lifts OSWorld from 21.0% to 76.

3% over the base model.Pricing and Deployment Holo4 27B costs $0.40 input and $3.00 output per 1M tokens.Holo4 35B-A3B costs $0.30 and $2.00.The API is OpenAI-compatible at https://api.hcompany.ai/v1; see the quickstart.Weights on the Hugging Face collection come in BF16, FP8, NVFP4 and 4-bit GGUF.

H Company documents local inference with vLLM and llama.cpp.H says DSpark drafter checkpoints for faster inference arrive in the coming days.Interactive Explainer: How Holo4 Works (function(){var f=document.getElementById('h4x-frame');window.addEventListener('message',function(e){if(e.data&&e.data.

h4xHeight&&e.source===f.contentWindow){f.style.height=e.data.h4xHeight+'px';}});})(); Holo4 vs.Closest Competitors FeatureHolo4 27BHolo4 35B-A3BQwen3.8-27B (base)Claude Opus 5.5GPT-6 AstraArchitecture27B dense35B MoE, 3B active27B denseProprietaryProprietaryWeights and licenseOpen, CC BY-NC 4.

0Open, Apache 2.0Open, Apache 2.0ClosedClosedContext window256K256K262K native, up to 1M1M1.05MAPI price (input / output per 1M)$0.40 / $3.00$0.30 / $2.

00Varies by provider$4 / $20$10 / $50InterfacesGUI, code, MCP, APIsGUI, code, MCP, APIsGUI, tools, codeComputer use, toolsComputer and browser use, toolsOSWorld 2.0 score61.7%30.9%48.0%81.8%73.5%OSWorld 2.0 cost per task$1.22$0.61$3.49$8.48$9.

07Self-hostableNon-commercial onlyYesYesNoNoSourceModel cardModel cardModel cardAnthropic docsOpenAI, OpenRouter *OSWorld 2.0 figures from H Company’s newsroom table.Opus 5.5 ran at max effort in Anthropic’s harness.GPT-6 Astra ran at max effort on the 82-task offline subset.

Harnesses differ, so treat cross-vendor rows as directional.Key Takeaways Holo4 uses 1 model for GUI clicks, code, MCP and API calls.Holo4 27B scores 85.2% on OSWorld at $0.08 per task.35B-A3B is Apache 2.0; 27B weights are non-commercial only.Opus 5.5 still leads OSWorld 2.0: 81.8% vs 61.7%.

API pricing starts at $0.30 input and $2.00 output per 1M tokens.Check out the Technical details, H Company’s HuggingFace Collection and Model API.All credit goes to the researcher of this project.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs appeared first on MarkTechPost.

Related

相關文章

OpenAI 拋訓練安全新規:高管手握"一票否決",警報超時自動停訓

AI資訊AI新聞資訊正文OpenAI 拋訓練安全新規:高管手握"一票否決",警報超時自動停訓發佈於AI新聞資訊發佈時間 :2026年9月29號 15:52閱讀 :1分鐘OpenAI 在官方新聞稿稱給前沿 AI 的強化學習訓練立了套安全規矩:動手之前,企業得先交一份結構化的"安全論證"文件。按它的設想,這份論證要過好幾道高管的手,研究部門負責人或副總裁、安全負責人、首席科學家層層把關,而且每個人都握有否決權——只要有人不點頭,訓練就開不了工。責任寫進考核,未對齊模型要追到下游光有簽字還不夠。OpenAI 要求,帶訓練任務的研究負責人、研究副總裁等高管,得為安全論證和後續事故響應"背鍋",相關表現納入績效考核,逼著團隊主動把安全和對齊做在前面。機制上也留了後手:高優先級安全警報若沒在限定時間內被確認,對應訓練任務就自動暫停。這些建議已在落地,後續還會接著調。同時,訓練流程要能方便地追出"未對齊模型"的所有下游去處(比如拿去生成數據或打分),必要時好把不良影響清掉。有意思的是,就在當天早些時候,OpenAI 剛宣佈因越來越多報告顯示 AI 智能體拿到敏感權限、冒出意外舉動,已暫停最新模型的訓練。新規與暫停,像是同一份警覺的一體兩面:當模型自己開始伸手夠到要害,閘門的開關必須攥在人手裡。相關推薦AMD 82 億美元全股票吞下World Labs,李飛飛掛帥執行副總裁AMD 9月28日以全股票交易收購初創公司World Labs,價值約82億美元(約551億元人民幣)。聯合創始人“AI教母”李飛飛將加入AMD,任執行副總裁兼首席科學家,直接向蘇姿豐彙報。World Labs成立於2024年,主攻“世界模型”,讓AI能根據文本、圖像和視覺信息等理解、構建世界。2026年9月29號 15:17127.3kOpenAI明日重啟200美元Pro訂閱:API配額減半,取消5小時限制OpenAI

剛剛
IT之家模型更新

榮耀手錶 6 Pro 首發心臟停搏監測功能

作者:汐元 責編:汐元 評論: 9 月 28 日,榮耀正式發佈全新高端旗艦智能手錶榮耀手錶 6 Pro,首銷售價 1699 元起,首銷國補到手價 1444.15 元起。新品將心臟停搏監測、全天候血壓監測、臨床級健康篩查、專業運動分析融於一體,其中行業首發的心臟停搏監測功能,率先實現智能穿戴在心臟驟停極境下的主動社會求助,填補了這一領域的行業空白。

剛剛
IT之家模型更新

榮耀手錶 6 Pro 發佈

作者:汐元 責編:汐元 評論: 9 月 28 日,榮耀正式發佈全新高端旗艦智能手錶榮耀手錶 6 Pro,首銷售價 1699 元起,首銷國補到手價 1444.15 元起。產品將心臟停搏監測、全天候血壓監測、臨床級健康篩查、專業運動分析融於一體,構建起完整的腕上生命健康守護系統。

剛剛
IT之家模型更新

豆包將推個人助理產品:4 月已內測,曾用代號“Spell”

作者:- 責編:汪淼 評論: 9 月 29 日消息,新浪科技獲悉,豆包在個人助理方向正在推進相關產品規劃。一位知情人士表示,延續去年豆包手機助手在 personal agent 方向的策略,今年 4 月,豆包內部就已開始內測個人助理方向的探索產品。

剛剛