Holo4: powering generalist computer-use agents

2026年9月28日 09:44
站內 AI 整理稿

Back to Articles Holo4: powering generalist computer-use agents Team Article Published September 28, 2026 Upvote 2 Tony Wu h-tonywu Follow Hcompany maxime h-maxime Follow Hcompany Frederic Renard frenard-h Follow Hcompany Vincent Coyette vincentcoyette Follow Hcompany Emrick Sinitambirivoutin emricksini-h Follow Hcompany Avshalom Manevich avshalom-h Follow Hcompany Antonio Loison antonioloison Follow Hcompany Antoine Bonnet ABonnetH Follow Hcompany Maxime Langevin maxime-hcompany Follow Hcompany Aleix Cambray (H-AI) h-aleixcambray Follow Hcompany Tony Wu tonywu71 Follow Hcompany Mats L.

Richter MatsLRichter Follow Hcompany Michael Eickenberg michaelhai Follow Hcompany Sławek Mucha smucha-h Follow Hcompany Matthias Brunel mbrunel-H Follow Hcompany Holo4 is our new series of agentic models.It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts.

Both are available on the H Models API.We are also releasing an updated version of Holotron 3: Holotron4 Nano.Holo4 builds on our previous model and interacts with software through any available interface: GUIs, code, MCP and APIs.

It scores well on academic benchmarks, but we built it for real business workflows.It was trained through supervised and reinforcement learning on a large set of environments and tasks, including those generated by our Agentic Task Factory.

Get started now: 🤖 Models: Holo4-27B | Holo4-35B-A3B | Holotron4 Nano 🗂️ Full collection (FP16, FP8, GGUF): Holo4 🎞️ Trajectories: viewer | dataset ⚡ H Models API: quickstart 📝 Full blog post: hcompany.

ai/newsroom/holo4 Models built for every interface Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools.It uses whichever fits the task.

Most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application that has no API.

Real work is not siloed that way, and a single business task can require combining these different approaches.Holo4 runs on desktops, on the web, on Android, in a code sandbox and against business APIs.It is the same model in each case and it is called the same way.

You do not need to select a different model for each platform.Holo4 models improve significantly over their Qwen base.Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%.

However, it does so with orders of magnitude fewer parameters and at a much lower cost.We open-source every trajectory behind our scores on public benchmarks: replay each step at trajectories.hcompany.ai or download them from Hugging Face.

Competitive with the frontier, at a fraction of the cost On the hardest academic benchmarks for desktop control (OSWorld 2.0) and API use (AutomationBench), Holo4 competes with frontier models at a much lower cost per task.Notes on the cost-performance charts OSWorld 2.0.

Costs are estimated from the input and output tokens of each agentic run.Holo4 is priced at H Models API rates (single run).Qwen3.8 27B: model card score, cost from the tokens of our run at Alibaba Cloud list prices.Qwen3.

6 35B-A3B: single run in our harness, at Alibaba Cloud list prices with cache hits at 20% of the input price.OpenAI launch data supplies the GPT and Opus effort sweeps; other closed and open-weight points use the official OSWorld 2.0 leaderboard.Releases, harnesses and task subsets differ.

The line connects non-dominated score and cost pairs among the closed models; Holo4 is excluded.AutomationBench.Holo4, Qwen3.8 27B and Qwen3.6 35B-A3B: AutomationBench v1.0.6, scores and costs measured in our internal harness.

Other models: public-set scores from the AutomationBench README, cost per task from the official leaderboard, which runs on the private set.We will report Holo4 on the private set once it is evaluated.

AI that does work Trained on environments and tasks from our Agentic Task Factory, Holo4 models excel on professional software.The examples below show Holo4 27B alongside Qwen3.8 27B, its base model.Same prompt and harness for both models.

3D modeling · Eiffel tower Build a 3D model of the Eiffel Tower in FreeCAD, at a scale of 1 mm to 1 metre, centred on the origin and aligned to the X and Y axes.Work to this design.The tower is square in plan at every height, never round.

Its half-width, measured from the central axis out to the corner, is 62.5 mm at ground level, 32.5 mm at height 57, 17.5 mm at height 115, and 9.35 mm at height 276.

Between those heights the half-width follows a smooth curve that falls steeply near the ground and gently higher up, never a straight line.Four identical legs, one per quadrant, each a square column whose outer corner follows that profile.

Each leg is 14 mm across at the ground and tapers to 4 mm at height 276.The legs stand apart from the ground up to the first platform, and converge as they rise.Nothing fills the space between them: the tower is open, and you can see straight through it from every side.

Three platforms, each a solid square slab centred on the axis: 72 mm across and 4 mm thick at height 57; 40 mm across and 3 mm thick at height 115; 22 mm across and 3 mm thick at height 276.A mast from height 276 to 324, square, 8 mm across at its base tapering to 2 mm at the tip.

Every part must be a closed solid with non-zero volume, and no part may fill the space between the legs.Holo4 27B (84 calls, 1.3M tokens) Qwen3.8 27B (60 calls, 1.

0M tokens) 3D modeling · H logo Build a 3D model in FreeCAD of the H company logo: a solid filled disc beside a blocky sans-serif capital letter H, both extruded to the same thickness, the two shapes of similar height and set apart so they do not overlap, with the centre of the disc level with the middle of the H.

Colour both shapes black.Holo4 27B (94 calls, 1.5M tokens) Qwen3.8 27B (118 calls, 1.9M tokens) Game design · Pac-Man Build a Pac-Man-style game in Godot and leave it running.A rectangular maze of walls laid out on a grid, with pellets filling every open corridor.

A player marker moves continuously along the corridors, eating each pellet it passes over and scoring a point for it.Three ghosts move through the same corridors and chase the player.If a ghost catches the player, the player loses a life and everything resets to its starting position.

Score and lives are drawn on screen.No one is going to play this.The player drives itself with a simple heuristic: at each junction it heads toward the nearest pellet, unless a ghost is close, in which case it moves away from the ghost.

The game must run unattended and indefinitely, with no keyboard input at all.When it works, start the game and leave it playing.Holo4 27B (68 calls, 2.4M tokens, 268 lines) Qwen3.8 27B (197 calls, 11.

4M tokens, 327 lines) How we built Holo4 Agentic task factory Our internal set of agentic pipelines builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software.

So far it has produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.

Training Harness Alongside training, we rebuilt our harness, the loop that executes the model's actions and manages its context over hundreds of steps, using feedback from agentic performance on OSWorld 2.0.Agents tagged why each task failed and engineers reviewed their fixes.

The largest changes were giving the agent a reliable memory that can keep track of hundreds of steps, and a shell on the desktop machine itself.Opus 5 (70.2%) and GPT-5.6 Sol (66.2%) use max-effort partial rewards on the v2026.08.

08 offline set from OpenAI's launch chart, as in the cost-performance plot.Other reference scores come from model cards and the official leaderboard.Task releases, subsets and harnesses vary.

Holotron4 Nano Our post-training stack is designed to adapt to new foundation models and produce agents that generalize across interfaces and environments.As a member of the NVIDIA Nemotron Coalition, we applied our latest stack to the Nemotron 3 Nano Omni model as a follow-up to Holotron 3.

The same recipe turns Nemotron 3 Nano Omni into Holotron4 Nano, a generalist agentic model that significantly improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes.Gains are absolute percentage-point improvements over Nemotron 3 Nano Omni.

These gains show that our recipe transfers well and can turn a generalist model into an agentic expert.Nothing in it is size-specific.Run it yourself Both sizes are available today on the H Models API.

Weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF, next to our small model, Holotron4 Nano.We will release optimized DSpark drafter checkpoints in the coming days to further accelerate inference.

Models mentioned in this article 3 Datasets mentioned in this article 1 Collections mentioned in this article 1 More from this author NeoMME: an efficient Multimodal-native and Multilingual Encoder 111 September 3, 2026 Holo3.

1: Fast & Local Computer Use Agents 40 June 2, 2026 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 2 Models mentioned in this article 3 Datasets mentioned in this article 1 Collections mentioned in this article 1

Related

相關文章

IT之家AI Agent

英偉達發佈 AI 智能體安全平臺,可實時隔離異常智能體

作者:潞源 責編:潞源 評論: 感謝網友 華南吳彥祖、麻辣清補涼 的線索投遞!9 月 28 日消息,英偉達今天發佈開放式 AI 智能體安全平臺(NVIDIA Open Agent Safety Platform),可對 AI 智能體、智能體硬件、機器人系統等進行全棧治理和控制。

剛剛
鈦媒體AI Agent

爆火了的Muse和Instinct們不能做的,OS3能?

字母AI2026.09.28 15:37 · 來自北京全文4120字00:00 / 11:42Personal Agent的下一程,Rabbit OS3已經開跑。文 | 字母AI小扎拿出Muse Charm之後,Rabbit r1又被翻了出來。不少人看到這款新設備時,都會幻視兩年多前的橙色小方盒。在X上,有人將小扎拿著Charm、呂騁拿著r1的照片放在一起玩梗,不少外媒在報道Charm時,也都提到了這款初代AI硬件。一臺可以隨身攜帶的小設備,通過語音聽懂用戶的要求,再讓AI替人辦事。這幅畫面,確實讓人覺得熟悉。

剛剛
IT之家AI Agent

微軟 CEO 納德拉如何使用智能體:動手操控、追問、繼續推理,而非盲目信任

作者:清源 責編:清源 評論: 9 月 28 日消息,在最新一期 Sources Podcast 節目中,微軟 CEO 薩提亞 · 納德拉披露了 Cowork 和 Excel 智能體在自己日常工作中的用途。納德拉會用這些模型追蹤美國證券交易委員會(SEC)的文件並分析電子表格,其中尤其看重 Excel 智能體模擬因果推理的能力。

剛剛
AIbaseAI Agent

比爾·蓋茨:再熬約 20 年調整期,AI 將把人類帶進"富足時代"

當地時間 27 日,他接受 NBC《Meet the Press》採訪時,把 AI 發展分為兩個階段。第一階段是調整期,第二階段才是"富足時代"蓋茨表示,人類目前開始進入第一階段:AI 雖具備解決生活成本高企、醫療費用昂貴等問題的潛力,但還無力真正解決這些難題。

39 分鐘前