2 min read

Holo4: Open Models for Computer Agents

Holo4компьютерные агентыоткрытые модели

H Company released Holo4, an open-weight family of computer-use models for web, desktop, and mobile interfaces. It includes a dense 27B model and a sparse 35B-A3B variant. The release matters because quantized weights enable local deployment, independent evaluation, and testing on real-world interfaces.

What H Company actually released

What stands out here is not just the model size, but the practical focus of the release: Holo4 is designed to operate websites, desktops, and mobile interfaces. In its announcement dated September 28, 2026, H Company presented the family as general-purpose computer-use models with open weights.

The lineup contains two distinct architectures. Holo4-27B is a dense vision-language model based on Alibaba Qwen3.8-27B, while Holo4-35B-A3B uses a sparse mixture-of-experts architecture based on Alibaba Qwen3.6-35B-A3B, with roughly 3B parameters active at a time. The model card lists a 262,144-token context window for the 27B version.

The weights are available in BF16, FP8, NVFP4, and four-bit GGUF. That is not merely a cosmetic selection of formats: H Company positions the models for local deployment, including single-GPU setups, although actual hardware requirements depend on the chosen weight format.

According to H Company’s official materials, the dense model scored 61.7% on OSWorld 2.0, while the MoE version reached 30.9%. For OSWorld, the reported scores are 85.2% and 80.8%, with stated per-task costs of $0.08 and $0.05 respectively. These figures should not be combined: benchmark versions and evaluation protocols differ.

Why open weights genuinely change the equation

The central shift is straightforward: computer agents can now be studied and deployed without mandatory dependence on a closed API. Open and quantized weights make it possible to reproduce evaluations, modify the execution stack, and test model behavior on proprietary interfaces.

But the results also reveal an uncomfortable engineering detail: fewer active parameters do not automatically mean comparable action quality. Holo4-35B-A3B may be more attractive in terms of runtime cost, yet it trails the dense model in the reported tests, especially on OSWorld 2.0. Lower compute costs only matter when agent mistakes do not erase those savings.

I would prioritize more than an appealing leaderboard row: evaluate the stability of long action sequences, visual-loop latency, and recovery after a mistaken click. Open weights make that validation possible, but they do not guarantee a good outcome. The key question is whether Holo4-27B’s advantage in official tables transfers to live, changing interfaces.

We previously covered Anthropic’s reversal of silent Claude query downgrades and what it revealed about model transparency. That debate provides useful context for assessing the model insights shared by Humanistic AI’s researcher.