Skip to main content
Holo4 27B AI model performance chart, showing 85.2% on OSWorld benchmarks at $0.08 per task.

Editorial illustration for H Company's Holo4 27B Hits 85.2% on OSWorld Benchmarks at USD 0.08 Per Task

Holo4 27B Hits 85.2% on OSWorld at Just 8 Cents

H Company's Holo4 27B Hits 85.2% on OSWorld Benchmarks at USD 0.08 Per Task

• 4 min read

H Company put a price tag on its new computer-use models this week, and it's a small one. Holo4 27B clears 85.2% on the OSWorld benchmark at 8 cents per task, according to the company's own published numbers, edging out its own base model, Qwen3.8-27B, which needs 22 cents to hit 84.3%.

The release comes in two sizes. Holo4 27B is a dense model built on Qwen3.8-27B. Holo4 35B-A3B is a Mixture of Experts setup with 3B active parameters, built on Qwen3.6-35B-A3B. Both handle a 256K context window through the H Models API and pair with H's open hai-agents harness, which feeds the model screenshots and tool outputs, then carries out whatever clicks, typing, code or tool calls come back.

Licensing splits the two apart. The 35B-A3B weights are Apache 2.0, so anyone can self-host them commercially. The 27B weights are CC BY-NC 4.0, which routes commercial use back through H's own API.

The pitch is coverage: one model working across desktop, web, Android, code sandboxes and business software, instead of separate agents for each.

GUI-only agents fail without a screen. Tool-calling agents stall when an application has no API. Holo4 runs on desktop, web, Android, code sandboxes and business APIs.

Why this matters

The licensing split here matters as much as the benchmark numbers. H Company put Apache 2.0 on the 35B-A3B Mixture of Experts model, meaning startups can self-host it commercially without touching H's API. The 27B dense model, the one posting 85.2% on OSWorld at eight cents a task, is CC BY-NC 4.0, so any commercial use routes back through H's Model API and its pricing. That's a deliberate wedge: give away the harder-to-serve MoE model, keep the cheap, high-performing dense model on a leash.

For developers building agents that click through desktop UIs, call APIs, and write code in one pass, the cost-per-task figures are the real story. Beating a Qwen3.8 27B base by a point on OSWorld while running at roughly a third the cost, and undercutting Claude Opus 5.5 by two orders of magnitude on price, changes the math for anyone running agents at volume. Worth watching whether OSWorld 2.0's tougher 61.7% score holds up in production, and whether H's licensing terms shift once competitors match the price.

Common Questions Answered

What is Holo4 27B's performance on the OSWorld benchmark and how does it compare to Qwen3.8-27B?

Holo4 27B achieves 85.2% accuracy on the OSWorld benchmark at a cost of USD 0.08 per task, outperforming H Company's base model Qwen3.8-27B, which reaches 84.3% accuracy but requires 22 cents per task. This represents a significant improvement in both performance and cost efficiency for computer-use models.

What are the key differences between the two Holo4 model releases?

Holo4 comes in two sizes: Holo4 27B is a dense model built on Qwen3.8-27B, while Holo4 35B-A3B is a Mixture of Experts setup with 3B active parameters built on Qwen3.6-35B-A3B. Both models support a 256K context window and can operate across multiple platforms including desktop, web, and Android.

How does Holo4 overcome the limitations of GUI-only and tool-calling agents?

Unlike GUI-only agents that fail without a screen and tool-calling agents that stall when applications lack APIs, Holo4 runs across multiple environments including desktop, web, Android, code sandboxes, and business APIs. This versatility allows Holo4 to function effectively in diverse computing contexts where traditional agent approaches would struggle.

What are the licensing differences between the two Holo4 models and what do they mean for commercial use?

The Holo4 35B-A3B Mixture of Experts model is licensed under Apache 2.0, allowing startups to self-host it commercially without using H Company's API. In contrast, the Holo4 27B dense model uses CC BY-NC 4.0 licensing, meaning any commercial use must route through H Company's Model API at the published pricing of 8 cents per task.

What context window size do both Holo4 models support?

Both the Holo4 27B dense model and the Holo4 35B-A3B Mixture of Experts model support a 256K context window, enabling them to process and understand longer sequences of information in their computer-use tasks.

LIVE09:59H Company's Holo4 27B Hits 85.2% on OSWorld Benchmarks at USD 0.08 Per Task