Hcompany Crafts Holo-3.1-35B-A3B for Private On-Device Screen Control

Holo-3.1-35B-A3B is the largest model in a new family of vision-language agents that can see, understand, and control computer interfaces across web browsers, desktops, and now mobile devices. It automates repetitive screen-based tasks, locates buttons and fields with high accuracy, and follows complex business workflows without needing cloud connectivity. The model comes with several ready-to-use quantized formats that shrink its size for local graphics cards while preserving strong performance.
H Company, a French AI lab, built this family by fine-tuning the Qwen 3.5 base models specifically for autonomous computer use. They focused on delivering capable agents that run on consumer hardware, not just expensive data center GPUs. The release gives privacy-minded professionals and small teams a practical way to keep sensitive automation tasks completely off the internet.
Multi-environment control and efficient local inference
- Operates web, desktop, and mobile interfaces natively.
- Native function-calling for easy agent framework integration.
- Quantized checkpoints for FP8, NVFP4, and Q4 GGUF.
- Scores well on UI grounding and business workflow benchmarks.
- Scales from lightweight 0.8B to 35B-A3B models.
- Licensed under permissive Apache 2.0 terms.
Prosumer GPU owners and agencies benefit because the entire agent runs on-premises, so no screenshots or login details ever leave the machine. The 35B-A3B variant with Q4 GGUF quantization squeezes into about 24 GB of VRAM, which fits high-end consumer cards like an RTX 4090. Privacy-conscious professionals gain a dependable computer-use assistant that avoids monthly API bills and data-handling risks.
Built on Qwen 3.5 with a mobile-first expansion
The models inherit the Qwen 3.5 architecture and are fine-tuned to handle mobile interfaces for the first time in this series, alongside continued desktop and browser support. Quantized versions unlock local deployment, though the largest model still demands a powerful GPU — expect a 24 GB card as the practical minimum for the Q4 GGUF edition. The team highlights strong cost-efficiency across the lineup, with the smaller models requiring far less compute while still handling everyday automation tasks well.
“Holo3.1 is our latest family of Vision-Language Models (VLMs) for computer use agents.” — Source: Hugging Face