Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#4Unlimited-OCRbaidu⬇2237kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#7Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#8krea2-identity-editconradlocke⬇0kTrending Model:#9Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#10OvisOCR2ATH-MaaS⬇17kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#4Unlimited-OCRbaidu⬇2237kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#7Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#8krea2-identity-editconradlocke⬇0kTrending Model:#9Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#10OvisOCR2ATH-MaaS⬇17k

InclusionAI Deploys VISTA-9B To Map Text Commands To Screen Clicks

Glowing computer mouse cursor that consists of a frosted acrylic material.

VISTA-9B is a visual model that understands screen layouts and translates natural language instructions into precise click coordinates. It looks at a screenshot and figures out exactly where to click based on what you ask it to do. The system normalizes these coordinates on a zero to one thousand scale to ensure accurate targeting.

The team at inclusionAI who are also behind the Ling-2.6-flash model created this tool by training it on the Qwen3.5 9B foundation. They developed a new training approach called view-consistent self-verified training to help the model handle different screen sizes and layouts. This method exposes the system to semantically similar but geometrically varied screenshots during the learning process.

Model features and target users

Key Features
  • Maps natural instructions to click coordinates.
  • Trained on Qwen3.5 nine billion backbone.
  • Handles geometrically different cropped screenshot views.
  • Uses deterministic decoding for accurate results.

People who build automated desktop agents will find this release highly useful for interfacing with visual applications. It allows systems to interact with software interfaces that lack traditional code hooks. Anyone needing to automate user interface interactions can rely on this model for precise visual navigation.

Development details and training requirements

Developers recommend having at least eight high memory GPUs to handle the training process for this architecture. The model avoids blindly copying failed coordinate generations by only adding oracle coordinates when a successful prediction occurs. Testing shows it outperforms standard GRPO methods by about one point on the ScreenSpot Pro benchmark.

"VISTA-9B is a GUI-grounding model that maps a screenshot and a natural-language instruction to a click coordinate in the normalized 0-1000 image frame." - Source: Hugging Face