Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10OvisOCR2ATH-MaaS⬇17kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10OvisOCR2ATH-MaaS⬇17k

NVIDIA Drops Cosmos3-Super-Text2Image for Pro Image Crafting

Close up object of a luminous cosmos orb with floating multilingual text.

Cosmos3-Super-Text2Image is a new open-source image generation model from NVIDIA that creates high-fidelity pictures directly from written prompts. It is a text-to-image variant of the broader Cosmos3 platform, which is built to help machines understand and simulate the physical world for robotics and autonomous vehicles. The model employs a 64-billion-parameter Mixture-of-Transformers architecture designed to produce visually detailed and text-aligned outputs.

NVIDIA along with Cosmos3-Super-Image2Video developed and released this model to give researchers and businesses a commercially usable tool for physical world simulation and visual content creation. The company has made the model available under a license that permits both non-commercial and commercial applications without extra payment. It is part of a larger family of models that also handle video, audio, and robotic action commands.

High-fidelity generation with local control

Key capabilities
  • 64-billion-parameter transformer model.
  • Uses prompt upsampling for better output quality.
  • Supports multiple popular aspect ratios.
  • Integrates with Hugging Face Diffusers library.
  • Served via OpenAI-compatible API with vLLM.
  • Released under permissive commercial license.
  • Runs on NVIDIA GPU hardware only.
  • Tested on H100 and GB200 systems.

Creators and small studios who need to generate images on their own hardware will find this model useful for keeping visual work private. The Diffusers integration means anyone already comfortable with community tools can load the model without learning entirely new software. Agencies and developers can use the permissive license to build the tool into paid client projects.

What developers should know

The model can struggle with out-of-distribution scenes, sometimes introducing unrealistic physics or causing objects to appear or disappear. NVIDIA notes that it lacks a real physics simulator, so outputs are approximations that should not be trusted as physically accurate for safety tasks. Longer or higher-resolution image requests may also expose imperfections, including imprecise spatial details.

"Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs." — Source: Hugging Face