Trending Model:#1Inklingthinkingmachines⬇13kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇339kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1263kTrending Model:#4Unlimited-OCRbaidu⬇2123kTrending Model:#5GLM-5.2zai-org⬇532kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2117kTrending Model:#7krea2-identity-editconradlocke⬇0kTrending Model:#8Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇2007kTrending Model:#9Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇17kTrending Model:#10OvisOCR2ATH-MaaS⬇15kTrending Model:#1Inklingthinkingmachines⬇13kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇339kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1263kTrending Model:#4Unlimited-OCRbaidu⬇2123kTrending Model:#5GLM-5.2zai-org⬇532kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2117kTrending Model:#7krea2-identity-editconradlocke⬇0kTrending Model:#8Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇2007kTrending Model:#9Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇17kTrending Model:#10OvisOCR2ATH-MaaS⬇15k

Multimodal

About multimodal releases

Discover new open‑source multimodal models. This archive covers models that can handle multiple functions such ah text, images, audio, and more.

Latest multimodal models

June 15, 2026
Google Drops Gemma-4-12B-It A Senses First Model That Runs Offline

Google just released Gemma-4-12B-It, an open-weights instruction-tuned model that handles text, images, video, and audio natively in one compact 12 billion parameter package. Instead of bolting on separate vision and […]

Read More
June 10, 2026
Google Drops Gemma-4-12B: One Model, Three Formats, Zero Encoders

Google has released Gemma-4-12B, a 12-billion-parameter open model that handles text, images, and audio in a single decoder-only system. The unified design ditches separate encoders, so all data goes straight […]

Read More
June 8, 2026
ByteDance Bernini Crafts Videos With Words, Not Pixel Paintbrushes

ByteDance has released Bernini, an open-source framework that unifies video generation and editing through a semantic planning approach. Instead of controlling pixels directly, the system uses a multimodal large language […]

Read More
June 6, 2026
Cosmos3-Nano Conjures Video, Audio, and Robot Commands from Any Input

Nvidia has released Cosmos3-Nano, a 16-billion-parameter omnimodal model that turns text, images, video, audio, or action data into dynamic video with synced sound, reasoning text, or robot movement commands. The […]

Read More
June 2, 2026
Gemma-4-Harmonia-31B-uncensored-heretic Slashes Refusals by 91%

Gemma-4-Harmonia-31B-uncensored-heretic is a decensored version of a 31-billion-parameter language model that dramatically cuts response refusals by 91%. The release uses an ablation technique to strip away content restrictions while keeping […]

Read More
June 2, 2026
PaddleOCR-VL-1.6 Smashes Document Parsing Accuracy At 96.33%

PaddleOCR-VL-1.6 is a compact document parsing model that reaches a new state-of-the-art accuracy of 96.33% on the OmniDocBench v1.6 benchmark. It boosts recognition of text, formulas, tables, ancient documents, rare […]

Read More
June 1, 2026
QwenLM Drops Qwen-Image-Bench to Grade AI Art Like a Pro

Qwen-Image-Bench is a new evaluation toolkit that scores images from any text-to-image model using a fine-tuned judge AI called Q-Judger. It checks generated images across five major quality dimensions, including […]

Read More
June 1, 2026
Qwen3.6-27B-pure-GGUF Squeezes Full 27B Model Onto One 16GB GPU

A new quantized version of Alibaba's coding model has been released to the community, offering a 27B parameter AI that can run entirely on a single 16GB graphics card. The […]

Read More
May 31, 2026
StepFun Delivers Step-3.7-Flash MoE Vision Model for Local AI Agents

Step-3.7-Flash is a 198-billion-parameter vision‑language model that uses a sparse mixture‑of‑experts design to activate only about 11 billion parameters per token. It handles images and text natively through a 1.8‑billion‑parameter […]

Read More