Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#4Unlimited-OCRbaidu⬇2237kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#7Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#8krea2-identity-editconradlocke⬇0kTrending Model:#9Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#10OvisOCR2ATH-MaaS⬇17kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#4Unlimited-OCRbaidu⬇2237kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#7Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#8krea2-identity-editconradlocke⬇0kTrending Model:#9Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#10OvisOCR2ATH-MaaS⬇17k

Google Drops Gemma-4-12B-It A Senses First Model That Runs Offline

A luminous gemstone representing the gemma-4-12B-it model.

Google just released Gemma-4-12B-It, an open-weights instruction-tuned model that handles text, images, video, and audio natively in one compact 12 billion parameter package. Instead of bolting on separate vision and audio encoders, it feeds raw pixels and sound waves straight into the transformer, cutting latency and making the entire model easier to fine-tune. This encoder-free design is optimized for consumer GPUs and workstations, letting you run sophisticated multimodality entirely offline.

Google DeepMind created the Gemma 4 family (along with assistant variants such as Gemma-4-31B-It-Assistant) to bring frontier-level AI to personal devices, and the 12B version focuses on removing the encoder bottleneck. The team trained the model on web documents, code, math, images, and audio—all with a January 2025 cutoff—so it can reason across more than 140 languages. By eliminating external preprocessors, they simplified local deployment while keeping strong reasoning, coding, and agentic abilities.

One model, many senses

Key features for local use
  • Processes images, video, and audio without separate encoders.
  • 256K token context window for long documents.
  • Built-in configurable thinking mode for step-by-step reasoning.
  • Native function calling to build autonomous agents.
  • Speech recognition and translation across multiple languages.
  • Runs comfortably on a single 24GB consumer GPU.
  • Handles variable image resolutions up to 1120 tokens.
  • Supports 35+ languages out of the box.

This model fits anyone who wants capable multimodality without sending data to the cloud—hobbyists with a single GPU, small agencies parsing invoices and transcripts, and privacy-sensitive professionals handling confidential documents. It scores highly on coding and reasoning benchmarks while staying lightweight enough to run in a home lab, so you can build offline assistants that truly understand your files. Because all modalities flow through one transformer, fine-tuning for a niche task becomes a single training pass instead of juggling multiple sub-models.

What the developers want you to know

Compared to the older Gemma 3 27B, this 12B model posts better scores on MMLU Pro, AIME, and LiveCodeBench despite being less than half the size. Google’s safety evaluations found minimal policy violations even with no safety filters active, though the model can still reflect biases from its web training data and shouldn’t be treated as a fact database. Developers should test outputs carefully and consider applying content safeguards for production use.

"Gemma 4 12B eliminates these encoders entirely, projecting raw image patches and audio waveforms directly into the LLM's embedding space through lightweight linear layers." — Source: Hugging Face