Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#4Unlimited-OCRbaidu⬇2237kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#7krea2-identity-editconradlocke⬇0kTrending Model:#8Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#9Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#10OvisOCR2ATH-MaaS⬇17kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#4Unlimited-OCRbaidu⬇2237kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#7krea2-identity-editconradlocke⬇0kTrending Model:#8Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#9Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#10OvisOCR2ATH-MaaS⬇17k

Unsloth Polishes Gemma-4-26B-A4B-It-Qat-GGUF With Speed Boosts

A multifaceted wireframe sloth holidng up a gemstone with tiny holographic circuitry traces visible within.

Unsloth has released gemma-4-26B-A4B-it-qat-GGUF, a new quantized version of Google DeepMind’s Gemma 4 26B Mixture-of-Experts model. It uses Quantization-Aware Training to preserve near-original quality while shrinking the model’s memory footprint. The release also ships a Multi-Token Prediction drafter to speed up text generation through speculative decoding.

The team at Unsloth, who've also quantized Gemma-4-12B-It-GGUF, converted the official QAT checkpoints into GGUF format, making it compatible with llama.cpp and other popular inference engines. This allows developers to run the 26B-parameter model with only 3.8 billion active parameters, dramatically reducing hardware requirements. A built-in MTP drafter uses the target model’s cache to predict multiple tokens at once, verified on the fly for lossless acceleration.

What’s inside the unsloth release

Key Features
  • Multi-Token Prediction drafter for faster generation.
  • Near-lossless Q4_K_XL quantization format.
  • QAT preserves quality comparable to bfloat16.
  • Only 3.8B active parameters during inference.
  • 256K token context window for long documents.
  • GGUF format works with llama.cpp and vLLM.
  • Native function calling for agentic workflows.
  • Supports image understanding and reasoning.

This release is for developers who need to run a powerful reasoning model on consumer GPUs or workstations. The MoE architecture and quantization let you handle long-context tasks like coding, document analysis, and agentic workflows without expensive hardware. The included MTP drafter further improves throughput, making interactive use smoother.

Developer notes and limitations

Unsloth notes that the MTP drafter shares the target model’s KV cache, so it does not alter the output. All drafted tokens are verified by the main model, ensuring lossless acceleration. The repository also includes other precision variants and explicit usage instructions in the MTP folder.

“The drafter shares the target’s KV cache and does not change the output (the target verifies every drafted token).” — Source: Hugging Face