Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10Laguna-S-2.1poolside⬇3kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10Laguna-S-2.1poolside⬇3k

Advanced-GGUF-Quantizer Slashes LLM Size While Preserving Quality

An advanced GGUF quantizer core is a transparent crystalline cube with internal geometric data layers.

The Advanced-GGUF-Quantizer toolkit is a new CUDA-powered utility for building highly optimized GGUF models, with special attention to NVIDIA’s NVFP4 and MXFP6 data types. It uses a refined scale fitting process and per-tensor error checks to shrink file sizes while keeping key quality metrics closer to the original model. The result is faster, smaller GGUF files that lose noticeably less accuracy than earlier conversion methods.

Michaelw9999 developed this project starting from NVFP4 kernel experiments in llama.cpp, then expanded it into a full quantization suite. The tool was created to fill a gap in open-source tooling for Blackwell’s native floating-point formats and to improve how GGUF models handle mixed precision. Testing shows it can produce NVFP4 models that perform better than some commercial alternatives while remaining entirely local and portable.

Refined fitting and smart tensor promotion

Key capabilities
  • RSF scale fitting improves NVFP4 and K-quants.
  • Promotes sensitive tensors to higher bit depths.
  • Supports mixed NVFP4/MXFP6 quantizations.
  • Evaluates with imatrix and saved-logit KLD data.
  • CUDA in-memory patching avoids extra disk writes.
  • Detailed reports, manifests, and tensor logs.
  • Fast, normal, and deep search modes available.

Owners of pro-level consumer GPUs like an RTX 5090 can use the tool to generate model files that run noticeably faster while staying faithful to the original. Small studios and local AI practitioners benefit from creating custom quantizations that match their hardware safety and quality targets. Privacy-minded professionals gain a way to run advanced models entirely on their own machine without any external dependencies.

Developer notes and early limitations

The author cautions that this is still an early-stage, rapidly changing codebase with many exposed options that can feel confusing. Deep search mode, while delivering excellent results, can take over 17 hours on a single 5090 for a 27‑billion‑parameter model and multi‑GPU testing has not happened yet. Future work aims to simplify the interface, reduce duplicated logic, and refine how the tool decides which error metric matters most.

“I started writing NVFP4 kernels for llama.cpp last year and needed the ability to quantize NVFP4 GGUFs, so this project started as an NVFP4 quantizer.” — Source: Reddit