Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10OvisOCR2ATH-MaaS⬇17kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10OvisOCR2ATH-MaaS⬇17k

Bonsai-Image-Binary-4B-Gemlite-1bit Packs Full Image AI in 0.93GB

Bonsai tree consists of cascading binary digits and intricate circuit traces.

Prism ML has released Bonsai-Image-Binary-4B-Gemlite-1bit, a text-to-image model that packs a full diffusion transformer into just 0.93 GB by using binary weights. It takes the FLUX.2 Klein 4B architecture and replaces most matrix-heavy layer parameters with simple {-1, +1} values plus a small scale per group. The result cuts the original 7.75 GB transformer down by 8.3× while still producing 1024×1024 images in under five seconds on a consumer RTX 3080.

The team at Prism ML, who also quantized Bonsai-8B-gguf, compressed every matmul-heavy linear layer in the 25-block transformer, applying binary representations to Q/K/V projections, MLP weights, and output projections. They kept a handful of precision-sensitive tensors in FP16 to maintain image quality. The entire CUDA deployment payload, including a 4-bit text encoder and FP16 VAE, sits at 4.09 GB, with the text encoder automatically offloaded after the prompt is processed.

What makes the compact binary design stand out

Key features
  • Weights stored as {-1, +1} with FP16 scaling.
  • Transformer shrunk from 7.75 GB to 0.93 GB.
  • 4-step FlowMatch-Euler sampler needs no CFG.
  • Runs natively on Linux and Windows.
  • Consumer RTX 3080 generates images in 4.5 seconds.
  • Text encoder offloaded after prompt for lower memory.
  • Total CUDA payload only 4.09 GB on disk.
  • Apache 2.0 license with Gemlite low-bit GEMM kernels.

Creatives and developers who want local image generation without a cloud dependency will find this release useful. Privacy-conscious professionals can keep prompts and assets entirely on their own CUDA-equipped workstations. Small teams can iterate faster without waiting in queues or paying per image, while still getting modern diffusion-transformer quality on everyday GPUs.

Developer notes and trade-offs

The binary model is not an exact clone of the FP16 original; it’s a compact deployment where quality depends on the prompt and how much fine detail matters. Small text, exact object counts, and strict compositional constraints may need extra evaluation. While the transformer is drastically reduced, other components like the VAE can become the memory bottleneck, so the runtime uses tiled decoding and encoder offloading to keep peak usage around 6.4 GiB on the RTX 3080.

"0.93 GB diffusion transformer, down from 7.75 GB for the FP16 FLUX.2 Klein 4B transformer" — Source: Hugging Face