Trending Model:#1Unlimited-OCRbaidu⬇2414kTrending Model:#2Inklingthinkingmachines⬇25kTrending Model:#3Laguna-S-2.1poolside⬇13kTrending Model:#4Ternary-Bonsai-27B-ggufprism-ml⬇576kTrending Model:#5Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇335kTrending Model:#6Solar-Open2-250Bupstage⬇0kTrending Model:#7Nanbeige4.2-3BNanbeige⬇5kTrending Model:#8GLM-5.2zai-org⬇596kTrending Model:#9Bonsai-27B-ggufprism-ml⬇1910kTrending Model:#10krea2-identity-editconradlocke⬇0kTrending Model:#1Unlimited-OCRbaidu⬇2414kTrending Model:#2Inklingthinkingmachines⬇25kTrending Model:#3Laguna-S-2.1poolside⬇13kTrending Model:#4Ternary-Bonsai-27B-ggufprism-ml⬇576kTrending Model:#5Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇335kTrending Model:#6Solar-Open2-250Bupstage⬇0kTrending Model:#7Nanbeige4.2-3BNanbeige⬇5kTrending Model:#8GLM-5.2zai-org⬇596kTrending Model:#9Bonsai-27B-ggufprism-ml⬇1910kTrending Model:#10krea2-identity-editconradlocke⬇0k

News

May 25, 2026
ThetaCursed's Anima-TrainFlow Corrals LoRA Training Into One Page

Anima-TrainFlow is a simple, single-page desktop tool for training LoRA adapters on the Anima 2B image generation model. It puts every setting you need right in front of you, skipping […]

Read More
May 24, 2026
Antirez Shrinks DeepSeek V4 Locally With Deepseek-V4-GGUF

A new quantized file for DeepSeek V4 Flash, called Deepseek-V4-GGUF, shrinks the massive AI model so it can run on high-end consumer hardware. It’s a set of GGUF format files […]

Read More
May 24, 2026
Emo’s Topic-Specialized Experts Cut Memory by 75% With 1% Loss

Emo is a new mixture-of-experts language model designed so groups of experts naturally specialize in specific topics during training, rather than requiring human labeling. The main release from the Allen […]

Read More
May 23, 2026
Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved Fewer Refusals

Llmfan46 has released Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved, a modified version of Qwen3.6-35B-A3B that cuts unwanted refusals by 88% while keeping all 19 multi-token prediction (MTP) layers fully intact. The model uses an abliteration […]

Read More
May 23, 2026
Anbeeld Supercharges Local AI With Beellama.cpp Speed Overhaul

Beellama.cpp is a fork of the popular llama.cpp project that squeezes extra speed and memory efficiency out of local GGUF model inference. It adds DFlash speculative decoding, TurboQuant KV‑cache compression, […]

Read More
May 23, 2026
ds4.pinokio Slots a Full DeepSeek V4 Brain Into Apple Silicon Macs With One Click

ds4.pinokio is a new launcher and browser interface that brings the massive DeepSeek V4 Flash AI model to Apple Silicon Macs. It builds on the ds4.c Metal-only inference engine created […]

Read More
May 23, 2026
NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-BF16 Unfolds Three Models

The NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-BF16 release packs three distinct reasoning model sizes — 30 billion, 23 billion, and 12 billion parameters — into a single checkpoint file. Rather than requiring separate training runs, […]

Read More
May 23, 2026
Tokenspeed Streams Fake Tokens To Let You Feel LLM Speed

Coming across tokens-per-second benchmarks is easy, but truly understanding what "47 tok/s" feels like while you work is much harder. A new open-source tool called Tokenspeed solves this problem by […]

Read More
May 19, 2026
ExLlamaV3 Supercharges Home AI with Triple-Speed DFlash Decoding

ExLlamaV3 is an inference library that lets you run large language models on consumer graphics cards. It introduces the EXL3 quantization format, which compresses models to very low bitrates while […]

Read More
1 31 32 33 34 35 81