Qwopus3.6-27B-Coder-MTP-GGUF is a new quantized coding model designed for fast, local agentic software development. It represents a specialized version of the Qwopus3.6-27B-Coder, packaged as a GGUF file for efficient single-GPU […]
News
Gemma-4-12B-OBLITERATED is the first language model to completely remove built-in safety refusals with absolutely zero loss in benchmark performance. This modified version of Google’s Gemma 4 12B scores identically to […]
North-Mini-Code-1.0 is a new open-source AI model from CohereLabs built specifically for code generation, autonomous software engineering, and terminal tasks. It packs a 30-billion-parameter architecture but only activates 3 billion […]
MiniMax-M3 is a new native multimodal AI model from MiniMaxAI that processes text, images, and video with a 1 million token context window. The model contains about 428 billion total […]
Moonshot AI has released Kimi-K2.7-Code, an open-source coding agentic model that significantly upgrades long-horizon software engineering performance. It builds directly on the Kimi K2.6 architecture while cutting thinking-token usage by […]
Google has released a new open-weights AI model called Diffusiongemma-26B-A4B-it, which uses a unique method to generate text significantly faster than traditional models. Instead of creating text one token at […]
Anvil is a new open-source model runner that wraps llama.cpp, providing transparent controls and fleet management for local AI on private hardware. It stores models as plain GGUF files you […]
MindLab Research has released Macaron-V1-Preview-749B, a 749-billion-parameter AI model built to serve as a personal agent that can use tools and generate dynamic user interfaces. The system combines a massive […]
Vllm-doctor is a new command-line tool that diagnoses performance bottlenecks in vLLM inference servers by reading live metrics. It turns raw data into clear explanations of what is wrong, why […]