Google just released Gemma-4-12B-It, an open-weights instruction-tuned model that handles text, images, video, and audio natively in one compact 12 billion parameter package. Instead of bolting on separate vision and […]
News
Google has released Gemma-4-12B, a 12-billion-parameter open model that handles text, images, and audio in a single decoder-only system. The unified design ditches separate encoders, so all data goes straight […]
Tapping into the mathematical beauty of fractals, Openmandel is a new tool that lets AI models generate and explore detailed Mandelbrot images on command. It functions as an MCP server, […]
VibeETL is a self-hosted, visual ETL platform that lets you build data processing pipelines by dragging and dropping nodes onto an interactive canvas. Instead of writing code to move and […]
The Realtime-Multilingual-Asr-Router is a lightweight coordinator that routes live audio between small, specialized speech recognition models instead of relying on a single large multilingual system. It detects speech boundaries with […]
The Comfyui-preset-gallery is a new visual extension for ComfyUI that saves, organizes, and reuses your go‑to prompt snippets and style templates. It adds a dedicated Preset Gallery node, letting you […]
img2imgVideoTransparency is a new ComfyUI workflow that creates short video clips with transparent backgrounds from just two images. It uses Wan 2.2 image-to-video models, LightX2V LoRAs for fast generation, and […]
Eisbach-Medium is a new LoRA adapter for the Stable Audio 3 Medium model that generates long-form instrumental music with clear narrative development. The adapter weighs only 41MB and uses a […]
TripoSplat is a new open-source model that turns a single image into a 3D Gaussian splat, with a variable point count up to 262,144. It runs locally on a consumer […]