Omni-Video 2 is a unified video editing and generation framework that combines a text-to-video diffusion model with vision-language understanding. The system can generate videos from text descriptions and edit existing […]
News
Voice Clone Studio is a modular Gradio-based web application that handles voice cloning, voice design, multi-speaker conversations, voice conversion, and sound effects generation. The tool consolidates multiple AI audio engines […]
ComfyUI-ZImageTurboProgressiveLockedUpscale is a new custom node for ComfyUI that handles progressive image upscaling through multiple stages. The node takes a different approach than traditional methods by using sigma slicing and […]
ComfyUI-Yedp-Action-Director is a custom node that has a fully interactive 3D viewport directly into ComfyUI. Users can load 3D character animations from .FBX, .GLB, or .BVH files, preview them in […]
ComfyUI-meancache-z is a new custom node that accelerates inference for Z-Image Flow Matching models without requiring any model fine-tuning. The tool, similar to the Z-Image loras, achieves speedups between 1.4x […]
Ming-flash-omni-2.0 is a unified multimodal model from inclusionAI that processes images, text, audio, and video while generating both speech and images. Built on a Mixture-of-Experts (MoE) architecture with 6 billion […]
MRS-Core is a deterministic reasoning engine built for large language models and autonomous agents. It provides a modular foundation constructed from a small set of reusable operators that execute in […]
WorldVQA is a new benchmark designed to test how well AI models can identify and name visual objects from memory. Created by MoonshotAI, it measures factual visual knowledge rather than […]
ComfyUI-wan-i2v-control is a custom node pack for ComfyUI that brings precise region control to WAN image-to-video generation. The tool intercepts the conditioning process and applies masks to specific areas, letting […]