SenseNova-U1-A3B-MoT is a new open-source vision-language model that handles image understanding, generation, and editing through a unified architecture without relying on separate visual encoders. This release belongs to the SenseNova […]
News
The Opendesk framework gives any AI agent direct control over a desktop computer—screenshots, mouse, keyboard, and app interaction—just like a real person. It works across macOS, Linux, and Windows, turning […]
The studiomi300 pipeline turns a single text prompt into a complete 30-second cinematic reel, complete with consistent characters, music, and voice-over. It strings together multiple large AI models—a director, image […]
MusiCue is an open-source tool that converts songs into typed, timeline-based cues for driving animation and show-control software. Developer cedarconnor built it to break audio into separated stems, beats, drum […]
The open-source project comfyui-node-canvas introduces a local GUI app called ComfyUI Node Builder that helps people create custom nodes for ComfyUI without manually writing all the repetitive code. It provides […]
OmniNFT is a set of LoRA adapters that fine-tune the open-source LTX Video model to produce better-aligned audio and video. Using reinforcement learning, it guides the generation process so that […]
DramaBox is a text-to-speech system that turns scene descriptions and dialogue into expressive speech, complete with laughs, sighs, and pauses. It can clone a speaker’s timbre from just a 10-second […]
ComfyUI-PlagueKind-Nodes is a custom node for ComfyUI that unifies image and mask resizing in a single step. It offers multiple scaling modes, preserves aspect ratios, and ensures masks stay perfectly […]
ComfyUI-DramaBox is a new custom node pack that brings ResembleAI’s expressive text-to-speech system directly into ComfyUI workflows. It turns text prompts into spoken audio using the LTX-2.3 audio diffusion model, […]