Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10Laguna-S-2.1poolside⬇3kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10Laguna-S-2.1poolside⬇3k

ByteDance Bernini Crafts Videos With Words, Not Pixel Paintbrushes

Impossibly smooth marble sculpture of a glowing translucent microchip circuits.

ByteDance has released Bernini, an open-source framework that unifies video generation and editing through a semantic planning approach. Instead of controlling pixels directly, the system uses a multimodal large language model to plan edits in a conceptual embedding space, then a diffusion transformer renders the final video. The inference code, model weights, and a Gradio demo are now available under an Apache 2.0 license.

The Bernini team separated training into two stages—a planner and a renderer—so each component retains its pretrained strengths while only needing light co-training. This method lets the MLLM’s language understanding, including chain-of-thought reasoning, guide complex edits like altering a subject’s motion or inserting new objects into a scene. The result is a flexible pipeline that handles text-to-video, image editing, and reference-image-guided video manipulation with high consistency.

Semantic planning meets diffusion rendering

Key Features
  • Unifies video generation and editing in one framework.
  • MLLM planner uses chain-of-thought reasoning for edits.
  • DiT-based renderer synthesizes photorealistic video.
  • Segment-Aware 3D RoPE handles multiple visual inputs.
  • Planner and renderer trained separately, then co-trained.
  • Supports text, image, and reference-image guided tasks.
  • Achieves first-tier results compared to closed-source tools.
  • Gradio UI for single-GPU or multi-GPU inference.

Researchers and developers who need a video model that understands nuanced editing instructions will benefit most from this release. The Apache 2.0 license also makes it accessible for commercial and academic projects. Hobbyists and small studios with powerful hardware can run the tool locally to produce quality edits without relying on cloud services.

What to know before running it

The current version is optimized for Hopper GPUs like the H100, using FlashAttention-3 for speed, though it will fall back to slower attention methods on other CUDA cards. Video tasks require multiple GPUs with Ulysses sequence parallelism, while single-image generation and editing can run on a single GPU. The team provides both a complete diffusers-format model and separate checkpoint files, along with detailed case files that keep long prompts out of the command line.

"On video editing, Bernini reaches the first tier among leading closed-source commercial models." — Source: Hugging Face