Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#8Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10OvisOCR2ATH-MaaS⬇17kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#8Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10OvisOCR2ATH-MaaS⬇17k

NAVA From ERNIE Research Syncs Video And Audio In One Pass

Merging film reel and speaker cone constructed from translucent glass and delicate neon orange lattice structure.

NAVA is a new open-source model that creates synchronized video and audio from a single text prompt. It generates 720p clips with stereo sound, supports multiple speaking voices with distinct timbres, and can even continue a scene from a still image. Despite packing only 6.3 billion parameters, it reaches top scores on benchmarks for both visual quality and audio-visual alignment.

The model comes from Baidu’s ERNIE research team and is released under an Apache 2.0 license. It was designed to overcome the lagging sync you get when large models bolt audio and video together after generation, or the muddled output when all modalities are mixed from the start. NAVA first carves out a dedicated alignment space for the sound and picture, then brings in the text prompt as a guide, which keeps everything tight and coherent while staying remarkably lightweight.

720p generation and multi-speaker voice control

Key Features
  • Synchronized 720p video in about one minute.
  • Stereo audio generated jointly with visuals.
  • Per-speaker timbre control via reference audio.
  • Text-driven camera motion and shot direction.
  • Landscape, portrait, and square aspect ratios.
  • Memory offloading for lower‑VRAM GPU setups.

This tool fits media makers who need to produce talking‑head scenes, cinematic clips, or audio‑visual storyboards on local hardware. It shines for small studios and privacy‑conscious professionals because everything runs in‑house without cloud dependencies. And because the entire system is contained in a single 6.3B‑parameter model, it stays within reach of a typical multi‑GPU workstation.

Developer notes and hardware requirements

The full inference pipeline is released end‑to‑end, including training scripts for fine‑tuning and a Gradio web demo. NAVA was trained primarily on dense Chinese captions, so users with short or English prompts are strongly encouraged to use the built‑in prompt rewriter for best results. The team also clearly flags ethical concerns: using the model to mimic real people’s faces or voices without consent is prohibited by the license and may be illegal.

“NAVA first establishes audio-video correspondence in a dedicated alignment space and then applies context as external conditioning to guide the aligned representation.” — Source: GitHub