The Nemotron-3.5-ASR-Streaming-0.6b model is NVIDIA’s latest open speech recognition release, designed to transcribe audio in real time across 40 language-locales from a single model. It can handle both low-latency streaming […]
Audio
About audio model releases
Latest audio models
MOSS-SoundEffect-v2.0 is an open model that creates high-fidelity sound effects straight from text prompts. It can generate rain, city noise, animal calls, human actions, and even short musical clips with […]
VTS (Voice To Sound) is a newly released open-source model that turns a short vocal imitation and a text description into a realistic sound effect. Instead of fumbling to describe […]
MOSS-TTS-v1.5 is an upgraded open-source text-to-speech model from the OpenMOSS team, building on their earlier 1.0 release. It keeps zero-shot voice cloning, long-form generation, and multilingual capabilities while delivering more […]
DramaBox is a text-to-speech system that turns scene descriptions and dialogue into expressive speech, complete with laughs, sighs, and pauses. It can clone a speaker’s timbre from just a 10-second […]
Scenema-Audio is a new open-source model that clones voices and generates speech with emotional acting, scene sounds, and zero-shot identity transfer. It doesn’t just read text aloud—it interprets stage directions […]
Supertonic-3 is a lightweight text-to-speech system that runs entirely on your device using ONNX Runtime, with no cloud calls needed for synthesis. This open-weight release expands language support from 5 […]
ControlFoley transforms video clips into synchronized soundtracks by combining visual scenes, written descriptions, and existing audio samples into a single generation system. This new framework produces matching sound effects and […]
Trelis recently released a specialized speech transcription model that handles overlapping conversations between two participants. The system processes audio clips locally without relying on external cloud servers. Built as an […]