MOSS-TTS-Nano-100M is a lightweight, open-source text-to-speech engine that generates natural audio directly on standard computers. The system converts typed prompts into clear speech while maintaining strict efficiency for daily use. […]
Audio
About audio model releases
Latest audio models
OmniVoice is an open source text-to-speech system that converts written words into spoken audio across more than six hundred languages. The software enables instant voice matching and allows users to […]
VoxCPM2 is an open text-to-speech system that generates studio-quality audio from written text. The tool reads words at standard clarity and outputs polished speech at a higher frequency without needing […]
ACE-Step recently published ACE-Step 1.5 XL, an open audio generation model that produces complete music tracks in just eight steps. This streamlined process significantly reduces rendering wait times while preserving […]
Foundation-1 is a text-to-sample model built for structured music production. It generates tempo-synced, key-aware loops that slot directly into production workflows instead of producing generic audio textures. RoyalCities developed this […]
LongCat-AudioDiT is a new text-to-speech model that generates high-fidelity audio directly from text inputs. It operates directly on the waveform latent space rather than relying on intermediate acoustic representations like […]
PrismAudio is a new framework that generates audio from video using reinforcement learning with Chain-of-Thought (CoT) planning. Developed by the FunAudioLLM team, it breaks down the complex task of video-to-audio […]
WAVe-1B-Multimodal-NL is a 1 billion parameter model that checks the quality of synthetic speech at the word level. It examines how well spoken audio matches its written transcript, catching errors […]
MOSS-TTS Family is an open-source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high-fidelity audio generation across complex real-world scenarios, including long-form […]