VTS Turns Your Hummed Imitation Into a Real Sound Effect
VTS (Voice To Sound) is a newly released open-source model that turns a short vocal imitation and a text description into a realistic sound effect. Instead of fumbling to describe…
Read more →VTS (Voice To Sound) is a newly released open-source model that turns a short vocal imitation and a text description into a realistic sound effect. Instead of fumbling to describe…
Read more →AudioX-Turbo is a new open source framework that generates audio and music from text, video, and existing audio signals. This release processes your inputs in just four steps to create […]
The newly released Dasheng-Audiogen is an open source artificial intelligence model that creates full audio scenes from text descriptions. Instead of producing just one type of sound, it can blend […]
The new release called MiMo-Audio-7B-Base is an open-source audio language model designed to learn new tasks from just a few examples. It processes over one hundred million hours of audio […]
Inflect-Nano-v1 is a tiny English text-to-speech model that turns written words into spoken audio. It includes its own audio generator and uses less than five million parameters to function. The […]
ZONOS2 is a new text-to-speech model designed to generate highly expressive and natural sounding audio. It predicts high quality audio tokens to create studio-grade sound at a 44.1 kHz sample […]
Google has released Magenta-Realtime-2, an open music generation model designed to create music on your own device with extremely low delay. This new system lets you steer musical output in […]
Dots.tts is a new 2-billion-parameter text-to-speech model that converts text directly into high-fidelity 48 kHz audio without relying on discrete audio codec tokens. The system operates fully end-to-end, using an […]
MisoTTS is a new 8 billion parameter text-to-speech model now available on Hugging Face. It converts written text into natural, conversational speech while maintaining voice consistency from short audio samples. […]
Boson AI has released higgs-audio-v3-tts-4b, a 4-billion-parameter text-to-speech model designed specifically for conversational voice AI. Rather than simply reading text aloud, the model produces expressive speech with emotional tone, natural […]