Nari Labs DiaDia is a 1.6B parameter open-source text-to-speech model that generates ultra-realistic dialogue with emotion control, audio conditioning, and nonverbal sounds in one pass.(0)0
S SadTalker AISadTalker generates realistic talking head videos from a single portrait and audio clip using 3D motion coefficients. Open-source CVPR 2023 research tool.(0)0
S StyleTTS 2StyleTTS 2 achieves human-level TTS synthesis using style diffusion and adversarial training with large speech language models. Open-source and free to use.(0)0
MusicGen MetaAudioCraft is Meta AI's open-source framework for generative audio, including MusicGen (text-to-music), AudioGen (text-to-sound), and EnCodec (neural audio codec).(0)0
Tortoise TTSTortoise TTS is a free, open-source multi-voice TTS system emphasizing realistic prosody and intonation. Clone voices and generate high-quality speech with this research-grade Python library.(0)0
StemRollerStemRoller is a free, open-source AI tool that splits any song into vocals, drums, bass, and instrumental stems using Facebook's Demucs model. Create karaoke and acapella tracks instantly.(0)0
Revocalize AI MasteringMaster your music tracks instantly and for free with Revocalize AI Mastering. Genre-aware AI processing delivers professional, release-ready sound in seconds.(0)0
Vocali.seSeparate vocals and music from any song in seconds with Vocali.se. Free AI-powered tool — no signup, no software, no cost required.(0)0
Coqui AICoqui AI was an open-source TTS and voice cloning platform featuring the XTTS model, supporting 17+ languages and voice cloning from just 3 seconds of audio.(0)0
EDGE DanceEDGE is an open-source AI model from Stanford that generates editable, physically plausible dance choreographies from any music. Features diffusion model, joint-wise constraints, and motion in-betweening.(0)0