Audio
Voice cloning, music generation, speech-to-text.
60 tools
AI audio has split cleanly into three lanes: speech synthesis (TTS + voice cloning), music generation, and speech-to-text — each with a clear leader.
Covers voice cloning and TTS (ElevenLabs, Resemble.ai, Murf), AI music generation (Suno, Udio), and speech-to-text (AssemblyAI, Whisper).
Pick ElevenLabs for voice quality. Pick Suno or Udio for AI music. Pick AssemblyAI when you need diarisation and timestamps; pick Whisper when you can self-host and want zero cost.
Otter.ai
AI meeting notetaker that transcribes calls, summarizes them, and pulls out action items in real time.
Soundful
Template-driven AI music generator that spits out royalty-free, commercially licensable tracks in seconds.
Scribbl
Bot-free AI meeting recorder, transcriber, and summarizer for Google Meet.
Bland AI
Enterprise voice AI for automated phone calls at scale
Fish Audio
Expressive, emotion-controllable text-to-speech and voice cloning with an open-model heritage
Horch
Privacy-first, on-device meeting assistant for macOS
Kyutai Moshi
Open-source, full-duplex speech-to-speech foundation model with sub-200ms latency
Retell AI
Build, test and deploy production-grade AI voice agents for inbound and outbound calls.
so-vits-svc
SoftVC VITS Singing Voice Conversion — open-source pipeline for training and running singing-voice models.
Threadfork
Private, local-first AI meeting notetaker for macOS
Vapi
Developer platform for building, deploying, and scaling production voice AI agents
VoxAI
Local macOS conversation recorder with live AI copilot and speaker-labelled transcripts