Skip to main content
📖 The AI Tool Bible

AI tools tagged Supports Audio

48 tools matching this tag.model

All tags →
EL

ElevenLabs

Featured
Audio · ElevenLabs Multilingual v2
9.4

The gold standard for AI voice cloning and TTS.

Freemium· Free: $0 · Starter: $6 · Creator: $11 · Pro: $99 · Scale: $299TTSvoice cloning
G4

GPT-4o

Featured
Writing · GPT-4o
9.4

OpenAI's multimodal flagship behind ChatGPT.

Freemium· Basic: $10 · Pro: $30 · Enterprise: Contact salesgeneral writingsummarization
AS

AssemblyAI

Audio · Universal / Slam-1
8.7

Speech-to-text API with diarisation, summarisation, and topic detection.

Freemium· Pre-recorded Speech-to-Text API: $0.21 /hr · Universal-2: $0.15 /hr · Realtime Speech-to-Text API: $0.45 /hr · Universal-Streaming: $0.15 /hr · Universal-Streaming Multilingual: $0.15 /hrtranscriptiondiarisation
WH

Whisper

Audio · Whisper large-v3
8.6

OpenAI's open-source speech-to-text — the de-facto baseline.

Free· Free open weights; $0.006/min via OpenAI APItranscriptionself-hosted
GO

Gong

Audio · In-house speech and language models, with additional agentic features reportedly built on frontier LLMs
8.5

Revenue AI platform that captures, transcribes, and analyzes customer conversations to drive sales outcomes.

Enterprise· Custom pricing based on per-user licenses plus a platform fee scaled to team size; no public tier pricing. Prospects request a quote via a demo form.Sales call recording and transcriptionDeal risk and pipeline inspection
RE

Replicate

Fine-tuning · Thousands of community + first-party models
8.5

One-API platform for running and fine-tuning open-source models.

Paid· Pay-per-second of GPU timemodel hostingfine-tuning
FA

Figma AI

Image Generation · Multi-model: routes to OpenAI, Anthropic Claude, Google Gemini, and GitHub-hosted models plus Figma fine-tuned models
8.3

AI workflows built into the design tool product teams already use

Freemium· Starter plan: Free · Professional plan: Contact sales · Organization plan: Contact sales · Enterprise plan: Contact salesAI-assisted UI generation from promptsDesign system-aware component search
GR

Grain

Audio · Claude, ChatGPT
8.3

AI meeting notes, coaching, and an MCP feed for your agents

Freemium· Free (45-min meeting cap, 90-day history) / Starter ~$15 per user/mo (unlimited meetings, advanced AI) / Business & Enterprise tiers with CRM sync and admin controlsAI meeting notes and summariesSales call transcription and coaching
HE

HeyGen

Video · HeyGen proprietary
8.3

Avatar video + lip-sync translation at scale.

Paid· Free: $0/mo · Creator: $29/mo · Pro: $49/mo · Business: $149/mo · Enterprise: Contact Saleslocalisationavatar video
MU

Mubert

Audio · Proprietary sample-based generative engine
8.3

AI music generator that spits out royalty-free background tracks for video, podcast, and app use.

Freemium· Ambassador: Free · Creator: $14 · Pro Most popular: $39 · Business: $199background-musicroyalty-free-soundtracks
AU

AudioCraft

Audio · MusicGen, AudioGen, EnCodec
8.2

Meta's open-source research toolkit for generating music and sound effects from text via a single autoregressive language model.

Free· Free and open source; self-hostedtext-to-musicsound-effects
FA

Fireflies.ai

Audio · Multi-model (proprietary ASR + LLM layer)
8.2

AI meeting assistant that joins calls, transcribes them, and turns the talk into searchable notes and action items.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesmeeting transcriptioncall summaries
GA

Google AI Studio

Coding · Gemini 2.5 Pro / Flash, Imagen, Veo
8.2

Browser-based playground and API console for prototyping with Google's Gemini models.

Freemium· Free tier with rate limits; paid via Gemini API usage-based pricingprompt-prototypinggemini-api-keys
GV

Google Veo

Video · Veo 3.1
8.2

Google DeepMind's flagship text-to-video model with native audio generation and cinematic camera control.

Paid· Metered via Gemini API; also bundled in Google AI and Workspace planstext-to-videoimage-to-video
HI

Higgsfield

Video · Multi-model (Sora 2, Veo 3.1, Kling 3.0, Seedance 2.0, Nano Banana Pro)
8.2

AI video and image generation suite that aggregates 30+ frontier models under one workflow.

Paid· Subscription tiers (pricing not clearly published on landing)text-to-videoimage-generation
LS

LTX Studio

Video · LTX-2, LTXV-13B, Veo 3, Flux.2
8.2

Storyboard-first AI video platform from Lightricks with shot-level camera and character control.

Freemium· Free: €0 · Lite: €15 · Standard: €35 · Pro: €125text-to-videostoryboarding
LU

Ludwig

Fine-tuning · Multi-model (PyTorch + HuggingFace Transformers)
8.2

Declarative, YAML-driven deep learning framework for fine-tuning LLMs and multi-modal models without writing training loops.

Free· Free, Apache 2.0 open sourcellm-fine-tuningmulti-modal-training
OP

OpenAI Playground

Writing · Multi-model (GPT-4o, GPT-4.1, o-series, DALL-E, Whisper, TTS)
8.2

OpenAI's official browser sandbox for prompting, tuning, and testing every model on the platform before you ship API code.

Paid· Basic: $10 · Pro: $20 · Enterprise: Contact salesprompt-engineeringmodel-comparison
PO

Poe

Writing · Multi-model (GPT, Claude, Gemini, DeepSeek, FLUX, Sora, others)
8.2

One subscription, every frontier model - Quora's multi-model chat hub with bots, image, video, and audio generation.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesmulti-model chatimage generation
SO

Soundraw

Audio · Proprietary in-house model
8.2

AI music generator that spits out royalty-free, customizable tracks by genre and mood.

Freemium· Creator: €5.83/mo · Artist Starter: €10.75/mo · Artist Pro: €12.42/mo · Artist Unlimited: €17.42/mo · Enterprise: Askbackground musicvideo soundtracks
UN

Unsloth

Fine-tuning · Llama, Mistral, Gemma, Qwen, GLM (multi-model)
8.2

Open-source LLM fine-tuning toolkit with custom kernels that train 2-30x faster and use up to 90% less VRAM.

Freemium· Free open-source; Pro and Enterprise contact saleslora-finetuningqlora
AI

AIVA

Audio · Proprietary (AIVA)
8.1

AI music composition tool that generates royalty-friendly tracks in 250+ styles with editable MIDI output.

Freemium· Free; Standard ~€11/mo, Pro ~€33/mo (billed yearly)music-generationsoundtrack-composition
KA

Kling AI

Video · Kling 3.0 (Omni One)
8.1

Kuaishou's flagship AI video generator, currently topping the ELO leaderboard for text-to-video and image-to-video.

Freemium· Free 66 daily credits; Standard $6.99/mo to Ultra $180/mo; API per-secondtext-to-videoimage-to-video
LO

LocalAI

Writing · Multi-model (llama.cpp, diffusers, whisper, etc.)
8.1

Self-hosted OpenAI-compatible API for running LLMs, image, and audio models on your own hardware.

Free· Free and open source (MIT)local-llm-inferenceopenai-api-replacement
NO

NotebookLM

RAG · Gemini 2.5
8.1

Google's source-grounded research notebook that turns your documents into chats, briefs, and AI-hosted podcasts.

Freemium· Free tier; Plus via Google One AI Premium ($19.99/mo) or Workspace add-ondocument Q&Aresearch synthesis
EI

Edge Impulse

Fine-tuning · Multi-model (TF Lite Micro, custom DSP blocks)
8.0

End-to-end platform for training and deploying ML models on microcontrollers, sensors, and other edge hardware.

Freemium· Developer: $0edge-aitinyml
FA

Fal.ai

Image Generation · Multi-model (Flux, Stable Diffusion, video/audio models)
8.0

Serverless GPU inference platform optimized for fast diffusion and generative media APIs.

Paid· Usage-based; serverless from ~$1.89/GPU-hour, per-output pricing on model APIstext-to-imagetext-to-video
SE

Sesame

Audio · Sesame CSM (1B / 3B / 8B)
8.0

Conversational voice AI aiming to cross the uncanny valley with context-aware, emotionally aware speech.

Free· Free research preview; consumer product pricing not announcedconversational-voicetext-to-speech
WE

WellSaid

Audio · Proprietary WellSaid TTS
8.0

Enterprise-grade AI text-to-speech built on licensed voice actor recordings.

Freemium· Free trial; paid plans for teams and enterprise (contact sales for API)text-to-speeche-learning narration
DE

Deepgram

Audio · Nova, Flux, Speak (proprietary)
7.9

Production-grade speech-to-text, text-to-speech, and voice-agent APIs for real-time and batch audio.

Freemium· Free credits on signup; usage-based pricing; enterprise contracts availablespeech-to-texttext-to-speech
KA

Kaiber

Video · Proprietary (undisclosed)
7.9

AI video generator built around music-reactive animation and image-to-video transforms.

Freemium· Free trial credits; paid plans typically start around $5-15/momusic-videotext-to-video
AR

Argil

Video · Proprietary avatar model; integrates VEO3 and Hailuo for AI Fictions
7.8

AI avatar video generator that clones your face and voice from a single photo and a minute of audio.

Paid· Classic: $39 · Pro: $149 · Scale: $499 · Enterprise: Custom pricingai-avatarstalking-head-video
HE

Hedra

Video · Character-3 (proprietary) + multi-model canvas
7.8

AI creative agent for character-driven video, image, and audio generation built around the Character-3 model.

Freemium· Basic: $15 · Creator: $30 · Professional: $75 · Teams: $75 · Enterprise: Customtalking-avatar-videocharacter-animation
LI

Limitless

Agents · Proprietary (multi-model)
7.8

AI wearable pendant that records and transcribes your conversations into a searchable personal memory.

Free· Free for existing customers post-Meta acquisition; no longer sold to new buyersmeeting-transcriptionpersonal-memory
MU

Murf

Audio · Murf Gen2
7.8

TTS aimed at corporate voiceover and e-learning.

Freemium· Free preview; from $19/mo Creator; $66/mo Businessvoiceovere-learning
VO

Voicebox

Audio · Multi-model (Chatterbox, Qwen TTS, Whisper, etc.)
7.4

Open-source desktop voice studio for local cloning, dictation, and giving MCP agents a voice.

Free· Free and open source; optional $VOICEBOX token donationsvoice-cloningtext-to-speech
AA

Azure AI Speech (Neural TTS)

Audio · Azure Neural TTS (plus HD and Azure OpenAI voices)
7.3

Microsoft's enterprise-grade neural text-to-speech with 100+ languages, custom brand voices, and SSML control.

Freemium· Free (F0): Free · Pay as You Go: Voice Live Prices: $- · Commitment Tiers – Standard: $- for 2,000 hourstext-to-speechvoice-cloning
DI

Dia

Audio · Dia-1.6B
7.3

Open-weights 1.6B text-to-dialogue model that generates ultra-realistic multi-speaker conversations in one pass.

Free· Free, open weights (Apache 2.0); hosted larger version waitlisteddialogue-generationvoice-cloning
GE

Gemini

Writing · Gemini 2.x (Flash / Pro / Ultra)
7.3

Google's flagship multimodal AI assistant with deep integration into Workspace and Android.

Freemium· Free tier; Google One AI Premium $19.99/mo includes Gemini Advancedchat-assistantresearch
GE

Geniusrise

Agents · Multi-model
7.2

Open-source framework for building, deploying, and scaling AI microservices across text, vision, and audio.

Free· Free, open source; self-hostedinference-servingfine-tuning
LB

LLM by Datasette

Coding · Multi-model
7.2

A CLI and Python library for running prompts against any LLM provider and logging everything to SQLite.

Free· Free and open source (Apache 2.0); pay underlying model providers separatelycli-promptingprompt-logging
LA

LOVO AI

Audio · Proprietary (LOVO Pro V2 voices)
7.2

Text-to-speech and voice cloning platform with 500+ voices, an integrated video editor, and a developer API.

Freemium· 14-day free Pro trial, no credit card; paid subscription tierstext-to-speechvoice-cloning
VA

Veritone Automate Studio

Agents · Multi-model (aiWARE engine ecosystem)
7.2

Low-code workflow builder for orchestrating multi-engine AI pipelines across audio, video, text, and images.

Enterprise· Contact sales; no public pricingmedia-monitoringmetadata-extraction
VV

Veritone Voice

Audio · Proprietary (Veritone aiWARE)
7.2

Enterprise-grade voice cloning and synthesis platform built for broadcasters, studios, and large media operations.

Enterprise· Contact sales / demo onlyvoice-cloningtext-to-speech
VI

Vibe

Audio · OpenAI Whisper (via whisper.cpp)
7.2

Offline desktop transcription app powered by Whisper, with diarization, batch processing, and an HTTP API.

Free· Free and open-source (MIT)transcriptionsubtitles
AA

AnkiDecks AI

Writing
7.1

AI flashcard generator that turns PDFs, slides, YouTube videos and handwritten notes into Anki-ready decks.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesflashcard-generationspaced-repetition
DR

DaVinci Resolve

Video · DaVinci Neural Engine (proprietary)
7.1

Hollywood-grade post-production suite with a Neural Engine that quietly automates the tedious parts of editing, color, and audio.

Freemium· Free Version: Free · DaVinci Resolve Studio: $295video-editingcolor-grading
NA

Nudge AI

Writing
7.1

AI scribe and clinical documentation platform that generates audit-ready notes within 30 seconds of a session.

Freemium· Free (5 notes); Pro $99/mo; Enterprise customclinical-notesai-scribe