Speech & Audio AI Terms — Whisper, TTS, Suno, Voice Cloning | AI Terms Guide
🎧
🎧 25 terms · AI that hears and speaks

Speech & audio — AI beyond text and pixels

Whisper, TTS, ASR, voice cloning, music generation. Every term you need to work with AI that hears, speaks, and makes music.

25
Terms
4
Sub-topics
Weekly
Updates

Speech and audio AI have been in a quiet renaissance. OpenAI's voice models made real-time conversational AI feel natural. ElevenLabs made voice cloning accessible. Suno and Udio can generate remarkably good music from text prompts. And Whisper quietly became the reference for speech recognition — still widely deployed for transcription, captioning, and voice interfaces.

The 25 terms in this category are grouped into four sub-topics: speech recognition (ASR), speech synthesis (TTS), music and audio generation, and the underlying audio processing concepts. Fewer terms than other categories because speech and audio are a smaller (though rapidly-growing) slice of AI activity.

Beyond terminology

Go deeper on Speech & Audio

📖 Related concept tutorials

Long-form guides that walk through how these concepts actually work.

🎯 Related models & tools

Real products and models where you'll encounter these terms.

Frequently Asked

Questions about Speech & Audio

For open-source and general-purpose transcription, yes. For real-time, low-latency, or highest-accuracy commercial needs, Deepgram and AssemblyAI often outperform. Whisper large-v3 is the current strongest Whisper variant.

Very good — indistinguishable from human speech in most cases with 30+ seconds of source audio. This has significant ethical implications: consent, misuse for scams, and impersonation are real concerns.

Contested and jurisdiction-dependent. In the US, purely AI-generated works generally cannot be copyrighted, but works with substantial human creative input can. This is an active legal question — this isn't legal advice.

ASR is speech-to-text (understanding). TTS is text-to-speech (generation). Voice assistants like ChatGPT Voice do both — ASR the input, LLM the reasoning, TTS the output.

Yes for many. Whisper runs on modest hardware. Local TTS is available (Coqui TTS, Piper). Voice cloning has open-source options. Music generation locally is limited but possible with MusicGen.

Share with