Speech & audio — AI beyond text and pixels
Whisper, TTS, ASR, voice cloning, music generation. Every term you need to work with AI that hears, speaks, and makes music.
Speech and audio AI have been in a quiet renaissance. OpenAI's voice models made real-time conversational AI feel natural. ElevenLabs made voice cloning accessible. Suno and Udio can generate remarkably good music from text prompts. And Whisper quietly became the reference for speech recognition — still widely deployed for transcription, captioning, and voice interfaces.
The 25 terms in this category are grouped into four sub-topics: speech recognition (ASR), speech synthesis (TTS), music and audio generation, and the underlying audio processing concepts. Fewer terms than other categories because speech and audio are a smaller (though rapidly-growing) slice of AI activity.
The most important terms in Speech & Audio
Start here if you're new. These entries explain the foundational vocabulary in depth.
Whisper
OpenAI's open-source speech recognition model. Multilingual, robust to accents and noise, still the reference implementation for ASR in 2026.
Read the full entry →SynthesisText-to-Speech (TTS)
Generating spoken audio from text. Modern TTS (ElevenLabs, OpenAI, Play.ht) is nearly indistinguishable from human speech.
Read the full entry →RecognitionAutomatic Speech Recognition (ASR)
Converting spoken audio to text. Also called speech-to-text (STT). Whisper is the dominant open model; commercial options include Deepgram, AssemblyAI, and Google.
Read the full entry →SynthesisVoice Cloning
Generating speech in a specific person's voice from a short audio sample. Powered by ElevenLabs, OpenAI Voice Engine, and open-source models. Raises significant ethical concerns.
Read the full entry →GenerationSuno
Text-to-music model that generates full songs with vocals from a text prompt. One of the most impressive multimodal capabilities of 2024-2026.
Read the full entry →ArchitectureAudio Diffusion
Diffusion models applied to audio waveforms or spectrograms. Foundation of modern music generation and voice synthesis.
Read the full entry →All Speech & Audio terms, organized
Grouped into sub-topics so you can find neighbors and prerequisites, not just alphabetical entries.
Speech recognition (ASR)
Speech synthesis (TTS)
Music & audio generation
Underlying concepts
Related term categories
These categories connect naturally to Speech & Audio — many terms cross-reference between them.
Go deeper on Speech & Audio
📖 Related concept tutorials
Long-form guides that walk through how these concepts actually work.
- How Whisper Works
ASR architecture explained.
- How Modern TTS Works
From text to natural speech.
- Music Generation Explained
How Suno and Udio actually work.
🎯 Related models & tools
Real products and models where you'll encounter these terms.
- ElevenLabs
Voice generation, cloning, TTS.
- Suno
Text-to-music generation.
- OpenAI Voice
OpenAI's voice models.
Questions about Speech & Audio
For open-source and general-purpose transcription, yes. For real-time, low-latency, or highest-accuracy commercial needs, Deepgram and AssemblyAI often outperform. Whisper large-v3 is the current strongest Whisper variant.
Very good — indistinguishable from human speech in most cases with 30+ seconds of source audio. This has significant ethical implications: consent, misuse for scams, and impersonation are real concerns.
Contested and jurisdiction-dependent. In the US, purely AI-generated works generally cannot be copyrighted, but works with substantial human creative input can. This is an active legal question — this isn't legal advice.
ASR is speech-to-text (understanding). TTS is text-to-speech (generation). Voice assistants like ChatGPT Voice do both — ASR the input, LLM the reasoning, TTS the output.
Yes for many. Whisper runs on modest hardware. Local TTS is available (Coqui TTS, Piper). Voice cloning has open-source options. Music generation locally is limited but possible with MusicGen.
Reviewed by the AI Terms Guide editorial team on August 6, 2026. Last updated: August 6, 2026. Spotted an issue? Let us know.
Explore our AI reference network
Six specialist sites, one shared editorial standard.
AI Terms Weekly
One deep term, three new models, one comparison, and the paper of the week — every Tuesday.
Free · No spam · Join 30,000+ AI professionals