The foundation model vocabulary
GPT, Claude, Gemini, Llama, Mistral — the concepts and terminology of the model families reshaping AI. Base vs instruction-tuned, sizes, tiers, and versioning.
Foundation model families are the "brands" of modern AI. When you say "GPT-5" you're referring to a specific model in OpenAI's GPT family; when you say "Claude Opus" you're referring to a tier within Anthropic's Claude family. Understanding how families work — what versions mean, what tiers indicate, how base and instruction-tuned variants differ — is fundamental literacy.
The 60 terms here are grouped into four sub-topics: the family concept itself and its history, individual family details (GPT, Claude, Gemini, Llama, Mistral, and others), model tiers and sizing (Opus/Sonnet/Haiku, Small/Medium/Large, distillations), and the development lifecycle vocabulary (pretraining, post-training, releases). For individual model specs, see the models section. For company backgrounds, see providers.
The most important terms in Foundation Model Families
Start here if you're new. These entries explain the foundational vocabulary in depth.
Foundation Model
A model trained on broad data at scale that can be adapted to many downstream tasks. Term coined by Stanford in 2021 to describe the paradigm shift from task-specific models.
Read the full entry →FoundationsBase Model
A pretrained model before instruction tuning. Completes text but does not follow instructions. Foundation for chat models via fine-tuning.
Read the full entry →FamiliesGPT Family
OpenAI's Generative Pretrained Transformer family — from GPT-1 (2018) to GPT-5 (2026). The family that defined the era of large language models.
Read the full entry →FamiliesClaude Family
Anthropic's family of frontier models named for Claude Shannon. Organized in tiers (Opus, Sonnet, Haiku) with numeric generations (3, 4, 4.5, 4.6, 4.7, 4.8).
Read the full entry →FamiliesLlama Family
Meta's open-weight family — from Llama 1 (2023) to Llama 4 (2026). The reference open-weight family that made high-quality LLMs accessible.
Read the full entry →SizingModel Tier
A size or capability level within a family. Claude has Opus/Sonnet/Haiku. OpenAI has GPT-5/mini/nano. Enables provider price discrimination and capability trade-offs.
Read the full entry →All Foundation Model Families terms, organized
Grouped into sub-topics so you can find neighbors and prerequisites, not just alphabetical entries.
Family concepts
Specific families
Model tiers & sizing
Development lifecycle
Related term categories
These categories connect naturally to Foundation Model Families — many terms cross-reference between them.
Go deeper on Foundation Model Families
📖 Related concept tutorials
Long-form guides that walk through how these concepts actually work.
- Understanding Model Families
What versions and tiers actually mean.
- Scaling Laws Explained
Why bigger models are better.
- Base vs Instruction-Tuned
The critical distinction.
🎯 Related models & tools
Real products and models where you'll encounter these terms.
- All AI models
Every model organized by family.
- All providers
The labs behind the families.
- Anthropic vs OpenAI
Two flagship families compared.
Questions about Foundation Model Families
Match model tier to task complexity. Simple tasks → smallest tier (Haiku, Flash, mini). Coding, reasoning → mid or flagship tier. See our model comparisons for concrete choices by use case.
Base model completes text — give it 'The capital of France is' and it says 'Paris.' Chat model responds to instructions and follows conversation format. Nearly all public APIs are chat models. Base models are used mainly for further fine-tuning.
Usually yes but not uniformly. Newer versions may excel at some tasks and regress on others. Always benchmark on your own use case before assuming 'newer = better.'
A capability that appears at scale but not at smaller sizes — sometimes suddenly. Chain-of-thought reasoning, instruction following, and some multilingual capabilities famously emerged as scale increased. Whether emergence is real or an artifact of measurement is debated.
The provider has scheduled it for removal. Existing integrations still work but won't receive updates and will eventually stop responding. Providers give notice periods (usually 3-6 months) before shutdown.
Reviewed by the AI Terms Guide editorial team on August 6, 2026. Last updated: August 6, 2026. Spotted an issue? Let us know.
Explore our AI reference network
Six specialist sites, one shared editorial standard.
AI Terms Weekly
One deep term, three new models, one comparison, and the paper of the week — every Tuesday.
Free · No spam · Join 30,000+ AI professionals