Foundation Model Families — GPT, Claude, Gemini, Llama, Mistral | AI Terms Guide
🎯
🎯 60 terms · Model families and lineages

The foundation model vocabulary

GPT, Claude, Gemini, Llama, Mistral — the concepts and terminology of the model families reshaping AI. Base vs instruction-tuned, sizes, tiers, and versioning.

60
Terms
4
Sub-topics
Weekly
Updates

Foundation model families are the "brands" of modern AI. When you say "GPT-5" you're referring to a specific model in OpenAI's GPT family; when you say "Claude Opus" you're referring to a tier within Anthropic's Claude family. Understanding how families work — what versions mean, what tiers indicate, how base and instruction-tuned variants differ — is fundamental literacy.

The 60 terms here are grouped into four sub-topics: the family concept itself and its history, individual family details (GPT, Claude, Gemini, Llama, Mistral, and others), model tiers and sizing (Opus/Sonnet/Haiku, Small/Medium/Large, distillations), and the development lifecycle vocabulary (pretraining, post-training, releases). For individual model specs, see the models section. For company backgrounds, see providers.

Full directory

All Foundation Model Families terms, organized

Grouped into sub-topics so you can find neighbors and prerequisites, not just alphabetical entries.

Beyond terminology

Go deeper on Foundation Model Families

📖 Related concept tutorials

Long-form guides that walk through how these concepts actually work.

🎯 Related models & tools

Real products and models where you'll encounter these terms.

Frequently Asked

Questions about Foundation Model Families

Match model tier to task complexity. Simple tasks → smallest tier (Haiku, Flash, mini). Coding, reasoning → mid or flagship tier. See our model comparisons for concrete choices by use case.

Base model completes text — give it 'The capital of France is' and it says 'Paris.' Chat model responds to instructions and follows conversation format. Nearly all public APIs are chat models. Base models are used mainly for further fine-tuning.

Usually yes but not uniformly. Newer versions may excel at some tasks and regress on others. Always benchmark on your own use case before assuming 'newer = better.'

A capability that appears at scale but not at smaller sizes — sometimes suddenly. Chain-of-thought reasoning, instruction following, and some multilingual capabilities famously emerged as scale increased. Whether emergence is real or an artifact of measurement is debated.

The provider has scheduled it for removal. Existing integrations still work but won't receive updates and will eventually stop responding. Providers give notice periods (usually 3-6 months) before shutdown.

Share with