The vocabulary of AI safety
Constitutional AI, red teaming, jailbreaks, prompt injection, interpretability, superalignment. Every term you need to reason about how AI is (and isn't) safe.
AI safety and alignment are the fields concerned with making AI systems do what humans want them to do, safely and reliably. Beyond the theoretical debates, safety vocabulary is now essential for anyone deploying AI in production — you cannot evaluate a security review, prompt injection defense, or content policy without knowing these terms.
The 40 terms are grouped into four sub-topics: alignment methods (how we shape model behavior), attack vectors (how models can be manipulated), defenses and evaluation (red teaming, guardrails), and interpretability (understanding what models actually do internally). Every technique here has real trade-offs — we don't oversimplify.
The most important terms in AI Safety & Alignment
Start here if you're new. These entries explain the foundational vocabulary in depth.
Alignment
The problem of making AI systems reliably do what humans want. Covers everything from RLHF for chat models to research on maintaining goals in more capable systems.
Read the full entry →Alignment methodsConstitutional AI
Anthropic's approach to alignment where the model is trained against a set of written principles (a 'constitution') rather than raw human preferences. Reduces the need for humans to label harmful outputs directly.
Read the full entry →EvaluationRed Teaming
Deliberately trying to make a model behave badly — output harmful content, leak information, produce jailbreaks. Standard practice before deploying any frontier model.
Read the full entry →AttacksJailbreak
A prompt or interaction pattern that gets a model to bypass its safety training. Named after phone jailbreaks; increasingly hardened against as models mature.
Read the full entry →AttacksPrompt Injection
Attacks where malicious instructions embedded in retrieved content, tool responses, or documents override the developer's original instructions. The top security issue for RAG and agent systems.
Read the full entry →UnderstandingInterpretability
The field of understanding what neural networks actually compute internally — not just what they output. Mechanistic interpretability is a growing subfield.
Read the full entry →All AI Safety & Alignment terms, organized
Grouped into sub-topics so you can find neighbors and prerequisites, not just alphabetical entries.
Alignment methods
Attacks & failure modes
Defenses & evaluation
Interpretability & understanding
Related term categories
These categories connect naturally to AI Safety & Alignment — many terms cross-reference between them.
Go deeper on AI Safety & Alignment
📖 Related concept tutorials
Long-form guides that walk through how these concepts actually work.
- How RLHF Works
The full technique explained.
- Understanding Prompt Injection
Why it's the #1 agent security issue.
- What is Interpretability
Anthropic's and OpenAI's work explained.
🎯 Related models & tools
Real products and models where you'll encounter these terms.
- Anthropic
Safety-focused frontier lab.
- Claude Opus 4.8
Trained with Constitutional AI.
- AI guardrail tools
Lakera, Rebuff, and prompt injection defenses.
Questions about AI Safety & Alignment
No. Bias mitigation is one piece. AI safety also covers robustness (models breaking on edge cases), security (jailbreaks, prompt injection), reliability (hallucinations), and long-term alignment research on more capable systems.
Overlapping. AI safety focuses on the model behaving as intended for users. AI security focuses on protecting against adversarial actors. Prompt injection sits at their intersection.
No. Every frontier model release brings new jailbreaks. The field has hardened significantly — trivial jailbreaks that worked in 2023 no longer work — but sophisticated attackers still find bypasses.
Yes, but it can be removed by anyone with the weights. This is one reason for the ongoing debate about open-weight release policies. See our provider profiles for how different labs approach this.
Different trade-offs. Constitutional AI is more transparent (you can read the principles), easier to iterate on, and less dependent on human labelers for harmful outputs. RLHF has been more battle-tested. Most frontier labs use hybrids.
Reviewed by the AI Terms Guide editorial team on August 6, 2026. Last updated: August 6, 2026. Spotted an issue? Let us know.
Explore our AI reference network
Six specialist sites, one shared editorial standard.
AI Terms Weekly
One deep term, three new models, one comparison, and the paper of the week — every Tuesday.
Free · No spam · Join 30,000+ AI professionals