AI Safety & Alignment Terms — Complete Reference | AI Terms Guide
🛡️
🛡️ 40 terms · Making AI safe and aligned

The vocabulary of AI safety

Constitutional AI, red teaming, jailbreaks, prompt injection, interpretability, superalignment. Every term you need to reason about how AI is (and isn't) safe.

40
Terms
4
Sub-topics
Weekly
Updates

AI safety and alignment are the fields concerned with making AI systems do what humans want them to do, safely and reliably. Beyond the theoretical debates, safety vocabulary is now essential for anyone deploying AI in production — you cannot evaluate a security review, prompt injection defense, or content policy without knowing these terms.

The 40 terms are grouped into four sub-topics: alignment methods (how we shape model behavior), attack vectors (how models can be manipulated), defenses and evaluation (red teaming, guardrails), and interpretability (understanding what models actually do internally). Every technique here has real trade-offs — we don't oversimplify.

Full directory

All AI Safety & Alignment terms, organized

Grouped into sub-topics so you can find neighbors and prerequisites, not just alphabetical entries.

Beyond terminology

Go deeper on AI Safety & Alignment

📖 Related concept tutorials

Long-form guides that walk through how these concepts actually work.

🎯 Related models & tools

Real products and models where you'll encounter these terms.

Frequently Asked

Questions about AI Safety & Alignment

No. Bias mitigation is one piece. AI safety also covers robustness (models breaking on edge cases), security (jailbreaks, prompt injection), reliability (hallucinations), and long-term alignment research on more capable systems.

Overlapping. AI safety focuses on the model behaving as intended for users. AI security focuses on protecting against adversarial actors. Prompt injection sits at their intersection.

No. Every frontier model release brings new jailbreaks. The field has hardened significantly — trivial jailbreaks that worked in 2023 no longer work — but sophisticated attackers still find bypasses.

Yes, but it can be removed by anyone with the weights. This is one reason for the ongoing debate about open-weight release policies. See our provider profiles for how different labs approach this.

Different trade-offs. Constitutional AI is more transparent (you can read the principles), easier to iterate on, and less dependent on human labelers for harmful outputs. RLHF has been more battle-tested. Most frontier labs use hybrids.

Share with