How to measure AI — every metric explained
MMLU, HumanEval, GSM8K, Chatbot Arena, LLM-as-Judge. Every benchmark, metric, and evaluation approach with what it actually measures — and what it doesn't.
Evaluation is the hardest problem in modern AI. Benchmark numbers can be misleading, popular benchmarks get 'gamed,' and public leaderboards don't always predict real-world quality. Understanding what each benchmark actually measures — and what it doesn't — is essential for making sensible decisions.
These 40 terms are grouped into four sub-topics: popular benchmarks (MMLU, HumanEval, GSM8K, and others you'll see in every model announcement), evaluation methods (LLM-as-Judge, human eval, automatic metrics), aggregators (Chatbot Arena, Artificial Analysis, HuggingFace leaderboards), and specialized evals for RAG, agents, and safety.
The most important terms in Evaluation & Benchmarks
Start here if you're new. These entries explain the foundational vocabulary in depth.
MMLU
Massive Multitask Language Understanding — 57 subjects from math to law. The most cited general-capability benchmark; heavily saturated by top models but still useful for tracking progress.
Read the full entry →BenchmarksHumanEval
OpenAI's coding benchmark: 164 Python programming problems judged by unit test pass rate. Saturated but still widely reported. Successor benchmarks: MBPP, LiveCodeBench, SWE-Bench.
Read the full entry →AggregatorsChatbot Arena (LMArena)
Crowdsourced blind pairwise voting between LLM responses. Elo rating system produces the most-cited public LLM leaderboard.
Read the full entry →MethodsLLM-as-Judge
Using a strong LLM to grade other model outputs. Cheaper and faster than human eval, though it has biases (position, verbosity) that require careful calibration.
Read the full entry →BenchmarksGSM8K
Grade School Math 8K — 8,000 word problems requiring multi-step arithmetic reasoning. Saturated by top models but a good CoT test.
Read the full entry →BenchmarksMT-Bench
Multi-turn conversation benchmark from LMSYS. 80 questions across categories, graded by GPT-4 pairwise. Standard for chat model quality.
Read the full entry →All Evaluation & Benchmarks terms, organized
Grouped into sub-topics so you can find neighbors and prerequisites, not just alphabetical entries.
Popular benchmarks
Evaluation methods
Leaderboards & aggregators
Specialized evals
Related term categories
These categories connect naturally to Evaluation & Benchmarks — many terms cross-reference between them.
Go deeper on Evaluation & Benchmarks
📖 Related concept tutorials
Long-form guides that walk through how these concepts actually work.
- How to Evaluate an LLM
End-to-end evaluation guide.
- Building an Eval Set
Practical guide for your use case.
- LLM-as-Judge — Full Guide
Setting it up without the pitfalls.
🎯 Related models & tools
Real products and models where you'll encounter these terms.
- All models
Every model page includes benchmark data.
- Best for coding
Ranked using coding benchmarks.
- Best for reasoning
Ranked using reasoning benchmarks.
Questions about Evaluation & Benchmarks
None single-handedly. For general capability: MMLU-Pro, GPQA, LiveCodeBench, and Chatbot Arena together give a reasonable picture. For your use case, build your own eval set — it's the only benchmark that matters for your decisions.
Different eval harnesses, different prompting, different sampling parameters, and different versions of the benchmark. Always check what version was used and how it was run before comparing numbers.
It reflects human preferences well but has biases — length bias, formatting bias, subject bias. Great signal but not the whole story. Combine with capability benchmarks and your own evals.
When benchmark test data appears in training data. Makes numbers look better than they are. Newer benchmarks (LiveCodeBench, MMLU-Pro) have contamination controls; older ones (MMLU, HumanEval) are heavily contaminated.
Task success rate on realistic benchmarks (SWE-Bench for coding, WebArena for browsing), plus operational metrics: tokens used, tools called, error recovery, human intervention rate. See our evaluation guide.
Reviewed by the AI Terms Guide editorial team on August 6, 2026. Last updated: August 6, 2026. Spotted an issue? Let us know.
Explore our AI reference network
Six specialist sites, one shared editorial standard.
AI Terms Weekly
One deep term, three new models, one comparison, and the paper of the week — every Tuesday.
Free · No spam · Join 30,000+ AI professionals