AI Evaluation & Benchmark Terms — Complete Reference | AI Terms Guide
📊
📊 40 terms · Measuring what matters

How to measure AI — every metric explained

MMLU, HumanEval, GSM8K, Chatbot Arena, LLM-as-Judge. Every benchmark, metric, and evaluation approach with what it actually measures — and what it doesn't.

40
Terms
4
Sub-topics
Weekly
Updates

Evaluation is the hardest problem in modern AI. Benchmark numbers can be misleading, popular benchmarks get 'gamed,' and public leaderboards don't always predict real-world quality. Understanding what each benchmark actually measures — and what it doesn't — is essential for making sensible decisions.

These 40 terms are grouped into four sub-topics: popular benchmarks (MMLU, HumanEval, GSM8K, and others you'll see in every model announcement), evaluation methods (LLM-as-Judge, human eval, automatic metrics), aggregators (Chatbot Arena, Artificial Analysis, HuggingFace leaderboards), and specialized evals for RAG, agents, and safety.

Full directory

All Evaluation & Benchmarks terms, organized

Grouped into sub-topics so you can find neighbors and prerequisites, not just alphabetical entries.

Adjacent categories

Related term categories

These categories connect naturally to Evaluation & Benchmarks — many terms cross-reference between them.

Beyond terminology

Go deeper on Evaluation & Benchmarks

📖 Related concept tutorials

Long-form guides that walk through how these concepts actually work.

🎯 Related models & tools

Real products and models where you'll encounter these terms.

Frequently Asked

Questions about Evaluation & Benchmarks

None single-handedly. For general capability: MMLU-Pro, GPQA, LiveCodeBench, and Chatbot Arena together give a reasonable picture. For your use case, build your own eval set — it's the only benchmark that matters for your decisions.

Different eval harnesses, different prompting, different sampling parameters, and different versions of the benchmark. Always check what version was used and how it was run before comparing numbers.

It reflects human preferences well but has biases — length bias, formatting bias, subject bias. Great signal but not the whole story. Combine with capability benchmarks and your own evals.

When benchmark test data appears in training data. Makes numbers look better than they are. Newer benchmarks (LiveCodeBench, MMLU-Pro) have contamination controls; older ones (MMLU, HumanEval) are heavily contaminated.

Task success rate on realistic benchmarks (SWE-Bench for coding, WebArena for browsing), plus operational metrics: tokens used, tools called, error recovery, human intervention rate. See our evaluation guide.

Share with