Every AI model, specs and benchmarks included
Frontier LLMs, open-weight releases, reasoning models, multimodal models — every major model with cited specs, real benchmarks, and current pricing.
AI Terms Guide tracks every major AI model with real specs, cited benchmarks, current pricing, and honest evaluations. When a model is released, we publish a full page within 48 hours. When pricing changes, we update within a week. When new benchmarks land, we refresh the numbers.
The models section covers frontier LLMs (Claude, GPT, Gemini), open-weight releases (Llama, Mistral, DeepSeek, Qwen), specialist reasoning models (o-series, DeepSeek-R1), multimodal models, image generators, and video and audio models. Every entry links to its provider profile, cross-references relevant concepts, and lists head-to-head comparisons.
Browse models by the lab that built them
Model families organized by provider. Click any provider name for the full company profile.
Claude family
- Claude Opus 4.8 · flagship
- Claude Opus 4.7
- Claude Sonnet 4.6 · workhorse
- Claude Haiku 4.5 · fast
Flagship models — updated weekly
Every model page has specs, cited benchmarks, current pricing, honest strengths and weaknesses, and links to head-to-head comparisons.
Claude Opus 4.8
Anthropic's most capable model with extended thinking, industry-leading coding, and strong agentic tool use. Optimized for hard reasoning and long-horizon tasks.
GPT-5
OpenAI's flagship model unifying reasoning, multimodal input, and long-context tool use in a single interface. Successor to GPT-4o and the o-series.
Gemini 3
Google's latest frontier model with strong long-context handling, native multimodal input, and deep Google-ecosystem integration.
Claude Sonnet 4.6
The best balance of capability and cost. Strong on coding, reasoning, and reliable tool use for production apps.
Llama 4
Meta's open-weight family with multiple sizes, competitive benchmarks, and permissive licensing for self-hosting and fine-tuning.
DeepSeek R1
Open-weight reasoning model trained with reinforcement learning. Matches top proprietary models on math and coding benchmarks at fraction of the cost.
Find the right model for your task
Different tasks demand different capabilities. These pages rank the best models for each.
Best for coding
Ranked coding models — SWE-bench scores, agentic ability, tool use quality.
Best for reasoning
Extended-thinking models compared on math, logic, and multi-step problems.
Best for long context
Models tested on 200K-1M+ token workloads. "Needle in a haystack" plus real long tasks.
Best for vision
Vision-language models compared on OCR, diagram understanding, and visual reasoning.
Best for cost
Cheap models that still deliver — cost per million tokens compared with quality.
Best open-weight
Self-hostable models with permissive licenses. Llama, Mistral, DeepSeek, Qwen.
Popular model comparisons
Claude Opus 4.8 vs GPT-5
Frontier flagships compared on coding, reasoning, and cost.
Gemini 3 vs Claude Sonnet 4.6
Long-context multimodal vs. balanced workhorse.
GPT-5 vs Gemini 3
OpenAI vs Google — reasoning, tool use, ecosystem.
Llama 4 vs Mistral Large 2
Two open-weight flagships compared for self-hosting.
DeepSeek R1 vs o-series
Two reasoning-model families compared on hard benchmarks.
Claude Haiku 4.5 vs GPT-5 mini
Cheap, fast models compared for high-volume workloads.
Questions about Models
Pricing and new releases are reviewed weekly. Benchmark numbers are refreshed monthly as new evaluation results are published. Every model page displays a "last updated" date. See our full update cadence.
Both — clearly labeled. Provider benchmarks appear alongside independent evaluations from Artificial Analysis, Chatbot Arena, and academic leaderboards when available.
Depends on the task. Our comparisons and "best for X" pages ({/models/best-for-coding/}, best-for-reasoning, best-for-long-context, etc.) walk through decisions with real trade-offs. Comparing named models is often the fastest path to a decision.
Yes — Llama, Mistral, DeepSeek, Qwen, and other open-weight families get full coverage including self-hosting notes, hardware requirements, and licensing terms.
Within 48 hours of a major release, typically. Rapid releases (like reasoning models or new API endpoints) may get a first-cut page within hours, then get updated as more information becomes available.
Reviewed by the AI Terms Guide editorial team on August 6, 2026. Last updated: August 6, 2026. Spotted an issue? Let us know.
Explore our AI reference network
Six specialist sites, one shared editorial standard.
AI Terms Weekly
One deep term, three new models, one comparison, and the paper of the week — every Tuesday.
Free · No spam · Join 30,000+ AI professionals