The infrastructure powering AI
Training datasets, GPUs, TPUs, HBM, data centers, cloud services, and the companies that make modern AI possible.
Modern AI isn't just algorithms — it's an enormous industrial infrastructure. Training a frontier model costs tens of millions and requires access to specialized hardware (H100 GPUs, TPU pods), massive datasets curated over years (Common Crawl, The Pile, FineWeb), specialized data centers, and cloud services built specifically for AI workloads.
The 110 terms here are grouped into four sub-topics: training datasets (Common Crawl, The Pile, FineWeb, and specialized corpora), hardware and chips (NVIDIA GPUs, TPUs, custom silicon, HBM), infrastructure and data centers (cooling, networking, power), and the ecosystem of companies (chip makers, cloud providers, dataset providers, tooling companies). Related content in training, inference, and providers.
The most important terms in Datasets, Infra & Companies
Start here if you're new. These entries explain the foundational vocabulary in depth.
NVIDIA H100
The GPU that trained most of modern AI. Announced 2022, still the most common training accelerator. Successor: H200, then Blackwell B100/B200.
Read the full entry →DatasetsCommon Crawl
The internet-scale web scrape underlying most LLM training. A nonprofit repository of petabytes of web pages; every major model has trained on some Common Crawl derivative.
Read the full entry →DatasetsThe Pile
EleutherAI's curated 800GB dataset from 22 diverse sources. Standard reference for open LLM pretraining before FineWeb.
Read the full entry →HardwareTPU
Google's Tensor Processing Unit — custom silicon for AI training and inference. Gemini and other Google models train on TPU pods.
Read the full entry →HardwareHBM
High Bandwidth Memory — the specialized DRAM stacked on top of GPUs that makes AI workloads fast. HBM3, HBM3e, and HBM4 are the current generations.
Read the full entry →InfrastructureAI Data Center
Purpose-built facilities for AI training and inference. Distinct from general cloud data centers in power density, cooling, and networking requirements.
Read the full entry →All Datasets, Infra & Companies terms, organized
Grouped into sub-topics so you can find neighbors and prerequisites, not just alphabetical entries.
Training datasets
Hardware & chips
Infrastructure
Ecosystem & companies
Related term categories
These categories connect naturally to Datasets, Infra & Companies — many terms cross-reference between them.
Go deeper on Datasets, Infra & Companies
📖 Related concept tutorials
Long-form guides that walk through how these concepts actually work.
- What is a Training Run
A full frontier training job explained.
- Understanding HBM
Why AI GPUs are memory-bound.
- AI Data Centers Explained
Why they differ from cloud data centers.
🎯 Related models & tools
Real products and models where you'll encounter these terms.
- AI providers
The companies using this infrastructure.
- Groq
Custom LPU chips.
- Cerebras
Wafer-scale AI chips.
Questions about Datasets, Infra & Companies
LLM inference and training are memory-bandwidth-bound — you spend more time moving data than doing math. HBM's very high bandwidth (5+ TB/s per stack) is why AI GPUs cost 10x normal GPUs. It's the biggest single bottleneck for AI hardware.
Different trade-offs. TPUs are optimized for large batch training with predictable workloads. GPUs are more flexible and dominate the open ecosystem. Both remain competitive; frontier models train on both.
Frontier models train on 15-20+ trillion tokens. That's roughly the entire internet (filtered), plus books, code, and specialized datasets. Chinchilla scaling laws initially suggested 20 tokens per parameter is optimal; recent work suggests much more.
Depends who you ask: (1) compute — access to enough H100/H200/Blackwell GPUs, (2) power — data centers need gigawatt-scale electricity, (3) high-quality training data — the internet has been picked over. All three are being pushed hard.
Common Crawl (public web scrape), Wikipedia, book datasets (often licensed or gray-legal), GitHub code, curated public datasets (The Pile, FineWeb), synthetic data generated by other models, and licensed private data (news archives, academic papers). Sourcing is increasingly a competitive advantage.
Reviewed by the AI Terms Guide editorial team on August 6, 2026. Last updated: August 6, 2026. Spotted an issue? Let us know.
Explore our AI reference network
Six specialist sites, one shared editorial standard.
AI Terms Weekly
One deep term, three new models, one comparison, and the paper of the week — every Tuesday.
Free · No spam · Join 30,000+ AI professionals