Logo

GPU Fundamentals for AI Inference

0:00 / 4:38
The Performance Loop
0:00 - 0:32

Learn the scientific approach to GPU optimization: fix workload, inspect, classify limiter, change small, and measure.

The GPU Execution Hierarchy
0:32 - 1:04

Kernels, grids, blocks, and warps of 32 threads—and why divergence stalls the group.

Hiding Latency with Occupancy
1:04 - 1:35

How GPUs swap ready warps to hide memory wait times and why occupancy matters.

Memory Coalescing
1:35 - 2:08

Efficient global-memory fetches when threads access contiguous addresses, and Structure of Arrays.

Shared Memory and Bank Conflicts
2:08 - 2:39

On-chip shared memory banks, conflicts when threads hit the same bank, and padding fixes.

The Roofline Model
2:39 - 3:08

Arithmetic intensity shows whether you are compute-bound or memory-bound on the roofline.

NIRC and Operator Fusion
3:08 - 3:40

How GPU-resident fused MLPs cut memory traffic in a neural rendering pipeline.

The Validation Judge
3:40 - 4:13

Quality gates with FLIP and exceedance-rate validation for LLM-proposed optimizations.

The Path to Mastery
4:13 - 4:38

Recap: performance loop, warps, memory ceilings, and evidence-based validation.