/
GPU Fundamentals for AI Inference
AI Generated
Export
Save to your account
Sign up
Report bug
Your browser does not support this video.
0:00 / 4:38
GPU Fundamentals for AI Inference
4:38 · AI Generated · 7/22/2026
Chapters
Summary
Transcript
Chapters
Summary
Transcript
Chapters
1
The Performance Loop
0:00
2
The GPU Execution Hierarchy
0:32
3
Hiding Latency with Occupancy
1:04
4
Memory Coalescing
1:35
5
Shared Memory and Bank Conflicts
2:08
6
The Roofline Model
2:39
7
NIRC and Operator Fusion
3:08
8
The Validation Judge
3:40
9
The Path to Mastery
4:13