GPU Scheduler
System Analysis
Normal Behavior
When an ML training job or inference pod is scheduled, the GPU scheduler inspects available accelerator VRAM and interconnect topology, binds specific GPU devices to the container cgroup, and enforces memory quotas.
Failure Behavior
If the scheduler assigns workloads without accounting for memory fragmentation or thermal throttling, jobs fail with CUDA out-of-memory errors, low-priority jobs starve production inference, or unallocated GPUs sit idle due to topology lockouts.
Business Consequence
Unknown
Known Aliases
Technical Terminology
Failure Indicators
System Architecture (Graph)
FAQ
How does it normally behave?
When an ML training job or inference pod is scheduled, the GPU scheduler inspects available accelerator VRAM and interconnect topology, binds specific GPU devices to the container cgroup, and enforces memory quotas.
How does it fail?
If the scheduler assigns workloads without accounting for memory fragmentation or thermal throttling, jobs fail with CUDA out-of-memory errors, low-priority jobs starve production inference, or unallocated GPUs sit idle due to topology lockouts.
What is the business consequence?
Unknown
What is the difference between GPU time-slicing and Multi-Instance GPU (MIG)?
Time-slicing allows multiple containers to share a GPU by rapidly interleaving execution, but provides no memory isolation—if one container over-allocates VRAM, all containers crash. Multi-Instance GPU (MIG) physically partitions the GPU into up to 7 hardware-isolated instances with dedicated memory, cache, and compute slices, guaranteeing complete fault isolation.
Explore the system
AI Summary
GPU Scheduler is a undefined system in TinyCTO.tv. When an ML training job or inference pod is scheduled, the GPU scheduler inspects available accelerator VRAM and interconnect topology, binds specific GPU devices to the container cgroup, and enforces memory quotas.
