Skip to main content

GPU Scheduler

System Analysis

Normal Behavior

When an ML training job or inference pod is scheduled, the GPU scheduler inspects available accelerator VRAM and interconnect topology, binds specific GPU devices to the container cgroup, and enforces memory quotas.

Failure Behavior

If the scheduler assigns workloads without accounting for memory fragmentation or thermal throttling, jobs fail with CUDA out-of-memory errors, low-priority jobs starve production inference, or unallocated GPUs sit idle due to topology lockouts.

Business Consequence

Unknown

Known Aliases

Accelerator OrchestratorGPU Workload SchedulerCluster GPU Manager

Technical Terminology

MIG (Multi-Instance GPU)NVLinkVRAM AllocationTopology AwarenessGang Scheduling

Failure Indicators

Allocation StarvationThermal ThrottlingCUDA OOMWorkload Preemption

System Architecture (Graph)

ARCHITECTURE FLOWCHARTCanonical Architecture Diagram
⚡ TinyCTO.tv

FAQ

How does it normally behave?

When an ML training job or inference pod is scheduled, the GPU scheduler inspects available accelerator VRAM and interconnect topology, binds specific GPU devices to the container cgroup, and enforces memory quotas.

How does it fail?

If the scheduler assigns workloads without accounting for memory fragmentation or thermal throttling, jobs fail with CUDA out-of-memory errors, low-priority jobs starve production inference, or unallocated GPUs sit idle due to topology lockouts.

What is the business consequence?

Unknown

What is the difference between GPU time-slicing and Multi-Instance GPU (MIG)?

Time-slicing allows multiple containers to share a GPU by rapidly interleaving execution, but provides no memory isolation—if one container over-allocates VRAM, all containers crash. Multi-Instance GPU (MIG) physically partitions the GPU into up to 7 hardware-isolated instances with dedicated memory, cache, and compute slices, guaranteeing complete fault isolation.

AI Summary

GPU Scheduler is a undefined system in TinyCTO.tv. When an ML training job or inference pod is scheduled, the GPU scheduler inspects available accelerator VRAM and interconnect topology, binds specific GPU devices to the container cgroup, and enforces memory quotas.