AI Roadmap/Layer 02 · 02. Compute, Network & Cloud Infrastructure
2.4

2.4 GPU Scheduling, Multi-tenancy & Cloud Control Plane

Kubernetes (GPU Operator/Kueue/Volcano), Slurm batch queues, Ray distributed execution, Gang scheduling, MIG multi-tenancy, and FinOps GPU cost governance.

Kubernetes & Slurm GPU Scheduling & Orchestration

NVIDIA GPU Operator automated management, Kueue/Volcano batch queues, Slurm HPC workload management, Gang scheduling, and topology-aware GPU placement.

🏢 Companies
Google Cloud (GKE)AWS (EKS)Microsoft Azure (AKS)CoreWeaveRed Hat
🛠️ Tech Stack
KubernetesSlurmGPU OperatorKueueVolcanoGang SchedulingTopology-Aware Placement
💼 Roles & Salary
Platform Engineer、Scheduler Engineer、Cloud Infrastructure Engineer、DevOps / SRE
💰 $185K - $390K / year (Cloud K8s & Scheduling) | ¥450K - ¥1.15M / year
📚 Prerequisites: Kubernetes Core & CRD Controllers • Slurm Resource Management & Queues • Gang Scheduling & Topology Placement • NVIDIA Container Toolkit & Runtime

Ray Distributed Execution, Multi-Tenancy & FinOps

Ray Core / KubeRay elastic execution graphs, Run:ai dynamic pooling, MIG hardware partitioning, Fair-share quotas, and Spot/preemptible FinOps cost optimization.

🏢 Companies
AnyscaleRun:aiCoreWeaveGoogle CloudAWS
🛠️ Tech Stack
Ray / KubeRayRun:aiMulti-tenancyFair-share QuotasFinOpsSpot InstancesGPU Pooling
💼 Roles & Salary
Platform Engineer、Scheduler Engineer、FinOps Engineer、DevOps / SRE
💰 $180K - $380K / year (Ray & Cloud FinOps) | ¥420K - ¥1.1M / year
📚 Prerequisites: Ray Distributed Framework & Actor Model • GPU Multi-tenancy Quotas & Preemption • Spot / Preemptible Instance Fault-tolerance • FinOps Cost Analytics & Utilization Metrics