Back to AI Roadmap
LAYER 02

02. Compute, Network & Cloud Infrastructure

Benchmark Companies:
AWSAWSAzureAzureGoogle CloudGoogle CloudCoreWeaveCoreWeaveLambda LabsLambda LabsOracle CloudOracle CloudCrusoeCrusoeNebiusNebiusNVIDIANVIDIAAristaAristaBroadcomBroadcomCiscoCiscoMarvellMarvellAstera LabsAstera LabsCredoCredoCoherentCoherentLumentumLumentumPure StoragePure StorageVAST DataVAST DataWekaWekaDDN StorageDDN StorageMinIOMinIORun:aiRun:aiAnyscaleAnyscaleRed HatRed HatAWSAWSGoogle CloudGoogle CloudAWSAWSAzureAzure

Orchestrating massive GPU clusters with non-blocking RDMA fabrics, high-throughput checkpoint storage, and scalable cloud scheduling.

📊Layer View Mode:
🌐3D CROSS-CORRELATION MATRIX

02. Compute, Network & Cloud Infrastructure · Cross-Correlation Ecosystem

Click any company, tech category, or career role to illuminate all cross-associations and dim unrelated entities.

🔍
🏢

Benchmark Enterprises

29
AWSAWS
6
AzureAzure
4
Google CloudGoogle Cloud
5
CoreWeaveCoreWeave
4
Lambda LabsLambda Labs
1
Oracle CloudOracle Cloud
1
CrusoeCrusoe
1
NebiusNebius
1
NVIDIANVIDIA
2
AristaArista
2
BroadcomBroadcom
2
CiscoCisco
2
MarvellMarvell
1
Astera LabsAstera Labs
1
CredoCredo
2
CoherentCoherent
1
LumentumLumentum
1
Pure StoragePure Storage
2
VAST DataVAST Data
2
WekaWeka
2
DDN StorageDDN Storage
1
MinIOMinIO
1
Run:aiRun:ai
1
AnyscaleAnyscale
1
Red HatRed Hat
1
AWSAWS
5
Google CloudGoogle Cloud
5
AWSAWS
5
AzureAzure
4
🛠️

Tech Stack Categories & Atomic Nodes

4 Tracks
2.1

2.1 AI Cluster Engineering & GPU Fleet Operations

10K+ GPU supercomputing clusters, rack-scale topologies, rail-optimized networks, DCGM telemetry, MIG partitioning, and automated fleet self-healing.
Sub-Page
10K+ GPU SuperPOD Clusters & Topologies
💰 $220K - $460K / year (AI Cloud Infra) | ¥600K - ¥1.5M / year
Rack-scale GPU systems, non-blocking fat-tree fabrics, rail-optimized network design, fault-domain isolation, and 10K+ GPU scaling efficiency optimization.
Benchmark:CoreWeaveCoreWeaveAWSAWSAzureAzureGoogle CloudGoogle CloudOracle CloudOracle CloudLambda LabsLambda LabsCrusoeCrusoeNebiusNebius
GPU ClusterSuperPODFat-TreeRail-OptimizedFault DomainsScaling EfficiencyBespoke AI Cloud
GPU Fleet Operations & Node Health Automation
💰 $195K - $420K / year (Fleet Reliability) | ¥500K - ¥1.3M / year
NVIDIA DCGM continuous telemetry, Multi-Instance GPU (MIG) slicing, bare-metal provisioning, firmware lifecycle automation, and self-healing node drain pipelines.
Benchmark:NVIDIANVIDIACoreWeaveCoreWeaveAWSAWSGoogle CloudGoogle CloudAzureAzure
DCGMMIG SlicingFleet ReliabilityBare-metal ProvisioningFirmware LifecycleNode RemediationSLO Monitoring
2.2

2.2 AI Fabric, Interconnect & Optical Networking

NVLink/NVSwitch fabrics, InfiniBand NDR/XDR, RoCEv2 lossless Ethernet, PFC/ECN congestion control, NCCL collectives tuning, and 800G/1.6T high-speed optics.
Sub-Page
NVLink & InfiniBand / RoCE RDMA Fabrics
💰 $210K - $440K / year (RDMA & AI Fabrics) | ¥550K - ¥1.4M / year
1.8TB/s inter-GPU NVLink, cross-node zero-copy RDMA, Priority Flow Control (PFC), ECN/DCQCN congestion control, NCCL collective tuning, and DPU/SmartNIC hardware offloading.
Benchmark:NVIDIANVIDIAAristaAristaBroadcomBroadcomCiscoCiscoMarvellMarvellAstera LabsAstera LabsCredoCredo
NVLinkInfiniBand NDRRoCEv2RDMA VerbsNCCLPFC / ECNDPU OffloadLossless Ethernet
800G/1.6T Optics & Network Telemetry
💰 $200K - $420K / year (Optical Networks) | ¥500K - ¥1.3M / year
800G/1.6T high-density optical transceivers (OSFP/QSFP-DD), Co-Packaged Optics (CPO), silicon photonics, In-band Network Telemetry (INT), and real-time BER link quality monitoring.
Benchmark:AristaAristaCiscoCiscoCoherentCoherentLumentumLumentumBroadcomBroadcomCredoCredo
800G/1.6T OpticsSilicon PhotonicsCPOOSFPIn-band TelemetryEye DiagramBER Monitoring
2.3

2.3 AI Data Plane, Storage & Checkpointing

Lustre, GPFS, Weka, VAST Data DASE flash architecture, NVMe-oF, GPUDirect Storage (GDS), and sub-minute multi-TB checkpoint persistence.
Sub-Page
Parallel File Systems & Distributed Storage
💰 $195K - $410K / year (AI Storage Systems) | ¥500K - ¥1.25M / year
High-concurrency parallel training I/O: Lustre, IBM GPFS (Storage Scale), Weka filesystem, VAST Data DASE architecture, NVMe-oF, and horizontal POSIX metadata scaling.
Benchmark:VAST DataVAST DataPure StoragePure StorageWekaWekaDDN StorageDDN StorageAWSAWSAzureAzure
LustreGPFS / Storage ScaleWekaVAST DataNVMe-oFPOSIX FilesystemMetadata Scaling
GPUDirect Storage, Tiering & Checkpoint IO
💰 $190K - $400K / year (I/O & Checkpointing) | ¥480K - ¥1.2M / year
NVIDIA GPUDirect Storage (GDS direct-to-GPU IO), multi-tier caching (DRAM/NVMe/Object), sub-minute multi-TB checkpointing, and non-blocking asynchronous persistence pipelines.
Benchmark:VAST DataVAST DataPure StoragePure StorageWekaWekaMinIOMinIOGoogle CloudGoogle CloudAWSAWS
GPUDirect Storage (GDS)Checkpoint IOMulti-tier CachingZero-Copy IOMinIOData Locality
2.4

2.4 GPU Scheduling, Multi-tenancy & Cloud Control Plane

Kubernetes (GPU Operator/Kueue/Volcano), Slurm batch queues, Ray distributed execution, Gang scheduling, MIG multi-tenancy, and FinOps GPU cost governance.
Sub-Page
Kubernetes & Slurm GPU Scheduling & Orchestration
💰 $185K - $390K / year (Cloud K8s & Scheduling) | ¥450K - ¥1.15M / year
NVIDIA GPU Operator automated management, Kueue/Volcano batch queues, Slurm HPC workload management, Gang scheduling, and topology-aware GPU placement.
Benchmark:Google CloudGoogle CloudAWSAWSAzureAzureCoreWeaveCoreWeaveRed HatRed Hat
KubernetesSlurmGPU OperatorKueueVolcanoGang SchedulingTopology-Aware Placement
Ray Distributed Execution, Multi-Tenancy & FinOps
💰 $180K - $380K / year (Ray & Cloud FinOps) | ¥420K - ¥1.1M / year
Ray Core / KubeRay elastic execution graphs, Run:ai dynamic pooling, MIG hardware partitioning, Fair-share quotas, and Spot/preemptible FinOps cost optimization.
Benchmark:AnyscaleAnyscaleRun:aiRun:aiCoreWeaveCoreWeaveGoogle CloudGoogle CloudAWSAWS
Ray / KubeRayRun:aiMulti-tenancyFair-share QuotasFinOpsSpot InstancesGPU Pooling
💼

Career Track Roles

17
💼GPU Infrastructure Engineer
2
💼HPC Engineer
1
💼Cluster Engineer
1
💼Fleet Reliability Engineer
1
💼RDMA Systems Engineer
1
💼Network Engineer
2
💼Network Performance Engineer
2
💼Optical Systems Engineer
1
💼Storage Engineer
2
💼Distributed Systems Engineer
2
💼I/O Performance Engineer
1
💼Platform Engineer
2
💼Scheduler Engineer
2
💼Cloud Infrastructure Engineer
1
💼DevOps / SRE
3
💼FinOps Engineer
1
💼Data Infrastructure Engineer
1
🔄Sub-Domain Sequential Path (4 stages):
2.1

2.1 AI Cluster Engineering & GPU Fleet Operations

Open Sub-Page

10K+ GPU supercomputing clusters, rack-scale topologies, rail-optimized networks, DCGM telemetry, MIG partitioning, and automated fleet self-healing.

10K+ GPU SuperPOD Clusters & Topologies

Rack-scale GPU systems, non-blocking fat-tree fabrics, rail-optimized network design, fault-domain isolation, and 10K+ GPU scaling efficiency optimization.

🏢 Companies
CoreWeaveCoreWeaveAWSAWSAzureAzureGoogle CloudGoogle CloudOracle CloudOracle CloudLambda LabsLambda LabsCrusoeCrusoeNebiusNebius
🛠️ Tech Stack
GPU ClusterSuperPODFat-TreeRail-OptimizedFault DomainsScaling EfficiencyBespoke AI Cloud
💼 Roles & Salary
GPU Infrastructure Engineer、HPC Engineer、Cluster Engineer
💰 $220K - $460K / year (AI Cloud Infra) | ¥600K - ¥1.5M / year

GPU Fleet Operations & Node Health Automation

NVIDIA DCGM continuous telemetry, Multi-Instance GPU (MIG) slicing, bare-metal provisioning, firmware lifecycle automation, and self-healing node drain pipelines.

🏢 Companies
NVIDIANVIDIACoreWeaveCoreWeaveAWSAWSGoogle CloudGoogle CloudAzureAzure
🛠️ Tech Stack
DCGMMIG SlicingFleet ReliabilityBare-metal ProvisioningFirmware LifecycleNode RemediationSLO Monitoring
💼 Roles & Salary
GPU Infrastructure Engineer、Fleet Reliability Engineer、DevOps / SRE
💰 $195K - $420K / year (Fleet Reliability) | ¥500K - ¥1.3M / year
2.2

2.2 AI Fabric, Interconnect & Optical Networking

Open Sub-Page

NVLink/NVSwitch fabrics, InfiniBand NDR/XDR, RoCEv2 lossless Ethernet, PFC/ECN congestion control, NCCL collectives tuning, and 800G/1.6T high-speed optics.

NVLink & InfiniBand / RoCE RDMA Fabrics

1.8TB/s inter-GPU NVLink, cross-node zero-copy RDMA, Priority Flow Control (PFC), ECN/DCQCN congestion control, NCCL collective tuning, and DPU/SmartNIC hardware offloading.

🏢 Companies
NVIDIANVIDIAAristaAristaBroadcomBroadcomCiscoCiscoMarvellMarvellAstera LabsAstera LabsCredoCredo
🛠️ Tech Stack
NVLinkInfiniBand NDRRoCEv2RDMA VerbsNCCLPFC / ECNDPU OffloadLossless Ethernet
💼 Roles & Salary
RDMA Systems Engineer、Network Engineer、Network Performance Engineer
💰 $210K - $440K / year (RDMA & AI Fabrics) | ¥550K - ¥1.4M / year

800G/1.6T Optics & Network Telemetry

800G/1.6T high-density optical transceivers (OSFP/QSFP-DD), Co-Packaged Optics (CPO), silicon photonics, In-band Network Telemetry (INT), and real-time BER link quality monitoring.

🏢 Companies
AristaAristaCiscoCiscoCoherentCoherentLumentumLumentumBroadcomBroadcomCredoCredo
🛠️ Tech Stack
800G/1.6T OpticsSilicon PhotonicsCPOOSFPIn-band TelemetryEye DiagramBER Monitoring
💼 Roles & Salary
Optical Systems Engineer、Network Performance Engineer、Network Engineer
💰 $200K - $420K / year (Optical Networks) | ¥500K - ¥1.3M / year
2.3

2.3 AI Data Plane, Storage & Checkpointing

Open Sub-Page

Lustre, GPFS, Weka, VAST Data DASE flash architecture, NVMe-oF, GPUDirect Storage (GDS), and sub-minute multi-TB checkpoint persistence.

Parallel File Systems & Distributed Storage

High-concurrency parallel training I/O: Lustre, IBM GPFS (Storage Scale), Weka filesystem, VAST Data DASE architecture, NVMe-oF, and horizontal POSIX metadata scaling.

🏢 Companies
VAST DataVAST DataPure StoragePure StorageWekaWekaDDN StorageDDN StorageAWSAWSAzureAzure
🛠️ Tech Stack
LustreGPFS / Storage ScaleWekaVAST DataNVMe-oFPOSIX FilesystemMetadata Scaling
💼 Roles & Salary
Storage Engineer、Distributed Systems Engineer、Data Infrastructure Engineer
💰 $195K - $410K / year (AI Storage Systems) | ¥500K - ¥1.25M / year

GPUDirect Storage, Tiering & Checkpoint IO

NVIDIA GPUDirect Storage (GDS direct-to-GPU IO), multi-tier caching (DRAM/NVMe/Object), sub-minute multi-TB checkpointing, and non-blocking asynchronous persistence pipelines.

🏢 Companies
VAST DataVAST DataPure StoragePure StorageWekaWekaMinIOMinIOGoogle CloudGoogle CloudAWSAWS
🛠️ Tech Stack
GPUDirect Storage (GDS)Checkpoint IOMulti-tier CachingZero-Copy IOMinIOData Locality
💼 Roles & Salary
Storage Engineer、I/O Performance Engineer、Distributed Systems Engineer
💰 $190K - $400K / year (I/O & Checkpointing) | ¥480K - ¥1.2M / year
2.4

2.4 GPU Scheduling, Multi-tenancy & Cloud Control Plane

Open Sub-Page

Kubernetes (GPU Operator/Kueue/Volcano), Slurm batch queues, Ray distributed execution, Gang scheduling, MIG multi-tenancy, and FinOps GPU cost governance.

Kubernetes & Slurm GPU Scheduling & Orchestration

NVIDIA GPU Operator automated management, Kueue/Volcano batch queues, Slurm HPC workload management, Gang scheduling, and topology-aware GPU placement.

🏢 Companies
Google CloudGoogle CloudAWSAWSAzureAzureCoreWeaveCoreWeaveRed HatRed Hat
🛠️ Tech Stack
KubernetesSlurmGPU OperatorKueueVolcanoGang SchedulingTopology-Aware Placement
💼 Roles & Salary
Platform Engineer、Scheduler Engineer、Cloud Infrastructure Engineer、DevOps / SRE
💰 $185K - $390K / year (Cloud K8s & Scheduling) | ¥450K - ¥1.15M / year

Ray Distributed Execution, Multi-Tenancy & FinOps

Ray Core / KubeRay elastic execution graphs, Run:ai dynamic pooling, MIG hardware partitioning, Fair-share quotas, and Spot/preemptible FinOps cost optimization.

🏢 Companies
AnyscaleAnyscaleRun:aiRun:aiCoreWeaveCoreWeaveGoogle CloudGoogle CloudAWSAWS
🛠️ Tech Stack
Ray / KubeRayRun:aiMulti-tenancyFair-share QuotasFinOpsSpot InstancesGPU Pooling
💼 Roles & Salary
Platform Engineer、Scheduler Engineer、FinOps Engineer、DevOps / SRE
💰 $180K - $380K / year (Ray & Cloud FinOps) | ¥420K - ¥1.1M / year