AI Roadmap/Layer 02 · 02. Compute, Network & Cloud Infrastructure
2.3

2.3 AI Data Plane, Storage & Checkpointing

Lustre, GPFS, Weka, VAST Data DASE flash architecture, NVMe-oF, GPUDirect Storage (GDS), and sub-minute multi-TB checkpoint persistence.

Parallel File Systems & Distributed Storage

High-concurrency parallel training I/O: Lustre, IBM GPFS (Storage Scale), Weka filesystem, VAST Data DASE architecture, NVMe-oF, and horizontal POSIX metadata scaling.

🏢 Companies
VAST DataPure StorageWekaIODDNAWS (FSx)Microsoft Azure
🛠️ Tech Stack
LustreGPFS / Storage ScaleWekaVAST DataNVMe-oFPOSIX FilesystemMetadata Scaling
💼 Roles & Salary
Storage Engineer、Distributed Systems Engineer、Data Infrastructure Engineer
💰 $195K - $410K / year (AI Storage Systems) | ¥500K - ¥1.25M / year
📚 Prerequisites: Distributed Consistency & Replication • Lustre / GPFS Parallel Architecture • NVMe-over-Fabrics (NVMe-oF) Protocols • Metadata Server (MDS) Performance

GPUDirect Storage, Tiering & Checkpoint IO

NVIDIA GPUDirect Storage (GDS direct-to-GPU IO), multi-tier caching (DRAM/NVMe/Object), sub-minute multi-TB checkpointing, and non-blocking asynchronous persistence pipelines.

🏢 Companies
VAST DataPure StorageWekaIOMinIOGoogle CloudAWS
🛠️ Tech Stack
GPUDirect Storage (GDS)Checkpoint IOMulti-tier CachingZero-Copy IOMinIOData Locality
💼 Roles & Salary
Storage Engineer、I/O Performance Engineer、Distributed Systems Engineer
💰 $190K - $400K / year (I/O & Checkpointing) | ¥480K - ¥1.2M / year
📚 Prerequisites: GPUDirect Storage (cuFile API) • Asynchronous Checkpoint Pipelines • Linux Page Cache & Direct I/O Tuning • Object Storage S3 & Data Tiering