Capacity Estimation & Latency Budgeting is the first-principles quantitative derivation of throughput, storage footprint, network bandwidth, hardware compute costs, and latency SLAs during System Design interviews and engineering capacity planning; it formalizes 5 core derivation equations: 1) Throughput:
Average QPS=86,400sDAU×Requests/User, with
Peak QPS=(2∼5)×Average QPS; 2) Storage:
Daily Storage=DAU×Events/User×Bytes/Event, scaled across 5-year retention with 3x replication factor; 3) Network Bandwidth:
Peak QPS×Payload Bytes×8 bits; 4) GPU Compute Fleet:
GPUs=⌈BatchConcurrency/GPUPeak QPS×Latency/Request⌉; 5) Latency SLA decomposition allocating strict millisecond slices across Network RTT, Feature Store IO, Neural Forward Passes, and Serialization.