pixel-space diffusion is computationally intractable above
256×256 — attention and feature maps grow quadratically with resolution; the
8× spatial reduction cuts feature-map area to
1/64, lowering training/sampling compute by roughly two orders of magnitude, while the decoder restores visual detail, the encoder absorbs semantic compression, and conditions (text, ControlNet) plug into the latent space uniformly.