GQA only shares KV heads, still caching
2Hd values per token; MLA cuts per-token cache to
dc, with
dc=512≪2Hd — DeepSeek-V2 drops its KV cache from about 41.5GB (GQA variant) to about 5.1GB (MLA), a reduction of 80%+; during training, recomputing K/V from
ctKV at prefill also saves substantial activation memory.