distribution discrepancy is the central metric of probabilistic ML — cross-entropy loss in classification, the
DKL(q∥p) term in VAEs, student-approximates-teacher in distillation, KL penalties constraining policy drift in RLHF; classic interview questions ("why cross-entropy loss", "prove KL non-negative") all start here.