upgrades imitation from 'passively fitting expert actions' (BC) to 'actively matching the expert trajectory distribution' — the discriminator supplies dense adversarial feedback at every step, so errors never compound quadratically as in BC; meanwhile max-entropy IRL assumes expert trajectories satisfy
p(τ)∝exp(R(τ)) — higher-reward trajectories are exponentially more likely to be chosen — giving a principled solution space for the reward and ruling out trivial solutions (e.g., constant reward) and IRL's ill-posedness.