1Derive the truth table for declaring Model A win, Model B win, or Tie across the 4 combinations of the Bidirectional Swap Test?
2Explain the underlying cause of Position Bias stemming from autoregressive causal attention masking and early-token attention anchoring?
3How does Random Permutation Averaging mitigate position bias when ranking lists of 3+ candidates simultaneously?
4Empirically evaluate the efficacy of prompt debiasing instructions ('Ignore presentation order') on reducing position bias?
5How to estimate a global position bias coefficient on a sub-sample to calibrate large batches without doubling evaluation costs?