two evaluation methods for long-context ability. (1) Needle-in-a-Haystack: insert one memorable fact (the needle) at a random position inside a long irrelevant text (the haystack), ask for it, and measure precise recall
Acc=total testsneedles retrieved, plotted as a recall heatmap over needle position and context length; (2) Lost-in-the-Middle (Liu et al., 2023): position experiments show models use content from the beginning and end far better than the middle — a U-shaped dip in the middle region, measured with multi-document KV retrieval tasks.