trillion-scale corpora without dedup cause repeated memorization, wasted compute and benchmark contamination. MinHash+LSH cuts near-duplicate search from
O(n2) to near-linear, scaling to trillions of tokens; PPL filtering removes ~30%-50% of low-quality docs; annealing raises the high-quality share above 50%, gaining several points on downstream benchmarks at the same ~1.4T-token compute budget.