TB/PB-scale pretraining curation, MinHash LSH deduplication, benchmark decontamination, human-in-the-loop annotation, and verifiable synthetic data engines.
MinHash/LSH fuzzy deduplication, FastText quality/language scoring, HTML boilerplate removal, benchmark decontamination (preventing test set leakage), and multimodal pair filtering.
Teacher-model Evol-Instruct evolutionary scaling, code execution sandbox validation, rejection sampling, human-in-the-loop QA gates, and data provenance/licensing tracking.