Can AI-generated data make language models worse? New study proposes screening before training
A research paper proposing early detection of synthetic data contamination in language-model training corpora has won the Best Paper Award at ASIACONF 2026 in Pune. Titled “SynthSentry,” the study introduces a corpus-level approach that requires neither access to the generating model nor synthetic labels. The researchers examine how contamination can contribute to model collapse and…
