Can AI-generated data make language models worse? New study proposes screening before training

A research paper proposing early detection of synthetic data contamination in language-model training corpora has won the Best Paper Award at ASIACONF 2026 in Pune. Titled “SynthSentry,” the study introduces a corpus-level approach that requires neither access to the generating model nor synthetic labels. The researchers examine how contamination can contribute to model collapse and argue for screening datasets before training begins. The conference received 4,231 submissions, with 263 papers selected through peer review.
Read more at the source

Disclaimer: The content of this post is sourced from external sites and is for informational purposes only. All rights and credits belong to the original authors and publishers.