the internet is full of ai generated text now, which is poison to training models. But it’s good at pretending.
This misconception shows up again and again. It’s wishful thinking from people who want to think AI researchers are idiots and AIs are going to kill themselves.
These models aren’t trained on “the internet”. They don’t just thoughtlessly rip everything that’s ever been posted every time they want to make an updated bot. The vast bulk of training data was scraped years ago, predating the current tide of generative muck, and additions are carefully curated to avoid the exact thing you’re talking about. A scrape of the 2018 internet is plenty, and will remain so for years and years.