Comment on ChatGPT bombs test on diagnosing kids’ medical cases with 83% error rate | It was bad at recognizing relationships and needs selective training, researchers say.

<- View Parent
Maven@lemmy.sdf.org ⁨10⁩ ⁨months⁩ ago

the internet is full of ai generated text now, which is poison to training models. But it’s good at pretending.

This misconception shows up again and again. It’s wishful thinking from people who want to think AI researchers are idiots and AIs are going to kill themselves.

These models aren’t trained on “the internet”. They don’t just thoughtlessly rip everything that’s ever been posted every time they want to make an updated bot. The vast bulk of training data was scraped years ago, predating the current tide of generative muck, and additions are carefully curated to avoid the exact thing you’re talking about. A scrape of the 2018 internet is plenty, and will remain so for years and years.

source
Sort:hotnewtop