Natanox@discuss.tchncs.de 3 days ago
My question would be if it uses stuff like CommonCrawl as training material. Given their size I assume they used anything incl. unethical training data, but at least admit it?
The only models I ever found that even tried to only resort to ethically obtained data, being FOSS etc. were the tiny ones from PleIAs. And as expected they’re completely useless.
So far I concluded that an “ideologically sound” LLM is impossible due to lack of training data. Unless your ideology allows to steal stuff.
melfie@lemmy.zip 3 days ago
I’m assuming a lot of its training data is synthetic and distilled from Chinese models that were themselves trained from pirated data and distilled from American models trained on pirated data. It would be quite remarkable if the training data involved no piracy whatsoever. Then again, it’s open source, so I suppose it would be essentially reversing a reverse Robin Hood.
I mean, they “plan” to publish that info, right?
I’d bet they used at least CommonCrawl (with it being the majority of data), arXiv, Wikipedia and the set containing all of Github. Probably not a lot of distillation.
CommonCrawl is one of the reasons small websites and social instances get DDoS’ed by rules-ignoring AI crawlers. Wikipedia data basically always gets used without paying them. Github… well, it’s a prime example of how they broke millions of licenses.
There’s also other stuff commonly used. Don’t get me started on the training sets for image generation. It’s beyond disgusting (and I’m not even talking just about theft and cultural destruction at this point).
This stuff is a bottomless pit, and if those (F)OSS models really aim to be up at the top they’ll have to break every possible moral, ethical and legal rule just like everyone else. Even more so if they omit closed-source training sets.