Comment on China is attempting to mirror the entire GitHub over to their own servers, users report

<- View Parent
sugar_in_your_tea@sh.itjust.works ⁨4⁩ ⁨months⁩ ago

I disagree that it needs to be explicit. The current law is the fair use doctrine, which generally has more to do with the intended use than specific amounts of the text. The point is that humans should know where that limit is and when they’ve crossed it, with motive being a huge part of it.

I think machines and algorithms should have to abide by a much narrower understanding of “fair use” because they don’t have motive or the ability to Intuit when they’ve crossed the line. So scraping copyrighted works to produce an LLM should probably generally be illegal, imo.

That said, our current copyright system is busted and desperately needs reform. We should be limiting copyright to 14 years (as in the original copyright act of 1790), with an option to explicitly extend for another 14 years. That way LLMs can scrape comment published >28 years ago with no concerns, and most content produced >14 years (esp. forums and social media where copyright extension is incredibly unlikely). That would be reasonable IMO and sidestep most of the issues people have with LLMs.

source
Sort:hotnewtop