Comment

Comment on In Cringe Video, OpenAI CTO Says She Doesn’t Know Where Sora’s Training Data Came From

PoliticallyIncorrect@lemmy.world ⁨11⁩ ⁨months⁩ ago

If you read a book or watch a movie and get inspired by it to create something new and different, it’s plagiarism and copyright infringement?

source

Sort:hotnew top

buffaloseven@fedia.io ⁨11⁩ ⁨months⁩ ago
There’s a long history of this and you might find some helpful information in looking at “transformative use” of copyrighted materials. Google Books is a famous case where the technology company won the lawsuit.

The real problem is that LLMs constantly spit out copyrighted material verbatim. That’s not transformative. And it’s a near-impossible problem to solve while maintaining the utility. Because these things aren’t actually AI, they’re just monstrous statistical correlation databases generated from an enormous data set.

Much of the utility from them will become targeted applications where the training comes from public/owned datasets. I don’t think the copyright case is going to end well for these companies…or at least they’re going to have to gradually chisel away parts of their training data, which will have an outsized impact as more and more AI generated material finds its way into the training data sets.

source
- stephen01king@lemmy.zip ⁨11⁩ ⁨months⁩ ago
  How constantly does it spit out copyrighted material? Is there data on that?
  
  source
  - buffaloseven@fedia.io ⁨11⁩ ⁨months⁩ ago
    There's more and more research starting to happen on it, but I've seen anywhere from 20% to 60% of responses. Here's a recent study where they explicitly try to coerce LLMs to break copyright: https://www.patronus.ai/blog/introducing-copyright-catcher
    
    I don't have the time to grab them right now, but in many of the lawsuits brought forward against companies developing LLMs, their openings contain some statistics gathered on how frequently they infringed by returning copyrighted material.
    
    source
potustheplant@feddit.nl ⁨11⁩ ⁨months⁩ ago
You do realize that AI is just a marketing term, right? None of these models learn, have intelligence or create truly original work. As a matter of fact, if people don’t continue to create original content, these models would stagnate or enter a poisonous feedback loop that would poison themselves with their own erroneous responses.

AIs don’t think. They copy with extra steps.

source
- PoliticallyIncorrect@lemmy.world ⁨11⁩ ⁨months⁩ ago
  I know AI it’s just a marketing term I usually use quotes when I write the AI term, but anyway it isn’t what real human intelillence does too?, you don’t create things from nowhere, usually people use different sources to accomplish a conclusion, I believe it’s exactly what “AI” does, just it speed up the process, instead of spending 30 minutes reading information about a random stuff, you just ask to the “AI” and it does it in 20 seconds, if you need instant answer to something I think it is pretty usable.
  
  I know it doesn’t think by itself but it speed up the process of searching objective stuff at the internet.
  
  source
  - potustheplant@feddit.nl ⁨11⁩ ⁨months⁩ ago
    Except that the information it gives you is often objectively incorrect and it makes up sources (this happened to me a lot of times). And no, it can’t do what a human can. It doesn’t interpret the information it gets and it can’t reach new conclusions based on what it “knows”.
    
    I honestly don’t know how you can even begin to compare an LLM to the human brain.
    
    source