The New York Times sues OpenAI and Microsoft for copyright infringement

Submitted ⁨⁨1⁩ ⁨year⁩ ago⁩ by ⁨L4s@lemmy.world [bot]⁩ to ⁨technology@lemmy.world⁩

https://edition.cnn.com/2023/12/27/tech/new-york-times-sues-openai-microsoft/index.html

The New York Times sues OpenAI and Microsoft for copyright infringement::The New York Times has sued OpenAI and Microsoft for copyright infringement, alleging that the companies’ artificial intelligence technology illegally copied millions of Times articles to train ChatGPT and other services to provide people with information – technology that now competes with the Times.

source

Comments

Sort:hotnew top

phoneymouse@lemmy.world ⁨1⁩ ⁨year⁩ ago
There is something wrong when search and AI companies extract all of the value produced by journalism for themselves. Sites like Reddit and Lemmy also have this issue. I’m not sure what the solution is. I don’t like the idea of a web full of paywalls, but I also don’t like the idea of all the profit going to the ones who didn’t create the product.

source
- kromem@lemmy.world ⁨1⁩ ⁨year⁩ ago
  What’s the value of old journalism?
  
  It’s a product where the value curve is heavily weighted towards recency.
  
  In theory, the greatest value theft is when the AP writes a piece and two dozen other ‘journalists’ copy the thing changing the text just enough not to get sued. Which is completely legal, but what effectively killed investigative journalism.
  
  A LLM taking years old articles and predicting them until it can effectively learn relationships between language itself and events described in those articles isn’t some inherent value theft.
  
  It’s not the training that’s the problem, it’s the application of the models that needs policing.
  
  Like if someone took a LLM, fed it recently published news stories, and had it rewrite them just differently enough that no one needed to visit the original publisher.
  
  Even if we have it legal for humans to do that (which really we might want to revisit, or at least create a special industry specific restriction regarding), maybe we should have different rules for the models.
  
  But to try to claim a LLM that’s allowing coma patients to communicate or to problem solve self-driving algorithms or to diagnose medical issues is stealing the value of old NYT articles in its doing so is not really an argument I see much value in.
  
  source
  - jacksilver@lemmy.world ⁨1⁩ ⁨year⁩ ago
    Except no one is claiming that LLMs are the problem, they’re claiming GPT, or more specifically GPTs training data, is the problem. Transformer models still have a lot of potential, but the question the NYT is asking is “can you just takes anyone else’s work to train them”.
    
    source
    -> View More Comments
  - ChucklesMacLeroy@lemmy.world ⁨1⁩ ⁨year⁩ ago
    Really gave me a whole new perspective. Thanks for that.
    
    source
- AllonzeeLV@lemmy.world ⁨1⁩ ⁨year⁩ ago
  
  but I also don’t like the idea of all the profit going to the ones who didn’t create the product.
  
  Should… should we tell him?
  
  source
  - kilgore_trout@feddit.it ⁨1⁩ ⁨year⁩ ago
    Tell them instead of mocking them.
    
    Yes, “that’s how the world works”. But doesn’t mean we should stop trying to change it.
    
    source
- Kecessa@sh.itjust.works ⁨1⁩ ⁨year⁩ ago
  The solution is imposing to these companies the responsibility of tracking the profit per media, tax them and redistribute that money based on the tracking info. They’re able to track all the pages you visit, it’s complete bullshit when they say they don’t know how much they make for each places their ads are displayed.
  
  source
- Boiglenoight@lemmy.world ⁨1⁩ ⁨year⁩ ago
  AI training is piracy by another name.
  
  source
  - uriel238@lemmy.blahaj.zone ⁨1⁩ ⁨year⁩ ago
    Elaborate. Consumption of copyrighted materials is normal use whether by a human or a machine.
    
    source
    -> View More Comments
- DogWater@lemmy.world ⁨1⁩ ⁨year⁩ ago
  Ai isn’t creating the product. It consumed it.
  
  source
LainOfTheWired@lemy.lol ⁨1⁩ ⁨year⁩ ago
My question is how is an AI reading a bunch of articles any different from a human doing it. With this logic no one would legally be able to write an article as they are using bits of other peoples work they read that they learnt to write a good article with.

They are both making money with parts of other peoples work.

source
- hansl@lemmy.world ⁨1⁩ ⁨year⁩ ago
  It was thought that the LLM wouldn’t keep the actual data internally verbatim. If you can memorize an article, and recite it to everyone free of charge, technically it’s plagiarism. Same if you sing a song to a crowd when you don’t have the rights.
  
  The Google research (and other discovery) proved that you can actually extract verbatim training data from a LLM. Which has a lot of implications for copyright.
  
  source
- MirthfulAlembic@lemmy.world ⁨1⁩ ⁨year⁩ ago
  The physical limitations are an important difference. A human can only read and remember so much material. With AI, you can scale that exponentially with more compute resources. Frankly, IP law was not written with this possibility in mind and needs to be updated to find a balance.
  
  source
- JonEFive@midwest.social ⁨1⁩ ⁨year⁩ ago
  Let me ask you this: when have you ever seen ChatGPT cite its sources and give appropriate credit to the original author?
  
  If I were to just read the NYT and make money by simply summarizing articles and posting those summaries on my own website without adding anything to it like my own commentary and without giving credit to the author, that would rightfully be considered plagiarism.
  
  This is a really interesting conundrum though. I would argue that AI isn’t capable of original thought the way that humans are and therefore AI creators must provide due compensation to the authors and artists whose data they used.
  
  AI is only giving back some amalgamation of words and concepts that it has been trained on. You might say that humans do the same, but that isn’t exactly true. The human brain is a funny thing. It can forget, it can misremember. It can manipulate. It can exaggerate. It can plan. It can have irrational or emotional responses. AI can’t really do those things on its own. It’s just mimicking human behavior at best.
  
  Most importantly to me though, AI is not capable of spontaneous thought. It is only capable of providing information that it has been trained on and only when prompted.
  
  source
  - thru_dangers_untold@lemm.ee ⁨1⁩ ⁨year⁩ ago
    There is evidence to suggest originality, such as DeepMind’s solution to the cap set problem.
    
    www.nature.com/articles/s41586-023-06924-6
    
    On the other hand LLM’s have some incredible text compression abilities
    
    arxiv.org/abs/2308.07633
    
    I’m pretty sure there is copyright infringement going on by the letter of the law. But I also think the world would be better off if copyright laws were a bit more loose. Not wild-west anything-goes libertarianism, but more open than the current state.
    
    source
    -> View More Comments
  - General_Effort@lemmy.world ⁨1⁩ ⁨year⁩ ago
    
    Let me ask you this: when have you ever seen ChatGPT cite its sources and give appropriate credit to the original author?
    
    Bing chat now does that by default. Normally you have to prompt that manually.
    
    If I were to just read the NYT and make money by simply summarizing articles and posting those summaries on my own website without adding anything to it like my own commentary and without giving credit to the author, that would rightfully be considered plagiarism.
    
    No. It would be considered journalism. If you read the news a bit, you will find that they reference the output of other news corporations quite a bit. If your preferred news source does not do that, then they simply don’t cite their sources.
    
    source
    -> View More Comments
- BURN@lemmy.world ⁨1⁩ ⁨year⁩ ago
  An AI does not learn like a human does. Therefore the same laws and principles can’t be applied to computer “learning” as can be to human learning.
  
  They’re fundamentally different uses of the material.
  
  source
- myfavouritename@lemmy.world ⁨1⁩ ⁨year⁩ ago
  I think the important difference in this case is like the difference between a human enjoying a song that they hear being performed vs a company recording a song that someone is performing and then replaying that song on demand for paying customers.
  
  source
  - d3Xt3r@lemmy.nz ⁨1⁩ ⁨year⁩ ago
    Except, it’s not replaying those song exactly,
    
    not even in their entirety. It’s taking a few notes from here and there and effectively playing a “new” song - which isn’t all that different from a human artist who is “inspired” by the works of other artists and produces a new work in the same genre.
    
    source
- topinambour_rex@lemmy.world ⁨1⁩ ⁨year⁩ ago
  The main difference being the volume. An example I like is how Google trained his gaming AI to starcraft 2. This AI was able to beat high ranked professional gamers. It was trained by watching a century of games.
  
  source
burliman@lemmy.today ⁨1⁩ ⁨year⁩ ago
Reminds me of Nokia suing Apple (two waves), Blockbuster suing Netflix, and Yahoo suing Facebook. Threatened, declining company suing a disruptor is what we can expect will always happen I guess. Will be nice to see this stuff finally tested in court though.

source