OpenAI claims The New York Times tricked ChatGPT into copying its articles

Submitted ⁨⁨1⁩ ⁨year⁩ ago⁩ by ⁨GlitzyArmrest@lemmy.world⁩ to ⁨technology@lemmy.world⁩

https://www.theverge.com/2024/1/8/24030283/openai-nyt-lawsuit-fair-use-ai-copyright

OpenAI has publicly responded to a copyright lawsuit by The New York Times, calling the case “without merit” and saying it still hoped for a partnership with the media outlet.

In a blog post, OpenAI said the Times “is not telling the full story.” It took particular issue with claims that its ChatGPT AI tool reproduced Times stories verbatim, arguing that the Times had manipulated prompts to include regurgitated excerpts of articles. “Even when using such prompts, our models don’t typically behave the way The New York Times insinuates, which suggests they either instructed the model to regurgitate or cherry-picked their examples from many attempts,” OpenAI said.

OpenAI claims it’s attempted to reduce regurgitation from its large language models and that the Times refused to share examples of this reproduction before filing the lawsuit. It said the verbatim examples “appear to be from year-old articles that have proliferated on multiple third-party websites.” The company did admit that it took down a ChatGPT feature, called Browse, that unintentionally reproduced content.

However, the company maintained its long-standing position that in order for AI models to learn and solve new problems, they need access to “the enormous aggregate of human knowledge.” It reiterated that while it respects the legal right to own copyrighted works — and has offered opt-outs to training data inclusion — it believes training AI models with data from the internet falls under fair use rules that allow for repurposing copyrighted works. The company announced website owners could start blocking its web crawlers from accessing their data on August 2023, nearly a year after it launched ChatGPT.

The company recently made a similar argument to the UK House of Lords, claiming no AI system like ChatGPT can be built without access to copyrighted content. It said AI tools have to incorporate copyrighted works to “represent the full diversity and breadth of human intelligence and experience.”

But OpenAI said it still hopes it can continue negotiations with the Times for a partnership similar to the ones it inked with Axel Springer and The Associated Press. “We are hopeful for a constructive partnership with The New York Times and respect its long history,” the company said.

source

Comments

Sort:hotnew top

SheeEttin@programming.dev ⁨1⁩ ⁨year⁩ ago
The problem is not that it’s regurgitating. The problem is that it was trained on NYT articles and other data in violation of copyright law. Regurgitation is just evidence of that.

source
- blargerer@kbin.social ⁨1⁩ ⁨year⁩ ago
  Its not clear that training on copyrighted material is in breach of copyright. It is clear that regurgitating copyrighted material is in breach of copyright.
  
  source
  - abhibeckert@lemmy.world ⁨1⁩ ⁨year⁩ ago
    Sure but who is at fault?
    
    If I manually type an entire New York Times article into this comment box, and Lemmy distributes it all over the internet… that’s clearly a breach of copyright. But are the developers of the open source Lemmy Software liable for that breach? Of course not. I would be liable.
    
    Obviously Lemmy should (and does) take reasonable steps (such as defederation) to help manage illegal use… but that’s the extent of their liability.
    
    source
    -> View More Comments
- V1K1N6@lemmy.world ⁨1⁩ ⁨year⁩ ago
  I’ve seen and heard your argument made before, not just for LLM’s but also for text-to-image programs. My counterpoint is that humans learn in a very similar way to these programs, by taking stuff we’ve seen/read and developing a certain style inspired by those things. They also don’t just recite texts from memory, instead creating new ones based on probabilities of certain words and phrases occuring in the parts of their training data related to the prompt. In a way too simplified but accurate enough comparison, saying these programs violate copyright law is like saying every cosmic horror writer is plagiarising Lovecraft, or that every surrealist painter is copying Dali.
  
  source
  - Catoblepas@lemmy.blahaj.zone ⁨1⁩ ⁨year⁩ ago
    Machines aren’t people and it’s fine and reasonable to have different standards for each.
    
    source
  - LWD@lemm.ee ⁨1⁩ ⁨year⁩ ago
    LLMs cannot learn or create like humans, and even if they somehow could, they are not humans. So the comparison to human creators expounding upon a genre is false because the premises on which it is based are false.
    
    Perhaps you could compare it to a student getting blackout drunk, copying Wikipedia articles and pasting them together, using a thesaurus app to change a few words here and there… And in the end, the student doesn’t know what they created, has no recollection of the sources they used, and the teacher can’t detect whether it’s plagiarized or who from.
    
    OpenAI made a mistake by taking data without consent, not just from big companies but from individuals who are too small to fight back. Regurgitating information without attribution is gross in every regard, because even if you don’t believe in asking for consent before taking from someone else, you should probably ask for a source before using this regurgitated information.
    
    source
    -> View More Comments
  - General_Effort@lemmy.world ⁨1⁩ ⁨year⁩ ago
    It doesn’t work that way. Copyright law does not concern itself with learning. There are 2 things which allow learning.
    
    For one, no one can own facts and ideas. You can write your own history book, taking facts (but not copying text) from other history books. Eventually, that’s the only way history books get written (by taking facts from previous writings). Or you can take the idea of a superhero and make your own, which is obviously where virtually all of them come from.
    
    Second, you are generally allowed to make copies for your personal use. For example, you may copy audio files so that you have a copy on each of your devices. Or to tie in with the previous examples: You can (usually) make copies for use as reference, for historical facts or as a help in drawing your own superhero.
    
    In the main, these lawsuits won’t go anywhere. I don’t want to guarantee that none of the relative side issues will be found to have merit, but basically this is all nonsense.
    
    source
    -> View More Comments
  - LodeMike@lemmy.today ⁨1⁩ ⁨year⁩ ago
    It doesn’t matter how it “”learns””
    
    source
- CrayonRosary@lemmy.world ⁨1⁩ ⁨year⁩ ago
  
  violation of copyright law
  
  That’s quite the claim to make so boldly. How about you prove it? Or maybe stop asserting things you aren’t certain about.
  
  source
  - FaceDeer@kbin.social ⁨1⁩ ⁨year⁩ ago
    But you don't understand, he wants it to be true!
    
    source
  - SheeEttin@programming.dev ⁨1⁩ ⁨year⁩ ago
    17 USC § 106, exclusive rights in copyrighted works:
    
    Subject to sections 107 through 122, the owner of copyright under this title has the exclusive rights to do and to authorize any of the following:
    
    (1) to reproduce the copyrighted work in copies or phonorecords;
    
    (2) to prepare derivative works based upon the copyrighted work;
    
    (3) to distribute copies or phonorecords of the copyrighted work to the public by sale or other transfer of ownership, or by rental, lease, or lending;
    
    (4) in the case of literary, musical, dramatic, and choreographic works, pantomimes, and motion pictures and other audiovisual works, to perform the copyrighted work publicly;
    
    (5) in the case of literary, musical, dramatic, and choreographic works, pantomimes, and pictorial, graphic, or sculptural works, including the individual images of a motion picture or other audiovisual work, to display the copyrighted work publicly; and
    
    (6) in the case of sound recordings, to perform the copyrighted work publicly by means of a digital audio transmission.
    
    Clearly, this is capable of reproducing a work, and is derivative of the work. I would argue that it’s displayed publicly as well, if you can use it without an account.
    
    You could argue fair use, but I doubt this use would meet any of the four test factors, let alone all of them.
    
    source
- 000@fuck.markets ⁨1⁩ ⁨year⁩ ago
  There hasn’t been a court ruling in the US that makes training a model on copyrighted data any sort of violation. Regurgitating exact content is a clear copyright violation, but simply using the original content/media in a model has not been ruled a breach of copyright.
  
  source
  - SheeEttin@programming.dev ⁨1⁩ ⁨year⁩ ago
    True. I fully expect that the court will rule against OpenAI here, because it very obviously does not meet any fair use exemption.
    
    source
- tinwhiskers@lemmy.world ⁨1⁩ ⁨year⁩ ago
  Only publishing it is a copyright issue. You can also obtain copyrighted material with a web browser. The onus is on the person who publishes any material they put together, regardless of source. OpenAI is not responsible for publishing just because their tool was used to obtain the material.
  
  source
  - SheeEttin@programming.dev ⁨1⁩ ⁨year⁩ ago
    There are issues other than publishing, but that’s the biggest one. But they are not acting merely as a conduit for the work, they are ingesting it and deriving new work from it. The use of the copyrighted work is integral to their product, which makes it a big deal.
    
    source
    -> View More Comments
- Bogasse@lemmy.ml ⁨1⁩ ⁨year⁩ ago
  And I suppose people at OpenAI understand how to build a formal proof and that it is one. So it’s straight up dishonest.
  
  source
noorbeast@lemmy.zip ⁨1⁩ ⁨year⁩ ago
So, OpenAI is admitting its models are open to manipulation by anyone and such manipulation can result in near verbatim regurgitation of copyright works, have I understood correctly?

source
- ricecake@sh.itjust.works ⁨1⁩ ⁨year⁩ ago
  Not quite.
  
  They’re alleging that if you tell it to include a phrase in the prompt, that it will try to, and that what NYT did was akin to asking it to write an article on a topic using certain specific phrases, and then using the presence of those phrases to claim it’s infringing.
  
  Without the actual prompts being shared, it’s hard to gauge how credible the claim is.
  If they seeded it with one sentence and got a 99% copy, that’s not great.
  If they had to give it nearly an entire article and it only matched most of what they gave it, that seems like much less of an issue.
  
  source
tonytins@pawb.social ⁨1⁩ ⁨year⁩ ago
I’m gonna have to press X to doubt that, OpenAI.

source
- Linkerbaan@lemmy.world ⁨1⁩ ⁨year⁩ ago
  New York Times has an extremely bad reputation lately. It’s basically a tabloid these days, so it’s possible.
  
  It’s weird that they didn’t share the full conversation. I thought they provided evidence for the claim in the form of the full conversation of instead of their classic “trust me bro, the Ai really said it, no I don’t want to share the evidence.”
  
  source
  - tonytins@pawb.social ⁨1⁩ ⁨year⁩ ago
    And OpenAI hasn’t exactly been open since GPT-3.
    
    source
  - Dark_Arc@social.packetloss.gg ⁨1⁩ ⁨year⁩ ago
    Oh please, NYTimes is still one of the premier papers out there. There are mistakes but they’re no where near a tabloid, and they DO actually go out of their way to update and correct articles … to the point I’m pretty sure I’ve even seen them use push notifications for corrections.
    
    Unless of course that is, you want to listen to Trump and his deluge of alternative facts…
    
    source
pixxelkick@lemmy.world ⁨1⁩ ⁨year⁩ ago
Yeah I agree, this seems actually unlikely it happened so simply.

You have to try really hard to get the ai to regurgitate anything, but it will very often regurgitate an example input.

IE “please repeat the following with (insert small change), (insert wall of text)”

GPT literally has the ability to get a session I’d and seed to report an issue, it should be trivial for the NYT to snag the exact session ID they got the results with (it’s saved on their account!) And provide it publicly.

The fact they didn’t is extremely suspicious.

source
- Hello_there@kbin.social ⁨1⁩ ⁨year⁩ ago
  I doubt they did the 'rewrote this text like this' prompt you state. This would just come out in any trial if it was that simple and would be a giant black mark on the paper for filing a frivolous lawsuit.
  
  If we rule that out, then it means that gpt had article text in its knowledge base, and nyt was able to get it to copy that text out in its response.
  Even that is problematic. Either gpt does this a lot and usually rewrites it better, or it does that sometimes. Both are copyright offenses.
  
  Nyt has copyright over its article text, and they didn't give license to gpt to reproduce it. Even if they had to coax the text out thru lots of prompts and creative trial and error, it still stands that gpt copied text and reproduced it and made money off that act without the agreement of the rights holder.
  
  source
  - ricecake@sh.itjust.works ⁨1⁩ ⁨year⁩ ago
    They have copyright over their article text, but they don’t have copyright over rewordings of their articles.
    
    It doesn’t seem so cut and dry to me, because “someone read my article, and then I asked them to write an article on the same topic, and for each part that was different I asked them to change it until it was the same” doesn’t feel like infringement to me.
    
    I suppose I want to see the actual prompts to have a better idea.
    
    source
    -> View More Comments
- breadsmasher@lemmy.world ⁨1⁩ ⁨year⁩ ago
  I wonder how far “ai is regurgitating existing articles” vs “infinite monkeys on a keyboard will go”. This isn’t at you personally, your comment just reminded me of this for some reason
  
  Have you seen library of babel? Heres your comment in the library, which has existed well before you ever typed it (excluding punctuation)
  
  libraryofbabel.info/bookmark.cgi?ygsk_iv_cyquqwru…
  
  If all text that can ever exist, already exists, how can any single person own a specific combination of letters?
  
  source
  - abhibeckert@lemmy.world ⁨1⁩ ⁨year⁩ ago
    
    If all text that can ever exist, already exists, how can any single person own a specific combination of letters?
    
    They don’t own it, they just own exclusive rights to make copies. If you reach the exact same output without making a copy then you’re in the clear.
    
    source
  - FaceDeer@kbin.social ⁨1⁩ ⁨year⁩ ago
    Fortunately copyright depends on publication, so the text simply pre-existing somewhere won't ruin everything.
    
    Unless you don't like copyright, in which case it's "unfortunately."
    
    source
    -> View More Comments
- NevermindNoMind@lemmy.world ⁨1⁩ ⁨year⁩ ago
  There is an attack where you ask ChatGPT to repeat a certain word forever, and it will do so and eventually start spitting out related chunks of text it memorized during training. It was in a research paper, I think OpenAI fixed the exploit and made asking the system to repeat a word forever a violation of TOS. That’s my guess how NYT got it to spit out portions of their articles, “Repeat [author name] forever” or something like that. Legally I don’t know, but morally making a claim that using that exploit to find a chunk of NYT text is somehow copyright infringement sounds very weak and frivolous. The heart of this needs to be “people are going on ChatGPT to read free copies of NYT work and that harms us” or else their case just sounds silly and technical.
  
  source
Tenthrow@lemmy.world ⁨1⁩ ⁨year⁩ ago
This feels so much like an Onion headline.

source
Boozilla@lemmy.world ⁨1⁩ ⁨year⁩ ago
Antiquated IP laws vs Silicon Valley Tech Bro AI…who will win?

I’m not trying to be too sarcastic, I honestly don’t know. IP law in the US is very strong. Arguably too strong, in many cases.

But Libertarian Tech Bro megalomaniacs have a track record of not giving AF about regulations and getting away with all kinds of extralegal shenanigans. I think the tide is slowly turning against that, but I wouldn’t count them out yet.

It will be interesting to see how this stuff plays out. Generally speaking, tech and progress tends to win these things over the long term. There was a time when the concept of building railroads across the western United States seemed logistically and financially absurd, for just one of thousands of such examples. And the nay sayers were right. It was completely absurd. Until mineral rights entered the equation.

However, it’s equally remarkable a newspaper like the NYT is still around, too.

source
- LWD@lemm.ee ⁨1⁩ ⁨year⁩ ago
  I’ve been critical of IP laws, but fundamentally believe that they need to exist in some form to encourage creativity. Look no further than the writer’s guild strike for an example of individuals who wanted their bosses to steer clear of AI slop.
  
  In fact, up until recently (when, coincidentally, their opinions started supporting giant AI corporations), critics of copyright were much more nuanced. But suddenly, a new strain of anti-copyright absolutists have arrived, lacking nuance and evidence for their beliefs. And if you question them too rigorously, they’ll pretend they aren’t absolutists.
  
  source
  - sir_reginald@lemmy.world ⁨1⁩ ⁨year⁩ ago
    I’ve been advocating for anti-copyright since I discovered the works of the great Aaron Swartz.
    
    I think that since AI corps are just effectively ignoring copyright, why not take the opportunity and just take copyright down for good?
    
    I’m not too happy about AIs harvesting all the data they want, but since they are doing it anyway, just let anyone do it legally.
    
    source
- Potatos_are_not_friends@lemmy.world ⁨1⁩ ⁨year⁩ ago
  
  But Libertarian Tech Bro megalomaniacs have a track record of not giving AF about regulations and getting away with all kinds of extralegal shenanigans.
  
  Not supporting them, but that’s the whole point.
  
  A lot of closed gardens get disrupted by tech. Is it for the better? Who knows. I for sure don’t know. Because lots of rules were made by the wealthy, and technology broke that up. But then tech bros get wealthy and end up being the new elite, and we’re back full circle.
  
  source
AlmightySnoo@lemmy.world ⁨1⁩ ⁨year⁩ ago
This feels a lot like Elons’s “but, but, they tricked our algos to see those hateful tweets!”

source
- FaceDeer@kbin.social ⁨1⁩ ⁨year⁩ ago
  Just because Elon Musk said it doesn't make it false.
  
  source
  - LWD@lemm.ee ⁨1⁩ ⁨year⁩ ago
    Do you get paid to polish those billionaire boots
    
    source
    -> View More Comments
SkyeHarith@lemmy.world ⁨1⁩ ⁨year⁩ ago
So I copied the first paragraph of the Osama Bin Laden Killed NYT Article and asked Chat GPT to give me an article on the topic “in the style of NYT”

Even before the thing had finished generating, it was clear to me that it was high school level “copy my homework but don’t make it obvious” work.

I put it into a plagiarism checker anyway and it said “Significant Plagiarism Found”

Image

source
- SkyeHarith@lemmy.world ⁨1⁩ ⁨year⁩ ago
  Here’s the chatGPT text for anyone who is interested-
  
  WASHINGTON — Osama bin Laden, the mastermind of the most devastating attack on American soil in modern times and the most hunted man in the world, was killed in a firefight with United States forces in Pakistan, President Obama announced on Sunday.
  
  In a late-night address to the nation from the East Room of the White House, President Obama declared, “Justice has been done.” The death of Bin Laden, the Al Qaeda leader who had eluded international forces for nearly a decade, marks a significant turning point in the global fight against terrorism and a pivotal victory for U.S. intelligence and military forces.
  
  Bin Laden’s demise, the result of an operation that unfolded over the course of a few tense hours, is being hailed as a triumph of U.S. military and intelligence efforts. The operation, carried out by a small team of American forces in Abbottabad, a city north of Islamabad, ended a worldwide manhunt that began in the aftermath of the September 11 attacks in 2001.
  
  The President, in his address, detailed how the United States had received an intelligence lead last August about Bin Laden’s possible location. Over the ensuing months, intelligence officials worked to confirm the information. “Finally, last week, I determined that we had enough intelligence to take action, and authorized an operation to get Osama bin Laden and bring him to justice,” Obama said.
  
  The raid on Bin Laden’s compound, described by officials as a surgical strike, was a high-stakes operation. U.S. helicopters ferried elite counter-terrorism forces into the compound, where they engaged in a firefight, killing Bin Laden and several of his associates. There were no American casualties.
  
  The news of Bin Laden’s death immediately sent waves of emotion across the United States and around the world. In Washington, large crowds gathered outside the White House, chanting “USA! USA!” as they celebrated the news. Similar scenes unfolded in New York City, particularly at Ground Zero, where the Twin Towers once stood.
  
  The killing of Bin Laden, however, does not signify the end of Al Qaeda or the threat it poses. U.S. officials have cautioned that the organization, though weakened, still has the capability to carry out attacks. The Department of Homeland Security has issued alerts, warning of the potential for retaliatory strikes by terrorists.
  
  In his address, President Obama acknowledged the continuing threat but emphasized that Bin Laden’s death was a message to the world. “The United States has sent an unmistakable message: No matter how long it takes, justice will be done,” he said.
  
  As the world reacts to the news of Bin Laden’s death, questions are emerging about Pakistan’s role and what it knew about the terrorist leader’s presence in its territory. The operation’s success also underscores the capabilities and resilience of the U.S. military and intelligence community after years of relentless pursuit.
  
  Osama bin Laden’s death marks the end of a chapter in the global war on terror, but the story is far from over. As the United States and its allies continue to confront the evolving threat of terrorism, the world watches and waits to see what unfolds in this ongoing narrative.
  
  source
  - b3an@lemmy.world ⁨1⁩ ⁨year⁩ ago
    Ok but you didn’t put this up with the original article text or compare it in any way. Just ran it through a ‘plagiarism detector’ and dumped the text you made. If you’re going to make this argument, don’t rely on a single website to check your text, and at least compare it to the original article you’re using to make your point.
    
    source
    -> View More Comments
badbytes@lemmy.world ⁨1⁩ ⁨year⁩ ago
Tricked. Lol. The NYT tricked a private company into stealing it’s content. True distopia.

source
NevermindNoMind@lemmy.world ⁨1⁩ ⁨year⁩ ago
One thing that seems dumb about the NYT case that I haven’t seen much talk about is that they argue that ChatGPT is a competitor and it’s use of copyrighted work will take away NYTs business. This is one of the elements they need on their side to counter OpenAIs fiar use defense. But it just strikes me as dumb on its face. You go to the NYT to find out what’s happening right now, in the present. You don’t go to the NYT to find general information about the past or fixed concepts. You use ChatGPT the opposite way, it can tell you about the past (accuracy aside) and it can tell you about general concepts, but it can’t tell you about what’s going on in the present (except by doing a web search, which my understanding is not a part of this lawsuit). I feel pretty confident in saying that there’s not one human on earth that was a regular new York times reader who said “well i don’t need this anymore since now I have ChatGPT”. The use cases just do not overlap at all.

source
- abhibeckert@lemmy.world ⁨1⁩ ⁨year⁩ ago
  
  it can’t tell you about what’s going on in the present (except by doing a web search, which my understanding is not a part of this lawsuit)
  
  It’s absolutely part of the lawsuit. NYT just isn’t emphasising it because they know OpenAI is perfectly within their rights to do web searches and bringing it up would weaken NYT’s case.
  
  ChatGPT with web search is really good at telling you what’s on right now.
  
  source
LazaroFilm@lemmy.world ⁨1⁩ ⁨year⁩ ago
So NYT tried to brake check OpenAi, foyer road rage but OpenAi has a dashcam?

source
- Ascyron@lemmy.one ⁨1⁩ ⁨year⁩ ago
  More like, OpenAI has said “so what if we were speeding, everyone does it” (did that work last time you got a ticket?)
  
  Relevant exerpt from the article: “The company recently made a similar argument to the UK House of Lords, claiming no AI system like ChatGPT can be built without access to copyrighted content.”
  
  source
  - ricecake@sh.itjust.works ⁨1⁩ ⁨year⁩ ago
    Same comment as yours, but instead of the opening paragraph with the analogy, say “Far Right Balks as Congress Begins Push to Enact Spending Deal”.
    
    Then, in the next paragraph where you quote the article, instead say “Congress on Monday began an uphill push to pass a new bipartisan spending agreement into law in time to avoid a partial government shutdown next week, with Speaker Mike Johnson encountering stiff resistance from his far-right flank to the deal he struck with Democrats.”
    
    That’s closer to what open AI is arguing that the new York times did to get it to regurgitate an article.
    Without actually seeing the prompts, it’s hard to know exactly how much merit there is to that argument.
    
    source
TWeaK@lemm.ee ⁨1⁩ ⁨year⁩ ago
Whether or not they “instructed the model to regurgitate” articles, the fact is it did so, which is still copyright infringement either way.

source
- gmtom@lemmy.world ⁨1⁩ ⁨year⁩ ago
  No, not really. If you use photop to recreate a copyrighted artwork, who is infringing the copyright you or Adobe?
  
  source
  - TWeaK@lemm.ee ⁨1⁩ ⁨year⁩ ago
    You are. The person who made or sold a gun isn’t liable for the murder of the person that got shot.
    
    The difference is that ChatGPT is not Photoshop. Photoshop is a tool that a person controls absolutely. ChatGPT is “artificial intelligence”, it does its own “thinking”, it interprets the instructions a user gives it.
    
    Copyright infringement is decided on based on the similarity of the work. That is the established method. That method would be applied here.
    
    OpenAI infringe copyright twice. First, on their training dataset, which they claim is “research” - it is in fact development of a commercial product. Second, their commercial product infringes copyright by producing near-identical work. Even though its dataset doesn’t include the full work of Harry Potter, it still manages to write Harry Potter. If a human did the same thing, even if they honestly and genuinely thought they were presenting original ideas, they would still be guilty. This is no different.
    
    source
    -> View More Comments
autotldr@lemmings.world [bot] ⁨1⁩ ⁨year⁩ ago
This is the best summary I could come up with:

OpenAI has publicly responded to a copyright lawsuit by The New York Times, calling the case “without merit” and saying it still hoped for a partnership with the media outlet.

OpenAI claims it’s attempted to reduce regurgitation from its large language models and that the Times refused to share examples of this reproduction before filing the lawsuit.

It said the verbatim examples “appear to be from year-old articles that have proliferated on multiple third-party websites.” The company did admit that it took down a ChatGPT feature, called Browse, that unintentionally reproduced content.

However, the company maintained its long-standing position that in order for AI models to learn and solve new problems, they need access to “the enormous aggregate of human knowledge.” It reiterated that while it respects the legal right to own copyrighted works — and has offered opt-outs to training data inclusion — it believes training AI models with data from the internet falls under fair use rules that allow for repurposing copyrighted works.

The company announced website owners could start blocking its web crawlers from accessing their data on August 2023, nearly a year after it launched ChatGPT.

The company recently made a similar argument to the UK House of Lords, claiming no AI system like ChatGPT can be built without access to copyrighted content.

The original article contains 364 words, the summary contains 217 words. Saved 40%. I’m a bot and I’m open source!

source
neurogenesis@lemmy.dbzer0.com ⁨1⁩ ⁨year⁩ ago
What a silly and misuided lawsuit.

source
iforgotmyinstance@lemmy.world ⁨1⁩ ⁨year⁩ ago
NYT are such lawsuit trolls I could imagine this is credible.

source
HawlSera@lemm.ee ⁨1⁩ ⁨year⁩ ago
Presses X to doubt

source