Comment

Comment on Largest study of its kind shows AI assistants misrepresent news content 45% of the time – regardless of language or territory

jordanlund@lemmy.world ⁨1⁩ ⁨month⁩ ago

I wish they had broke it out by AI. The article states:

“Gemini performed worst with significant issues in 76% of responses, more than double the other assistants, largely due to its poor sourcing performance.”

But I don’t see that anywhere in the linked PDF of the “full results”.

This sort of study should also be re-done from time to time to track AI version numbers.

source

Sort:hotnew top

Rothe@piefed.social ⁨1⁩ ⁨month⁩ ago
It doesn’t really matter, “AI” is being asked to do a task it was never meant to do. It isn’t good at it, and it will never be good at it.

source
- snooggums@piefed.world ⁨1⁩ ⁨month⁩ ago
  Using an LLM to return accurate information is like using a shoe to hammer a nail.
  
  source
  - athatet@lemmy.zip ⁨1⁩ ⁨month⁩ ago
    Except that a shoe is vaguely hammer ish. More like pounding a screw in with your forehead.
    
    source
  - Rooster326@programming.dev ⁨1⁩ ⁨month⁩ ago
    We’ve all done it?
    
    source
    snooggums@piefed.world ⁨1⁩ ⁨month⁩ ago
    Nope, my soles are too soft.
    
    source
- Cocodapuf@lemmy.world ⁨1⁩ ⁨month⁩ ago
  Wow, way to completely ignore the content of the comment you’re replying to. Clearly, some are better than others… so, how do the others perform? It’s worth knowing before we make assertions.
  
  The excerpt they quoted said:
  
  “Gemini performed worst with significant issues in 76% of responses, more than double the other assistants, largely due to its poor sourcing performance.”
  
  So that implies that “the other assistants” performed more than twice as well, so presumably that means encountering serious issues less than 38% of the time (still not great, but better). But they said “more than double the other assistants”, does that mean double the rate of one of the others or double the average of the others? If it’s an average it would mean that some models probably performed better, while others performed worse.
  
  This was the point, what was reported was insufficient information.
  
  source
  - Rothe@piefed.social ⁨5⁩ ⁨weeks⁩ ago
    Yes, you are a techbro. You suck because your ideas doesn’t take into consideration actual real life. Fuck you.
    
    source
    Cocodapuf@lemmy.world ⁨5⁩ ⁨weeks⁩ ago
    Wow, that’s just incredibly dismissive and rude. And in response to a completely reasonable comment!
    
    Look, forget the whole AI discussion, I don’t care. Here’s the thing, I really like Lemmy. I really like this community and I want to continue using it as a way to have discussions with people about interesting topics. What I don’t want to see is people yelling insults and swearing at any user they disagree with.
    
    Frankly, that behavior is unwelcome. That’s reddit behavior, you can go there if that’s what you want to do.
    
    source
nick@campfyre.nickwebster.dev ⁨1⁩ ⁨month⁩ ago
And also which version of the models. Gemini 2.5 Flash is a completely different experience to 2.5 Pro.

source