Comment

Comment on How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms [TLDR: 25%]

unpossum@sh.itjust.works ⁨2⁩ ⁨months⁩ ago

GLM 4.5 is from August. Isn’t the real tl;dr that a seven month old open model, which was behind proprietary models at the time, did better than most humans would?

source

Sort:hotnew top

MHard@lemmy.world ⁨2⁩ ⁨months⁩ ago
The task described in this article is asking questions about a document that was provided to the llm in the context.

I would hope that if you give a human a text and ask them to cite facts from it they would do better than 99% correct.

Also, when the tokens exceeded 200k, the llm error rate was higher than 10%

source
- unpossum@sh.itjust.works ⁨2⁩ ⁨months⁩ ago
  
  I would hope that if you give a human a text and ask them to cite facts from it they would do better than 99% correct.
  
  That’s literally what school exams are about, isn’t it?
  
  Token window is a problem for all llms though, that’s not easily solved, but it can be worked around to a certain extent.
  
  source