The task described in this article is asking questions about a document that was provided to the llm in the context.
I would hope that if you give a human a text and ask them to cite facts from it they would do better than 99% correct.
Also, when the tokens exceeded 200k, the llm error rate was higher than 10%
unpossum@sh.itjust.works 1 week ago
That’s literally what school exams are about, isn’t it?
Token window is a problem for all llms though, that’s not easily solved, but it can be worked around to a certain extent.