Comment on ChatGPT 'got absolutely wrecked' by Atari 2600 in beginner's chess match — OpenAI's newest model bamboozled by 1970s logic

<- View Parent
nednobbins@lemm.ee ⁨2⁩ ⁨weeks⁩ ago

I wouldn’t either but that’s exactly what lmsys.org found.

That blog post had ratings between 858 and 1169. Those are slightly higher than the average rating of human users on popular chess sites. Their latest leaderboard shows them doing even better.

lmarena.ai/leaderboard has one of the Gemini models with a rating of 1470. That’s pretty good.

source
Sort:hotnewtop