CogGym Measures the Gap Between Models and Human Judgments
A unified evaluation of 50 language models across 258 cognitive experiments finds improving model--human fit, but a substantial gap from human reliability remains.
Sep 22, 20263 min2609.21259