This perspective piece examines why improvements in LLMs' next-word prediction ability have *reduced* their fit as models of human reading behavior, focusing on the role of LLMs' 'superhumanness' in linguistic prediction.
As LLMs have become better at predicting upcoming words, their alignment with human reading behavior has declined — driven by LLMs' vastly larger training data, stronger long-term memory of training examples, and stronger short-term memory compared to humans.
This is a theoretical/opinion piece without new empirical data; claims about memory and training data as drivers of misalignment are proposed but not experimentally verified in this paper.
Researchers using LLM surprisal scores to model human reading should be cautious: more capable LLMs may be *worse* proxies for human linguistic prediction. Building or selecting LLMs with human-like memory constraints may better explain reading behavior.