We just released a new research paper: “Language Models Agree With Each Other, Not With Readers”
Many studies have found that language models produce more homogeneous outputs than humans. But the human comparison groups in those studies are usually recruited, instructed, and paid to perform the same task.
We wanted to compare models with people reading naturally, without being told what to find important.
Using Glasp’s public highlighting data, we analyzed:
• 2,523 reader highlight sets
• 120 web documents
• 18 language models
• 11 model vendors
• Models spanning 2024 to 2026
We measured how often two readers or two models selected the same sentences, after controlling for sentence position and length.
The median model pair agreed 2.3 times more than two human readers.
GPT-5.4 and Claude Opus 5, despite coming from rival labs, agreed 5.1 times more than two readers.
We also tested four models released after the original analysis was completed:
• GPT-5.5
• GPT-5.6 Luna
• GPT-5.6 Sol
• GPT-5.6 Terra
This gave us an out-of-sample test of the paper’s main finding.
None of the four models agreed with readers significantly more than human readers agreed with each other. None surpassed the best model in the original panel or approached the estimated crowd-consensus ceiling.
But their agreement with the original panel’s frontier models was extremely high.
Across 12 comparisons, the new models reached a median agreement of +0.224 with the panel’s frontier incumbents. That is higher than the +0.203 agreement between GPT-5.4 and Claude Opus 5.
The newer models did not move clearly beyond the human agreement level.
They did, however, continue to converge strongly with other frontier models.
Another result surprised us.
Newer models are becoming more aligned with human readers, but they are becoming even more aligned with one another. From 2024 to 2026, agreement with readers increased 2.9 times, while agreement between models increased 3.2 times.
Models are becoming more human-like and more alike at the same time.
We also tested whether models simply preferred a different style of sentence. After controlling for sentence length and position, none of the surface features we measured reliably separated model-selected sentences from reader-selected sentences.
Models and readers choose sentences that look similar, but they choose different ones.
This matters for search, summarization, recommendations, research, AI evaluation, and any system that uses multiple models as if they represent independent perspectives.
Using several frontier models may create the appearance of diversity without providing genuinely different views of what matters.
This work was conducted with
@KeiWatanabe17 at
@_Glasp