Model leaderboards look authoritative, a single ranked list, but understanding how they’re actually built explains why the “best” model changes depending on which leaderboard you check.
Some rankings use human preference votes between anonymous model pairs, others automated benchmark tests with objectively right answers, others real-world usage data. A model ranking first on human preference and mid-pack on a coding benchmark isn’t contradictory, they measure different things.
Our full guide to reading benchmark scores covers a real risk: scores climbing rapidly toward the ceiling on an older benchmark often means it’s become too easy or is showing up in training data, not that the capability gap has closed.
A lab publishing its own model’s results has an obvious incentive to present favorable numbers, whether through selective reporting or a tuned benchmark configuration. Independently run leaderboards carry more weight for this reason.
Check what a leaderboard actually measures before trusting its ranking, prefer independent benchmarks over vendor-published numbers, and test the top candidates against your own task. See Papers With Code for results across many benchmarks.




