Every model release comes with a wall of benchmark percentages. Here’s how to actually read them instead of trusting whichever number a company leads with.
A coding benchmark like Terminal-Bench measures a specific set of coding tasks, not general intelligence. A model scoring higher there says nothing about performance on a different task type. Check what a benchmark actually tests before treating a score as a general ranking.
Watch for saturation. Our coverage of computer-use agents on OSWorld 2.0 is the clearest example: scores climbed from roughly 12% to 85% in about two years, a suspiciously steep curve that turned out to mean the benchmark had become too easy, not that models had solved the underlying problem.
A company reporting its own benchmark results has an obvious incentive to present them favorably. Treat vendor-reported numbers as claims, not verified facts, until independent testing confirms them.
Test a model against your own task, not a demo, and track your real acceptance or edit rate. See Papers With Code for independently tracked leaderboards across most major benchmarks.




