Ask who is winning the AI race and you will get an answer about benchmarks. It is the wrong scoreboard, and it is wrong in a specific way that anyone who has ever graded a forecast will recognise immediately.
There are at least four scoreboards being kept simultaneously. They disagree with each other, they are measuring genuinely different things, and only one of them has historically predicted who ends up in control.
Scoreboard one: benchmarks
The most cited and the least informative.
A benchmark measures performance on a fixed set of problems with known answers. That is useful early, when models fail most of them and the differences are enormous. It stops being useful at exactly the point it becomes interesting, for three reasons.
**Saturation.** When every serious model scores above ninety per cent, the remaining spread is mostly noise plus whatever quirks the test has. A test that nobody fails cannot rank anyone. This is the same defect as a forecast interval so wide that every outcome falls inside it: it can no longer be wrong, and something that cannot be wrong is not telling you anything.
**Contamination.** These systems are trained on enormous crawls of the public internet. Benchmark questions and their answers live on the public internet. Nobody has a clean way to prove a given test set was excluded, and the incentive to look closely is weak.
**Optimisation pressure.** Once a number is public and compared, effort flows toward it. This is not cheating; it is what measurement does to any organisation. But it means the score drifts away from the general capability it was meant to proxy.
None of that makes benchmarks useless. It makes them a floor rather than a ranking - good for "can this system do the thing at all", weak for "which of these is better".
Scoreboard two: revenue
Harder to game and harder to read.
Revenue tells you someone is paying, which is more than a benchmark tells you. What it does not separate is whether the payment reflects durable value or an experimental budget. A large share of enterprise AI spending over the past two years has been pilot spending, and pilots renew at a much lower rate than production systems.
The number worth watching is not revenue. It is net revenue retention - what an existing customer spends this year against last. Growth from new logos measures curiosity. Growth from existing accounts measures whether anything worked.
Scoreboard three: distribution
Historically, the one that decides.
The lesson of every previous platform contest is that the best technology loses to the technology that is already there. Being the default in an operating system, a browser, an office suite or a phone has repeatedly beaten being better, because most users do not choose a tool - they use the one in front of them.
Applied here: a company with a modestly worse model and a billion existing users is in a stronger position than a company with a modestly better model and a signup page. The gap in model quality between serious contenders is currently small and shrinking. The gap in distribution is enormous and widening.
This scoreboard also explains behaviour that looks irrational on the others - free tiers that lose money, aggressive bundling, giving capability away inside products people already open every morning.
Scoreboard four: compute
The least discussed and the most capital-intensive.
Access to chips, to power and to the physical buildings that hold them is a constraint that money alone cannot resolve quickly. Lead times are measured in years. This is the one axis where an incumbent's balance sheet converts directly into position, and it is why several of the largest players are effectively in the infrastructure business regardless of what their marketing says.
It is also the scoreboard most likely to produce a surprise, because capital commitments made against a forecast of demand become extremely awkward if that forecast is wrong.
Why the public conversation uses the worst one
Benchmarks dominate because they are cheap, public, comparable and produce a single number. Revenue is disclosed selectively. Distribution is hard to summarise. Compute is opaque.
So the most visible scoreboard is the least predictive, which is a familiar pattern to anyone who has watched a metric get chosen because it was available rather than because it was informative.
What a scoreboard has to have to mean anything
Three properties, and benchmarks satisfy roughly one and a half of them.
**It has to be possible to fail.** If the test cannot separate good from bad, it is not a test. This is the saturation problem, and it is why a rising score can indicate a broken measure rather than an improving system.
**The claim has to be fixed before the result.** A number reported after the fact by the party being measured is a claim, not evidence. The order matters more than the honesty of the person reporting.
**Failures have to appear alongside successes.** A record that only shows what worked is marketing regardless of how accurate each individual entry is.
We hold ourselves to this, which is the only reason we are entitled to write it down. Our published forecasts state a 50% range; across 56 resolved forecasts the outcome landed inside 47 times. Eighty-four per cent against a target of fifty - a failure, because a range that catches almost everything carries no information. It sits on our public scoring page next to the results that look good, and we published the diagnosis before we had the fix.
The short answer
Nobody is winning on the scoreboard everyone is watching, because that scoreboard has stopped being able to distinguish the players. The contest is being decided on distribution and compute, both of which favour incumbents, and it will become legible on the revenue scoreboard about two years after it is actually settled.
Educational content - not financial advice.