Who gets to keep score
- Every measurement layer the web adopted got a public scoreboard before it got an agreed standard. The scoreboard is what forced the definition.
- Agent readiness has audits now and no scoreboard, which means no company can tell whether it is behind.
- A score deserves trust when the method is published, the evidence is per check, the rerun is free, and the conflicts are stated.
- A vendor scoring you on criteria only that vendor can satisfy is running a sales tool, not a benchmark.
- I am building one, and I sell the remediation. You should know that before you read a number I publish.
Five posts ago I said the browser had started grading sites for agents and nobody was grading the industry. This is the end of that argument, and it is the part I have a stake in, so I am going to say what the stake is up front: I sell the work that fixes a bad score. Read what follows knowing that.
Here is what I keep coming back to. Every measurement layer the web has adopted arrived in the same order, and it is not the order anybody plans. Somebody publishes a comparison. The comparison is partial and people argue with it. The arguing produces the definition. The definition becomes the standard. Accessibility went that way. Page speed went that way, badly, twice. Security went that way when browsers began labeling sites insecure and the industry discovered it had opinions about certificates.
The takeaway
The scoreboard forces the standard, not the other way around. Whoever publishes first defines the category, so the question worth asking now is what a scoreboard has to do to deserve being believed.
Four conditions, and why each one exists
The method is published before any result is. Not a methodology page written after the fact to justify a ranking. The full rubric, in advance, so a firm can score itself and get the same answer. I published mine in the second post of this series for exactly that reason.
Every check records its evidence. Not “fails structured data” but the URL and the literal thing found or not found there. The first time somebody calls to say the score is wrong, the evidence line is the only thing standing between a correction and a fight. Twice now the caller has been right.
Rescoring is free. A benchmark that charges to acknowledge a fix is a protection racket with a spreadsheet. Fix it, ask, get rescored, and the new number replaces the old one with both dates visible.
The conflicts are stated. Anyone publishing a score is selling something. Say what. A score is not disqualified by its author having a business, it is disqualified by hiding one.
The vendor scoreboard problem
Watch for the benchmark whose criteria happen to be that vendor's feature list. It scores you low, the remedy is their product, and the score rises when you buy it. That is not measurement, it is a quote with a chart on top.
The tell is checkable. Ask whether a firm could score full marks using nothing but open standards and its own engineering team. If the answer is no, the benchmark is a sales tool. My rubric passes that test on purpose: every point on it can be earned with published standards, a text editor and someone who knows what they are doing.
What a first round honestly looks like
Partial. One industry, a hundred or so firms, conformance measured from the public web, published with the date and the rubric version stamped on it. Not a market study, not a prediction, and no claim about anybody's business quality.
It will be argued with. Some of the arguments will be right and the results will get corrected in public, because the alternative is pretending a first attempt was perfect, which is how a benchmark loses the only thing it has.
And one round proves nothing on its own. A single score is a snapshot. Four rounds is a trend, and the trend is the part nobody can reconstruct after the fact, which is why the only version of this worth doing is the one that keeps a schedule.
Where this leaves you
Score yourself before somebody else does. The rubric is published, the checks are mechanical, and the exercise costs an afternoon. Then fix the failures rather than the number, since the failures are specific and the number is only a summary of them.
And keep the two measurements apart in your head. Conformance is not citation. One says software can read you. The other says a model recommends you. A scoreboard can settle the first honestly. Nobody should claim to settle the second with a single number, and when someone tries, that is the moment to ask what they sell.
I am building the first one. When it publishes, it will carry the method, the evidence and this disclosure attached to it.
Common questions
Why does a public benchmark matter more than a private audit?
Because a number with nothing to compare it against tells you nothing. Scoring 62 is meaningless until you know the field runs from 20 to 80 and your closest competitor is at 71. Private audits give firms a number and no context, which is comfortable and useless. A published comparison is uncomfortable and actionable, and that discomfort is most of the value.
What stops a benchmark from being a marketing weapon?
Four things, and they are all structural rather than promises. The method is published in full before any result is. Every check records the evidence that produced it. Any firm can have its site rescored for free after it fixes something. And whoever runs it states what they sell. Take away any one of those and you have a lead-generation asset wearing a lab coat.
Is it fair to score companies without asking them?
It is if you only measure what is already public and reproducible from the open web, which is what the conformance half of this is. Anyone can rerun it. Judgment about a company's quality, ethics or competence would be a different matter, and none of that belongs in a mechanical score.
Sources
Want to know where you sit before the list exists? That is the head to head read, scored against a real competitor field.