AI Optimization
What a deterministic score can and cannot carry
A stable number is a precondition for measuring change and it is not a judgement about a business. The difference decides how a score should be used.
What determinism buys
One thing, and it is the necessary thing: the ability to tell whether something changed. If the same inputs produce the same output, then a movement in the number is a movement in the world. Without that, improvement and noise are indistinguishable and every subsequent decision is guesswork.
What it does not buy
Correctness. A deterministic model can be consistently wrong, and stability makes a wrong model more convincing rather than less, because it produces reassuringly steady numbers.
Determinism is a floor. Validity is a separate argument, and it has to be made in public by publishing the weighting so it can be disagreed with.
The distinction in one line: determinism tells you the ruler does not move. It says nothing about whether you are measuring the right thing.
What a score is for
Tracking movement over time on a fixed method. Its primary and best use.
Comparing a business against a defined field on identical criteria. Legitimate when the field and criteria are stated.
Prioritising work. Only via its components, never via the total. A composite tells you something is wrong; the four measures underneath tell you what.
What a score is not for
Comparing across tools. Different definitions, different fields, different weights. A business scoring 72 on one and 41 on another has learned nothing about itself.
Predicting revenue. Several unmeasured steps sit between visibility and money.
Judging quality. It measures whether systems can perceive a business accurately, not whether the business is good.
Why the composite is published at all
Because a single number is the only thing that can be tracked in a line, and because businesses need one figure to answer whether things are moving. It is a summary of a finding, not the finding, and every report puts the four measures next to it for that reason.
The version problem
Improving a scoring model breaks comparability with everything measured before it. So the method carries a version, results record the version that produced them, and comparisons across a change are withheld. The cost is a gap in the trend line. The alternative is a trend line that lies, which is worse and much harder to notice.