How this measures.
Compared with the real model
The same requests go to the provider under test and to the real model. We report the difference.
What a provider says about itself does not count as a result.
The normal spread first
The same provider answers differently from call to call. A difference counts only when it goes beyond that spread.
Each check has its own threshold.
How sure we are
- PROVEN
- Proven — two independent checks agree.
- STRONG
- Strong — one check found a deviation.
- SUSPECTED
- Possible — the deviation is inside the spread.
- INCONCLUSIVE
- Not checked — the check could not run. It does not affect the score.
The score is arithmetic
The trust score is a function of the stored findings. The same findings give the same number.
A judge model writes the summary. It never produces the number.
A proven hidden instruction caps the overall score.
So does a mismatch in who is answering. If the replies do not come from the claimed model, the other checks describe a different model, and the score cannot rise above that finding.
What we do not know
- —Whether the provider answers everyone this way. We measure from one place.
- —What it does right now. Every result carries a date.
- —Anything about what was not checked.
- —Intent. We report the numbers only.
What we do not publish
The checks themselves, the requests we send and the thresholds.
Part of every check group is not published at all.