Get started
Core concepts
Oloproof measures an AI system the way a study measures anything: the same cases, the same judges, the same seeds, run on every change. What it adds to a test suite is that every number arrives with its uncertainty, and a release decision is made on the interval rather than on the estimate.
The chain
Seven links, and each one is a record you can read on its own.
| Link | What it is |
|---|---|
| Observation | What the system did, captured as an execution artifact |
| Measurement | What an evaluator judged about it |
| Inference | A metric with an interval, over a stated denominator |
| Diagnosis | Failures clustered into factors, with counts and assumptions |
| Decision | PASS, FAIL, INSUFFICIENT_EVIDENCE or MANUAL_REVIEW |
| Recommendation | A release action, which is a field and not a fifth state |
| Evidence | The provenance that lets somebody else check all of it |
Each link is stored separately and is independently reusable. Re-running a suite against an unchanged candidate reuses the executions and the judgments it already has, which is why a repeated run costs almost nothing.
The four decision states
There are exactly four, and one of them is the reason this product exists.
- PASS — the interval's lower bound clears the minimum threshold.
- FAIL — the interval's upper bound is below it.
- INSUFFICIENT_EVIDENCE — neither, so the evidence does not decide.
- MANUAL_REVIEW — a person is asked, because the method declined to answer.
For a minimum threshold T and an interval [L, U]: pass iff L >= T, fail iff U < T, otherwise insufficient. Maximum-threshold rules are the mirror image.
INSUFFICIENT_EVIDENCE is not a soft failure. It is the honest answer when a suite is too small or an effect too close to the threshold to separate from noise, and treating it as a pass is the single most common way an evaluation misleads the team running it.
What is not a decision state
RUN_ERROR, PARTIAL and CANCELLED describe what happened to a run, not what was decided about a system. A run that errored carries no quality result at all, and presenting one as a failure tells a developer their change is bad when the truth is that the harness fell over.
The distinction matters most when it is inconvenient. A metric that counts an ungraded case as a zero charges a system's own outages to the model as quality failures, which is a real finding from running this engine against a live assistant.
Denominators
A rate is a number over a denominator, and the denominator is where evaluations go wrong quietly. Oloproof records four counts for every metric — total, eligible, observed and missing — and a case that could not be measured is bounded rather than dropped.
A denominator that shrinks when a system starts failing is how a failing system reports a rising score.
Where to go next
- Writing a suite: what a case, a system and an evaluator are, and how you declare them.
- Gating CI: turning a decision into an exit code.
- Judges: what an LLM judge must clear before it may gate a release.
- Clustered cases: what changes when cases are not independent.