Skip to content

Get started

Core concepts

Oloproof measures an AI system the way a study measures anything: the same cases, the same judges, the same seeds, run on every change. What it adds to a test suite is that every number arrives with its uncertainty, and a release decision is made on the interval rather than on the estimate.

The chain

Seven links, and each one is a record you can read on its own.

LinkWhat it is
ObservationWhat the system did, captured as an execution artifact
MeasurementWhat an evaluator judged about it
InferenceA metric with an interval, over a stated denominator
DiagnosisFailures clustered into factors, with counts and assumptions
DecisionPASS, FAIL, INSUFFICIENT_EVIDENCE or MANUAL_REVIEW
RecommendationA release action, which is a field and not a fifth state
EvidenceThe provenance that lets somebody else check all of it

Each link is stored separately and is independently reusable. Re-running a suite against an unchanged candidate reuses the executions and the judgments it already has, which is why a repeated run costs almost nothing.

The four decision states

There are exactly four, and one of them is the reason this product exists.

  • PASS — the interval's lower bound clears the minimum threshold.
  • FAIL — the interval's upper bound is below it.
  • INSUFFICIENT_EVIDENCE — neither, so the evidence does not decide.
  • MANUAL_REVIEW — a person is asked, because the method declined to answer.

For a minimum threshold T and an interval [L, U]: pass iff L >= T, fail iff U < T, otherwise insufficient. Maximum-threshold rules are the mirror image.

INSUFFICIENT_EVIDENCE is not a soft failure. It is the honest answer when a suite is too small or an effect too close to the threshold to separate from noise, and treating it as a pass is the single most common way an evaluation misleads the team running it.

What is not a decision state

RUN_ERROR, PARTIAL and CANCELLED describe what happened to a run, not what was decided about a system. A run that errored carries no quality result at all, and presenting one as a failure tells a developer their change is bad when the truth is that the harness fell over.

The distinction matters most when it is inconvenient. A metric that counts an ungraded case as a zero charges a system's own outages to the model as quality failures, which is a real finding from running this engine against a live assistant.

Denominators

A rate is a number over a denominator, and the denominator is where evaluations go wrong quietly. Oloproof records four counts for every metric — total, eligible, observed and missing — and a case that could not be measured is bounded rather than dropped.

A denominator that shrinks when a system starts failing is how a failing system reports a rising score.

Where to go next

  • Writing a suite: what a case, a system and an evaluator are, and how you declare them.
  • Gating CI: turning a decision into an exit code.
  • Judges: what an LLM judge must clear before it may gate a release.
  • Clustered cases: what changes when cases are not independent.