Guides
Comparison rules
Comparing a candidate to a baseline shows the workflow. This page is the reference for the rules a comparison is decided by: which kinds exist, what each asks, what a margin is measured in, and what moves the interval they are read from.
Three kinds
A comparison rule names a kind and a metric. It never takes min or max: a difference is not inferred from a threshold.
| Kind | Asks | Takes |
|---|---|---|
| superiority | Is the candidate better than the baseline? | no margin |
| non_inferiority | Is the candidate no worse than the baseline by more than the margin? | margin, and optionally direction |
| equivalence | Is the candidate within the margin of the baseline, in both directions? | margin |
A margin is in the metric's own units. For a rate, 0.05 is five percentage points; for a latency quantile, 5 is five milliseconds.
version: 1
rules:
- id: not-worse
kind: non_inferiority
metric: exact_label
margin: 0.05
- id: same-quality
kind: equivalence
metric: exact_label
margin: 0.05
- id: better
kind: superiority
metric: exact_label
- id: not-slower
kind: non_inferiority
metric: latency_p50
direction: max
margin: 5The comparison rules live in their own file, passed with --policy, because release.yaml decides a single run and a threshold rule and a comparison rule read different evidence.
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlThe same rules against two candidates
Two hundred labelled questions. The baseline mislabels twenty of them. Each candidate below was run over the same suite and compared against that baseline with the file above.
A candidate that behaves exactly like the baseline:
Comparison sha256:7d2396c475b94028925930d4e52ae070d9a216a4558e22646169db492757c4de of run_01M3C44TVGT0X067W2H0BZ6GC9 against run_01M3C44SWMXMSQJE17SYHV87PW · 200 paired cases
exact_label: +0.0 points [-2.7, +2.7] · 200 paired · 0 missing · 0 excluded
latency_p50: -0.067 ms [-0.113, -0.001] ms · p50 of per-case differences · 200 paired · 0 missing
Decisions
not-worse exact_label non-inferiority, margin 5.0 points PASS lower_bound_above_margin
same-quality exact_label equivalence, margin ±5.0 points PASS interval_within_margins
better exact_label superiority INSUFFICIENT_EVIDENCE interval_overlaps_zero
not-slower latency_p50 maximum increase 5 ms PASS lower_bound_above_margin
Gate: BLOCK (exit 3)A candidate that fixes the twenty:
Comparison sha256:1e36f48662cbbc0af20217a2c69248f90b473bbf0873dc6dd222ccb1895bb33a of run_01M3C44VQ0TFFNW4FRM5Z2D12R against run_01M3C44SWMXMSQJE17SYHV87PW · 200 paired cases
exact_label: +10.0 points [+4.6, +17.9] · 200 paired · 0 missing · 0 excluded
latency_p50: -0.038 ms [-0.089, +0.009] ms · p50 of per-case differences · 200 paired · 0 missing
Decisions
not-worse exact_label non-inferiority, margin 5.0 points PASS lower_bound_above_margin
same-quality exact_label equivalence, margin ±5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margins
better exact_label superiority PASS difference_above_zero
not-slower latency_p50 maximum increase 5 ms PASS lower_bound_above_margin
Gate: BLOCK (exit 3)Read the two gates together. Both block, for opposite reasons. The first candidate is established as equivalent and not established as better; the second is established as better and therefore not established as equivalent. A policy that holds all three kinds for one metric is asking two incompatible questions, and one of them will always be unanswered. Choose the rule that states the claim you are making.
Direction
non_inferiority defaults to direction: min: higher is better, and the rule asks that the candidate not fall more than the margin below the baseline. direction: max is for metrics where lower is better, such as latency, tokens or cost. The rule then reads as a ceiling on an increase, which is how the terminal prints it: maximum increase 5 ms.
superiority and equivalence take no direction. Equivalence bounds both sides already, and superiority asks whether the difference is above zero in the metric's own sense.
What moves the interval
A comparison is paired: each case's candidate result is set against its own baseline result. Two things widen the interval that results, and neither is the change you made.
Cases missing from either side
A case that errored in either run is missing from the pair, and it is bounded rather than dropped: the interval allows every missing case to have gone either way. A smaller suite, with a criterion named ok: the same four disagreements over thirty-four paired cases, first with no missing cases and then with four more that errored in both runs:
ok: +11.8 points [-7.2, +34.3] · 34 paired · 0 missing · 0 excludedok: +11.8 points [-24.8, +44.7] · 34 paired · 4 missing · 0 excludedThe estimate is the same, because it is computed from the cases observed on both sides. The interval is not, because four cases nobody saw could each have moved the difference by a whole case. Fix the errors before you add cases: more cases at the same error rate do not close the gap.
A paired ranking difference, such as ROC-AUC between two model versions, cannot bound missing rows this way: it compares the rows both sides scored. A rule on one reads missingness_unbounded until it declares max_missing_fraction, the share of missing rows it accepts as missing at random.
The method
The default interval for a rate difference is a bounded-mean interval over the per-case differences. A policy may name the conditional exact method instead:
version: 1
difference_method: conditional_exact_paired_difference@1
rules:
- {id: not-worse, metric: exact_label, kind: non_inferiority, margin: 0.05}The candidate above that is ten points better reads [+4.6, +17.9] under the default and this under the exact method:
exact_label: +10.0 points [+3.5, +15.8] · 200 paired · 0 missing · 0 excludedNeither is uniformly narrower, which is why the method is chosen in the policy, before any evidence is read, and never by the engine from the data. It applies to rate differences only.
When a rule cannot decide
An INSUFFICIENT_EVIDENCE comparison rule may carry a line underneath it saying what would decide it. The thirty-four-case comparison with nothing missing:
Decisions
not-worse ok non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 5 more paired cases would decide it, if the difference holds (39 in total at 12% discordance)
Gate: BLOCK (exit 3)That estimate assumes the difference and the share of cases that disagreed both hold as more cases arrive. oloproof plan COMPARISON_ID prints the same line with the time the extra cases would take.
When pairs are missing, no size is offered, because the missing pairs are bounded at their worst and more cases would not resolve them. The same comparison with four missing pairs prints no advice line, and the plan says why:
oloproof plan COMPARISON_IDComparison sha256:22c1aa4fddc4e96ca77e0509ac9835c2a377ed237ca80556bbf1a54a172239bd: no sample size can be computed for a rule that did not decide.
not-worse: 4 missing pairs are bounded at their worst, and the advice does not size under missingness; resolve them and compare againAnd when the difference itself sits on the wrong side of the margin, more cases would only confirm it, so the advice says that instead:
overall ok non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
no sample size would make this PASS: the difference itself (-5.0 points) is outside the margin, so more cases would move it toward FAILWhere to go next
- Comparing a candidate to a baseline covers the workflow and the first reading
of an interval.
- Clustered cases covers suites whose cases are not independent, which a
comparison does not yet decide.
- Slices covers rules scoped to one part of the suite.