Skip to content

Guides

Comparison rules

Comparing a candidate to a baseline shows the workflow. This page is the reference for the rules a comparison is decided by: which kinds exist, what each asks, what a margin is measured in, and what moves the interval they are read from.

Three kinds

A comparison rule names a kind and a metric. It never takes min or max: a difference is not inferred from a threshold.

KindAsksTakes
superiorityIs the candidate better than the baseline?no margin
non_inferiorityIs the candidate no worse than the baseline by more than the margin?margin, and optionally direction
equivalenceIs the candidate within the margin of the baseline, in both directions?margin

A margin is in the metric's own units. For a rate, 0.05 is five percentage points; for a latency quantile, 5 is five milliseconds.

version: 1
rules:
  - id: not-worse
    kind: non_inferiority
    metric: exact_label
    margin: 0.05
  - id: same-quality
    kind: equivalence
    metric: exact_label
    margin: 0.05
  - id: better
    kind: superiority
    metric: exact_label
  - id: not-slower
    kind: non_inferiority
    metric: latency_p50
    direction: max
    margin: 5

The comparison rules live in their own file, passed with --policy, because release.yaml decides a single run and a threshold rule and a comparison rule read different evidence.

oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

The same rules against two candidates

Two hundred labelled questions. The baseline mislabels twenty of them. Each candidate below was run over the same suite and compared against that baseline with the file above.

A candidate that behaves exactly like the baseline:

Comparison sha256:7d2396c475b94028925930d4e52ae070d9a216a4558e22646169db492757c4de of run_01M3C44TVGT0X067W2H0BZ6GC9 against run_01M3C44SWMXMSQJE17SYHV87PW · 200 paired cases
exact_label: +0.0 points [-2.7, +2.7] · 200 paired · 0 missing · 0 excluded
latency_p50: -0.067 ms [-0.113, -0.001] ms · p50 of per-case differences · 200 paired · 0 missing
Decisions
  not-worse  exact_label  non-inferiority, margin 5.0 points  PASS  lower_bound_above_margin
  same-quality  exact_label  equivalence, margin ±5.0 points  PASS  interval_within_margins
  better  exact_label  superiority  INSUFFICIENT_EVIDENCE  interval_overlaps_zero
  not-slower  latency_p50  maximum increase 5 ms  PASS  lower_bound_above_margin
Gate: BLOCK (exit 3)

A candidate that fixes the twenty:

Comparison sha256:1e36f48662cbbc0af20217a2c69248f90b473bbf0873dc6dd222ccb1895bb33a of run_01M3C44VQ0TFFNW4FRM5Z2D12R against run_01M3C44SWMXMSQJE17SYHV87PW · 200 paired cases
exact_label: +10.0 points [+4.6, +17.9] · 200 paired · 0 missing · 0 excluded
latency_p50: -0.038 ms [-0.089, +0.009] ms · p50 of per-case differences · 200 paired · 0 missing
Decisions
  not-worse  exact_label  non-inferiority, margin 5.0 points  PASS  lower_bound_above_margin
  same-quality  exact_label  equivalence, margin ±5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margins
  better  exact_label  superiority  PASS  difference_above_zero
  not-slower  latency_p50  maximum increase 5 ms  PASS  lower_bound_above_margin
Gate: BLOCK (exit 3)

Read the two gates together. Both block, for opposite reasons. The first candidate is established as equivalent and not established as better; the second is established as better and therefore not established as equivalent. A policy that holds all three kinds for one metric is asking two incompatible questions, and one of them will always be unanswered. Choose the rule that states the claim you are making.

Direction

non_inferiority defaults to direction: min: higher is better, and the rule asks that the candidate not fall more than the margin below the baseline. direction: max is for metrics where lower is better, such as latency, tokens or cost. The rule then reads as a ceiling on an increase, which is how the terminal prints it: maximum increase 5 ms.

superiority and equivalence take no direction. Equivalence bounds both sides already, and superiority asks whether the difference is above zero in the metric's own sense.

What moves the interval

A comparison is paired: each case's candidate result is set against its own baseline result. Two things widen the interval that results, and neither is the change you made.

Cases missing from either side

A case that errored in either run is missing from the pair, and it is bounded rather than dropped: the interval allows every missing case to have gone either way. A smaller suite, with a criterion named ok: the same four disagreements over thirty-four paired cases, first with no missing cases and then with four more that errored in both runs:

ok: +11.8 points [-7.2, +34.3] · 34 paired · 0 missing · 0 excluded
ok: +11.8 points [-24.8, +44.7] · 34 paired · 4 missing · 0 excluded

The estimate is the same, because it is computed from the cases observed on both sides. The interval is not, because four cases nobody saw could each have moved the difference by a whole case. Fix the errors before you add cases: more cases at the same error rate do not close the gap.

A paired ranking difference, such as ROC-AUC between two model versions, cannot bound missing rows this way: it compares the rows both sides scored. A rule on one reads missingness_unbounded until it declares max_missing_fraction, the share of missing rows it accepts as missing at random.

The method

The default interval for a rate difference is a bounded-mean interval over the per-case differences. A policy may name the conditional exact method instead:

version: 1
difference_method: conditional_exact_paired_difference@1
rules:
  - {id: not-worse, metric: exact_label, kind: non_inferiority, margin: 0.05}

The candidate above that is ten points better reads [+4.6, +17.9] under the default and this under the exact method:

exact_label: +10.0 points [+3.5, +15.8] · 200 paired · 0 missing · 0 excluded

Neither is uniformly narrower, which is why the method is chosen in the policy, before any evidence is read, and never by the engine from the data. It applies to rate differences only.

When a rule cannot decide

An INSUFFICIENT_EVIDENCE comparison rule may carry a line underneath it saying what would decide it. The thirty-four-case comparison with nothing missing:

Decisions
  not-worse  ok  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 5 more paired cases would decide it, if the difference holds (39 in total at 12% discordance)
Gate: BLOCK (exit 3)

That estimate assumes the difference and the share of cases that disagreed both hold as more cases arrive. oloproof plan COMPARISON_ID prints the same line with the time the extra cases would take.

When pairs are missing, no size is offered, because the missing pairs are bounded at their worst and more cases would not resolve them. The same comparison with four missing pairs prints no advice line, and the plan says why:

oloproof plan COMPARISON_ID
Comparison sha256:22c1aa4fddc4e96ca77e0509ac9835c2a377ed237ca80556bbf1a54a172239bd: no sample size can be computed for a rule that did not decide.
  not-worse: 4 missing pairs are bounded at their worst, and the advice does not size under missingness; resolve them and compare again

And when the difference itself sits on the wrong side of the margin, more cases would only confirm it, so the advice says that instead:

  overall  ok  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    no sample size would make this PASS: the difference itself (-5.0 points) is outside the margin, so more cases would move it toward FAIL

Where to go next

of an interval.

comparison does not yet decide.

  • Slices covers rules scoped to one part of the suite.