Skip to content

Hướng dẫn

Comparing a candidate to a baseline

Bản dịch của trang này đã lỗi thời, vì vậy nó được hiển thị bằng tiếng Anh.

The workflow the rest of this product exists for: you changed something, and you want to know whether it helped. A comparison pairs two runs case by case and reports the difference with an interval, so a two-point move is never mistaken for a win.

Two runs, paired

oloproof run
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID

Both runs must score the same cases, because the comparison is paired: each case's candidate result is set against its own baseline result, not against the baseline's average. Pairing is what makes a small difference detectable at all — it removes the variance between cases and leaves the variance from the change.

A comparison of two runs over the scaffolded example prints:

Comparison sha256:1c8429df... of run_01M3B2BCNV3GHD3Z8HKDQCV3XK against run_01M3B2BDH8FSN1MMG3MJ92Q0G2 · 30 paired cases
exact_label: +0.0 points [-16.5, +16.5] · 30 paired · 0 missing · 0 excluded

Those two runs are the same system twice, so the true difference is zero and the estimate says so. Read the interval instead: over thirty cases, this suite cannot detect a change smaller than about sixteen points in either direction. That number is a property of the suite, not of the change, and it is knowable before you make any change at all.

Gating on a comparison

A release policy can decide a comparison as well as a run. The rule most teams want is non-inferiority: not "did it get better" but "is it no worse than the baseline by more than a margin I can live with".

rules:
  - id: no-regression
    kind: non_inferiority
    metric: exact_label
    margin: 0.05
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy release.yaml

Against the same two identical runs:

exact_label: +0.0 points [-16.5, +16.5] · 30 paired · 0 missing · 0 excluded
Decisions
  no-regression  exact_label  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

Read that carefully: it is the whole argument. The two systems are identical. The estimate is exactly zero. And the gate blocks, because thirty cases cannot establish that the candidate is within five points of the baseline — the interval reaches to sixteen points worse.

A tool that compared the estimate to the margin would have passed this. It would have passed it for a candidate that was genuinely ten points worse, too, because thirty cases cannot tell those two situations apart.

exit 3 is INSUFFICIENT_EVIDENCE, not failure. The change is not rejected; the evidence has not decided. Gating CI lists the codes.

Equivalence, when you need both directions

Non-inferiority asks whether the candidate is no worse. Equivalence asks whether it is neither worse nor better by more than the margin — which is what you want when replacing a model with a cheaper one and the claim is that nothing changed.

  - id: same-quality
    kind: equivalence
    metric: exact_label
    margin: 0.05

What decides the interval's width

Three things, and only one of them is the change you made.

  • How many cases are paired. Cases missing from either side cannot pair, and the comparison

reports them separately rather than dropping them.

  • How much the cases disagree with each other. A suite whose cases all behave alike gives a

tighter interval than one with a few wild outliers.

  • The confidence level, which the policy sets and defaults to 0.95.

A comparison that cannot decide is a suite that is too small for the margin you chose, and the answer is more cases or a wider margin — not a different reading of the same evidence.

Slices

A suite that declares slices gets a difference for every slice and every metric. They are exploratory, so oloproof compare counts them in one line rather than printing them before the decision:

45 exploratory slice differences not shown; add --slices to list them

--slices prints them all. A slice that a rule decides on is not exploratory and is always shown, and oloproof inspect and oloproof export keep every slice either way.

Asking a judge which answer is better

A comparison measures each run against the suite. Sometimes the question is simply which of two answers is better, and a pairwise judge asks exactly that. Declare one in a file of its own:

criterion: helpfulness
provider: openai
model: gpt-4o-2024-08-06
rubric_text: Prefer the answer that is more correct and more complete.

Then compare two runs over the same suite:

oloproof prefer CANDIDATE_RUN_ID BASELINE_RUN_ID --judge pairwise.yaml

Every pair is asked twice, with the candidate's answer shown first and then second. A case counts for the candidate only when the judge prefers its answer both times, so a judge that simply favours whichever answer comes first produces no winner. How often its verdict followed the slot is printed as its position bias. The net preference carries an interval, and nothing is gated on it yet. Asking again costs nothing: every verdict is cached.

Where to go next

equivalence, direction, missing pairs and the interval method.

  • Clustered cases covers suites whose cases are not independent.
  • Slices covers gating one part of the suite inside a Holm family.
  • Gating CI covers the release policy and the exit codes.
  • Core concepts covers why a denominator that shrinks is how a failing system

reports a rising score.