Guides
Slices
A global rate can hold steady while one part of the suite collapses. A slice is a declared part of the suite, measured on its own, so the collapse is visible. Slices come with two disciplines, because looking at many parts of a suite is how an evaluation finds noise and calls it a finding.
Declaring slices
A slice is named in oloproof.yaml. The common kind reads a key from each case's metadata:
slices: [metadata.lang, metadata.topic]
min_slice_support: 30{"id": "c000", "input": {"lang": "en", "topic": "billing", "hard": false, "n": 0}, "expected": {"label": "yes"}, "metadata": {"lang": "en", "topic": "billing"}}Every slice kind the file accepts:
| Slice | Groups cases by | Needs |
|---|---|---|
| metadata.<key> | the value of that key in the case's metadata | the key on the cases |
| relevant_position | where the first relevant passage was retrieved: top_k, beyond_top_k or not_retrieved | a RAG system and expected.relevant |
| context_truncated | whether a token budget dropped passages from the context | a RAG system with a token_budget |
| first_tool | the tool an agent called first | an agent_trajectory/v1 record |
| repeated_action | whether an agent repeated an identical call | an agent_trajectory/v1 record |
| route | the agents that held control, written triage>billing | an agent_trajectory/v1 record whose steps name their agents |
| trajectory_length:4,8 | step count, bucketed at bounds you declare | an agent_trajectory/v1 record |
| confidence:0.5 | the score a model gave the row, banded at bounds you declare | a predictive: block |
A name the file does not recognise is refused before anything runs:
Configuration error: unknown slice 'lang'; declare metadata.<key>, relevant_position, context_truncated, first_tool, repeated_action, route, trajectory_length:<bounds> or confidence:<bounds>Slices report and never gate
A run prints its slices under their own heading. A candidate that fixed English and broke French:
│ metadata.lang=de │ ok │ 100.0% │ [95.4%, 100.0%] │ 80 / 80 observed · 0 missing · 0 excluded · exploratory · support 80 │
│ metadata.lang=en │ ok │ 100.0% │ [95.4%, 100.0%] │ 80 / 80 observed · 0 missing · 0 excluded · exploratory · support 80 │
│ metadata.lang=fr │ ok │ 32.5% │ [22.4%, 43.9%] │ 26 / 80 observed · 0 missing · 0 excluded · exploratory · support 80 │
│ metadata.topic=account │ ok │ 88.3% │ [81.2%, 93.5%] │ 106 / 120 observed · 0 missing · 0 excluded · exploratory · support 120 │
│ metadata.topic=billing │ ok │ 66.7% │ [57.4%, 75.1%] │ 80 / 120 observed · 0 missing · 0 excluded · exploratory · support 120 │The table's title says Slices (exploratory; never gated), and it means it. That run's global rate was 77.5% and its gate allowed the release, with French at a third of that. A release rule on a run names a global metric, and a rule scoped to a slice of a single run is refused:
Configuration error: release rule 'fr-floor' targets slice/metadata.lang=fr; a run's slice results are exploratory, and only a comparison rule gates a slice, inside a family with a correctionThat is the first discipline. A suite sliced ten ways has ten chances to show a difference by accident, and a gate that read slices would block releases on noise. Slices are for finding where to look.
Which slices stand out
The run also ranks its slices against the run's own rate and marks the ones that differ, with the Benjamini-Yekutieli procedure so the share of false marks stays controlled however the slices overlap. The terminal's slices table shows each one in its BY mark column — marked, not marked, or blank for a slice too thin to test — with the p-value it was ranked by. They are also stored in the run and exported in its bundle's run.json:
jq -c '.run.slice_metrics[] | {scope, estimate, fdr_marked, fdr_p_value}' .oloproof/bundles/RUN_ID/run.json{"scope":"slice/metadata.lang=de","estimate":1.0,"fdr_marked":true,"fdr_p_value":0.0001}
{"scope":"slice/metadata.lang=en","estimate":1.0,"fdr_marked":true,"fdr_p_value":0.0001}
{"scope":"slice/metadata.lang=fr","estimate":0.325,"fdr_marked":true,"fdr_p_value":0.0001}
{"scope":"slice/metadata.topic=account","estimate":0.8833333333333333,"fdr_marked":true,"fdr_p_value":0.0037003540039062493}
{"scope":"slice/metadata.topic=billing","estimate":0.6666666666666666,"fdr_marked":true,"fdr_p_value":0.008582189941406249}An unmarked slice is not shown to match the run; it is only not flagged. A mark directs attention and reaches no gate.
Support
min_slice_support is the fewest cases a slice needs before it gets an interval. Below it, the estimate is shown and the interval is withheld. From a project that set it to 4:
│ metadata.topic=account │ answer_correct │ 100.0% │ │ 1 / 1 observed · 0 missing · 0 excluded · exploratory · support 1 │
│ │ │ │ │ < 4, no interval │The default is 30. A small suite that lowers it gets intervals on thin slices, and those intervals are wide enough to say so themselves.
Gating on a slice, in a comparison
Where a slice genuinely must not regress, a comparison rule can be scoped to it. The scope is the slice's name prefixed with slice/, and the rule states the support it needs before any evidence is read:
version: 1
rules:
- {id: overall, metric: ok, kind: non_inferiority, margin: 0.05}
- {id: fr-not-worse, metric: ok, kind: non_inferiority, margin: 0.05, scope: slice/metadata.lang=fr, min_support: 30}
- {id: en-not-worse, metric: ok, kind: non_inferiority, margin: 0.05, scope: slice/metadata.lang=en, min_support: 30}
- {id: de-not-worse, metric: ok, kind: non_inferiority, margin: 0.05, scope: slice/metadata.lang=de, min_support: 30}
families:
- {id: languages, correction: holm, rules: [fr-not-worse, en-not-worse, de-not-worse]}Written without the prefix, the scope is refused:
Configuration error: rule 'fr-not-worse' has unknown scope 'metadata.lang=fr'; a comparison rule decides on the global scope or on slice/<name>=<value>With it, the candidate above against its baseline:
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlComparison sha256:cd2693b505e645919d397e8d7a87eaf0a841ab0c11c0b95484912363381b72e0 of run_01M3C46CYA163JSMSA5QZ6KAZA against run_01M3C46B3PSNP1ZTC8X05AG0GE · 240 paired cases
ok: -5.0 points [-15.2, +4.9] · 240 paired · 0 missing · 0 excluded
ok [slice/metadata.lang=de]: +8.8 points [-0.5, +21.5] · 80 paired · 0 missing · 0 excluded · exploratory · support 80
ok [slice/metadata.lang=en]: +18.8 points [+7.0, +34.4] · 80 paired · 0 missing · 0 excluded · exploratory · support 80
ok [slice/metadata.lang=fr]: -42.5 points [-60.3, -26.8] · 80 paired · 0 missing · 0 excluded · exploratory · support 80
ok [slice/metadata.topic=account]: +7.5 points [+0.9, +17.1] · 120 paired · 0 missing · 0 excluded · exploratory · support 120
ok [slice/metadata.topic=billing]: -17.5 points [-35.0, -0.1] · 120 paired · 0 missing · 0 excluded · exploratory · support 120
Decisions
overall ok non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
no sample size would make this PASS: the difference itself (-5.0 points) is outside the margin, so more cases would move it toward FAIL
fr-not-worse ok non-inferiority, margin 5.0 points FAIL upper_bound_below_margin
en-not-worse ok non-inferiority, margin 5.0 points PASS lower_bound_above_margin
de-not-worse ok non-inferiority, margin 5.0 points PASS lower_bound_above_margin
Family languages: Holm over 3 rules; adjusted levels 0.01667
Gate: BLOCK (exit 1)That is the second discipline. A slice-scoped rule gates only inside a families: entry, which lists the rules whose false failures are controlled together. holm is the correction: a rule in the family that fails is re-tested at a stricter level, sized by how many rules the family holds, so three chances to fail by accident do not add up to three times the risk. At a confidence level of 0.95 and three rules, the strongest failure is re-tested at 0.05 / 3, the next at 0.05 / 2, and the last at 0.05. Only failures are re-tested: a pass is not a rejection and needs no correction. Here one rule failed, it was re-tested at 0.01667, and French still fails. The global rule could not decide, and the advice under it says more cases would only move it toward FAIL.
A slice rule outside any family is refused:
Configuration error: slice rule 'fr-not-worse' belongs to no family; slice rules multiply the chances of a false FAIL, so they gate only inside a family with a correctionWhere to go next
- Comparison rules covers the kinds of rule a scope can be put on.
- RAG evaluation, Agents and tools and
Classifiers and regressors use the non-metadata slices.