Skip to content

Guides

Clustered cases

Most interval arithmetic assumes every case is independent of every other. Three turns of one conversation are not: when the conversation goes wrong, all three tend to. A suite that treats them as independent reports an interval narrower than the evidence supports, and a gate can pass on it.

Declaring a cluster

A case names its cluster with group_id, at the top level of the row beside id:

{"id": "conv00_t0", "group_id": "conv00", "input": {"question": "Conversation 0 turn 0: refund please (garbled)"}, "expected": {"label": "refund"}}
{"id": "conv00_t1", "group_id": "conv00", "input": {"question": "Conversation 0 turn 1: where is my order (garbled)"}, "expected": {"label": "other"}}
{"id": "conv00_t2", "group_id": "conv00", "input": {"question": "Conversation 0 turn 2: refund please (garbled)"}, "expected": {"label": "refund"}}

Once any case declares one, the whole suite is analysed by cluster. There is no per-metric switch back: treating grouped cases as independent is the unsafe direction, so it is not offered.

What changes

The same 108 turns from 36 conversations, one conversation in six garbled from its first turn to its last. Without group_id:

│ exact_label │ 83.3%    │ [74.9%, 89.9%] │ 90 / 108 observed · 0 missing · 0 excluded │

With it:

│ exact_label │ 83.3%    │ [65.2%, 94.6%] │ 90 / 108 observed · 0 missing · 0 excluded · 36 clusters · approximate │

The estimate is identical. The interval is about twice as wide, because the suite holds 36 independent observations rather than 108. Against a floor of 0.70, with the opt-in described below, the first run passes and the second does not:

label-floor: PASS (lower_bound_meets_minimum)
label-floor: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)

The second answer is the right one for this data.

The opt-in

The clustered interval is a studentized cluster bootstrap. It is an approximate method, and a policy has to say it accepts one before a rule may decide on it. Without that, the same grouped run reads:

label-floor: MANUAL_REVIEW (approximate_method_not_permitted)

and exits 4. To let the rule decide, add this to release.yaml:

allow_approximate_methods: true

Two more settings bound what the method is trusted with.

SettingDefaultWhat it does
min_clusters20Fewer clusters than this and a rule reads INSUFFICIENT_EVIDENCE with insufficient_clusters. It may be lowered to 10 and no further.
max_missing_fractionnoneSet on a rule. A clustered interval cannot bound missing cases the way an independent one does, so a clustered rule over a metric with any missing case reads missingness_unbounded until the rule states how much missingness it accepts.

A policy that sets min_clusters: 5 is refused before anything is decided:

Configuration error: p5.yaml: min_clusters: Input should be greater than or equal to 10

The same conversations with one of them erroring on every turn:

│ exact_label │ 82.9%    │ [64.3%, 94.5%] │ 87 / 105 observed · 3 missing · 0 excluded · 35 clusters · approximate │
label-floor: INSUFFICIENT_EVIDENCE (missingness_unbounded)

A rule that states the missingness it accepts:

rules:
  - id: label-floor
    metric: exact_label
    min: 0.60
    max_missing_fraction: 0.15
label-floor: PASS (lower_bound_meets_minimum)

Stating a fraction records an assumption: that the missing cases are missing at random. The decision is only as good as that assumption, which is why the engine will not make it for you.

What a clustered suite cannot do yet

Only pass/fail rates have a clustered interval. On a suite that declares group_id:

  • a mean, a quantile or a ranking metric has no interval, and a rule on one reads

MANUAL_REVIEW with unsupported_dependence_structure;

  • with replicates: above 1, no metric has an interval, pass/fail rates included, because the

two dependence structures compose and nothing models both;

  • a comparison of two runs has no interval for any metric.

A latency quantile on the grouped suite:

│ latency_p50 │ 1.365 ms │ no interval: unsupported_dependence_structure │ p50 of 108 observed · 0 missing · 0 excluded │

The last limitation matters most. Comparing two runs of the conversation suite, the candidate fixing every garbled conversation:

oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Comparison sha256:ed7820e6471f70ce2b5e16ba618fdd48ea92a5ed39d23cce0b719ab31fcb16af of run_01M3C43MWTH8127R7S5M4N9DQ9 against run_01M3C43J2DVM93EWWWRBMRW22E · 108 paired cases
exact_label: +16.7 points · 108 paired · 0 missing · 0 excluded
  no interval: unsupported_dependence_structure
Decisions
  no-regression  exact_label  non-inferiority, margin 5.0 points  MANUAL_REVIEW  unsupported_dependence_structure
Gate: BLOCK (exit 4)

No paired procedure for clustered suites has passed its validation grid, so the comparison reports the difference and asks a person rather than decide approximately. MANUAL_REVIEW rather than INSUFFICIENT_EVIDENCE, because more cases of the same kind would not fix it.

Where to go next