指南
聚类用例
大多数区间计算都假设每个用例与其他所有用例相互独立。同一段对话中的三轮并非如此:当对话出了问题,这三轮往往会一起出问题。把它们当作相互独立的测试套件,报告的区间会比证据所能支持的更窄,而门控可能因此通过。
声明聚类
用例通过 group_id 指定它所属的聚类,该字段位于数据行的顶层,与 id 并列:
{"id": "conv00_t0", "group_id": "conv00", "input": {"question": "Conversation 0 turn 0: refund please (garbled)"}, "expected": {"label": "refund"}}
{"id": "conv00_t1", "group_id": "conv00", "input": {"question": "Conversation 0 turn 1: where is my order (garbled)"}, "expected": {"label": "other"}}
{"id": "conv00_t2", "group_id": "conv00", "input": {"question": "Conversation 0 turn 2: refund please (garbled)"}, "expected": {"label": "refund"}}只要有任何一个用例声明了聚类,整个测试套件就会按聚类进行分析。没有按指标切换回去的开关:把分组的用例当作相互独立是不安全的方向,因此不提供这种选项。
有什么变化
同样来自 36 段对话的 108 轮,其中每六段对话中有一段从第一轮到最后一轮都是乱码。不使用 group_id 时:
│ exact_label │ 83.3% │ [74.9%, 89.9%] │ 90 / 108 observed · 0 missing · 0 excluded │使用它时:
│ exact_label │ 83.3% │ [65.2%, 94.6%] │ 90 / 108 observed · 0 missing · 0 excluded · 36 clusters · approximate │估计值完全相同。区间宽了大约一倍,因为这个测试套件包含的是 36 个独立观测,而不是 108 个。在最低阈值为 0.70、并启用下文所述的显式许可时,第一次运行通过,第二次则没有:
label-floor: PASS (lower_bound_meets_minimum)label-floor: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)对于这份数据,第二个答案才是正确的。
显式许可
聚类区间采用学生化聚类自助法(studentized cluster bootstrap)。它是一种近似方法,策略必须声明接受近似方法,规则才能依据它做出决策。否则,同一次分组运行的结果是:
label-floor: MANUAL_REVIEW (approximate_method_not_permitted)并以 4 退出。要让规则做出决策,请在 release.yaml 中加入:
allow_approximate_methods: true另有两项设置限定了该方法可被信任的范围。
| 设置 | 默认值 | 作用 |
|---|---|---|
| min_clusters | 20 | 聚类数少于该值时,规则结果为 INSUFFICIENT_EVIDENCE,原因为 insufficient_clusters。它最低可以调到 10,不能再低。 |
| max_missing_fraction | 无 | 在规则上设置。聚类区间无法像独立区间那样对缺失用例进行界定,因此对于存在任何缺失用例的指标,聚类规则的结果都是 missingness_unbounded,直到该规则声明它能接受多少缺失。 |
设置了 min_clusters: 5 的策略会在做出任何决策之前被拒绝:
Configuration error: p5.yaml: min_clusters: Input should be greater than or equal to 10同样的对话,其中一段在每一轮都出错时:
│ exact_label │ 82.9% │ [64.3%, 94.5%] │ 87 / 105 observed · 3 missing · 0 excluded · 35 clusters · approximate │label-floor: INSUFFICIENT_EVIDENCE (missingness_unbounded)声明了可接受缺失比例的规则:
rules:
- id: label-floor
metric: exact_label
min: 0.60
max_missing_fraction: 0.15label-floor: PASS (lower_bound_meets_minimum)声明一个比例就是记录一个假设:缺失的用例是随机缺失的。决策的可靠程度取决于这个假设,这也是引擎不会替你做出这个假设的原因。
聚类测试套件目前还不能做什么
只有通过/失败率有聚类区间。在声明了 group_id 的测试套件上:
- 均值、分位数或排序指标没有区间,基于它们的规则结果为 MANUAL_REVIEW,原因为 unsupported_dependence_structure;
- 当 replicates: 大于 1 时,任何指标都没有区间,包括通过/失败率,因为两种依赖结构会叠加,而没有任何方法能同时对两者建模;
- 两次运行之间的比较,对任何指标都没有区间。
分组测试套件上的一个延迟分位数:
│ latency_p50 │ 1.365 ms │ no interval: unsupported_dependence_structure │ p50 of 108 observed · 0 missing · 0 excluded │最后一项限制最为重要。比较对话测试套件的两次运行,其中候选修复了所有乱码对话:
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlComparison sha256:ed7820e6471f70ce2b5e16ba618fdd48ea92a5ed39d23cce0b719ab31fcb16af of run_01M3C43MWTH8127R7S5M4N9DQ9 against run_01M3C43J2DVM93EWWWRBMRW22E · 108 paired cases
exact_label: +16.7 points · 108 paired · 0 missing · 0 excluded
no interval: unsupported_dependence_structure
Decisions
no-regression exact_label non-inferiority, margin 5.0 points MANUAL_REVIEW unsupported_dependence_structure
Gate: BLOCK (exit 4)目前还没有任何适用于聚类测试套件的配对方法通过其验证网格,因此比较只报告差值并交由人来判断,而不是做出近似决策。结果是 MANUAL_REVIEW 而不是 INSUFFICIENT_EVIDENCE,因为增加更多同类用例并不能解决这个问题。