Skip to content

指南

进度与并发

针对线上模型运行一个测试套件需要几分钟,而其中大部分时间都花在等待模型上。本页介绍 Oloproof 会同时保持多少个进行中的调用、在调用运行期间向你展示什么,以及如何得知还需要多少次运行才能让一条未做出决策的规则得出结论。

并发

oloproof.yaml 中的 concurrency: 限制同时进行中的调用数量:

concurrency:
  system: 4
  judge: 4

system 限制对被测系统的调用,默认为 8。judge 限制对 LLM 评判模型的调用,默认为 4。两者分开设置,是因为它们通常受不同的速率限制约束。

针对每次调用耗时十分之一秒的系统运行一百二十个用例:

system:实际耗时
113.4s
161.7s
未变更,重新运行0.4s

最后一行体现的是缓存,而不是并发:每个执行结果都被复用了。并发不是系统身份的一部分,因此修改它永远不会使已存储的内容失效。

每个用例都是一次单独的调用。Oloproof 不会把用例合并到服务商的批处理 API 中,因此无法通过它享受服务商的批处理折扣;让运行更快的是并发和缓存。

重试

当评判器服务商或 HTTP 系统返回 429 或 5xx、超时或断开连接时,会以退避方式重试,最多尝试四次,并且绝不会早于 Retry-After 响应头所要求的时间。可调用系统需要主动选择启用:从 oloproof 中抛出带 retryable=True 的 TransientError:

from oloproof import TransientError, system


@system(name="example-support-bot", version="1")
def answer(case):
    ...
    raise TransientError("provider timed out", retryable=True)

retryable 默认为 False。不带它时,错误会记录在用例上,该用例随后被记为缺失,而不是被重试。一个在三十个用例的每一个上都在第一次调用时以这种方式抛出异常、第二次调用时才正常应答的系统:

│ exact_label │ 100.0%   │ [88.4%, 100.0%] │ 30 / 30 observed · 0 missing · 0 excluded │

同一个系统去掉 retryable=True 之后:

│ exact_label │          │ [0.0%, 100.0%] │ 0 / 0 observed · 30 missing · 0 excluded │

运行期间显示什么

在终端中,oloproof run 会在标准错误上重绘一个实时视图:已完成的用例数、缓存命中数、错误数,以及每个二元指标的临时估计值。其中一帧:

                    56/120 cases · 0 cached · 0 errors
┏━━━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Metric      ┃ Estimate ┃ Provisional interval ┃ Cases                   ┃
┡━━━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ exact_label │ 89.8%    │ [78.2%, 95.6%]       │ 49 observed · 0 missing │
└─────────────┴──────────┴──────────────────────┴─────────────────────────┘
     Provisional Wilson estimates over finished cases; not a decision.

标题说明的就是规则。临时区间是基于已完成部分计算的 Wilson 区间,值得观察,但不应据此做出决策:它会随着用例的到来不断重新计算,而一个被反复检查、直到看起来满意为止的区间,就不再是 95% 区间了。没有任何东西会因为它而提前停止运行。决策只做一次,基于完成后的证据,使用策略所指定的方法。

在终端之外,例如在 CI 中,不会绘制实时视图,表格会在结束时一次性打印。

事件流

--json 会把同样的进度以每行一个 JSON 对象的形式写到标准输出,并把表格写到标准错误:

oloproof run --json

一次包含 120 个用例的运行,保存为 run.ndjson 并用 jq -r '.type' run.ndjson | sort | uniq -c 计数,会产生:

 120 case_executed
 120 case_judged
  11 provisional_metrics
   1 run_finished
   2 run_phase_changed
   1 run_started

每个事件都带有运行 id 和时间戳:

{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.782628Z","type":"run_started","suite_digest":"sha256:c6ae32f25d38ddac175f688c15c40991c1e0ec5348f32bfabd9c493a3f688c28","cases":120}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.885640Z","type":"case_executed","scenario_id":"q000","status":"OK","from_cache":false,"latency_ms":102.11420899941004}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.885667Z","type":"case_judged","scenario_id":"q000","criterion":"exact_label","status":"OK","passed":true,"score":null,"from_cache":false}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:41.362779Z","type":"run_finished","status":"DECIDED","completeness":"COMPLETE","exit_code":0}

provisional_metrics 携带每个二元指标的实时 Wilson 区间:

jq -c 'select(.type == "provisional_metrics") | [.cases_done, .metrics[0].estimate, .metrics[0].wilson_lower, .metrics[0].wilson_upper]' run.ndjson
[1,1.0,0.20654931411298355,1.0]
[13,0.8461538461538461,0.5776536895684791,0.9567418216820717]
[25,0.88,0.7004420606159933,0.9583318285288502]
[37,0.8918918918918919,0.7529146844205937,0.9571481006263428]

不要提前关闭管道。在读取前几行后就停止的读取方(例如 head)会在运行存储其最后几个用例之前结束该运行,而该运行会被记录为 RUN_ERROR/PARTIAL。

还需要多少才能做出决策

结果为 INSUFFICIENT_EVIDENCE 的规则并没有失败;只是测试套件太小,无法将结果与阈值区分开。oloproof plan 会说明测试套件需要扩大多少,并根据该运行已经花费的成本估价。对于包含十八个用例的 examples/support_bot/:

oloproof plan RUN_ID --run
Run run_01M3C3WS0SBTFAG55M7ECM1EZ4
  observed  18 cases

exact-label-floor: about 1614 more cases would decide it, if the observed rate holds (1632 in total)
  time      <1s – 27s
  tokens    none reported by this run's providers
  assuming  the cases to come resemble the 18 already run
            cases run one after another; concurrency divides the time and not the cost

对运行规则进行规模估算,仅支持通过/失败率。对于均值,plan 会直接说明这一点,而不是去猜测:

Run run_01M3C3W5YWJY65X9YM6N02F3W4: no sample size can be computed for a rule that did not decide.
  error-budget: sizing a run rule is admitted for binary rates only, and days_error is a MEAN metric: a run stores its summary, not the per-case values sizing one would need

如果给的是比较 id,它会为该比较的规则做规划,并且不需要额外参数。

下一步