指南
進度與並行
針對線上模型執行一個套件需要數分鐘,其中大部分時間都花在等待模型。本頁說明 Oloproof 同時保持多少個進行中的呼叫、在呼叫執行期間向你顯示什麼,以及如何得知還需要多執行多少,才能讓一條未能判定的規則得出結論。
並行
oloproof.yaml 中的 concurrency: 限制同時進行中的呼叫數量:
concurrency:
system: 4
judge: 4system 限制對受測系統的呼叫,預設為 8。judge 限制對 LLM 評審的呼叫,預設為 4。兩者分開設定,是因為它們通常位於不同的速率限制之後。
一百二十個案例,針對一個每次呼叫需要十分之一秒的系統:
| system: | 實際耗時 |
|---|---|
| 1 | 13.4s |
| 16 | 1.7s |
| 未變更,重新執行 | 0.4s |
最後一列是快取的效果,而不是並行:每一次執行都被重複使用。並行設定不屬於系統的身分,因此變更它絕不會讓已儲存的內容失效。
每個案例都是一次獨立的呼叫。Oloproof 不會將案例分組送進提供者的批次 API,因此無法透過它取得提供者的批次折扣;讓執行變快的是並行與快取。
重試
當評審提供者或 HTTP 系統回應 429 或 5xx、逾時或中斷連線時,會以退避方式重試,最多嘗試四次,而且絕不早於 Retry-After 標頭所要求的時間。可呼叫系統若要選擇加入,需從 oloproof 引發帶有 retryable=True 的 TransientError:
from oloproof import TransientError, system
@system(name="example-support-bot", version="1")
def answer(case):
...
raise TransientError("provider timed out", retryable=True)retryable 預設為 False。若沒有它,錯誤會記錄在案例上,該案例隨即算作缺失,而不會重試。一個系統在三十個案例中每一個的第一次呼叫都如此引發錯誤,並在第二次呼叫時回應:
│ exact_label │ 100.0% │ [88.4%, 100.0%] │ 30 / 30 observed · 0 missing · 0 excluded │同一個系統移除 retryable=True 後:
│ exact_label │ │ [0.0%, 100.0%] │ 0 / 0 observed · 30 missing · 0 excluded │執行期間顯示的內容
在終端機中,oloproof run 會在標準錯誤輸出上重繪即時檢視:已完成的案例、快取命中、錯誤,以及每個二元指標的暫定估計值。其中一個畫面:
56/120 cases · 0 cached · 0 errors
┏━━━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Metric ┃ Estimate ┃ Provisional interval ┃ Cases ┃
┡━━━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ exact_label │ 89.8% │ [78.2%, 95.6%] │ 49 observed · 0 missing │
└─────────────┴──────────┴──────────────────────┴─────────────────────────┘
Provisional Wilson estimates over finished cases; not a decision.說明文字就是規則。暫定區間是針對已完成部分計算的 Wilson 區間,值得觀察,但不能據以做出決策:它會隨著案例陸續到達而重新計算,而一個反覆檢查直到看起來不錯為止的區間,就不再是 95% 區間了。任何執行都不會因它而提前停止。決策只做一次,依據完成後的證據,並採用政策所指定的方法。
在終端機以外的環境(例如 CI)中,不會繪製即時檢視,表格只在結束時印出一次。
事件串流
--json 會在標準輸出上以每行一個 JSON 物件的形式寫出相同的進度,並將表格寫到標準錯誤輸出:
oloproof run --json一次 120 個案例的執行,儲存為 run.ndjson,並以 jq -r '.type' run.ndjson | sort | uniq -c 計數,會產生:
120 case_executed
120 case_judged
11 provisional_metrics
1 run_finished
2 run_phase_changed
1 run_started每一個事件都帶有執行 ID 與時間戳記:
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.782628Z","type":"run_started","suite_digest":"sha256:c6ae32f25d38ddac175f688c15c40991c1e0ec5348f32bfabd9c493a3f688c28","cases":120}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.885640Z","type":"case_executed","scenario_id":"q000","status":"OK","from_cache":false,"latency_ms":102.11420899941004}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.885667Z","type":"case_judged","scenario_id":"q000","criterion":"exact_label","status":"OK","passed":true,"score":null,"from_cache":false}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:41.362779Z","type":"run_finished","status":"DECIDED","completeness":"COMPLETE","exit_code":0}provisional_metrics 帶有每個二元指標的即時 Wilson 區間:
jq -c 'select(.type == "provisional_metrics") | [.cases_done, .metrics[0].estimate, .metrics[0].wilson_lower, .metrics[0].wilson_upper]' run.ndjson[1,1.0,0.20654931411298355,1.0]
[13,0.8461538461538461,0.5776536895684791,0.9567418216820717]
[25,0.88,0.7004420606159933,0.9583318285288502]
[37,0.8918918918918919,0.7529146844205937,0.9571481006263428]不要提前關閉管線。在讀取前幾行後就停止的讀取端(例如 head)會在執行儲存最後幾個案例之前就結束它,而該執行會被記錄為 RUN_ERROR/PARTIAL。
還需要多少才能判定
顯示為 INSUFFICIENT_EVIDENCE 的規則並不是失敗;而是套件太小,無法將結果與門檻區分開來。oloproof plan 會說明套件需要擴大多少,並依據該次執行已經花費的資源估算成本。以十八個案例的 examples/support_bot/ 為例:
oloproof plan RUN_ID --runRun run_01M3C3WS0SBTFAG55M7ECM1EZ4
observed 18 cases
exact-label-floor: about 1614 more cases would decide it, if the observed rate holds (1632 in total)
time <1s – 27s
tokens none reported by this run's providers
assuming the cases to come resemble the 18 already run
cases run one after another; concurrency divides the time and not the cost為執行規則估算樣本數,僅適用於通過/失敗比率。對於平均數,計畫會直接說明這一點,而不是妄加猜測:
Run run_01M3C3W5YWJY65X9YM6N02F3W4: no sample size can be computed for a rule that did not decide.
error-budget: sizing a run rule is admitted for binary rates only, and days_error is a MEAN metric: a run stores its summary, not the per-case values sizing one would need若改為提供比較 ID,它會為該比較的規則制定計畫,且不需要任何旗標。