Skip to content

指南

進度與並行

針對線上模型執行一個套件需要數分鐘,其中大部分時間都花在等待模型。本頁說明 Oloproof 同時保持多少個進行中的呼叫、在呼叫執行期間向你顯示什麼,以及如何得知還需要多執行多少,才能讓一條未能判定的規則得出結論。

並行

oloproof.yaml 中的 concurrency: 限制同時進行中的呼叫數量:

concurrency:
  system: 4
  judge: 4

system 限制對受測系統的呼叫,預設為 8。judge 限制對 LLM 評審的呼叫,預設為 4。兩者分開設定,是因為它們通常位於不同的速率限制之後。

一百二十個案例,針對一個每次呼叫需要十分之一秒的系統:

system:實際耗時
113.4s
161.7s
未變更,重新執行0.4s

最後一列是快取的效果,而不是並行:每一次執行都被重複使用。並行設定不屬於系統的身分,因此變更它絕不會讓已儲存的內容失效。

每個案例都是一次獨立的呼叫。Oloproof 不會將案例分組送進提供者的批次 API,因此無法透過它取得提供者的批次折扣;讓執行變快的是並行與快取。

重試

當評審提供者或 HTTP 系統回應 429 或 5xx、逾時或中斷連線時,會以退避方式重試,最多嘗試四次,而且絕不早於 Retry-After 標頭所要求的時間。可呼叫系統若要選擇加入,需從 oloproof 引發帶有 retryable=True 的 TransientError:

from oloproof import TransientError, system


@system(name="example-support-bot", version="1")
def answer(case):
    ...
    raise TransientError("provider timed out", retryable=True)

retryable 預設為 False。若沒有它,錯誤會記錄在案例上,該案例隨即算作缺失,而不會重試。一個系統在三十個案例中每一個的第一次呼叫都如此引發錯誤,並在第二次呼叫時回應:

│ exact_label │ 100.0%   │ [88.4%, 100.0%] │ 30 / 30 observed · 0 missing · 0 excluded │

同一個系統移除 retryable=True 後:

│ exact_label │          │ [0.0%, 100.0%] │ 0 / 0 observed · 30 missing · 0 excluded │

執行期間顯示的內容

在終端機中,oloproof run 會在標準錯誤輸出上重繪即時檢視:已完成的案例、快取命中、錯誤,以及每個二元指標的暫定估計值。其中一個畫面:

                    56/120 cases · 0 cached · 0 errors
┏━━━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Metric      ┃ Estimate ┃ Provisional interval ┃ Cases                   ┃
┡━━━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ exact_label │ 89.8%    │ [78.2%, 95.6%]       │ 49 observed · 0 missing │
└─────────────┴──────────┴──────────────────────┴─────────────────────────┘
     Provisional Wilson estimates over finished cases; not a decision.

說明文字就是規則。暫定區間是針對已完成部分計算的 Wilson 區間,值得觀察,但不能據以做出決策:它會隨著案例陸續到達而重新計算,而一個反覆檢查直到看起來不錯為止的區間,就不再是 95% 區間了。任何執行都不會因它而提前停止。決策只做一次,依據完成後的證據,並採用政策所指定的方法。

在終端機以外的環境(例如 CI)中,不會繪製即時檢視,表格只在結束時印出一次。

事件串流

--json 會在標準輸出上以每行一個 JSON 物件的形式寫出相同的進度,並將表格寫到標準錯誤輸出:

oloproof run --json

一次 120 個案例的執行,儲存為 run.ndjson,並以 jq -r '.type' run.ndjson | sort | uniq -c 計數,會產生:

 120 case_executed
 120 case_judged
  11 provisional_metrics
   1 run_finished
   2 run_phase_changed
   1 run_started

每一個事件都帶有執行 ID 與時間戳記:

{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.782628Z","type":"run_started","suite_digest":"sha256:c6ae32f25d38ddac175f688c15c40991c1e0ec5348f32bfabd9c493a3f688c28","cases":120}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.885640Z","type":"case_executed","scenario_id":"q000","status":"OK","from_cache":false,"latency_ms":102.11420899941004}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.885667Z","type":"case_judged","scenario_id":"q000","criterion":"exact_label","status":"OK","passed":true,"score":null,"from_cache":false}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:41.362779Z","type":"run_finished","status":"DECIDED","completeness":"COMPLETE","exit_code":0}

provisional_metrics 帶有每個二元指標的即時 Wilson 區間:

jq -c 'select(.type == "provisional_metrics") | [.cases_done, .metrics[0].estimate, .metrics[0].wilson_lower, .metrics[0].wilson_upper]' run.ndjson
[1,1.0,0.20654931411298355,1.0]
[13,0.8461538461538461,0.5776536895684791,0.9567418216820717]
[25,0.88,0.7004420606159933,0.9583318285288502]
[37,0.8918918918918919,0.7529146844205937,0.9571481006263428]

不要提前關閉管線。在讀取前幾行後就停止的讀取端(例如 head)會在執行儲存最後幾個案例之前就結束它,而該執行會被記錄為 RUN_ERROR/PARTIAL。

還需要多少才能判定

顯示為 INSUFFICIENT_EVIDENCE 的規則並不是失敗;而是套件太小,無法將結果與門檻區分開來。oloproof plan 會說明套件需要擴大多少,並依據該次執行已經花費的資源估算成本。以十八個案例的 examples/support_bot/ 為例:

oloproof plan RUN_ID --run
Run run_01M3C3WS0SBTFAG55M7ECM1EZ4
  observed  18 cases

exact-label-floor: about 1614 more cases would decide it, if the observed rate holds (1632 in total)
  time      <1s – 27s
  tokens    none reported by this run's providers
  assuming  the cases to come resemble the 18 already run
            cases run one after another; concurrency divides the time and not the cost

為執行規則估算樣本數,僅適用於通過/失敗比率。對於平均數,計畫會直接說明這一點,而不是妄加猜測:

Run run_01M3C3W5YWJY65X9YM6N02F3W4: no sample size can be computed for a rule that did not decide.
  error-budget: sizing a run rule is admitted for binary rates only, and days_error is a MEAN metric: a run stores its summary, not the per-case values sizing one would need

若改為提供比較 ID,它會為該比較的規則制定計畫,且不需要任何旗標。

下一步