Skip to content

Guides

Progress and concurrency

A suite against a live model takes minutes, and most of them are spent waiting on the model. This page covers how many calls Oloproof keeps in flight, what it shows you while they run, and how to find out how much more running would settle a rule that did not decide.

Concurrency

concurrency: in oloproof.yaml bounds how many calls are in flight at once:

concurrency:
  system: 4
  judge: 4

system bounds calls to the system under test and defaults to 8. judge bounds calls to LLM judges and defaults to 4. They are separate because the two usually sit behind different rate limits.

A hundred and twenty cases against a system that takes a tenth of a second per call:

system:Wall time
113.4s
161.7s
unchanged, rerun0.4s

The last row is the cache, not concurrency: every execution was reused. Concurrency is not part of a system's identity, so changing it never invalidates what is stored.

Every case is its own call. Oloproof does not group cases into a provider's batch API, so a provider's batch discount is not available through it; concurrency and the cache are what make a run faster.

Retries

A judge provider or an HTTP system that answers 429 or a 5xx, times out, or drops the connection is retried with backoff, up to four attempts, and never sooner than a Retry-After header asks. A callable system opts in by raising TransientError from oloproof with retryable=True:

from oloproof import TransientError, system


@system(name="example-support-bot", version="1")
def answer(case):
    ...
    raise TransientError("provider timed out", retryable=True)

retryable defaults to False. Without it the error is recorded on the case, which is then missing rather than retried. A system that raised that way on the first call for each of thirty cases, and answered on the second:

│ exact_label │ 100.0%   │ [88.4%, 100.0%] │ 30 / 30 observed · 0 missing · 0 excluded │

The same system with retryable=True removed:

│ exact_label │          │ [0.0%, 100.0%] │ 0 / 0 observed · 30 missing · 0 excluded │

What a run shows while it runs

In a terminal, oloproof run redraws a live view on standard error: cases done, cache hits, errors, and a provisional estimate for each binary metric. One frame:

                    56/120 cases · 0 cached · 0 errors
┏━━━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Metric      ┃ Estimate ┃ Provisional interval ┃ Cases                   ┃
┡━━━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ exact_label │ 89.8%    │ [78.2%, 95.6%]       │ 49 observed · 0 missing │
└─────────────┴──────────┴──────────────────────┴─────────────────────────┘
     Provisional Wilson estimates over finished cases; not a decision.

The caption is the rule. A provisional interval is a Wilson interval over whatever has finished, which is a useful thing to watch and not a thing to decide on: it is recomputed as cases arrive, and an interval checked repeatedly until it looks good is not a 95% interval any more. Nothing stops a run early on it. The decision is taken once, on the finished evidence, with the method the policy names.

Outside a terminal, such as in CI, the live view is not drawn and the tables are printed once, at the end.

The event stream

--json writes the same progress as one JSON object per line on standard output, and the tables on standard error:

oloproof run --json

A run of 120 cases, saved to run.ndjson and counted with jq -r '.type' run.ndjson | sort | uniq -c, emits:

 120 case_executed
 120 case_judged
  11 provisional_metrics
   1 run_finished
   2 run_phase_changed
   1 run_started

Each one carries the run id and a timestamp:

{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.782628Z","type":"run_started","suite_digest":"sha256:c6ae32f25d38ddac175f688c15c40991c1e0ec5348f32bfabd9c493a3f688c28","cases":120}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.885640Z","type":"case_executed","scenario_id":"q000","status":"OK","from_cache":false,"latency_ms":102.11420899941004}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:37.885667Z","type":"case_judged","scenario_id":"q000","criterion":"exact_label","status":"OK","passed":true,"score":null,"from_cache":false}
{"run_id":"run_01M3C3ZNXPVQJCDF31DJSBXBHC","timestamp":"2026-09-25T11:07:41.362779Z","type":"run_finished","status":"DECIDED","completeness":"COMPLETE","exit_code":0}

provisional_metrics carries the running Wilson interval for each binary metric:

jq -c 'select(.type == "provisional_metrics") | [.cases_done, .metrics[0].estimate, .metrics[0].wilson_lower, .metrics[0].wilson_upper]' run.ndjson
[1,1.0,0.20654931411298355,1.0]
[13,0.8461538461538461,0.5776536895684791,0.9567418216820717]
[25,0.88,0.7004420606159933,0.9583318285288502]
[37,0.8918918918918919,0.7529146844205937,0.9571481006263428]

Do not close the pipe early. A reader that stops after the first few lines, such as head, ends the run before it stores its last cases, and the run is recorded as RUN_ERROR/PARTIAL.

How much more would decide it

A rule that reads INSUFFICIENT_EVIDENCE has not failed; the suite was too small to separate the result from the threshold. oloproof plan says how much larger it would need to be, priced from what the run already spent. For the eighteen-case examples/support_bot/:

oloproof plan RUN_ID --run
Run run_01M3C3WS0SBTFAG55M7ECM1EZ4
  observed  18 cases

exact-label-floor: about 1614 more cases would decide it, if the observed rate holds (1632 in total)
  time      <1s – 27s
  tokens    none reported by this run's providers
  assuming  the cases to come resemble the 18 already run
            cases run one after another; concurrency divides the time and not the cost

Sizing a run rule is admitted for pass/fail rates only. On a mean, the plan says so rather than guess:

Run run_01M3C3W5YWJY65X9YM6N02F3W4: no sample size can be computed for a rule that did not decide.
  error-budget: sizing a run rule is admitted for binary rates only, and days_error is a MEAN metric: a run stores its summary, not the per-case values sizing one would need

Given a comparison id instead, it plans the comparison's rules, and needs no flag.

Where to go next

estimates are priced from.