Guias
Gating CI
A tradução desta página está desatualizada, por isso ela é exibida em inglês.
A gate compares measured evidence against a threshold you declared, and exits with a code your CI understands. The decision is made on the interval, not on the estimate.
The release policy
release.yaml is the scaffolded second file.
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
warn_on: []
require_validated_evaluators: false
rules:
- id: label-floor
metric: exact_label
min: 0.80metric: names a criterion from oloproof.yaml. min: is a minimum threshold, so the rule passes when the interval's lower bound clears it and fails when the upper bound is below it.
block_on lists the decision states that stop a release. The default blocks on INSUFFICIENT_EVIDENCE as well as on FAIL, which is the point: a suite too small to separate an effect from noise has not shown the change is safe, and treating that as a pass is the common way an evaluation misleads the team running it.
Running the gate
oloproof gate RUN_ID --policy release.yamlRUN_ID is the id oloproof run printed. Gating is a separate command from running because a decision can be re-taken against an edited policy without re-executing anything: the evidence is stored and the threshold is not part of it.
Early stopping
Add early_stopping: true to a policy to let a run stop after a seeded batch once every rule has already decided PASS or FAIL. Runs without the flag execute every case exactly as before. The run records the seed, batch size, cases run, cases not needed and the avoided work; the terminal, bundle and workbench show those fields beside the decision.
Comparisons stop too. evaluate_comparison() runs the candidate and the baseline on the same seeded batches, so every case either side ran is a pair, and it stops both once every run rule and comparison rule has decided. A gate that decides only on the last batch stopped nothing, and the run is recorded as an ordinary completed run.
A stopped run keeps its methods: oloproof gate, plan and signoff re-decide it with the same intervals it stopped on.
Cases that never ran count at their worst possible value, so a stop is valid whenever it happens, but that also limits how early a PASS can come. Passing a minimum threshold T needs at least T × N of the N cases to have run, and a superiority rule needs more than half the suite. Early stopping saves the most on clear failures.
Some policies still run every case and record why in disabled_reasons: Holm families (holm_family_declared), quantiles, AUC or other ranking metrics, clustered metrics, replicated runs, and comparisons whose policy names conditional_exact_paired_difference@1 (difference_method:<id>) have no admitted optional-stopping path yet.
Exit codes
| Code | Meaning |
|---|---|
| 0 | Every rule passed |
| 1 | A rule failed |
| 2 | The configuration or invocation was wrong, before anything ran |
| 3 | The evidence did not decide |
| 4 | A person is asked to decide |
| 5 | The run did not complete, so there is no quality result |
When more than one gate is blocked, the code reported is the first of 1, 5, 4, 3 that applies. A failure outranks an incomplete run, which outranks a request for review, which outranks insufficient evidence.
Code 2 is not a quality result. It means the invocation was wrong and nothing was measured, so CI should treat it as a broken build rather than a failing system.
Overriding a blocked gate
Shipping against a blocked gate takes two named approvers and a written reason.
oloproof signoff RUN_ID --by alice --by bob --reason "..."The decisions still say what the evidence said. Only the release action moves, and the override is recorded in the audit log, chained to the entry before it, pinned to the exact gate it approved. New evidence or an edited threshold needs approving again.
In a project that requires verified evidence, the workspace gate decides the release, and a sign-off from the terminal overrides only the gate your run pushed. Sign off against the workspace gate on the run page instead: the Owner gives the reason, and a different Owner or Admin approves it. It takes effect once both have, for that gate alone; if the workspace decides the gate again, ask again.
Verified evidence
A gate decided on your machine says what your evidence shows. A workspace can ask for more, per project, on its settings page, where an Owner or Admin chooses what the project requires:
- Nothing beyond the gate: Decisions stand as pushed.
- Verified human sample: Every decision resting on a judge (an LLM judge, a model or custom code) must rest on a PPI interval over a human sample the workspace drew and checked itself; a judge-free metric stands. The project can also require two or more labellers per sampled case, counted by the account that gave each verdict, in the review queue or with oloproof review --sample, not by the name typed beside it. Where they disagree on a case, an Owner or Admin who labelled none of it resolves it on the queue page or the run page.
- Verified end to end: As above, and every decision must rest on outputs a registered runner produced: a runner your platform team operates, or the managed worker.
With a requirement in force, an Owner or Admin also pins what the workspace decides. They pin the evaluation: a finished run whose suite, evaluators and metrics become the project's evaluation, and which is itself the baseline every comparison is decided against. The settings page lists runs a registered runner signed first, and under verified end to end only a signed run can be pinned; under the human sample level, open an unsigned run before pinning it, since its records are whatever its pusher sent. And they pin the release policy, choosing from the policies pushed with the project's runs that the workspace has checked. With either missing the workspace decides nothing. The workspace then decides every rule of that policy itself, for runs of exactly the pinned evaluation, from the evidence it holds and the samples it drew, and the page of the run a version is decided on leads with that workspace gate, each rule shown with its terms; the comparison page shows none. What it cannot decide on verified evidence is manual review, with the reason. A run with another suite, other evaluators or other metrics is another evaluation, and is not decided. To grow or change the evaluation, an Owner pins a newer run. What the workspace decides is a system version, on its first run of the evaluation: under verified end to end, the first run a registered runner started for that version, whether or not it was pushed or finished; under the human sample level, the first the workspace received. Running the same version again until one run passes changes nothing, and every version evaluated is listed on the run page with its result. Comparisons are the workspace's too: it compares each version's deciding run with the pinned run, so you need not push a comparison at all. The policy your push carries still decides your local gate and CI; in the workspace it only shows beside the result. Under the verified human sample level, which evaluator produced a metric is what your push says; only verified end to end binds it.
To run a suite on a registered runner:
- On the machine that will run it, oloproof runner keygen --out runner.key writes the key and prints its public half.
- An Owner or Admin registers that public key for the project (POST /api/workspaces/<id>/runners), which returns its key id.
- Log the runner in to the project with oloproof login, then run with OLOPROOF_RUNNER_KEY_FILE (or the key itself in OLOPROOF_RUNNER_KEY) and OLOPROOF_RUNNER_KEY_ID set.
Before it executes anything, the runner records the run's suite with the workspace, signed, and refuses to run if it cannot. The runner signs what it executed, including the suite, the system and evaluator versions, and oloproof push sends that signature first; the workspace then checks every record it holds against it, including copies your egress policy redacted. A push that writes a verified run must carry that signature every time.
Verified execution attests which system version the runner executed and what it produced. It does not attest that this is the version in production, and it cannot see outputs anyone obtained by calling the system some other way. A runner holds its signing key in the process that runs your system and evaluators, so it must only run code its operator controls: a provider-backed system, a deployed endpoint or the platform team's own modules, never code the person being evaluated can change. Revoking a runner key for rotation keeps what it signed verified; revoking it as compromised (DELETE ...?compromised=true) voids every run it signed.
Where to go next
- Running in CI covers the job itself: installing, the event stream, and comparing a
pull request against its base.
- Writing a suite covers the file the metrics come from.
- Errors covers what a run that did not complete reports.