Guides
Running in CI
Gating CI covers the release policy and what each exit code means. This page is the part that happens inside a CI job: how to install Oloproof there, where the evidence lives while the job runs, how to read a result without parsing a table, and how to compare a pull request against the branch it targets.
The exit code is the gate
oloproof run exits on its gate. A CI step that runs it fails when the gate blocks, which is usually what you want and needs no extra wiring:
| Code | A CI step should |
|---|---|
| 0 | pass |
| 1 | fail: a rule failed |
| 2 | fail as a broken build: the configuration was wrong and nothing was measured |
| 3 | fail, or warn: the suite could not decide |
| 4 | fail and route to a person |
| 5 | fail as a broken build: the run did not complete |
Code 3 is the one teams argue about. The default policy blocks on it, because a suite too small to decide has not shown the change is safe. A team that wants the job to go green while it grows its suite can say so in the policy, rather than by ignoring the exit code:
version: 1
block_on: [FAIL, MANUAL_REVIEW]
warn_on: [INSUFFICIENT_EVIDENCE]
rules:
- id: exact-label-floor
metric: exact_label
min: 0.70The same stored run of examples/support_bot/, which blocks with 3 under its own policy, under this one:
oloproof gate RUN_ID --policy advisory.yamlexact-label-floor: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)
about 1614 more cases would decide it, if the observed rate holds (1632 in total)It exits 0. The decision is unchanged. Only the release action moved, and it moved because a file under review says so.
Reading the result as data
--json writes one JSON event per line to standard output and the human tables to standard error, so a CI log keeps the tables and a script reads the events. Saved to run.ndjson, the run's id is on every line, and the last line carries the exit code:
jq -r 'select(.type == "run_started") | .run_id' run.ndjson
jq -c 'select(.type == "run_finished")' run.ndjsonrun_01M3C3WS0SBTFAG55M7ECM1EZ4
{"run_id":"run_01M3C3WS0SBTFAG55M7ECM1EZ4","timestamp":"2026-09-25T11:06:03.646415Z","type":"run_finished","status":"DECIDED","completeness":"COMPLETE","exit_code":3}Progress and concurrency lists every event type.
Where the evidence lives
A run is stored in .oloproof/store.sqlite beside the oloproof.yaml it was run from, and oloproof init puts .oloproof/ in .gitignore. A CI job starts with an empty store, so every job re-executes every case, and nothing from one job is visible to the next.
That matters most for a comparison, which needs both runs in one store. Two checkouts of a project each get their own store, so a baseline run from one cannot be found from the other:
Configuration error: unknown run 'run_01M3C3XSX9KY5T8VFZXA1CVBES'OLOPROOF_HOME points every command at one store, wherever its oloproof.yaml is. The directory it names holds .oloproof/store.sqlite.
Installing Oloproof in a job
Oloproof is not published to a package index, so there is no pip install line that fetches it by name. A job installs it from a checkout of the Oloproof repository, the same way a developer does: actions/checkout with repository: naming wherever your team's copy of that repository lives, and a token: that can read it if it is private. The repository root builds the package, and the package provides the oloproof command.
Comparing a pull request against its base
Check out the base branch and the pull request side by side, run each into the same store, and compare:
name: oloproof
on: pull_request
jobs:
evaluate:
runs-on: ubuntu-latest
env:
OLOPROOF_HOME: ${{ github.workspace }}/evidence
steps:
- uses: actions/checkout@v4
with:
path: pr
- uses: actions/checkout@v4
with:
ref: ${{ github.base_ref }}
path: main
- uses: actions/checkout@v4
with:
repository: YOUR_ORG/Oloproof
token: ${{ secrets.OLOPROOF_REPO_TOKEN }}
path: oloproof-src
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: python -m pip install ./oloproof-src
- run: mkdir -p "$OLOPROOF_HOME"
- run: oloproof run --config main/oloproof.yaml --json > base.ndjson || true
- run: oloproof run --config pr/oloproof.yaml --json > cand.ndjson || true
- name: compare
run: |
candidate=$(jq -r 'select(.type == "run_started") | .run_id' cand.ndjson)
baseline=$(jq -r 'select(.type == "run_started") | .run_id' base.ndjson)
oloproof compare "$candidate" "$baseline" --config pr/oloproof.yaml --policy pr/compare.yamlYOUR_ORG/Oloproof and OLOPROOF_REPO_TOKEN are placeholders for your copy of the repository and a secret that can read it. The two runs carry || true because their own gates are not the question here; the comparison's exit code is.
The same steps on a laptop, with the scaffolded project checked out twice:
Comparison sha256:31ac104779bda1726ba55b661107cbe256fae5f1d831c79b1283a07651b58cbf of run_01M3C3XX4TM0WSZW8K81PHRGAW against run_01M3C3XW7X64B0Z0ACZV27WSW6 · 30 paired cases
exact_label: +0.0 points [-16.5, +16.5] · 30 paired · 0 missing · 0 excluded
Decisions
no-regression exact_label non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
Gate: BLOCK (exit 3)Both runs must be over the same suite. A pull request that edits a case changes the suite's digest, and the comparison refuses rather than pair cases that are no longer the same case:
Configuration error: runs 'run_01M3C3XXVZ1X0J2PR1ZW1WVGVZ' and 'run_01M3C3XW7X64B0Z0ACZV27WSW6' used different suites (sha256:baff4f101901d9a37cd440f99b9a70032f9488891b4f590f18a81017c26ba794 and sha256:2011286a7ec00c8c31560bb6d037b54a501a15f6251ac7c2d578d0c40e226d72); comparisons pair scenario by scenario over one suiteA comparison of a run that did not finish exits 5, the same as a gate over one: there is no paired result to decide on.
Where to go next
- Gating CI covers the policy and the order exit codes are reported in.
- Comparison rules covers what a comparison policy can ask.