Reference
CLI reference
Every command Oloproof registers, with what it does. Each one takes --help, which prints the full text these summaries come from.
This page is generated from the CLI by scripts/generate_cli_reference.py. Editing it by hand is pointless: the next run overwrites it, and make check fails while the committed copy is stale.
Commands
| Command | What it does |
|---|---|
| oloproof audit | Every act a person took in this project, and what each one proves. |
| oloproof collect | Receive OpenTelemetry traces, and serve case content to reviewers, here. |
| oloproof compare | Report the paired counts, differences and decisions of two stored runs. |
| oloproof credentials list | The names this workspace holds. Never the values. |
| oloproof credentials set | Store a provider key for this workspace, read from the environment or standard input. |
| oloproof diagnose | Re-execute a run's failed cases with gold context, under an intervention. |
| oloproof doctor | Report the Python, platform and provider environment this machine offers. |
| oloproof evaluators calibrate | Measure a probability judge's stated probabilities against people, and fit a correction. |
| oloproof evaluators list | List every stored evaluator version with the status a decision would read. |
| oloproof evaluators probe | Probe a judge without labels: how often its verdict moves when nothing that matters did. |
| oloproof evaluators retire | Withdraw trust in an evaluator version. Terminal: record a new version to replace it. |
| oloproof evaluators try | Measure a draft evaluator against the labels already recorded, and validate nothing. |
| oloproof evaluators validate | Measure the evaluator against stored human labels and validate it on that evidence. |
| oloproof export | Write a portable bundle for a run, a comparison or an evaluator version. |
| oloproof gate | Decide the release against the policy, and exit on the decision. |
| oloproof init | Scaffold a project: an example system, oloproof.yaml and release.yaml. |
| oloproof inspect | Read a stored run, case, comparison or diagnosis. |
| oloproof job | What became of a managed job, exiting on the gate of the run it produced. |
| oloproof labels export | Write cases to a CSV for people to label: a blind random sample, or cases to review. |
| oloproof labels import | Record every filled-in row of a labelling file as a label, or none of them. |
| oloproof login | Connect this terminal to a workspace, by approving it in a browser. |
| oloproof plan | What it would take to decide what this run or comparison could not. |
| oloproof prefer | Which of two runs answered better, by a pairwise judge asked in both orders. |
| oloproof push | Send a run's evidence to the workspace. |
| oloproof review | Label cases one at a time in the terminal, storing each verdict as it is given. |
| oloproof run | Execute every case against the system and store the evidence. |
| oloproof runner keygen | Make a runner's Ed25519 key pair; print the public key to register with the workspace. |
| oloproof signoff | Record shipping against a blocked gate, with two approvers and a written reason. |
| oloproof traces evaluate | Evaluate a traffic sample the workspace drew, on the outputs production recorded. |
| oloproof traces list | The newest traces in this project's store. |
| oloproof traces promote | Add a trace to a dataset as a test case, with where it came from. |
| oloproof traces show | One trace, with its input, output, spans and artifacts, as JSON. |
| oloproof usage | What this project consumed, measured from its own evidence. |
Exit codes
Every command uses the same codes. Gating CI lists what each means and the order they are reported in when more than one gate is blocked.