Skip to content

Documentation

Write a suite. Gate on it.

Oloproof decides releases on confidence intervals rather than point estimates. These pages cover declaring what to measure, turning a decision into an exit code, and what an LLM judge must clear before it may gate anything.

QuickstartInstall the engine, scaffold a project, and get a first run and a gate decision on your own machine. Nothing leaves it unless you push.Core conceptsOloproof measures an AI system the way a study measures anything: the same cases, the same judges, the same seeds, run on every change. What it adds to a test suite is that every number arrives with its uncertainty, and a release decision is made on the interval rather than on the estimate.Writing a suiteA suite is two files and a dataset. oloproof init scaffolds both, and the scaffold is commented because the first onboarding walk lost four attempts to a missing string rather than a missing feature.The Python APIEverything the CLI does, the library does. Use it when the evaluation belongs inside a script, a notebook or a test suite rather than beside a config file.Recording what a system didA system's output is what an evaluator reads by default. Everything else it did — the prompt it built, the passages it retrieved, the tools it called, the tokens it spent — is recorded alongside the output as artifacts and usage, and every artifact is stored content-addressed with the execution that produced it.Progress and concurrencyA suite against a live model takes minutes, and most of them are spent waiting on the model. This page covers how many calls Oloproof keeps in flight, what it shows you while they run, and how to find out how much more running would settle a rule that did not decide.Gating CIA gate compares measured evidence against a threshold you declared, and exits with a code your CI understands. The decision is made on the interval, not on the estimate.Running in CIGating CI covers the release policy and what each exit code means. This page is the part that happens inside a CI job: how to install Oloproof there, where the evidence lives while the job runs, how to read a result without parsing a table, and how to compare a pull request against the branch it targets.Comparing a candidate to a baselineThe workflow the rest of this product exists for: you changed something, and you want to know whether it helped. A comparison pairs two runs case by case and reports the difference with an interval, so a two-point move is never mistaken for a win.Comparison rulesComparing a candidate to a baseline shows the workflow. This page is the reference for the rules a comparison is decided by: which kinds exist, what each asks, what a margin is measured in, and what moves the interval they are read from.Clustered casesMost interval arithmetic assumes every case is independent of every other. Three turns of one conversation are not: when the conversation goes wrong, all three tend to. A suite that treats them as independent reports an interval narrower than the evidence supports, and a gate can pass on it.SlicesA global rate can hold steady while one part of the suite collapses. A slice is a declared part of the suite, measured on its own, so the collapse is visible. Slices come with two disciplines, because looking at many parts of a suite is how an evaluation finds noise and calls it a finding.JudgesAn LLM judge is one kind of evaluator, not all of them. Oloproof has three kinds — deterministic, LLM judge, and custom — and the rules on this page are about the second, because a model's verdict is the one that needs measuring against a human's.RAG evaluationA retrieval-augmented answer can be wrong for four different reasons: the right passage was never retrieved, it was retrieved and ranked too low, it was ranked high enough and then dropped from the context, or it reached the model and the model got it wrong anyway. A single accuracy number cannot tell them apart. Oloproof runs a RAG system as two stages it can see, measures each, and re-executes failed cases under controlled changes to find out which reason applies.Agents and toolsOloproof does not drive an agent. Your agent runs its own loop, calls its own tools, and records what happened as an agent_trajectory/v1 artifact. Every agent metric is read out of that record, so the record is the whole integration.Classifiers and regressorsA predictive model is evaluated like any other system: a callable returns a prediction for each case, and evaluators read it. What changes is the denominators. Accuracy, recall and precision are three rates over three different sets of rows, and a model can look good on one while failing the question the business is asking.NotificationsA passing run notifies nobody. When Oloproof sends a notification it means that something needs a decision or is about to stop working: a release a gate blocked, a managed run that ended, an allowance running out.ErrorsThe distinction this page exists for: a run that failed and a run that errored are different things, and presenting one as the other tells a developer their change is bad when the truth is that the harness fell over.CLI referenceEvery command Oloproof registers, with what it does. Each one takes --help, which prints the full text these summaries come from.