Get started
Writing a suite
A suite is two files and a dataset. oloproof init scaffolds both, and the scaffold is commented because the first onboarding walk lost four attempts to a missing string rather than a missing feature.
The project file
oloproof.yaml declares what is measured and what does the measuring.
version: 1
project: example
dataset: data/example.jsonl
system:
name: example-support-bot
version: "1"
callable: app:answer
config: {}
evaluators:
- type: exact_match
criterion: exact_label
field: labelsystem.callable is an import path to the thing under test. criterion is the name a metric, a threshold and a report will all use for this evaluator's result, so it is worth choosing a word you will recognise in a gate.
The evaluator types
Every type: the file accepts, by family. A rejected field answers with the ones that evaluator takes, so a wrong guess is one run away from the right one.
| Family | Types |
|---|---|
| Deterministic | exact_match, contains, regex, json_schema |
| LLM judge | rubric_judge, groundedness_judge, citation_support_judge, probability_judge, cascade |
| model | model_classifier |
| Retrieval | hit_rate, recall, mrr, ndcg, citation_validity |
| Agent | agent_max_steps, agent_tool_called, agent_no_tool_loop, agent_tool_sequence, agent_no_undeclared_tool, agent_constraints_satisfied |
| Multi-agent | agent_route, agent_tool_permissions, agent_max_handoffs |
| Predictive | predictive_correct, predictive_precision, predictive_recall, predictive_ranking, predictive_brier, predictive_log_loss, predictive_absolute_error |
An LLM judge on an endpoint that is not OpenAI's needs two more fields: base_url: and api_key_env:. The key is read from the environment at call time and never stored.
Running it
oloproof runA run executes every case, judges each one with every evaluator, and stores the results as content-addressed records. Running it again against an unchanged system and an unchanged dataset reuses what it already has, which is why a repeated run costs almost nothing and is not metered.
oloproof plan RUN_ID --run reports what settling an undecided rule would take, priced from what that run actually spent. Given a comparison id instead, it needs no flag.
Replicates
A case measured once tells you what happened once. replicates: measures each case more than once and reports how many changed their verdict between identical measurements.
That number is worth knowing before you trust any comparison: against a live assistant, about one observed case in nine changed its verdict between identical runs.
Where to go next
- Recording what a system did covers artifacts, usage and latency metrics.
- RAG evaluation, Agents and tools and
Classifiers and regressors cover the retrieval, agent and predictive types.
- Gating CI turns a decision into an exit code.
- Judges covers what an LLM judge must clear before it may gate a release.
- Core concepts explains the four decision states and why denominators matter.