Skip to content

Get started

Writing a suite

A suite is two files and a dataset. oloproof init scaffolds both, and the scaffold is commented because the first onboarding walk lost four attempts to a missing string rather than a missing feature.

The project file

oloproof.yaml declares what is measured and what does the measuring.

version: 1
project: example
dataset: data/example.jsonl
system:
  name: example-support-bot
  version: "1"
  callable: app:answer
  config: {}
evaluators:
  - type: exact_match
    criterion: exact_label
    field: label

system.callable is an import path to the thing under test. criterion is the name a metric, a threshold and a report will all use for this evaluator's result, so it is worth choosing a word you will recognise in a gate.

The evaluator types

Every type: the file accepts, by family. A rejected field answers with the ones that evaluator takes, so a wrong guess is one run away from the right one.

FamilyTypes
Deterministicexact_match, contains, regex, json_schema
LLM judgerubric_judge, groundedness_judge, citation_support_judge, probability_judge, cascade
modelmodel_classifier
Retrievalhit_rate, recall, mrr, ndcg, citation_validity
Agentagent_max_steps, agent_tool_called, agent_no_tool_loop, agent_tool_sequence, agent_no_undeclared_tool, agent_constraints_satisfied
Multi-agentagent_route, agent_tool_permissions, agent_max_handoffs
Predictivepredictive_correct, predictive_precision, predictive_recall, predictive_ranking, predictive_brier, predictive_log_loss, predictive_absolute_error

An LLM judge on an endpoint that is not OpenAI's needs two more fields: base_url: and api_key_env:. The key is read from the environment at call time and never stored.

Running it

oloproof run

A run executes every case, judges each one with every evaluator, and stores the results as content-addressed records. Running it again against an unchanged system and an unchanged dataset reuses what it already has, which is why a repeated run costs almost nothing and is not metered.

oloproof plan RUN_ID --run reports what settling an undecided rule would take, priced from what that run actually spent. Given a comparison id instead, it needs no flag.

Replicates

A case measured once tells you what happened once. replicates: measures each case more than once and reports how many changed their verdict between identical measurements.

That number is worth knowing before you trust any comparison: against a live assistant, about one observed case in nine changed its verdict between identical runs.

Where to go next

Classifiers and regressors cover the retrieval, agent and predictive types.

  • Gating CI turns a decision into an exit code.
  • Judges covers what an LLM judge must clear before it may gate a release.
  • Core concepts explains the four decision states and why denominators matter.