Get started
Quickstart
Install the engine, scaffold a project, and get a first run and a gate decision on your own machine. Nothing leaves it unless you push.
These steps are the README's, rendered here rather than retyped: README.md is the copy a developer meets on GitHub and the one the onboarding tests hold against the code, so it stays the source and this page is generated from it by scripts/generate_quickstart.py.
Install For Development
From the repository root:
python3 -m venv .venv
.venv/bin/python -m pip install -e '.[dev]'
pnpm install --frozen-lockfile
make check
pnpm --filter @oloproof/web checkThe public Python surface is:
from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, Contains, JsonSchema, Regex, RubricJudge, evaluatorEverything under oloproof_core is engine internals.
Quick Start
Generate a small project:
. .venv/bin/activate
oloproof init /tmp/oloproof-demo
cd /tmp/oloproof-demo
oloproof runEvaluators
The type: values oloproof.yaml accepts:
| deterministic | exact_match, contains, regex, json_schema |
| LLM judge | rubric_judge, groundedness_judge, citation_support_judge; probability_judge: a yes/no, choice or score question answered by one token's probabilities; cascade: a probability judge first, a second judge only where it is unsure |
| model | model_classifier: a trained classifier or NLI model on a TEI server, scored against a threshold |
| retrieval | hit_rate, recall, mrr, ndcg, citation_validity |
| agent | agent_max_steps, agent_tool_called, agent_no_tool_loop, agent_tool_sequence, agent_no_undeclared_tool, agent_constraints_satisfied |
| multi-agent | agent_route, agent_tool_permissions, agent_max_handoffs |
| predictive | predictive_correct, predictive_precision, predictive_recall, predictive_ranking, predictive_brier, predictive_log_loss, predictive_absolute_error |
A judge names its provider (anthropic, openai or openai_compatible), its model, and its rubric as rubric_text or rubric_file. One on a provider that speaks the OpenAI API but is not OpenAI — Groq, Together, a local server — also needs base_url and api_key_env:
evaluators:
- type: rubric_judge
criterion: answer_correct
provider: openai_compatible
model: openai/gpt-oss-20b
base_url: https://api.groq.com/openai/v1
api_key_env: GROQ_API_KEY
rubric_text: "PASS if the answer conveys the same fact as the reference."Naming a field an evaluator does not take is answered with the ones it does.
Re-decide an existing run without rerunning the system or evaluators:
oloproof gate RUN_ID --policy release.yamlInspect failures and one case:
oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case refund_00Export a bundle and open it in the web workbench:
oloproof export RUN_ID
OLOPROOF_BUNDLE_DIR="$PWD/.oloproof/bundles" pnpm --dir /path/to/Oloproof --filter @oloproof/web devThe bundle path must be absolute, because pnpm starts the app from apps/web. The workbench opens on its workspace index; a lone OLOPROOF_BUNDLE_DIR appears there as the local/bundles project. To hold several projects on one machine, point OLOPROOF_WORKBENCH_DIR at a directory containing a workbench.json instead (docs/API_CONTRACTS.md).