Skip to content

Get started

Quickstart

Install the engine, scaffold a project, and get a first run and a gate decision on your own machine. Nothing leaves it unless you push.

These steps are the README's, rendered here rather than retyped: README.md is the copy a developer meets on GitHub and the one the onboarding tests hold against the code, so it stays the source and this page is generated from it by scripts/generate_quickstart.py.

Install For Development

From the repository root:

python3 -m venv .venv
.venv/bin/python -m pip install -e '.[dev]'
pnpm install --frozen-lockfile
make check
pnpm --filter @oloproof/web check

The public Python surface is:

from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, Contains, JsonSchema, Regex, RubricJudge, evaluator

Everything under oloproof_core is engine internals.

Quick Start

Generate a small project:

. .venv/bin/activate
oloproof init /tmp/oloproof-demo
cd /tmp/oloproof-demo
oloproof run

Evaluators

The type: values oloproof.yaml accepts:

deterministicexact_match, contains, regex, json_schema
LLM judgerubric_judge, groundedness_judge, citation_support_judge; probability_judge: a yes/no, choice or score question answered by one token's probabilities; cascade: a probability judge first, a second judge only where it is unsure
modelmodel_classifier: a trained classifier or NLI model on a TEI server, scored against a threshold
retrievalhit_rate, recall, mrr, ndcg, citation_validity
agentagent_max_steps, agent_tool_called, agent_no_tool_loop, agent_tool_sequence, agent_no_undeclared_tool, agent_constraints_satisfied
multi-agentagent_route, agent_tool_permissions, agent_max_handoffs
predictivepredictive_correct, predictive_precision, predictive_recall, predictive_ranking, predictive_brier, predictive_log_loss, predictive_absolute_error

A judge names its provider (anthropic, openai or openai_compatible), its model, and its rubric as rubric_text or rubric_file. One on a provider that speaks the OpenAI API but is not OpenAI — Groq, Together, a local server — also needs base_url and api_key_env:

evaluators:
  - type: rubric_judge
    criterion: answer_correct
    provider: openai_compatible
    model: openai/gpt-oss-20b
    base_url: https://api.groq.com/openai/v1
    api_key_env: GROQ_API_KEY
    rubric_text: "PASS if the answer conveys the same fact as the reference."

Naming a field an evaluator does not take is answered with the ones it does.

Re-decide an existing run without rerunning the system or evaluators:

oloproof gate RUN_ID --policy release.yaml

Inspect failures and one case:

oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case refund_00

Export a bundle and open it in the web workbench:

oloproof export RUN_ID
OLOPROOF_BUNDLE_DIR="$PWD/.oloproof/bundles" pnpm --dir /path/to/Oloproof --filter @oloproof/web dev

The bundle path must be absolute, because pnpm starts the app from apps/web. The workbench opens on its workspace index; a lone OLOPROOF_BUNDLE_DIR appears there as the local/bundles project. To hold several projects on one machine, point OLOPROOF_WORKBENCH_DIR at a directory containing a workbench.json instead (docs/API_CONTRACTS.md).