Skip to content

Get started

The Python API

Everything the CLI does, the library does. Use it when the evaluation belongs inside a script, a notebook or a test suite rather than beside a config file.

from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, Contains, JsonSchema, Regex, RubricJudge, evaluator

Everything under oloproof_core is engine internals and is not part of this surface.

A system takes a case's input, not the case

This is the one thing worth getting right first, because getting it wrong fails quietly.

@system(name="support-bot", version="1")
def answer(case):
    return {"label": "refund" if "refund" in case["question"].lower() else "other"}

A dataset row looks like this:

{"id":"refund_00","input":{"question":"Can I get a refund? #0"},"expected":{"label":"refund"}}

The function receives the input object, so case["question"] is the question and there is no case["expected"]. What it returns is the output an evaluator reads, so ExactMatch(field="label") compares the returned label against the case's expected label.

A system can also be a method of an object, such as evaluate(system=model.answer, ...), or a callable object. It must then declare a version, as in evaluate(system=system(model.answer, version="v2"), ...): what it does depends on the object's state, which Oloproof cannot see, and two objects with the same method would otherwise share cached results. Change the version whenever that state changes. @system(version=...) written on the method inside the class, or on a callable class, is refused, since it would be one version for every object; use system(model, version="v2") for a callable object.

Running one

result = evaluate(
    system=answer,
    dataset="data/example.jsonl",
    evaluators=[ExactMatch(criterion="exact_label", field="label")],
)

for metric in result.metrics:
    print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)
exact_label 1.0 lower=0.8842966917779722 upper=1.0 30 0

Thirty cases, all correct, and the interval still reaches down to 88.4%. Thirty cases cannot establish more than that, whatever the estimate says.

Reading a case that failed

If the system raises, that case is missing rather than wrong, and the metric reports it:

exact_label None lower=0.0 upper=1.0 0 30

An estimate of None with an interval of the whole range means nothing was observed. Check n_missing before trusting an estimate: a system whose every case raised produces a result object that is structurally fine, and the count is what tells you it is empty.

That is the denominator contract doing its job rather than a failure of it — a case that errored is bounded, not dropped — but nothing raises on your behalf, so the check is yours.

Writing your own evaluator

@evaluator turns a function into an evaluator. The function takes one argument, the case, and reads what it needs from it: case.output is what the system returned and case.expected is the row's expected object.

from oloproof.evaluators import evaluator

@evaluator(criterion="known_label", reads=("output",))
def known_label(case):
    return case.output["label"] in {"refund", "other"}

result = evaluate(system=answer, dataset="data/example.jsonl", evaluators=[known_label])
known_label 1.0 lower=0.8842966917779722 upper=1.0 30 0

reads declares which parts of the case it reads, for provenance; the default is ("output", "expected"). It returns True or False, or a score if it declares value_type="score" and the score_range its scores lie in. Its version includes a digest of the file that defines it, so editing it invalidates its cached judgments.

A function written as def f(output, expected) is refused when it is declared, with the form that works, rather than failing on every case.

Applying a policy

evaluate takes a policy and decides the run, the same rules the CLI reads from release.yaml:

result = evaluate(
    system=answer,
    dataset="data/example.jsonl",
    evaluators=[ExactMatch(criterion="exact_label", field="label")],
    policy="release.yaml",
)
print(result.gate)

Without a policy, result.gate is None: there is no rule to fall short of, so there is nothing to decide.

Where to go next

evaluator.