Skip to content

Przewodniki

Judges

Tłumaczenie tej strony jest nieaktualne, dlatego jest wyświetlana po angielsku.

An LLM judge is one kind of evaluator, not all of them. Oloproof has three kinds — deterministic, LLM judge, and custom — and the rules on this page are about the second, because a model's verdict is the one that needs measuring against a human's.

Why a judge is measured

A deterministic check needs no agreement measurement: it is deterministic, and running it twice gives the same answer. A judge is a model, and a model's error rate is a fact about it that nobody knows until somebody measures it.

So a judge that has never been compared with human labels stays UNVALIDATED, and a policy may refuse to let an unvalidated judge decide a release.

Getting to VALIDATED

Promotion needs evidence, not a command.

oloproof evaluators list
oloproof evaluators validate EVALUATOR_ID --by alice

oloproof evaluators list prints each stored version with its id and status; validate takes that id and the name of whoever is standing behind the measurement.

VALIDATED requires a measured agreement against human labels.

Recording labels

A label is one person's verdict on what a system produced for one case of a run, for one criterion. The people who know the right answers rarely write Python, so a run's cases go out as a file anyone can fill in, a spreadsheet included, and come back as labels:

oloproof labels export RUN_ID --criterion answer_correct --sample 50 --out sample.csv
oloproof labels import sample.csv

--sample 50 draws fifty cases at random, prints the seed it used and records the draw on the run, and leaves the judge's verdict out of the file: a person shown a verdict tends to agree with it. Fill in passed (pass or fail) and labelled_by on each row you judge; a row left blank is skipped. The whole file is checked before anything is stored, so one wrong row stores nothing and names the row.

In a terminal, oloproof review asks case by case and stores each verdict as it is given:

oloproof review RUN_ID --criterion answer_correct --by alice --sample 20

To settle particular cases rather than measure the judge, name them with --cases or take every case the judge failed with --failures. Those show the judge's verdict and are recorded as review labels, which never count toward agreement.

In a hosted workspace, an Owner or Admin can open a review queue from the project's Review page and assign reviewers. A queue that measures the judge has the workspace draw its cases and hides the judge's verdict; a queue of flagged cases shows it and records review labels. Reviewers label in the browser with Pass, Fail or Can't tell, or the keys P, F and C. Each verdict is recorded under the reviewer's own account, which is how the workspace counts independent labellers, and becomes a label the workspace verifies like one pushed from a terminal. Can't tell is recorded and makes no label. Giving a second verdict on a case keeps the first.

A workspace counts only verdicts it recorded itself toward its check of a sample it drew: labels in its review queue, and oloproof review --sample from a terminal, which sends each verdict to the workspace under your account as you give it. Labels imported from a file and pushed still count on your machine, but not in the workspace's check. When a project requires two or more labellers and they disagree on a case, the workspace decides nothing until an Owner or Admin who labelled none of that case resolves it, on the queue page or the run page.

A queue can also ask for a score on a scale you choose, or for a preference between two runs: reviewers see the two answers as A and B, in an order fixed for each reviewer and case that only the workspace can compute, and the page does not say which run is which. That stops a reviewer anchoring on the run they expect to win; it does not stop one who goes looking, since every member can open both runs. Scores and preferences are recorded, and each reviewer sees their own in the queue, but nothing on the run page and no metric or gate reads them yet; a metric of type human_score or human_preference is refused in oloproof.yaml until a method for them is validated.

If your project keeps inputs and outputs on your machines, the workspace has nothing to show reviewers, so run oloproof collect beside the project's store and give its address in the workspace settings. Each reviewer's browser then fetches the case from it directly, with a five-minute ticket the workspace signs for that reviewer and that case; the content never passes through the workspace. Use --certfile and --keyfile to serve HTTPS on your network; http://localhost works for one machine.

From Python, record_label records one:

from oloproof import record_label

record_label(
    run_id="run_01M3C15NADYQ8S96WCPD8HHPGQ",
    scenario_id="refund_00",
    criterion="answer_correct",
    passed=True,
    labelled_by="alice",
)

The run id is the one oloproof run printed. The label names the answer that run produced for the case, so it is compared only with the judge's verdict on that same answer, never with the verdict on a later run's answer to the same question. Refused rather than stored, because a label that matches nothing would be kept and never counted:

  • a case the run does not hold;
  • a criterion it did not judge;
  • a case whose execution produced no output.

Measurement or review: a label is purpose="measurement" by default, which is what measures a judge. Pass purpose="review" for a case you picked because of what its answer or its verdict looked like, such as a failure you are settling. Review labels are kept and never counted toward agreement: a case chosen for how it looked is not a sample, and a reviewer shown the judge's verdict tends to agree with it.

Agreement counts one verdict per person per case. Re-running a labelling script, with or without a different note, does not narrow the interval below, and a person who changes their mind is counted by their latest verdict. Two people labelling one case are two labels, and both count, but the case is counted once: agreement becomes the average, over cases, of the share of labellers who agreed with the judge. Five people on one case add five opinions and one case of evidence.

Then oloproof evaluators validate measures agreement over every measurement label whose answer the judge judged, in any run.

Judge-corrected gates

When a judge-produced pass/fail metric has current-run measurement labels, Oloproof can report ppi_judge_rate@1: the judge-only interval, the human-only interval and the PPI interval shown side by side. The gate decides on the PPI interval, whose estimand is the human-defined pass rate.

Only the random, seeded, blind measurement sample enters the correction. The labels must carry the exported sample's seed, requested size, blind flag and selected execution ids; hand-entered purpose="measurement" labels without that provenance still measure agreement, but they cannot decide a PPI gate. Only the run's first sample for the criterion counts, and only if Oloproof chose its seed. Each case's place in the draw depends only on its own answer and content, so exporting again, a cached rerun, or renaming, reordering or dropping other cases gives the same cases, and a larger sample adds cases to a smaller one. If the first sample you draw uses a --seed of your own, that run stays out of PPI; rerun it from the cache and draw again without --seed.

A sample drawn on your machine is marked local: it is good-faith, because its seed is known before anyone labels, so someone reading the outputs first could arrange the suite to make it favourable. For a PPI gate other people rely on, push the run and let the workspace draw the sample: once you are logged in (oloproof login) and the run is pushed, oloproof labels export --sample and oloproof review --sample ask the workspace to draw, wait for its cases and mark the sample workspace. The workspace checks the result from the verdicts it receives: those given in its review queue and those oloproof review --sample sends. A file filled in from labels export --sample and imported is a local round trip, and counts on your machine only. The workspace never tells you the order it drew in, only the cases. From the moment it draws, the suite and the judge are frozen for verification: you may add cases, but a run that removes, edits, reorders or re-judges any case the workspace drew over, or adds, renames or changes any judge, is not verified. Settle your suite and your judges before you draw. Which cases a suite contains is still your choice, as it is for every metric. The workspace can check a case's reference answer only on a push that carries its content; what a push carries is set by oloproof.yaml on the machine that pushes, so wherever the content is withheld the reference answers people label against rest on what was pushed. Running the suite on a registered runner binds the reference answers to what the runner scored against, and a pinned evaluation fixes them for every verified run: see "Verified evidence" in the gating guide. oloproof push sends your labels with the run as local evidence; the workspace does not count them toward its check. A verdict oloproof review --sample could not send is kept and sent again by the next oloproof review or oloproof push. Drawing a workspace sample needs a key that can write runs, so Reviewers label in the review queue. Pass --local to draw on your machine anyway.

A later draw, or a --seed you chose, could have been picked after reading the outputs, so its labels measure the judge without entering PPI. Draw the sample once the run has finished. Review labels are still stored for case review, but they are excluded from the PPI rectifier and shown as excluded beside the interval. Each case counts its first recorded measurement verdict only: a re-check or a second opinion recorded later is kept and measures agreement, but it does not change the PPI interval, because people re-check the cases that failed. That only works if the first verdict recorded is your first look. Don't revise verdicts in the file before importing it; oloproof review --sample records each verdict as you give it, and is the safer way to label a measurement sample. A file with two measurement verdicts for one case is refused: import a second opinion as its own file.

The agreement bar

A policy can declare the agreement a judge must clear before it may gate a release.

minimum_evaluator_agreement: 0.90
require_validated_evaluators: true

The comparison uses the lower bound of the measured agreement, not the estimate — the same rule the product applies to everything else it decides on. A judge measured at 92% agreement on thirty labelled cases has an interval wide enough that its lower bound may sit under a 90% bar, and in that case it does not gate.

Agreement is a binary rate over the labelled cases a run judged, so it takes a Clopper-Pearson interval and bounds uncomparable labels by worst-case substitution.

A policy can also bound the judge's bias:

maximum_evaluator_bias: 0.03

Then the whole interval of the judge's pass rate minus the people's must lie within three points either way. A judge whose errors all run one way can clear an agreement bar and still pass more than people do, and every pass rate it measures is off by that much. If the interval straddles the margin, the judge is refused as unmeasured rather than as biased: record more labels.

Cohen's kappa is reported beside it as a point estimate with no interval, because the usual interval for it is a normal approximation and unvalidated methods are not admitted. Kappa is absent rather than zero when a marginal is degenerate. Absent is not zero: a kappa of zero says the judge agreed no better than chance, which is a measurement.

What else validate reports

Agreement alone can hide which way a judge is wrong. A judge that agrees with people 95% of the time, and whose every error is passing something they failed, runs several points above them on every pass rate it measures. So validate and evaluators list also print:

  • bias: the judge's pass rate minus the people's, in points, with an interval;
  • passes what people pass, and fails what people fail: which way it errs, when each case has

one label;

  • people agree with each other: how often two people labelling the same case agreed, where two

or more did. If people disagree with each other, the rubric needs work before any judge does.

Trying a judge before adopting it

A judge's rubric takes a few attempts. Write the draft as it would appear under evaluators: in oloproof.yaml, in a file of its own, and try it against the answers people have already labelled:

oloproof evaluators try draft.yaml

It prints the draft's agreement, bias and hit rates as validate would, and validates nothing. The system is not run again, and the draft's judgments are cached, so trying the same draft twice costs nothing and a draft you adopt finds its judgments already made.

Probing a judge without labels

Some ways a judge goes wrong need no person to find. Ask it the same question twice, and a verdict that changes is noise. Show it the same answer with a paragraph of neutral filler appended, and a verdict that changes read length rather than content:

oloproof evaluators probe RUN_ID --criterion answer_correct --probe repeat
oloproof evaluators probe RUN_ID --criterion answer_correct --probe padding --field answer

Each prints how often the verdict moved, with an interval, and padding says which way. Both are recorded beside the agreement and shown on the evaluator page. They change nothing a run stored, probing again is free, and nothing gates on them yet.

When the judge's model changes

A judge configured as gpt-4o is answered by whatever the provider currently serves under that name. Each verdict records the model that actually answered, and a validation records the ones it was measured on. If a later run is answered by a different model, that judge's rules report evaluator_recalibration_required instead of deciding: label a sample of the new run and validate again. A pinned, dated model name avoids the surprise.

Why the scaffold turns the bar off

A new project has no human labels, so it has no measured agreement, so a judge in it cannot be VALIDATED. With the bar on, the first run of any project that swaps the scaffolded deterministic evaluator for a judge returns INSUFFICIENT_EVIDENCE with evaluator_not_validated and no way forward that day.

So require_validated_evaluators is scaffolded as false and says why. Turn it on once you have labelled some cases: a release decided by a judge nobody has measured is a decision whose error rate nobody can state.

Judges on other endpoints

A judge on an endpoint that is not OpenAI's takes base_url: and api_key_env:. The key is read from the environment at call time and is never written to a stored record, a bundle, an export or a log.

A model served on this machine needs no key. Ollama, LM Studio and llama.cpp serve the OpenAI API on a loopback address, and a judge pointed at one runs with no credential, no network and no account:

evaluators:
  - type: rubric_judge
    criterion: answer_correct
    provider: openai_compatible
    model: llama3.1
    base_url: http://localhost:11434/v1
    rubric_text: "PASS if the answer conveys the same fact as the reference."

Only localhost and loopback addresses count as this machine. Any other endpoint, including one on your network, still needs its key, and a missing key fails before any request is sent. A key that is set is sent to a local server too, for servers such as vLLM that can require one.

Ollama and llama.cpp answer one request at a time by default, and Oloproof sends judge calls 4 at a time and system calls 8 at a time. The rest wait in the server's queue, and on a slow model they wait past the time limit. Lower both when the judge, or the system you evaluate, runs on such a server:

concurrency: {system: 2, judge: 2}

A run that loses calls to timeouts says so under its metrics, with this setting named.

Model evaluators

An LLM judge writes its verdict. A model evaluator scores text with a trained model and never writes anything: a classifier over one text, such as a toxicity model, or a cross-encoder over a pair, such as a natural-language inference model asked whether the answer follows from the reference. It is faster, cheaper and steadier than an LLM judge, and it runs on a laptop.

model_classifier speaks the protocol of Hugging Face Text Embeddings Inference. The verdict is one label's score against one threshold: min_score passes at or above it, max_score passes below it.

evaluators:
  - type: model_classifier
    criterion: answer_entailed
    model: cross-encoder/nli-deberta-v3-base
    base_url: http://localhost:8080
    label: entailment
    min_score: 0.5
    text: output.answer
    premise: expected.answer

text: and premise: are paths starting at input, expected or output. With a premise: the model is asked about the pair, premise first. A case whose reference lacks the premise is not applicable; an answer whose text field is missing fails the case. Every label's score is kept on the judgment, so a verdict can be read at another threshold without asking again.

A model evaluator is held to what an LLM judge is held to. It is its own kind, so it must be validated against human labels before it may gate a release. Its version names its model: before its first case it asks the server which model it serves and refuses a different one, so swapping the model behind an address cannot borrow the old version's validation. Its verdicts are cached, and the judge probes run on it unchanged. On this machine it needs no key; any other address needs api_key_env:, sent as a bearer token.

Probability judges

A rubric judge writes a verdict and Oloproof reads the text. A probability judge asks a typed question, lets the model produce exactly one token, and reads the probability the model gave each possible answer. There is no text to parse and nothing is sampled, so a call is a single forward pass: on a laptop with a local model, a tenth of a second once the model is loaded.

Three forms of question, each with the answers that pass and one threshold on them:

evaluators:
  - type: probability_judge
    criterion: answer_supported
    provider: openai_compatible
    model: qwen2.5:3b
    base_url: http://localhost:11434/v1
    form: yes_no
    question: Is every claim in the output supported by the expected answer?
    min_probability: 0.8

  - type: probability_judge
    criterion: routed_well
    provider: openai_compatible
    model: qwen2.5:3b
    base_url: http://localhost:11434/v1
    form: choice
    question: Which team should the output send the customer to?
    options:
      billing: Payments, invoices, refunds
      technical: Bugs, outages, integrations
      sales: null
    pass_options: [billing]
    min_probability: 0.7

  - type: probability_judge
    criterion: helpful
    provider: openai_compatible
    model: qwen2.5:3b
    base_url: http://localhost:11434/v1
    form: score
    question: How helpful is the output to the person who asked?
    levels:          # lowest first
      poor: null
      adequate: Answers the question
      excellent: Answers it and anticipates the next one
    pass_at_least: adequate
    min_probability: 0.75

A case passes when the probability on the passing answers — yes, the pass_options, or every level from pass_at_least up — is at least min_probability. Options and levels are shown to the model as letters A to T, so there are at most twenty.

Every answer's probability is kept on the judgment, with the most probable answer, the *coverage* (how much of the model's probability landed on an answer at all) and a *confidence* from 0, spread evenly, to 1, all on one answer. When the coverage is under one half, the model did not answer the question, and the case records an error rather than a verdict.

It needs a server that reports token probabilities: Ollama, llama.cpp, vLLM, LM Studio and OpenAI do; a server that does not is refused, and so is provider: anthropic, whose API reports none. It is never quietly switched to reading text.

A general model's probabilities are not calibrated — 90% from a small model does not mean right nine times in ten — so the confidence is description, not evidence. The judge is an LLM judge in every other respect: it must be validated against human labels before it may gate, its verdicts are cached, and the probes run on it.

Calibrating a probability judge

Once people have labelled a random sample of cases the judge stated a probability for — at least forty, and more is better — ask how far its probabilities are from theirs:

oloproof evaluators calibrate EVALUATOR_ID

It reports, with intervals:

  • the Brier score: the mean squared gap between the stated probability and the share of people

who passed the case (0 is perfect);

  • stated minus people: the mean gap with its sign: positive means the judge passes more

readily than people do;

  • a reliability table in five fixed bins of stated probability, beside how often people passed

those cases. It is descriptive: nothing decides on it.

Then it fits a two-parameter correction (Platt scaling) on half the labelled cases and checks it on the other half, reporting the held-out Brier score before and after and the difference with its interval. The halves are fixed by each case's identity, so running it again gives the same split. With a few dozen labels a mild correction is usually not shown to help: the interval says so rather than overstating it.

Nothing is validated. The command records the result on the evaluator's registry entry and prints the lines that declare the correction:

    calibration:
      slope: 0.61
      intercept: -0.18
      from_version: sha256:…

Added to the judge's entry, the correction makes a new evaluator version: it judges again, applies min_probability to the corrected probability, and keeps the uncorrected one beside it. Validate the new version as any judge; its agreement leaves out the cases the correction was fitted on, because those labels chose its threshold. A correction moves the threshold and never the order of cases, so it makes the stated number mean what it says — it does not make the judge more accurate.

Cascading to a stronger judge

A cascade asks a probability judge first and a second judge only about the cases the first is unsure of. Most cases cost one fast call; the uncertain ones get the careful look.

evaluators:
  - type: cascade
    criterion: answer_supported
    escalate_between: [0.2, 0.8]
    first:
      type: probability_judge
      provider: openai_compatible
      model: qwen2.5:3b
      base_url: http://localhost:11434/v1
      form: yes_no
      question: Is every claim in the output supported by the expected answer?
      min_probability: 0.5
    then:
      type: rubric_judge
      provider: openai
      model: gpt-4.1
      rubric_text: PASS if every claim in the output is supported by the expected answer.

A case whose first-stage probability is inside escalate_between, ends included, goes to the second stage, whose verdict and rationale are the cascade's. Outside the band the first stage's verdict stands. The band must contain the first stage's min_probability, so the second stage only ever settles what the first was unsure of and never overrules a confident verdict. The stages take the cascade's criterion; a calibration declared on the first stage applies before the band is read.

The cascade is one evaluator and is validated as one: its agreement with people is the agreement of what actually decides. The stages' own statuses do not matter to a gate. Each judgment records its route, both stages' outcomes and what they used together, and oloproof evaluators validate adds how many cases were escalated and the agreement on each route. Those two are description: the escalated cases are the harder ones by construction, so agreement on them is not the second stage's accuracy.

Choose the band before looking at the labels, from what you are willing to pay: a band tuned on the labels and then validated on the same labels looks better than it is. A first-stage error is an error, never an escalation.

Where to go next

  • Writing a suite lists the three judge types and the fields each takes.
  • Gating CI covers require_validated_evaluators in the release policy.