Skip to content

Guides

RAG evaluation

A retrieval-augmented answer can be wrong for four different reasons: the right passage was never retrieved, it was retrieved and ranked too low, it was ranked high enough and then dropped from the context, or it reached the model and the model got it wrong anyway. A single accuracy number cannot tell them apart. Oloproof runs a RAG system as two stages it can see, measures each, and re-executes failed cases under controlled changes to find out which reason applies.

examples/support_rag/ is the project this page runs. It needs no provider credentials.

A staged system

The system is a class with a retrieval stage and a generation stage, decorated with @rag_system:

from oloproof import Passage, Retrieval, rag_system


@rag_system(
    name="support-rag",
    version="slice-b-example",
    depth=6,
    top_k=2,
    token_budget=40,
    index_version="kb-2026-09-16",
)
class SupportRag:
    def retrieve(self, input, depth):
        ...
        return Retrieval(query=input["question"], depth=depth, candidates=tuple(passages))

    def generate(self, input, context):
        ...
        return {"answer": answer, "citations": [best.doc_id]}

    def count_tokens(self, passage):
        return len((passage.text or "").split())

retrieve(input, depth) returns up to depth candidates, as Passage(doc_id=..., score=..., text=...), in the order your retriever produced them. Oloproof records the positions and never re-ranks. It then keeps the first top_k, drops passages past token_budget, and passes what is left to generate(input, context). count_tokens(passage) is needed only with a token_budget; Oloproof never estimates tokens.

oloproof.yaml points at the class and may override any of its settings:

system:
  name: support-rag
  version: slice-b-example
  rag:
    object: app:SupportRag
    depth: 6
    top_k: 2
    token_budget: 40
    index_version: kb-2026-09-16

index_version is part of the retrieval's identity. Change it when the index changes, or cached retrievals will be reused against an index that no longer returns them.

The code of retrieve is part of that identity too, down to the class that defines it. Two prompt versions written as two subclasses therefore search again for every case, even when they inherit the same retrieve, because retrieve could read an attribute the subclass changed. To compare prompts over one set of cached retrievals, keep one class and vary the prompt through its configuration.

What a case declares

A RAG case carries two fields under expected that no other kind of case needs:

{"id":"refund_annual","input":{"question":"How long do refunds take for annual plans?"},"expected":{"answer":"14 days","relevant":[{"doc_id":"kb-01"}],"gold_context":[{"doc_id":"kb-01","text":"Refunds for annual plans are issued within 14 days of an approved request. Support reviews each request on the same business day."}]},"metadata":{"topic":"billing"}}
  • expected.relevant lists the documents that answer the question. The retrieval metrics read

it, and a case without it is excluded from them with no_relevance_labels.

  • expected.gold_context is the passage text itself. Diagnosis substitutes it for the retrieved

context, and a failed case without it cannot be diagnosed.

The evaluators

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: hit_rate, k: 2}
  - {type: recall, k: 2}
  - {type: ndcg, k: 6}
  - {type: citation_validity}
slices: [metadata.topic, relevant_position, context_truncated]
min_slice_support: 4

hit_rate, recall, mrr and ndcg take k and name their own criterion from it: hit_rate_at_2. An entry in expected.relevant may also carry a chunk_id and a grade, which defaults to 1; relevance_unit: chunk then counts each chunk as its own unit rather than whole documents. citation_validity checks that every id an answer cites names a passage in the context it was given, and with require_citations: true an answer that cites nothing fails. groundedness_judge and citation_support_judge are LLM judges that read the assembled context, and take a provider and model like any other judge.

oloproof run

Three rows of the metrics table:

│ answer_correct  │ 69.2%    │ [38.5%, 91.0%]  │ 9 / 13 observed · 0 missing · 0 excluded     │
│ ndcg_at_6       │ 0.866    │ [0.506, 0.990]  │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0%   │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded    │

Between them, hit_rate_at_2 and recall_at_2 both read 92.3%, 12 of 13 observed, with a lower bound of 63.9%: one question's relevant passage was never retrieved.

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 miss

The Stages line is the staged system's own cache. Retrieval and generation are cached separately, so a change to generation never re-runs retrieval.

Finding out why

Four questions failed. diagnose re-executes them with the gold passage in place of the retrieved context, beside a control that re-executes them unchanged:

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE
Cases: oloproof inspect sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69

A case that passes with the gold passage and fails without it failed upstream of the model. A case that fails with the gold passage in hand is the model's. The control is what makes that reading safe: a case that recovers on a plain re-run was flaky, not diagnosed.

The diagnosis id lists the case behind every count:

oloproof inspect DIAGNOSIS_ID
money_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2

It names a factor, never a cause. "Implicated" and "candidate experiment" are the strongest words it uses, because four cases recovering under one intervention do not establish why they failed.

Testing a fix before making it

Two other interventions replay the recorded retrieval with a different setting, so no retriever call is made. top-k widens the cut:

oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69): 3 of 4 recovered
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Diagnosis sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE, top-k 4 run_01M3C3RKHTGGV37079V11JJ78M, replay control run_01M3C3RKJ076K7MSVQ94J4ZE2P
Cases: oloproof inspect sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085

Nothing recovered, which is what the labels predicted: no failure here was a passage ranked just below the cut. The replay carries the gold-context labels forward, so the two diagnoses read together. reranker takes --reranker with a module:function and replays the retrieval through your reranker instead.

Running the experiment

The diagnosis named a larger token budget. Raise token_budget to 120 and run again:

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 13 hit/0 miss; generate 7 hit/6 miss

Every retrieval was reused, because top_k and the token budget are outside the retrieval's identity; only the six cases whose context changed were generated again. Then compare, with the example's comparison policy:

oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Comparison sha256:d801ed897f1fa1b196ca1e3f73a1ef4ffe6845d6ee8bddaea8a44da94d27a75b of run_01M3C3RYJZ4PAJ8P4P9NT71360 against run_01M3C3R5Y9D18WGSZQFA7N4XRT · 13 paired cases
answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
hit_rate_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
recall_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
ndcg_at_6: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
citations_valid: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded

One exploratory line follows for each slice and metric, and then the decisions:

Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

The experiment did not help: not one of the thirteen changed its verdict. And thirteen paired cases could not have established a change of any size a release cares about, which is the other thing the interval says.

Where to go next

  • Slices covers relevant_position and context_truncated.
  • Comparison rules covers the rules the last step was decided by.
  • Judges covers what a groundedness judge must clear before it may gate.