Skip to content

指南

RAG 评估

本页的英文版本在翻译之后已有更改。英文页面为最新版本。 阅读英文版

检索增强生成的回答出错可能有四种不同的原因:正确的段落从未被检索到;它被检索到了,但排名太低;它的排名足够高,却随后被从上下文中丢弃;或者它到达了模型,而模型依然答错了。单一的准确率数字无法区分这些情况。Oloproof 将 RAG 系统作为两个它能看到的阶段来运行,分别度量每个阶段,并在受控的变更下重新执行失败的用例,以查明适用的是哪种原因。

examples/support_rag/ 是本页所运行的项目。它不需要任何提供方凭据。

分阶段的系统

系统是一个包含检索阶段和生成阶段的类,并使用 @rag_system 装饰:

from oloproof import Passage, Retrieval, rag_system


@rag_system(
    name="support-rag",
    version="slice-b-example",
    depth=6,
    top_k=2,
    token_budget=40,
    index_version="kb-2026-09-16",
)
class SupportRag:
    def retrieve(self, input, depth):
        ...
        return Retrieval(query=input["question"], depth=depth, candidates=tuple(passages))

    def generate(self, input, context):
        ...
        return {"answer": answer, "citations": [best.doc_id]}

    def count_tokens(self, passage):
        return len((passage.text or "").split())

retrieve(input, depth) 最多返回 depth 个候选段落,形式为 Passage(doc_id=..., score=..., text=...),顺序与你的检索器产生它们的顺序一致。Oloproof 记录这些位置,从不重新排序。随后它保留前 top_k 个,丢弃超出 token_budget 的段落,并将剩余部分传给 generate(input, context)。只有设置了 token_budget 时才需要 count_tokens(passage);Oloproof 从不估算 token 数。

oloproof.yaml 指向该类,并且可以覆盖它的任何设置:

system:
  name: support-rag
  version: slice-b-example
  rag:
    object: app:SupportRag
    depth: 6
    top_k: 2
    token_budget: 40
    index_version: kb-2026-09-16

index_version 是检索身份的一部分。索引变化时请修改它,否则已缓存的检索结果会被用于一个已不再返回这些结果的索引。

用例声明的内容

RAG 用例在 expected 下带有两个其他类型用例不需要的字段:

{"id":"refund_annual","input":{"question":"How long do refunds take for annual plans?"},"expected":{"answer":"14 days","relevant":[{"doc_id":"kb-01"}],"gold_context":[{"doc_id":"kb-01","text":"Refunds for annual plans are issued within 14 days of an approved request. Support reviews each request on the same business day."}]},"metadata":{"topic":"billing"}}
  • expected.relevant 列出能够回答该问题的文档。检索指标会读取它,没有它的用例会以 no_relevance_labels 被排除在这些指标之外。
  • expected.gold_context 是段落文本本身。诊断会用它替换检索到的上下文,没有它的失败用例无法被诊断。

评估器

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: hit_rate, k: 2}
  - {type: recall, k: 2}
  - {type: ndcg, k: 6}
  - {type: citation_validity}
slices: [metadata.topic, relevant_position, context_truncated]
min_slice_support: 4

hit_rate、recall、mrr 和 ndcg 接受 k,并据此为自己的判据命名:hit_rate_at_2。expected.relevant 中的条目还可以带有 chunk_id 和 grade,后者默认为 1;此时 relevance_unit: chunk 会把每个分块作为独立单位计数,而不是按整篇文档计数。citation_validity 检查回答引用的每个 id 是否都指向它所获得的上下文中的某个段落;设置 require_citations: true 时,没有任何引用的回答判为失败。groundedness_judge 和 citation_support_judge 是读取组装后上下文的 LLM 评判模型,与其他评判模型一样接受 provider 和 model。

oloproof run

指标表中的三行:

│ answer_correct  │ 69.2%    │ [38.5%, 91.0%]  │ 9 / 13 observed · 0 missing · 0 excluded     │
│ ndcg_at_6       │ 0.866    │ [0.506, 0.990]  │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0%   │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded    │

在它们之间,hit_rate_at_2 和 recall_at_2 都显示为 92.3%,13 个观测中有 12 个,下限为 63.9%:有一个问题的相关段落从未被检索到。

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 miss

Stages 行是分阶段系统自身的缓存。检索和生成分别缓存,因此对生成的修改永远不会重新运行检索。

查明原因

有四个问题失败了。diagnose 用黄金段落替换检索到的上下文来重新执行它们,同时设置一个对照组,原样重新执行它们:

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE
Cases: oloproof inspect sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69

一个在有黄金段落时通过、没有时失败的用例,其失败发生在模型的上游。一个手握黄金段落仍然失败的用例,问题在模型。对照组让这种解读变得可靠:在普通重跑中就恢复的用例属于不稳定,而不是被诊断出来的。

诊断 id 列出了每个计数背后的用例:

oloproof inspect DIAGNOSIS_ID
money_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2

它指出的是一个因素,而绝不是原因。"Implicated"(牵涉)和 "candidate experiment"(候选实验)是它使用的最强措辞,因为四个用例在一次干预下恢复,并不能确立它们失败的原因。

在实施修复之前先检验它

另外两种干预以不同的设置重放已记录的检索,因此不会调用检索器。top-k 放宽截断:

oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69): 3 of 4 recovered
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Diagnosis sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE, top-k 4 run_01M3C3RKHTGGV37079V11JJ78M, replay control run_01M3C3RKJ076K7MSVQ94J4ZE2P
Cases: oloproof inspect sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085

没有任何用例恢复,这与标签的预测一致:这里没有哪次失败是因为段落排名恰好落在截断线之下。重放会沿用黄金上下文的标签,因此两次诊断可以放在一起解读。reranker 接受 --reranker 以及一个 module:function,改为通过你的重排序器重放检索。

运行实验

诊断指出了更大的 token 预算。将 token_budget 提高到 120 并再次运行:

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 13 hit/0 miss; generate 7 hit/6 miss

每一次检索都被复用了,因为 top_k 和 token 预算不属于检索的身份;只有上下文发生变化的六个用例被重新生成。然后使用示例的比较策略进行比较:

oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Comparison sha256:d801ed897f1fa1b196ca1e3f73a1ef4ffe6845d6ee8bddaea8a44da94d27a75b of run_01M3C3RYJZ4PAJ8P4P9NT71360 against run_01M3C3R5Y9D18WGSZQFA7N4XRT · 13 paired cases
answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
hit_rate_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
recall_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
ndcg_at_6: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
citations_valid: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded

接下来每个切片和指标各有一行探索性输出,然后是判定:

Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

这个实验没有带来帮助:十三个用例中没有一个改变了结论。而且十三个配对用例也无法确立任何发布所关心幅度的变化,这正是区间传达的另一层信息。

下一步

  • 切片 介绍 relevant_position 和 context_truncated。
  • 比较规则 介绍最后一步所依据的规则。
  • 评判模型 介绍 groundedness 评判模型在参与门控之前必须满足的条件。