Skip to content

指南

RAG 評估

本頁的英文版本在翻譯之後已有變更。英文頁面為最新版本。 閱讀英文版

一個檢索增強的回答可能因為四種不同的原因而出錯:正確的段落從未被檢索到;它被檢索到了,但排名太低;它的排名夠高,之後卻從脈絡中被捨棄;或者它送到了模型面前,而模型仍然答錯。單一的準確率數字無法區分這些情況。Oloproof 將 RAG 系統當作兩個它看得見的階段來執行,分別量測每個階段,並在受控的變更下重新執行失敗的案例,以找出適用的是哪一種原因。

examples/support_rag/ 是本頁所執行的專案。它不需要任何供應商憑證。

分階段的系統

系統是一個具有檢索階段與生成階段的類別,並以 @rag_system 裝飾:

from oloproof import Passage, Retrieval, rag_system


@rag_system(
    name="support-rag",
    version="slice-b-example",
    depth=6,
    top_k=2,
    token_budget=40,
    index_version="kb-2026-09-16",
)
class SupportRag:
    def retrieve(self, input, depth):
        ...
        return Retrieval(query=input["question"], depth=depth, candidates=tuple(passages))

    def generate(self, input, context):
        ...
        return {"answer": answer, "citations": [best.doc_id]}

    def count_tokens(self, passage):
        return len((passage.text or "").split())

retrieve(input, depth) 最多回傳 depth 個候選段落,形式為 Passage(doc_id=..., score=..., text=...),順序與你的檢索器產生它們的順序相同。Oloproof 會記錄這些位置,且絕不重新排序。接著它保留前 top_k 個,捨棄超出 token_budget 的段落,並將剩下的傳給 generate(input, context)。只有設定了 token_budget 時才需要 count_tokens(passage);Oloproof 絕不估算 token 數。

oloproof.yaml 指向這個類別,並可以覆寫它的任何設定:

system:
  name: support-rag
  version: slice-b-example
  rag:
    object: app:SupportRag
    depth: 6
    top_k: 2
    token_budget: 40
    index_version: kb-2026-09-16

index_version 是檢索身分的一部分。索引改變時請修改它,否則快取的檢索結果會被重複用在一個已不再回傳它們的索引上。

案例宣告的內容

RAG 案例在 expected 之下帶有兩個其他種類的案例都不需要的欄位:

{"id":"refund_annual","input":{"question":"How long do refunds take for annual plans?"},"expected":{"answer":"14 days","relevant":[{"doc_id":"kb-01"}],"gold_context":[{"doc_id":"kb-01","text":"Refunds for annual plans are issued within 14 days of an approved request. Support reviews each request on the same business day."}]},"metadata":{"topic":"billing"}}
  • expected.relevant 列出能回答該問題的文件。檢索指標會讀取它,而沒有它的案例會以 no_relevance_labels 被排除在這些指標之外。
  • expected.gold_context 就是段落文字本身。診斷會以它取代檢索到的脈絡,而沒有它的失敗案例無法被診斷。

評估器

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: hit_rate, k: 2}
  - {type: recall, k: 2}
  - {type: ndcg, k: 6}
  - {type: citation_validity}
slices: [metadata.topic, relevant_position, context_truncated]
min_slice_support: 4

hit_rate、recall、mrr 與 ndcg 接受 k,並據此為自己的判準命名:hit_rate_at_2。expected.relevant 中的項目也可以帶有 chunk_id 與 grade,後者預設為 1;此時 relevance_unit: chunk 會把每個區塊當成獨立單位計數,而不是以整份文件計數。citation_validity 檢查回答所引用的每個 id 都指向它所拿到的脈絡中的某個段落;設定 require_citations: true 時,沒有任何引用的回答即為失敗。groundedness_judge 與 citation_support_judge 是讀取組裝後脈絡的 LLM 評審,並像其他評審一樣接受 provider 與 model。

oloproof run

指標表中的三列:

│ answer_correct  │ 69.2%    │ [38.5%, 91.0%]  │ 9 / 13 observed · 0 missing · 0 excluded     │
│ ndcg_at_6       │ 0.866    │ [0.506, 0.990]  │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0%   │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded    │

在這兩者之間,hit_rate_at_2 與 recall_at_2 都顯示 92.3%,13 個已觀察中有 12 個,下界為 63.9%:有一個問題的相關段落從未被檢索到。

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 miss

Stages 這一行是分階段系統自己的快取。檢索與生成分開快取,因此對生成的變更絕不會重新執行檢索。

找出原因

有四個問題失敗了。diagnose 以黃金段落取代檢索到的脈絡來重新執行它們,並在旁邊設置一個原封不動重新執行它們的對照組:

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE
Cases: oloproof inspect sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69

一個有黃金段落時通過、沒有時失敗的案例,是在模型的上游失敗的。一個手握黃金段落仍然失敗的案例,則是模型的問題。對照組讓這種解讀可靠:在單純重新執行時就恢復的案例是不穩定,而不是被診斷出來的。

診斷 id 會列出每個計數背後的案例:

oloproof inspect DIAGNOSIS_ID
money_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2

它指出的是一個因素,而絕不是原因。「Implicated」與「candidate experiment」是它使用的最強字眼,因為四個案例在一項介入下恢復,並不能確立它們失敗的原因。

在動手修正之前先測試修正

另外兩種介入會以不同的設定重播已記錄的檢索,因此不會呼叫檢索器。top-k 放寬截斷:

oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69): 3 of 4 recovered
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Diagnosis sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE, top-k 4 run_01M3C3RKHTGGV37079V11JJ78M, replay control run_01M3C3RKJ076K7MSVQ94J4ZE2P
Cases: oloproof inspect sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085

沒有任何案例恢復,這正是標註所預測的:這裡沒有任何失敗是段落排名恰好落在截斷線之下。重播會沿用黃金脈絡的標註,因此兩份診斷可以一起解讀。reranker 接受 --reranker 與一個 module:function,改為透過你的重新排序器重播檢索。

執行實驗

診斷指出了更大的 token 預算。將 token_budget 提高到 120 並再執行一次:

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 13 hit/0 miss; generate 7 hit/6 miss

每一次檢索都被重複使用了,因為 top_k 與 token 預算不屬於檢索的身分;只有脈絡改變的六個案例被重新生成。接著使用範例的比較政策進行比較:

oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Comparison sha256:d801ed897f1fa1b196ca1e3f73a1ef4ffe6845d6ee8bddaea8a44da94d27a75b of run_01M3C3RYJZ4PAJ8P4P9NT71360 against run_01M3C3R5Y9D18WGSZQFA7N4XRT · 13 paired cases
answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
hit_rate_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
recall_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
ndcg_at_6: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
citations_valid: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded

之後每個切片與指標各有一行探索性輸出,然後是決策:

Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

這個實驗沒有幫助:十三個案例中沒有一個改變了判定。而且十三個成對案例也不可能確立任何發布所在意之幅度的變化,這是區間所說的另一件事。

下一步

  • 切片 說明 relevant_position 與 context_truncated。
  • 比較規則 說明最後一步所依據的規則。
  • 評審 說明 groundedness 評審在可以設閘之前必須通過什麼。