Skip to content

ガイド

RAG の評価

このページの英語版は、翻訳後に変更されています。最新の内容は英語版です。 英語で読む

検索拡張 (retrieval-augmented) による回答が誤る理由は 4 つあります。正しいパッセージがそもそも取得されなかった、取得されたが順位が低すぎた、十分高い順位だったがコンテキストから落とされた、あるいはモデルに届いたのにモデルがそれでも誤った、のいずれかです。単一の正解率の数値では、これらを区別できません。Oloproof は RAG システムを、中身が見える 2 つのステージとして実行し、それぞれを測定し、失敗したケースを制御された変更のもとで再実行して、どの理由に当てはまるかを突き止めます。

このページで実行するプロジェクトは examples/support_rag/ です。プロバイダーの認証情報は不要です。

ステージ化されたシステム

システムは、検索ステージと生成ステージを持つクラスで、@rag_system でデコレートします。

from oloproof import Passage, Retrieval, rag_system


@rag_system(
    name="support-rag",
    version="slice-b-example",
    depth=6,
    top_k=2,
    token_budget=40,
    index_version="kb-2026-09-16",
)
class SupportRag:
    def retrieve(self, input, depth):
        ...
        return Retrieval(query=input["question"], depth=depth, candidates=tuple(passages))

    def generate(self, input, context):
        ...
        return {"answer": answer, "citations": [best.doc_id]}

    def count_tokens(self, passage):
        return len((passage.text or "").split())

retrieve(input, depth) は最大 depth 個の候補を Passage(doc_id=..., score=..., text=...) として、リトリーバーが生成した順序で返します。Oloproof は位置を記録し、並べ替え直すことは決してありません。その後、先頭の top_k 個を残し、token_budget を超えたパッセージを落とし、残ったものを generate(input, context) に渡します。count_tokens(passage) が必要なのは token_budget を使う場合だけです。Oloproof がトークン数を推定することはありません。

oloproof.yaml はそのクラスを指し、その設定のどれでも上書きできます。

system:
  name: support-rag
  version: slice-b-example
  rag:
    object: app:SupportRag
    depth: 6
    top_k: 2
    token_budget: 40
    index_version: kb-2026-09-16

index_version は検索の同一性の一部です。インデックスが変わったらこれを変更してください。そうしないと、キャッシュされた検索結果が、もはやそれを返さないインデックスに対して再利用されてしまいます。

ケースが宣言するもの

RAG のケースは、他の種類のケースには不要な 2 つのフィールドを expected の下に持ちます。

{"id":"refund_annual","input":{"question":"How long do refunds take for annual plans?"},"expected":{"answer":"14 days","relevant":[{"doc_id":"kb-01"}],"gold_context":[{"doc_id":"kb-01","text":"Refunds for annual plans are issued within 14 days of an approved request. Support reviews each request on the same business day."}]},"metadata":{"topic":"billing"}}
  • expected.relevant は、質問に答えるドキュメントを列挙します。検索メトリクスはこれを読み、これを持たないケースは no_relevance_labels としてそれらから除外されます。
  • expected.gold_context はパッセージのテキストそのものです。診断はこれを取得されたコンテキストの代わりに差し込み、これを持たない失敗ケースは診断できません。

評価器

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: hit_rate, k: 2}
  - {type: recall, k: 2}
  - {type: ndcg, k: 6}
  - {type: citation_validity}
slices: [metadata.topic, relevant_position, context_truncated]
min_slice_support: 4

hit_rate、recall、mrr、ndcg は k をとり、それをもとに自身の基準名を付けます: hit_rate_at_2。expected.relevant のエントリは chunk_id と grade を持つこともでき、grade のデフォルトは 1 です。このとき relevance_unit: chunk を指定すると、ドキュメント全体ではなく各チャンクを 1 つの単位として数えます。citation_validity は、回答が引用するすべての ID が、与えられたコンテキスト内のパッセージを指していることを確認します。require_citations: true を指定すると、何も引用しない回答は失敗します。groundedness_judge と citation_support_judge は組み立てられたコンテキストを読む LLM ジャッジであり、他のジャッジと同様に provider と model をとります。

oloproof run

メトリクス表のうち 3 行です。

│ answer_correct  │ 69.2%    │ [38.5%, 91.0%]  │ 9 / 13 observed · 0 missing · 0 excluded     │
│ ndcg_at_6       │ 0.866    │ [0.506, 0.990]  │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0%   │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded    │

その間にある hit_rate_at_2 と recall_at_2 はどちらも 92.3%(観測 13 件中 12 件)、下限 63.9% です。1 つの質問の関連パッセージが一度も取得されなかったのです。

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 miss

Stages の行は、ステージ化されたシステム自身のキャッシュです。検索と生成は別々にキャッシュされるため、生成を変更しても検索が再実行されることはありません。

理由を突き止める

4 つの質問が失敗しました。diagnose は、取得されたコンテキストの代わりにゴールドパッセージを使ってそれらを再実行し、それと並べて、変更なしで再実行する対照を実行します。

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE
Cases: oloproof inspect sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69

ゴールドパッセージがあれば合格し、なければ失敗するケースは、モデルより上流で失敗しています。ゴールドパッセージを手にしていても失敗するケースは、モデルの失敗です。この読み方を安全にしているのが対照です。単に再実行しただけで回復するケースは不安定だったのであり、診断されたわけではありません。

診断 ID は、各件数の裏にあるケースを列挙します。

oloproof inspect DIAGNOSIS_ID
money_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2

診断が名指しするのは要因であり、原因ではありません。使う言葉で最も強いのは「Implicated」と「candidate experiment」です。1 つの介入のもとで 4 件のケースが回復しても、それらがなぜ失敗したかは立証されないからです。

修正を加える前に試す

他の 2 つの介入は、記録された検索を別の設定で再生するため、リトリーバーは呼び出されません。top-k は切り捨て位置を広げます。

oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69): 3 of 4 recovered
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Diagnosis sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE, top-k 4 run_01M3C3RKHTGGV37079V11JJ78M, replay control run_01M3C3RKJ076K7MSVQ94J4ZE2P
Cases: oloproof inspect sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085

何も回復しませんでした。これはラベルが予測したとおりです。ここでの失敗に、切り捨て位置のすぐ下に順位付けされたパッセージによるものはありませんでした。再生はゴールドコンテキストのラベルを引き継ぐため、2 つの診断は合わせて読めます。reranker は --reranker で module:function をとり、代わりにあなたのリランカーを通して検索を再生します。

実験の実行

診断はトークン予算の拡大を挙げていました。token_budget を 120 に上げて、もう一度実行します。

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 13 hit/0 miss; generate 7 hit/6 miss

すべての検索が再利用されました。top_k とトークン予算は検索の同一性に含まれないからです。コンテキストが変わった 6 件のケースだけが再び生成されました。次に、この例の比較ポリシーで比較します。

oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Comparison sha256:d801ed897f1fa1b196ca1e3f73a1ef4ffe6845d6ee8bddaea8a44da94d27a75b of run_01M3C3RYJZ4PAJ8P4P9NT71360 against run_01M3C3R5Y9D18WGSZQFA7N4XRT · 13 paired cases
answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
hit_rate_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
recall_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
ndcg_at_6: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
citations_valid: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded

続いて、各スライスとメトリクスについて探索的な行が 1 行ずつあり、その後に判定が続きます。

Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

この実験は役に立ちませんでした。13 件のうち判定が変わったものは 1 件もありません。そして 13 件の対応するケースでは、リリースが気にかけるどんな大きさの変化も立証できなかったはずです。それが区間の語るもう 1 つのことです。

次に読むページ

  • スライス では、relevant_position と context_truncated を説明しています。
  • 比較ルール では、最後のステップの判定に使われたルールを説明しています。
  • ジャッジ では、根拠性 (groundedness) ジャッジがゲートとして使えるようになるまでにクリアすべきことを説明しています。