Skip to content

가이드

RAG 평가

이 페이지의 영어판이 번역 이후 변경되었습니다. 최신 페이지는 영어판입니다. 영어로 읽기

검색 증강 답변이 틀리는 이유는 네 가지로 서로 다를 수 있습니다. 올바른 구절이 아예 검색되지 않았거나, 검색되었지만 순위가 너무 낮았거나, 순위는 충분히 높았지만 컨텍스트에서 빠졌거나, 모델에 도달했는데도 모델이 틀렸을 수 있습니다. 하나의 정확도 수치로는 이들을 구별할 수 없습니다. Oloproof는 RAG 시스템을 볼 수 있는 두 단계로 실행하고, 각 단계를 측정하며, 실패한 케이스를 통제된 변경 아래에서 다시 실행해 어떤 이유에 해당하는지 알아냅니다.

이 페이지에서 실행하는 프로젝트는 examples/support_rag/입니다. 제공자 자격 증명은 필요 없습니다.

단계로 나뉜 시스템

시스템은 검색 단계와 생성 단계를 가진 클래스이며, @rag_system으로 데코레이트합니다.

from oloproof import Passage, Retrieval, rag_system


@rag_system(
    name="support-rag",
    version="slice-b-example",
    depth=6,
    top_k=2,
    token_budget=40,
    index_version="kb-2026-09-16",
)
class SupportRag:
    def retrieve(self, input, depth):
        ...
        return Retrieval(query=input["question"], depth=depth, candidates=tuple(passages))

    def generate(self, input, context):
        ...
        return {"answer": answer, "citations": [best.doc_id]}

    def count_tokens(self, passage):
        return len((passage.text or "").split())

retrieve(input, depth)는 최대 depth개의 후보를 Passage(doc_id=..., score=..., text=...) 형태로, 검색기가 만든 순서대로 반환합니다. Oloproof는 위치를 기록할 뿐 순위를 다시 매기지 않습니다. 그런 다음 처음 top_k개를 유지하고, token_budget을 넘는 구절을 빼고, 남은 것을 generate(input, context)에 전달합니다. count_tokens(passage)는 token_budget이 있을 때만 필요합니다. Oloproof는 토큰 수를 추정하지 않습니다.

oloproof.yaml은 이 클래스를 가리키며, 클래스의 어떤 설정이든 재정의할 수 있습니다.

system:
  name: support-rag
  version: slice-b-example
  rag:
    object: app:SupportRag
    depth: 6
    top_k: 2
    token_budget: 40
    index_version: kb-2026-09-16

index_version은 검색의 식별 정보에 포함됩니다. 인덱스가 바뀌면 이 값을 바꾸십시오. 그렇지 않으면 캐시된 검색 결과가 더 이상 그 결과를 반환하지 않는 인덱스에 대해 재사용됩니다.

케이스가 선언하는 것

RAG 케이스는 다른 종류의 케이스에는 필요 없는 두 필드를 expected 아래에 가집니다.

{"id":"refund_annual","input":{"question":"How long do refunds take for annual plans?"},"expected":{"answer":"14 days","relevant":[{"doc_id":"kb-01"}],"gold_context":[{"doc_id":"kb-01","text":"Refunds for annual plans are issued within 14 days of an approved request. Support reviews each request on the same business day."}]},"metadata":{"topic":"billing"}}
  • expected.relevant는 질문에 답하는 문서를 나열합니다. 검색 지표가 이를 읽으며, 이 필드가 없는

케이스는 no_relevance_labels로 검색 지표에서 제외됩니다.

  • expected.gold_context는 구절 텍스트 자체입니다. 진단은 검색된 컨텍스트 대신 이것을 넣으며, 이

필드가 없는 실패 케이스는 진단할 수 없습니다.

평가기

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: hit_rate, k: 2}
  - {type: recall, k: 2}
  - {type: ndcg, k: 6}
  - {type: citation_validity}
slices: [metadata.topic, relevant_position, context_truncated]
min_slice_support: 4

hit_rate, recall, mrr, ndcg는 k를 받고 그것으로 자신의 기준 이름을 정합니다: hit_rate_at_2. expected.relevant의 항목에는 chunk_id와 grade도 넣을 수 있으며, grade의 기본값은 1입니다. 이때 relevance_unit: chunk는 문서 전체가 아니라 각 청크를 하나의 단위로 셉니다. citation_validity는 답변이 인용한 모든 id가 그 답변에 주어진 컨텍스트의 구절을 가리키는지 확인하며, require_citations: true이면 아무것도 인용하지 않은 답변은 실패합니다. groundedness_judge와 citation_support_judge는 조립된 컨텍스트를 읽는 LLM 심사 모델이며, 다른 심사 모델과 마찬가지로 provider와 model을 받습니다.

oloproof run

지표 표의 세 행입니다.

│ answer_correct  │ 69.2%    │ [38.5%, 91.0%]  │ 9 / 13 observed · 0 missing · 0 excluded     │
│ ndcg_at_6       │ 0.866    │ [0.506, 0.990]  │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0%   │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded    │

그 사이에서 hit_rate_at_2와 recall_at_2는 모두 92.3%, 관찰된 13개 중 12개이며 하한은 63.9%입니다. 한 질문의 관련 구절이 아예 검색되지 않았습니다.

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 miss

Stages 줄은 단계로 나뉜 시스템 자체의 캐시입니다. 검색과 생성은 따로 캐시되므로, 생성을 바꿔도 검색은 다시 실행되지 않습니다.

원인 찾기

네 질문이 실패했습니다. diagnose는 검색된 컨텍스트 대신 정답 구절을 넣어 이들을 다시 실행하고, 그 옆에서 아무것도 바꾸지 않고 다시 실행하는 대조군을 함께 실행합니다.

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE
Cases: oloproof inspect sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69

정답 구절이 있으면 통과하고 없으면 실패하는 케이스는 모델보다 앞 단계에서 실패한 것입니다. 정답 구절을 받고도 실패하는 케이스는 모델의 몫입니다. 이렇게 읽는 것을 안전하게 만드는 것이 대조군입니다. 단순히 다시 실행해서 회복된 케이스는 진단된 것이 아니라 불안정했던 것입니다.

진단 id는 모든 개수 뒤에 있는 케이스를 나열합니다.

oloproof inspect DIAGNOSIS_ID
money_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2

진단은 요인을 지목할 뿐 원인을 지목하지 않습니다. 가장 강한 표현은 "Implicated"와 "candidate experiment"입니다. 하나의 개입 아래에서 네 케이스가 회복되었다고 해서 그들이 왜 실패했는지가 입증되지는 않기 때문입니다.

수정하기 전에 시험하기

다른 두 개입은 기록된 검색을 다른 설정으로 재생하므로, 검색기를 호출하지 않습니다. top-k는 잘라내는 지점을 넓힙니다.

oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:b2846cbe04099225bda4bf21beb1738f54d6cc0e9a7ce210bb3eb467619cbd69): 3 of 4 recovered
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Diagnosis sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085
Child runs: gold context run_01M3C3REAAZDF9FD8N2C8BG17P, control run_01M3C3REAFVXK2VM4MYAQYH6KE, top-k 4 run_01M3C3RKHTGGV37079V11JJ78M, replay control run_01M3C3RKJ076K7MSVQ94J4ZE2P
Cases: oloproof inspect sha256:668084c3e99df84f7b63f138792aa96f48e5cb8251fee3acae8e1cc0500ba085

아무것도 회복되지 않았으며, 이는 레이블이 예측한 그대로입니다. 여기의 실패 중 어느 것도 잘라내는 지점 바로 아래에 순위가 매겨진 구절 때문이 아니었습니다. 재생은 정답 컨텍스트 레이블을 이어받으므로 두 진단을 함께 읽을 수 있습니다. reranker는 module:function 형식의 --reranker를 받아, 검색 결과를 대신 여러분의 재순위 모델로 재생합니다.

실험 실행하기

진단은 더 큰 토큰 예산을 지목했습니다. token_budget을 120으로 올리고 다시 실행합니다.

Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 13 hit/0 miss; generate 7 hit/6 miss

모든 검색이 재사용되었습니다. top_k와 토큰 예산은 검색의 식별 정보 밖에 있기 때문입니다. 컨텍스트가 바뀐 여섯 케이스만 다시 생성되었습니다. 그런 다음 예제의 비교 정책으로 비교합니다.

oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Comparison sha256:d801ed897f1fa1b196ca1e3f73a1ef4ffe6845d6ee8bddaea8a44da94d27a75b of run_01M3C3RYJZ4PAJ8P4P9NT71360 against run_01M3C3R5Y9D18WGSZQFA7N4XRT · 13 paired cases
answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
hit_rate_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
recall_at_2: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
ndcg_at_6: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
citations_valid: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded

슬라이스와 지표마다 탐색용 줄이 하나씩 이어지고, 그다음에 결정이 나옵니다.

Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

실험은 도움이 되지 않았습니다. 열세 케이스 중 판정이 바뀐 것은 하나도 없습니다. 그리고 짝지은 케이스 열세 개로는 릴리스가 신경 쓸 만한 어떤 크기의 변화도 입증할 수 없었을 것이며, 이것이 구간이 말하는 또 다른 사실입니다.

다음 단계

  • 슬라이스는 relevant_position과 context_truncated를 다룹니다.
  • 비교 규칙은 마지막 단계의 결정 기준이 된 규칙을 다룹니다.
  • 심사 모델은 근거성 심사 모델이 게이트하기 전에 통과해야 하는 기준을 다룹니다.