Skip to content

가이드

에이전트와 도구

Oloproof는 에이전트를 구동하지 않습니다. 에이전트는 자체 루프를 실행하고, 자체 도구를 호출하며, 일어난 일을 agent_trajectory/v1 아티팩트로 기록합니다. 모든 에이전트 지표는 이 기록에서 읽어 내므로, 기록이 곧 통합의 전부입니다.

이 페이지에서 실행하는 프로젝트는 examples/support_agent/입니다. 환불 요청 마흔 건, 결정론적 에이전트로 구성되며 제공자 자격 증명은 필요 없습니다.

궤적 기록하기

시스템 안에서 단계가 일어날 때마다 만들어 케이스 기록기에 넘깁니다.

from oloproof import AgentConstraintCheck, AgentStep, AgentTrajectory, current_case, system


@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
    steps = []
    for name, arguments in plan(case):
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
        result = getattr(tools, name)(**arguments)
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))

    current_case().agent_trajectory(
        AgentTrajectory(
            steps=tuple(steps),
            terminal_status="success" if refunded else "failure",
            truncated=truncated,
            step_limit=STEP_LIMIT if truncated else None,
            constraints=(
                AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),
            ),
        )
    )
    return {"answer": "refunded" if refunded else "unresolved"}

각 부분의 용도는 다음과 같습니다.

필드기록하는 것
AgentStep.kindmessage, tool_call, tool_result, observation, decision 또는 final
AgentStep.tool_name, arguments, result호출과 돌아온 결과
terminal_statussuccess, failure 또는 unknown
truncated, step_limit에이전트가 한도에 도달해 추적이 도중에 끝났다는 사실
constraints환경이 수행한 검사, 각각 관찰한 단계와 함께
checkpoints재생을 다시 시작할 수 있는 지점, AgentCheckpoint 형태

시스템은 궤적을 기록한다는 사실을 데코레이터와 oloproof.yaml에서 선언합니다.

system:
  name: support-agent
  version: slice-e-example
  callable: app:run
  records: [agent_trajectory/v1]

records:가 없으면 에이전트 평가기는 모든 케이스를 누락으로 세는 대신 실행을 거부합니다.

케이스가 선언하는 것

특정 도구를 특정 순서로 호출해야 하는 케이스는 expected.tools 아래에 그렇게 명시합니다.

{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}

expected.tools: []는 케이스가 도구 호출을 전혀 기대하지 않는다는 뜻이며, 그 기준으로 판정됩니다. 키를 생략하면 도구 선택이 그 케이스에 해당하지 않는다는 뜻입니다. agent_tool_sequence와 agent_no_undeclared_tool은 이런 케이스를 no_declared_tools로 제외합니다. 이 케이스는 통과로 세어지는 대신 분모에서 빠집니다. 한 번도 측정되지 않은 케이스로 부풀린 비율은 비율이 아니기 때문입니다. 도구 이름 목록이 아닌 값은 실행 전에 거부됩니다.

평가기

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_tool_called, tool_name: lookup_order}
  - {type: agent_no_tool_loop, max_repeats: 2}
  - {type: agent_tool_sequence}
  - {type: agent_constraints_satisfied, constraints: [no_deletion]}
  - {type: agent_max_steps, max_steps: 10}
metrics:
  - {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
  - {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3
유형통과 조건받는 값
agent_tool_called도구가 최소 min_calls번 호출되었을 때tool_name, min_calls (기본값 1)
agent_no_tool_loop같은 도구와 같은 인수의 동일한 호출이 연속으로 max_repeats번을 넘게 반복되지 않았을 때max_repeats (기본값 2)
agent_tool_sequence호출된 도구가 expected.tools와 일치할 때ordered (기본값 true)
agent_no_undeclared_toolexpected.tools 밖의 도구가 호출되지 않았을 때없음
agent_constraints_satisfied환경이 기록한, 이름이 지정된 모든 제약이 통과했을 때constraints
agent_max_steps추적이 최대 max_steps 단계 이내였을 때max_steps

agent_steps와 agent_tool_calls는 마지막 평가기 뒤에 있는 분포이며, 분위수 지표로 쓰입니다.

oloproof run
│ answer_correct                 │ 82.5%    │ [67.2%, 92.7%]       │ 33 / 40 observed · 0 missing · 0 excluded   │
│ agent_tool_lookup_order_called │ 100.0%   │ [86.8%, 100.0%]      │ 39 / 39 observed · 1 missing · 0 excluded   │
│ agent_no_tool_loop             │ 74.4%    │ [56.1%, 87.4%]       │ 29 / 39 observed · 1 missing · 0 excluded   │
│ agent_tool_sequence            │ 60.0%    │ [43.3%, 75.2%]       │ 24 / 40 observed · 0 missing · 0 excluded   │
│ agent_constraints_satisfied    │ 92.5%    │ [79.6%, 98.5%]       │ 37 / 40 observed · 0 missing · 0 excluded   │
│ agent_steps_le_10              │ 97.5%    │ [86.8%, 100.0%]      │ 39 / 40 observed · 0 missing · 0 excluded   │
│ steps_p95                      │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50                 │ 2 calls  │ [2, 3] calls         │ p50 of 39 observed · 1 missing · 0 excluded │
│ no-deletion         │ agent_constraints_satisfied │ FAIL                  │ observed_failures_exceed_limit │

제약은 다른 기준과 똑같이 게이트합니다. no-deletion은 max_failures: 0인 관찰 개수 규칙입니다. 세 케이스가 delete_customer를 호출했으며, "우리가 실행한 스위트에서 이런 일은 일어나서는 안 된다"는 결정에는 구간이 필요 없습니다.

도중에 끝난 추적

한 케이스가 에이전트의 단계 한도에 도달했으므로 그 추적은 truncated로 기록됩니다. 잘린 추적에 대한 모든 개수는 하한이며, 하한은 어떤 질문은 해결하고 어떤 질문은 해결하지 못합니다.

기준잘린 케이스이유
agent_steps_le_10실패기록된 단계만으로 이미 한도 10을 넘었음이 입증됩니다
agent_tool_lookup_order_called누락기록되지 않은 부분에 호출이 있을 수 있습니다
agent_no_tool_loop누락추적 끝에 이른 반복은 그 너머로 이어질 수 있습니다
steps_p95, tool_calls_p50누락하한들에 대한 분위수는 분위수가 아닙니다

제외가 아니라 누락입니다. 케이스는 관찰되지 않은 채 분모에 남고, 구간은 그 케이스가 어느 쪽으로든 갔을 가능성을 허용합니다. 그래서 agent_tool_lookup_order_called는 관찰된 케이스 39개에 대해, 케이스 39개만으로 얻을 더 좁은 구간 대신 [86.8%, 100.0%]로 표시됩니다.

궤적에 대한 슬라이스

first_tool은 실행이 어떤 도구를 호출했든, 에이전트가 처음 손을 뻗은 도구로 케이스를 묶습니다. repeated_action은 동일한 호출을 반복한 케이스와 그렇지 않은 케이스를 나눕니다. trajectory_length:4,8은 케이스를 1-4, 5-8, 9 이상의 단계로 나누며, 경계를 직접 선언하게 하는 것은 구간 경계가 슬라이스가 말하는 내용을 바꾸기 때문입니다.

│ first_tool=search          │ agent_no_tool_loop             │ 0.0%     │ [0.0%, 57.9%]        │ 0 / 6 observed · 1 missing · 0 excluded ·          │
│ first_tool=search          │ agent_tool_sequence            │ 0.0%     │ [0.0%, 41.0%]        │ 0 / 7 observed · 0 missing · 0 excluded ·          │
│ repeated_action=false      │ answer_correct                 │ 93.1%    │ [77.2%, 99.2%]       │ 27 / 29 observed · 0 missing · 0 excluded ·        │
│ repeated_action=true       │ answer_correct                 │ 60.0%    │ [26.2%, 87.9%]       │ 6 / 10 observed · 0 missing · 0 excluded ·         │

루프는 신호일 뿐 설명이 아닙니다. 출력의 어디에도 반복 호출이 케이스 실패의 이유라고 말하지 않습니다. 그것을 말할 수 있는 것은 반복 호출을 제거한 재생뿐입니다.

여러 에이전트

examples/triage_agents/는 세 에이전트로 이루어진 팀을 실행합니다. triage가 각 요청을 billing 또는 tech에 넘기고, billing이 발행할 수 없는 환불은 사람에게 넘겨집니다. 여러 에이전트로 이루어진 시스템에서는 각 단계가 그 단계를 수행한 에이전트의 이름을 가지며, 제어권 이전은 handoff 단계입니다.

steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))

궤적은 모든 단계의 에이전트 이름을 지정하거나, 어느 단계도 지정하지 않습니다. 일부만 지정한 기록은 거부되며, 전혀 지정하지 않은 기록은 아래의 모든 기준에서 통과가 아니라 누락이 됩니다. 핸드오프 단계는 선택 사항입니다. 다음 단계를 다른 에이전트가 수행할 때도 제어권이 넘어가기 때문입니다. 핸드오프 단계는 사람처럼 아무 단계도 수행하지 않는 대상에게 제어권이 넘어간 것을 추적이 기록하는 방법입니다.

케이스는 거쳐야 할 에이전트를 expected.route 아래에 선언합니다.

{"id": "case_012", "input": {"topic": "billing", "request": "I was charged twice for order ord-012", "order_id": "ord-012", "behaviour": "escalated"}, "expected": {"answer": "escalated", "route": ["triage", "billing", "person"]}, "metadata": {"topic": "billing", "behaviour": "escalated"}}
evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_route}
  - type: agent_tool_permissions
    permissions:
      triage: []
      billing: [lookup_order, issue_refund]
      tech: [search_kb]
  - {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
유형통과 조건받는 값
agent_route제어권을 가졌던 에이전트들(반복은 합침)이 케이스의 expected.route와 같을 때없음
agent_tool_permissions모든 도구 호출이 권한 맵에서 그 도구 호출이 허용된 에이전트에 의해 이루어졌을 때permissions
agent_max_handoffs제어권이 최대 max_handoffs번 넘어갔을 때max_handoffs

경로에는 핸드오프의 수신자가 포함되므로, 사람에게 에스컬레이션하며 끝나는 추적의 경로는 그 사람에게 이릅니다. expected.route가 없는 케이스는 agent_route에서 세지 않습니다. 권한 맵은 닫혀 있습니다. 맵에 없는 에이전트는 어떤 도구도 호출할 수 없습니다.

│ answer_correct         │ 96.7%    │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route            │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1%    │ [73.4%, 99.2%]  │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2    │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │

두 요청에서 tech가 환불을 발행했습니다. 둘 다 answer_correct는 통과하고 agent_tool_permissions는 실패합니다. 고객은 자기 권한을 어긴 시스템으로부터 올바른 답을 받은 것입니다. 한 요청은 루프의 한도에 이를 때까지 billing과 tech 사이를 오갑니다. 그 추적은 잘렸으므로, 앞부분만으로 이미 결정되는 경로와 핸드오프 한도에서는 실패하고, 그렇지 않은 권한에서는 누락됩니다.

route는 에이전트들이 거친 경로로 케이스를 묶으며, triage>billing 형식으로 표기합니다. 잘린 추적은 어느 경로 그룹에도 들어가지 않습니다. 그 경로는 에이전트들이 그다음에 간 곳의 앞부분일 뿐이기 때문입니다.

출력의 어디에도 어느 에이전트에게 책임이 있는지 말하지 않습니다. 경로 분기는 두 경로가 어디서 갈라지는지를 말할 뿐입니다. 어떤 에이전트가 실패를 일으켰다는 것은 그 에이전트가 달리 행동했다면 무슨 일이 일어났을지에 대한 주장이며, 그것은 그 에이전트를 대체한 재생으로만 보여 줄 수 있습니다. Oloproof는 그런 재생을 실행하지 않습니다.

다음 단계