Skip to content

指南

代理與工具

Oloproof 不會驅動代理。你的代理執行自己的迴圈、呼叫自己的工具,並將發生的事情記錄為一個 agent_trajectory/v1 產出物。每一項代理指標都是從這份記錄中讀出的,因此這份記錄就是整個整合。

examples/support_agent/ 是本頁所執行的專案:四十個退款請求、一個確定性的代理,而且不需要任何提供者憑證。

記錄軌跡

在系統內部,隨著事情發生逐步建立各個步驟,並交給案例記錄器:

from oloproof import AgentConstraintCheck, AgentStep, AgentTrajectory, current_case, system


@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
    steps = []
    for name, arguments in plan(case):
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
        result = getattr(tools, name)(**arguments)
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))

    current_case().agent_trajectory(
        AgentTrajectory(
            steps=tuple(steps),
            terminal_status="success" if refunded else "failure",
            truncated=truncated,
            step_limit=STEP_LIMIT if truncated else None,
            constraints=(
                AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),
            ),
        )
    )
    return {"answer": "refunded" if refunded else "unresolved"}

各部分的用途:

欄位記錄的內容
AgentStep.kindmessage、tool_call、tool_result、observation、decision 或 final
AgentStep.tool_name、arguments、result該次呼叫及其回傳內容
terminal_statussuccess、failure 或 unknown
truncated、step_limit代理已達到其上限,且追蹤提前中止
constraints你的環境所做的檢查,每項都附上它所觀察的步驟
checkpoints重播可以從中恢復的位置,以 AgentCheckpoint 表示

系統會在裝飾器上以及 oloproof.yaml 中宣告它會記錄軌跡:

system:
  name: support-agent
  version: slice-e-example
  callable: app:run
  records: [agent_trajectory/v1]

若沒有 records:,代理評估器會拒絕執行,而不是把每個案例都算作缺失。

案例宣告的內容

應該依特定順序呼叫特定工具的案例,會在 expected.tools 下說明:

{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}

expected.tools: [] 表示該案例預期完全不呼叫任何工具,並據此評判。省略這個鍵則表示工具選擇不適用於該案例:agent_tool_sequence 和 agent_no_undeclared_tool 會以 no_declared_tools 將其排除。它會離開分母,而不是算作通過,因為用從未量測過的案例灌水的比率並不是比率。不是工具名稱清單的值會在執行前被拒絕。

評估器

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_tool_called, tool_name: lookup_order}
  - {type: agent_no_tool_loop, max_repeats: 2}
  - {type: agent_tool_sequence}
  - {type: agent_constraints_satisfied, constraints: [no_deletion]}
  - {type: agent_max_steps, max_steps: 10}
metrics:
  - {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
  - {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3
類型通過條件參數
agent_tool_called該工具至少被呼叫 min_calls 次tool_name、min_calls(預設 1)
agent_no_tool_loop沒有任何相同的呼叫(相同工具與引數)連續重複超過 max_repeats 次max_repeats(預設 2)
agent_tool_sequence所呼叫的工具符合 expected.toolsordered(預設 true)
agent_no_undeclared_tool沒有呼叫 expected.tools 以外的工具無
agent_constraints_satisfied環境所記錄的每一項具名約束都通過constraints
agent_max_steps追蹤最多只用了 max_steps 個步驟max_steps

agent_steps 和 agent_tool_calls 是最後一項背後的分布,以分位數指標呈現。

oloproof run
│ answer_correct                 │ 82.5%    │ [67.2%, 92.7%]       │ 33 / 40 observed · 0 missing · 0 excluded   │
│ agent_tool_lookup_order_called │ 100.0%   │ [86.8%, 100.0%]      │ 39 / 39 observed · 1 missing · 0 excluded   │
│ agent_no_tool_loop             │ 74.4%    │ [56.1%, 87.4%]       │ 29 / 39 observed · 1 missing · 0 excluded   │
│ agent_tool_sequence            │ 60.0%    │ [43.3%, 75.2%]       │ 24 / 40 observed · 0 missing · 0 excluded   │
│ agent_constraints_satisfied    │ 92.5%    │ [79.6%, 98.5%]       │ 37 / 40 observed · 0 missing · 0 excluded   │
│ agent_steps_le_10              │ 97.5%    │ [86.8%, 100.0%]      │ 39 / 40 observed · 0 missing · 0 excluded   │
│ steps_p95                      │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50                 │ 2 calls  │ [2, 3] calls         │ p50 of 39 observed · 1 missing · 0 excluded │
│ no-deletion         │ agent_constraints_satisfied │ FAIL                  │ observed_failures_exceed_limit │

約束就像任何其他準則一樣可以作為關卡。no-deletion 是一條觀察計數規則,設定為 max_failures: 0:有三個案例呼叫了 delete_customer,而「這在我們執行的套件中絕不能發生」不需要區間就能判定。

提前中止的追蹤

有一個案例達到了代理的步數上限,因此其追蹤被記錄為 truncated。對截斷追蹤所做的每項計數都是下界,而下界能解決某些問題,卻無法解決其他問題:

準則截斷的案例原因
agent_steps_le_10失敗已記錄的步驟就足以證明超過了十步的上限
agent_tool_lookup_order_called缺失該呼叫可能位於未被記錄的部分
agent_no_tool_loop缺失延續到追蹤末尾的重複可能會繼續下去
steps_p95、tool_calls_p50缺失對下界取分位數並不是分位數

是缺失而不是排除:該案例以未觀察的狀態留在分母中,而區間允許它朝任一方向發展。這就是為什麼 agent_tool_lookup_order_called 在 39 個觀察案例上顯示為 [86.8%, 100.0%],而不是僅憑 39 個案例所能得到的較窄區間。

以軌跡切分

first_tool 依代理最先使用的工具將案例分組,不論該次執行呼叫了哪些工具。repeated_action 將重複了相同呼叫的案例與沒有重複的案例分開。trajectory_length:4,8 將案例依 1-4、5-8 及 9 步以上分桶,分界由你宣告,因為分桶邊界會改變切片所表達的內容。

│ first_tool=search          │ agent_no_tool_loop             │ 0.0%     │ [0.0%, 57.9%]        │ 0 / 6 observed · 1 missing · 0 excluded ·          │
│ first_tool=search          │ agent_tool_sequence            │ 0.0%     │ [0.0%, 41.0%]        │ 0 / 7 observed · 0 missing · 0 excluded ·          │
│ repeated_action=false      │ answer_correct                 │ 93.1%    │ [77.2%, 99.2%]       │ 27 / 29 observed · 0 missing · 0 excluded ·        │
│ repeated_action=true       │ answer_correct                 │ 60.0%    │ [26.2%, 87.9%]       │ 6 / 10 observed · 0 missing · 0 excluded ·         │

迴圈是一個訊號,而不是解釋。輸出中沒有任何內容表示重複呼叫就是案例失敗的原因;只有移除該呼叫的重播才能說明這一點。

多個代理

examples/triage_agents/ 執行一個由三個代理組成的團隊:triage 將每個請求交給 billing 或 tech,而 billing 無法核發的退款則交給真人處理。在由多個代理組成的系統中,每個步驟都會標明執行它的代理,而控制權的轉移則是一個 handoff 步驟:

steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))

軌跡要嘛為每個步驟標明代理,要嘛全都不標明。只標明部分步驟的記錄會被拒絕,而全都不標明的記錄在以下每項準則中都算作缺失,而不是通過。交接步驟是選用的,因為當下一個步驟由另一個代理執行時,控制權同樣會轉手。它是追蹤用來記錄轉移給不執行任何步驟之代理(例如真人)的方式。

案例會在 expected.route 下宣告它應該經過的代理:

{"id": "case_012", "input": {"topic": "billing", "request": "I was charged twice for order ord-012", "order_id": "ord-012", "behaviour": "escalated"}, "expected": {"answer": "escalated", "route": ["triage", "billing", "person"]}, "metadata": {"topic": "billing", "behaviour": "escalated"}}
evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_route}
  - type: agent_tool_permissions
    permissions:
      triage: []
      billing: [lookup_order, issue_refund]
      tech: [search_kb]
  - {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
類型通過條件參數
agent_route握有控制權的代理(合併重複後)即為該案例的 expected.route無
agent_tool_permissions每次工具呼叫都是由對應表允許呼叫該工具的代理所發出permissions
agent_max_handoffs控制權轉手最多 max_handoffs 次max_handoffs

路徑包含交接的接收者,因此以升級給真人作結的追蹤,其路徑會通往該真人。沒有 expected.route 的案例不會被 agent_route 計入。權限對應表是封閉的:表中未列出的代理不得呼叫任何工具。

│ answer_correct         │ 96.7%    │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route            │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1%    │ [73.4%, 99.2%]  │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2    │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │

有兩個請求中是 tech 核發了退款。兩者都通過 answer_correct,卻未通過 agent_tool_permissions:客戶從一個違反自身權限的系統得到了正確答案。有一個請求在 billing 與 tech 之間來回交接,直到達到迴圈的上限。它的追蹤被截斷,因此在路徑與交接上限這兩項上失敗(其前段已足以判定),而在權限這一項上則為缺失(其前段無法判定)。

route 依案例中代理所走的路徑將案例分組,寫作 triage>billing。截斷的追蹤不會加入任何路徑群組,因為它的路徑只是其代理後續去向的前段。

輸出中沒有任何內容表示該歸咎於哪個代理。路徑分歧說明的是兩條路徑在何處分開。某個代理導致失敗,是關於「如果它採取不同行動會發生什麼」的主張,只有替換該代理的重播才能證明,而 Oloproof 不會執行這種重播。

下一步