指南
代理與工具
Oloproof 不會驅動代理。你的代理執行自己的迴圈、呼叫自己的工具,並將發生的事情記錄為一個 agent_trajectory/v1 產出物。每一項代理指標都是從這份記錄中讀出的,因此這份記錄就是整個整合。
examples/support_agent/ 是本頁所執行的專案:四十個退款請求、一個確定性的代理,而且不需要任何提供者憑證。
記錄軌跡
在系統內部,隨著事情發生逐步建立各個步驟,並交給案例記錄器:
from oloproof import AgentConstraintCheck, AgentStep, AgentTrajectory, current_case, system
@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
steps = []
for name, arguments in plan(case):
steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
result = getattr(tools, name)(**arguments)
steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))
current_case().agent_trajectory(
AgentTrajectory(
steps=tuple(steps),
terminal_status="success" if refunded else "failure",
truncated=truncated,
step_limit=STEP_LIMIT if truncated else None,
constraints=(
AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),
),
)
)
return {"answer": "refunded" if refunded else "unresolved"}各部分的用途:
| 欄位 | 記錄的內容 |
|---|---|
| AgentStep.kind | message、tool_call、tool_result、observation、decision 或 final |
| AgentStep.tool_name、arguments、result | 該次呼叫及其回傳內容 |
| terminal_status | success、failure 或 unknown |
| truncated、step_limit | 代理已達到其上限,且追蹤提前中止 |
| constraints | 你的環境所做的檢查,每項都附上它所觀察的步驟 |
| checkpoints | 重播可以從中恢復的位置,以 AgentCheckpoint 表示 |
系統會在裝飾器上以及 oloproof.yaml 中宣告它會記錄軌跡:
system:
name: support-agent
version: slice-e-example
callable: app:run
records: [agent_trajectory/v1]若沒有 records:,代理評估器會拒絕執行,而不是把每個案例都算作缺失。
案例宣告的內容
應該依特定順序呼叫特定工具的案例,會在 expected.tools 下說明:
{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}expected.tools: [] 表示該案例預期完全不呼叫任何工具,並據此評判。省略這個鍵則表示工具選擇不適用於該案例:agent_tool_sequence 和 agent_no_undeclared_tool 會以 no_declared_tools 將其排除。它會離開分母,而不是算作通過,因為用從未量測過的案例灌水的比率並不是比率。不是工具名稱清單的值會在執行前被拒絕。
評估器
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_tool_called, tool_name: lookup_order}
- {type: agent_no_tool_loop, max_repeats: 2}
- {type: agent_tool_sequence}
- {type: agent_constraints_satisfied, constraints: [no_deletion]}
- {type: agent_max_steps, max_steps: 10}
metrics:
- {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
- {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3| 類型 | 通過條件 | 參數 |
|---|---|---|
| agent_tool_called | 該工具至少被呼叫 min_calls 次 | tool_name、min_calls(預設 1) |
| agent_no_tool_loop | 沒有任何相同的呼叫(相同工具與引數)連續重複超過 max_repeats 次 | max_repeats(預設 2) |
| agent_tool_sequence | 所呼叫的工具符合 expected.tools | ordered(預設 true) |
| agent_no_undeclared_tool | 沒有呼叫 expected.tools 以外的工具 | 無 |
| agent_constraints_satisfied | 環境所記錄的每一項具名約束都通過 | constraints |
| agent_max_steps | 追蹤最多只用了 max_steps 個步驟 | max_steps |
agent_steps 和 agent_tool_calls 是最後一項背後的分布,以分位數指標呈現。
oloproof run│ answer_correct │ 82.5% │ [67.2%, 92.7%] │ 33 / 40 observed · 0 missing · 0 excluded │
│ agent_tool_lookup_order_called │ 100.0% │ [86.8%, 100.0%] │ 39 / 39 observed · 1 missing · 0 excluded │
│ agent_no_tool_loop │ 74.4% │ [56.1%, 87.4%] │ 29 / 39 observed · 1 missing · 0 excluded │
│ agent_tool_sequence │ 60.0% │ [43.3%, 75.2%] │ 24 / 40 observed · 0 missing · 0 excluded │
│ agent_constraints_satisfied │ 92.5% │ [79.6%, 98.5%] │ 37 / 40 observed · 0 missing · 0 excluded │
│ agent_steps_le_10 │ 97.5% │ [86.8%, 100.0%] │ 39 / 40 observed · 0 missing · 0 excluded │
│ steps_p95 │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50 │ 2 calls │ [2, 3] calls │ p50 of 39 observed · 1 missing · 0 excluded ││ no-deletion │ agent_constraints_satisfied │ FAIL │ observed_failures_exceed_limit │約束就像任何其他準則一樣可以作為關卡。no-deletion 是一條觀察計數規則,設定為 max_failures: 0:有三個案例呼叫了 delete_customer,而「這在我們執行的套件中絕不能發生」不需要區間就能判定。
提前中止的追蹤
有一個案例達到了代理的步數上限,因此其追蹤被記錄為 truncated。對截斷追蹤所做的每項計數都是下界,而下界能解決某些問題,卻無法解決其他問題:
| 準則 | 截斷的案例 | 原因 |
|---|---|---|
| agent_steps_le_10 | 失敗 | 已記錄的步驟就足以證明超過了十步的上限 |
| agent_tool_lookup_order_called | 缺失 | 該呼叫可能位於未被記錄的部分 |
| agent_no_tool_loop | 缺失 | 延續到追蹤末尾的重複可能會繼續下去 |
| steps_p95、tool_calls_p50 | 缺失 | 對下界取分位數並不是分位數 |
是缺失而不是排除:該案例以未觀察的狀態留在分母中,而區間允許它朝任一方向發展。這就是為什麼 agent_tool_lookup_order_called 在 39 個觀察案例上顯示為 [86.8%, 100.0%],而不是僅憑 39 個案例所能得到的較窄區間。
以軌跡切分
first_tool 依代理最先使用的工具將案例分組,不論該次執行呼叫了哪些工具。repeated_action 將重複了相同呼叫的案例與沒有重複的案例分開。trajectory_length:4,8 將案例依 1-4、5-8 及 9 步以上分桶,分界由你宣告,因為分桶邊界會改變切片所表達的內容。
│ first_tool=search │ agent_no_tool_loop │ 0.0% │ [0.0%, 57.9%] │ 0 / 6 observed · 1 missing · 0 excluded · │
│ first_tool=search │ agent_tool_sequence │ 0.0% │ [0.0%, 41.0%] │ 0 / 7 observed · 0 missing · 0 excluded · │
│ repeated_action=false │ answer_correct │ 93.1% │ [77.2%, 99.2%] │ 27 / 29 observed · 0 missing · 0 excluded · │
│ repeated_action=true │ answer_correct │ 60.0% │ [26.2%, 87.9%] │ 6 / 10 observed · 0 missing · 0 excluded · │迴圈是一個訊號,而不是解釋。輸出中沒有任何內容表示重複呼叫就是案例失敗的原因;只有移除該呼叫的重播才能說明這一點。
多個代理
examples/triage_agents/ 執行一個由三個代理組成的團隊:triage 將每個請求交給 billing 或 tech,而 billing 無法核發的退款則交給真人處理。在由多個代理組成的系統中,每個步驟都會標明執行它的代理,而控制權的轉移則是一個 handoff 步驟:
steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))軌跡要嘛為每個步驟標明代理,要嘛全都不標明。只標明部分步驟的記錄會被拒絕,而全都不標明的記錄在以下每項準則中都算作缺失,而不是通過。交接步驟是選用的,因為當下一個步驟由另一個代理執行時,控制權同樣會轉手。它是追蹤用來記錄轉移給不執行任何步驟之代理(例如真人)的方式。
案例會在 expected.route 下宣告它應該經過的代理:
{"id": "case_012", "input": {"topic": "billing", "request": "I was charged twice for order ord-012", "order_id": "ord-012", "behaviour": "escalated"}, "expected": {"answer": "escalated", "route": ["triage", "billing", "person"]}, "metadata": {"topic": "billing", "behaviour": "escalated"}}evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_route}
- type: agent_tool_permissions
permissions:
triage: []
billing: [lookup_order, issue_refund]
tech: [search_kb]
- {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]| 類型 | 通過條件 | 參數 |
|---|---|---|
| agent_route | 握有控制權的代理(合併重複後)即為該案例的 expected.route | 無 |
| agent_tool_permissions | 每次工具呼叫都是由對應表允許呼叫該工具的代理所發出 | permissions |
| agent_max_handoffs | 控制權轉手最多 max_handoffs 次 | max_handoffs |
路徑包含交接的接收者,因此以升級給真人作結的追蹤,其路徑會通往該真人。沒有 expected.route 的案例不會被 agent_route 計入。權限對應表是封閉的:表中未列出的代理不得呼叫任何工具。
│ answer_correct │ 96.7% │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1% │ [73.4%, 99.2%] │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2 │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │有兩個請求中是 tech 核發了退款。兩者都通過 answer_correct,卻未通過 agent_tool_permissions:客戶從一個違反自身權限的系統得到了正確答案。有一個請求在 billing 與 tech 之間來回交接,直到達到迴圈的上限。它的追蹤被截斷,因此在路徑與交接上限這兩項上失敗(其前段已足以判定),而在權限這一項上則為缺失(其前段無法判定)。
route 依案例中代理所走的路徑將案例分組,寫作 triage>billing。截斷的追蹤不會加入任何路徑群組,因為它的路徑只是其代理後續去向的前段。
輸出中沒有任何內容表示該歸咎於哪個代理。路徑分歧說明的是兩條路徑在何處分開。某個代理導致失敗,是關於「如果它採取不同行動會發生什麼」的主張,只有替換該代理的重播才能證明,而 Oloproof 不會執行這種重播。