指南
智能体与工具
Oloproof 不驱动智能体。你的智能体运行自己的循环,调用自己的工具,并将所发生的事情记录为一个 agent_trajectory/v1 制品。每个智能体指标都从这条记录中读取,因此这条记录就是全部的集成工作。
examples/support_agent/ 是本页所运行的项目:四十个退款请求、一个确定性的智能体,无需任何提供方凭据。
记录轨迹
在系统内部,随着事情发生构建各个步骤,并将它们交给用例记录器:
from oloproof import AgentConstraintCheck, AgentStep, AgentTrajectory, current_case, system
@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
steps = []
for name, arguments in plan(case):
steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
result = getattr(tools, name)(**arguments)
steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))
current_case().agent_trajectory(
AgentTrajectory(
steps=tuple(steps),
terminal_status="success" if refunded else "failure",
truncated=truncated,
step_limit=STEP_LIMIT if truncated else None,
constraints=(
AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),
),
)
)
return {"answer": "refunded" if refunded else "unresolved"}各部分的用途:
| 字段 | 记录的内容 |
|---|---|
| AgentStep.kind | message、tool_call、tool_result、observation、decision 或 final |
| AgentStep.tool_name、arguments、result | 调用以及返回的内容 |
| terminal_status | success、failure 或 unknown |
| truncated、step_limit | 智能体触及了步数上限,轨迹提前截止 |
| constraints | 你的环境所做的检查,每项都附带它所观察的步骤 |
| checkpoints | 重放可以从中恢复的位置,形式为 AgentCheckpoint |
系统需要在装饰器上和 oloproof.yaml 中声明它会记录轨迹:
system:
name: support-agent
version: slice-e-example
callable: app:run
records: [agent_trajectory/v1]没有 records: 时,智能体评估器会拒绝运行,而不是把每个用例都计为缺失。
用例声明的内容
应当按特定顺序调用特定工具的用例,在 expected.tools 下声明:
{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}expected.tools: [] 表示该用例预期完全不调用工具,并据此评判。省略该键表示工具选择不适用于该用例:agent_tool_sequence 和 agent_no_undeclared_tool 会以 no_declared_tools 将其排除。它会离开分母,而不是计为通过,因为混入了从未被度量的用例而虚高的比率并不是真正的比率。不是工具名列表的值会在运行之前被拒绝。
评估器
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_tool_called, tool_name: lookup_order}
- {type: agent_no_tool_loop, max_repeats: 2}
- {type: agent_tool_sequence}
- {type: agent_constraints_satisfied, constraints: [no_deletion]}
- {type: agent_max_steps, max_steps: 10}
metrics:
- {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
- {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3| 类型 | 通过条件 | 参数 |
|---|---|---|
| agent_tool_called | 该工具至少被调用 min_calls 次 | tool_name、min_calls(默认 1) |
| agent_no_tool_loop | 没有任何完全相同的调用(相同工具和参数)连续重复超过 max_repeats 次 | max_repeats(默认 2) |
| agent_tool_sequence | 调用的工具与 expected.tools 一致 | ordered(默认 true) |
| agent_no_undeclared_tool | 没有调用 expected.tools 之外的工具 | 无 |
| agent_constraints_satisfied | 环境记录的每个指定约束都已通过 | constraints |
| agent_max_steps | 轨迹最多用了 max_steps 步 | max_steps |
agent_steps 和 agent_tool_calls 是最后一项背后的分布,以分位数指标的形式提供。
oloproof run│ answer_correct │ 82.5% │ [67.2%, 92.7%] │ 33 / 40 observed · 0 missing · 0 excluded │
│ agent_tool_lookup_order_called │ 100.0% │ [86.8%, 100.0%] │ 39 / 39 observed · 1 missing · 0 excluded │
│ agent_no_tool_loop │ 74.4% │ [56.1%, 87.4%] │ 29 / 39 observed · 1 missing · 0 excluded │
│ agent_tool_sequence │ 60.0% │ [43.3%, 75.2%] │ 24 / 40 observed · 0 missing · 0 excluded │
│ agent_constraints_satisfied │ 92.5% │ [79.6%, 98.5%] │ 37 / 40 observed · 0 missing · 0 excluded │
│ agent_steps_le_10 │ 97.5% │ [86.8%, 100.0%] │ 39 / 40 observed · 0 missing · 0 excluded │
│ steps_p95 │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50 │ 2 calls │ [2, 3] calls │ p50 of 39 observed · 1 missing · 0 excluded ││ no-deletion │ agent_constraints_satisfied │ FAIL │ observed_failures_exceed_limit │约束与其他判据一样参与门控。no-deletion 是一条观测计数规则,设置了 max_failures: 0:有三个用例调用了 delete_customer,而"这在我们运行的测试套件中绝不能发生"无需区间即可判定。
提前截止的轨迹
有一个用例触及了智能体的步数上限,因此其轨迹被记录为 truncated。截断轨迹上的每个计数都是下限,而下限能解决某些问题,却解决不了另一些:
| 判据 | 截断的用例 | 原因 |
|---|---|---|
| agent_steps_le_10 | 失败 | 已记录的步骤已经证明超出了十步的上限 |
| agent_tool_lookup_order_called | 缺失 | 该调用可能位于未被记录的部分 |
| agent_no_tool_loop | 缺失 | 延续到轨迹末尾的重复可能在其后继续 |
| steps_p95、tool_calls_p50 | 缺失 | 基于下限计算的分位数不是分位数 |
是缺失而不是排除:该用例以未观测状态留在分母中,区间允许它朝任一方向变化。这就是为什么 agent_tool_lookup_order_called 基于 39 个观测用例显示为 [86.8%, 100.0%],而不是仅凭 39 个用例会给出的更窄区间。
基于轨迹的切片
first_tool 按智能体首先使用的工具对用例分组,无论该运行调用了哪些工具。repeated_action 将重复了完全相同调用的用例与没有重复的用例分开。trajectory_length:4,8 将用例按 1-4 步、5-8 步以及 9 步及以上分桶,边界由你声明,因为桶的边界会改变切片所表达的内容。
│ first_tool=search │ agent_no_tool_loop │ 0.0% │ [0.0%, 57.9%] │ 0 / 6 observed · 1 missing · 0 excluded · │
│ first_tool=search │ agent_tool_sequence │ 0.0% │ [0.0%, 41.0%] │ 0 / 7 observed · 0 missing · 0 excluded · │
│ repeated_action=false │ answer_correct │ 93.1% │ [77.2%, 99.2%] │ 27 / 29 observed · 0 missing · 0 excluded · │
│ repeated_action=true │ answer_correct │ 60.0% │ [26.2%, 87.9%] │ 6 / 10 observed · 0 missing · 0 excluded · │循环是一个信号,而不是解释。输出中没有任何内容说重复调用就是用例失败的原因;只有移除了该重复的重放才能说明这一点。
多个智能体
examples/triage_agents/ 运行一个三人团队:triage 将每个请求交给 billing 或 tech,而 billing 无权发放的退款会被转交给人。在由多个智能体组成的系统中,每个步骤都注明执行它的智能体,控制权的转移是一个 handoff 步骤:
steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))一条轨迹要么为每个步骤注明智能体,要么全都不注明。只注明了部分步骤的记录会被拒绝;全都不注明的记录对下面的每个判据都计为缺失,而不是通过。转交步骤是可选的,因为当下一步由另一个智能体执行时,控制权同样发生了转移。它是轨迹记录向一个不执行任何步骤的对象(例如人)转移控制权的方式。
用例在 expected.route 下声明它应当经过的智能体:
{"id": "case_012", "input": {"topic": "billing", "request": "I was charged twice for order ord-012", "order_id": "ord-012", "behaviour": "escalated"}, "expected": {"answer": "escalated", "route": ["triage", "billing", "person"]}, "metadata": {"topic": "billing", "behaviour": "escalated"}}evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_route}
- type: agent_tool_permissions
permissions:
triage: []
billing: [lookup_order, issue_refund]
tech: [search_kb]
- {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]| 类型 | 通过条件 | 参数 |
|---|---|---|
| agent_route | 掌握过控制权的智能体(合并连续重复后)与用例的 expected.route 一致 | 无 |
| agent_tool_permissions | 每次工具调用都由映射允许调用该工具的智能体发起 | permissions |
| agent_max_handoffs | 控制权转移不超过 max_handoffs 次 | max_handoffs |
路线包含转交的接收方,因此以升级给人而结束的轨迹,其路线终点就是这个人。没有 expected.route 的用例不计入 agent_route。权限映射是封闭的:映射中未列出的智能体不能调用任何工具。
│ answer_correct │ 96.7% │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1% │ [73.4%, 99.2%] │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2 │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │有两个请求中 tech 发放了退款。两者都通过了 answer_correct,却未通过 agent_tool_permissions:客户从一个违反了自身权限的系统那里得到了正确答案。有一个请求在 billing 和 tech 之间来回转交,直到达到循环上限。它的轨迹被截断,因此它未通过路线和转交上限这两项(其前缀已足以判定),而在权限这一项上计为缺失(其前缀无法判定)。
route 按用例中智能体所走的路线对用例分组,写作 triage>billing。截断的轨迹不加入任何路线分组,因为它的路线只是其智能体后续去向的一个前缀。
输出中没有任何内容说明应当归咎于哪个智能体。路线分歧说明的是两条路线在何处分开。说某个智能体导致了失败,是在断言如果它采取了不同的行动会发生什么,这只有替换该智能体的重放才能证明,而 Oloproof 不会运行这样的重放。