Skip to content

指南

智能体与工具

Oloproof 不驱动智能体。你的智能体运行自己的循环,调用自己的工具,并将所发生的事情记录为一个 agent_trajectory/v1 制品。每个智能体指标都从这条记录中读取,因此这条记录就是全部的集成工作。

examples/support_agent/ 是本页所运行的项目:四十个退款请求、一个确定性的智能体,无需任何提供方凭据。

记录轨迹

在系统内部,随着事情发生构建各个步骤,并将它们交给用例记录器:

from oloproof import AgentConstraintCheck, AgentStep, AgentTrajectory, current_case, system


@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
    steps = []
    for name, arguments in plan(case):
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
        result = getattr(tools, name)(**arguments)
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))

    current_case().agent_trajectory(
        AgentTrajectory(
            steps=tuple(steps),
            terminal_status="success" if refunded else "failure",
            truncated=truncated,
            step_limit=STEP_LIMIT if truncated else None,
            constraints=(
                AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),
            ),
        )
    )
    return {"answer": "refunded" if refunded else "unresolved"}

各部分的用途:

字段记录的内容
AgentStep.kindmessage、tool_call、tool_result、observation、decision 或 final
AgentStep.tool_name、arguments、result调用以及返回的内容
terminal_statussuccess、failure 或 unknown
truncated、step_limit智能体触及了步数上限,轨迹提前截止
constraints你的环境所做的检查,每项都附带它所观察的步骤
checkpoints重放可以从中恢复的位置,形式为 AgentCheckpoint

系统需要在装饰器上和 oloproof.yaml 中声明它会记录轨迹:

system:
  name: support-agent
  version: slice-e-example
  callable: app:run
  records: [agent_trajectory/v1]

没有 records: 时,智能体评估器会拒绝运行,而不是把每个用例都计为缺失。

用例声明的内容

应当按特定顺序调用特定工具的用例,在 expected.tools 下声明:

{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}

expected.tools: [] 表示该用例预期完全不调用工具,并据此评判。省略该键表示工具选择不适用于该用例:agent_tool_sequence 和 agent_no_undeclared_tool 会以 no_declared_tools 将其排除。它会离开分母,而不是计为通过,因为混入了从未被度量的用例而虚高的比率并不是真正的比率。不是工具名列表的值会在运行之前被拒绝。

评估器

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_tool_called, tool_name: lookup_order}
  - {type: agent_no_tool_loop, max_repeats: 2}
  - {type: agent_tool_sequence}
  - {type: agent_constraints_satisfied, constraints: [no_deletion]}
  - {type: agent_max_steps, max_steps: 10}
metrics:
  - {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
  - {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3
类型通过条件参数
agent_tool_called该工具至少被调用 min_calls 次tool_name、min_calls(默认 1)
agent_no_tool_loop没有任何完全相同的调用(相同工具和参数)连续重复超过 max_repeats 次max_repeats(默认 2)
agent_tool_sequence调用的工具与 expected.tools 一致ordered(默认 true)
agent_no_undeclared_tool没有调用 expected.tools 之外的工具无
agent_constraints_satisfied环境记录的每个指定约束都已通过constraints
agent_max_steps轨迹最多用了 max_steps 步max_steps

agent_steps 和 agent_tool_calls 是最后一项背后的分布,以分位数指标的形式提供。

oloproof run
│ answer_correct                 │ 82.5%    │ [67.2%, 92.7%]       │ 33 / 40 observed · 0 missing · 0 excluded   │
│ agent_tool_lookup_order_called │ 100.0%   │ [86.8%, 100.0%]      │ 39 / 39 observed · 1 missing · 0 excluded   │
│ agent_no_tool_loop             │ 74.4%    │ [56.1%, 87.4%]       │ 29 / 39 observed · 1 missing · 0 excluded   │
│ agent_tool_sequence            │ 60.0%    │ [43.3%, 75.2%]       │ 24 / 40 observed · 0 missing · 0 excluded   │
│ agent_constraints_satisfied    │ 92.5%    │ [79.6%, 98.5%]       │ 37 / 40 observed · 0 missing · 0 excluded   │
│ agent_steps_le_10              │ 97.5%    │ [86.8%, 100.0%]      │ 39 / 40 observed · 0 missing · 0 excluded   │
│ steps_p95                      │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50                 │ 2 calls  │ [2, 3] calls         │ p50 of 39 observed · 1 missing · 0 excluded │
│ no-deletion         │ agent_constraints_satisfied │ FAIL                  │ observed_failures_exceed_limit │

约束与其他判据一样参与门控。no-deletion 是一条观测计数规则,设置了 max_failures: 0:有三个用例调用了 delete_customer,而"这在我们运行的测试套件中绝不能发生"无需区间即可判定。

提前截止的轨迹

有一个用例触及了智能体的步数上限,因此其轨迹被记录为 truncated。截断轨迹上的每个计数都是下限,而下限能解决某些问题,却解决不了另一些:

判据截断的用例原因
agent_steps_le_10失败已记录的步骤已经证明超出了十步的上限
agent_tool_lookup_order_called缺失该调用可能位于未被记录的部分
agent_no_tool_loop缺失延续到轨迹末尾的重复可能在其后继续
steps_p95、tool_calls_p50缺失基于下限计算的分位数不是分位数

是缺失而不是排除:该用例以未观测状态留在分母中,区间允许它朝任一方向变化。这就是为什么 agent_tool_lookup_order_called 基于 39 个观测用例显示为 [86.8%, 100.0%],而不是仅凭 39 个用例会给出的更窄区间。

基于轨迹的切片

first_tool 按智能体首先使用的工具对用例分组,无论该运行调用了哪些工具。repeated_action 将重复了完全相同调用的用例与没有重复的用例分开。trajectory_length:4,8 将用例按 1-4 步、5-8 步以及 9 步及以上分桶,边界由你声明,因为桶的边界会改变切片所表达的内容。

│ first_tool=search          │ agent_no_tool_loop             │ 0.0%     │ [0.0%, 57.9%]        │ 0 / 6 observed · 1 missing · 0 excluded ·          │
│ first_tool=search          │ agent_tool_sequence            │ 0.0%     │ [0.0%, 41.0%]        │ 0 / 7 observed · 0 missing · 0 excluded ·          │
│ repeated_action=false      │ answer_correct                 │ 93.1%    │ [77.2%, 99.2%]       │ 27 / 29 observed · 0 missing · 0 excluded ·        │
│ repeated_action=true       │ answer_correct                 │ 60.0%    │ [26.2%, 87.9%]       │ 6 / 10 observed · 0 missing · 0 excluded ·         │

循环是一个信号,而不是解释。输出中没有任何内容说重复调用就是用例失败的原因;只有移除了该重复的重放才能说明这一点。

多个智能体

examples/triage_agents/ 运行一个三人团队:triage 将每个请求交给 billing 或 tech,而 billing 无权发放的退款会被转交给人。在由多个智能体组成的系统中,每个步骤都注明执行它的智能体,控制权的转移是一个 handoff 步骤:

steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))

一条轨迹要么为每个步骤注明智能体,要么全都不注明。只注明了部分步骤的记录会被拒绝;全都不注明的记录对下面的每个判据都计为缺失,而不是通过。转交步骤是可选的,因为当下一步由另一个智能体执行时,控制权同样发生了转移。它是轨迹记录向一个不执行任何步骤的对象(例如人)转移控制权的方式。

用例在 expected.route 下声明它应当经过的智能体:

{"id": "case_012", "input": {"topic": "billing", "request": "I was charged twice for order ord-012", "order_id": "ord-012", "behaviour": "escalated"}, "expected": {"answer": "escalated", "route": ["triage", "billing", "person"]}, "metadata": {"topic": "billing", "behaviour": "escalated"}}
evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_route}
  - type: agent_tool_permissions
    permissions:
      triage: []
      billing: [lookup_order, issue_refund]
      tech: [search_kb]
  - {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
类型通过条件参数
agent_route掌握过控制权的智能体(合并连续重复后)与用例的 expected.route 一致无
agent_tool_permissions每次工具调用都由映射允许调用该工具的智能体发起permissions
agent_max_handoffs控制权转移不超过 max_handoffs 次max_handoffs

路线包含转交的接收方,因此以升级给人而结束的轨迹,其路线终点就是这个人。没有 expected.route 的用例不计入 agent_route。权限映射是封闭的:映射中未列出的智能体不能调用任何工具。

│ answer_correct         │ 96.7%    │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route            │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1%    │ [73.4%, 99.2%]  │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2    │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │

有两个请求中 tech 发放了退款。两者都通过了 answer_correct,却未通过 agent_tool_permissions:客户从一个违反了自身权限的系统那里得到了正确答案。有一个请求在 billing 和 tech 之间来回转交,直到达到循环上限。它的轨迹被截断,因此它未通过路线和转交上限这两项(其前缀已足以判定),而在权限这一项上计为缺失(其前缀无法判定)。

route 按用例中智能体所走的路线对用例分组,写作 triage>billing。截断的轨迹不加入任何路线分组,因为它的路线只是其智能体后续去向的一个前缀。

输出中没有任何内容说明应当归咎于哪个智能体。路线分歧说明的是两条路线在何处分开。说某个智能体导致了失败,是在断言如果它采取了不同的行动会发生什么,这只有替换该智能体的重放才能证明,而 Oloproof 不会运行这样的重放。

下一步