Skip to content

Guides

Agents and tools

Oloproof does not drive an agent. Your agent runs its own loop, calls its own tools, and records what happened as an agent_trajectory/v1 artifact. Every agent metric is read out of that record, so the record is the whole integration.

examples/support_agent/ is the project this page runs: forty refund requests, a deterministic agent, and no provider credentials.

Recording a trajectory

Inside the system, build the steps as they happen and hand them to the case recorder:

from oloproof import AgentConstraintCheck, AgentStep, AgentTrajectory, current_case, system


@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
    steps = []
    for name, arguments in plan(case):
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
        result = getattr(tools, name)(**arguments)
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))

    current_case().agent_trajectory(
        AgentTrajectory(
            steps=tuple(steps),
            terminal_status="success" if refunded else "failure",
            truncated=truncated,
            step_limit=STEP_LIMIT if truncated else None,
            constraints=(
                AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),
            ),
        )
    )
    return {"answer": "refunded" if refunded else "unresolved"}

What each part is for:

FieldWhat it records
AgentStep.kindmessage, tool_call, tool_result, observation, decision or final
AgentStep.tool_name, arguments, resultthe call and what came back
terminal_statussuccess, failure or unknown
truncated, step_limitthat the agent hit its bound and the trace stops short
constraintschecks your environment made, each with the step it observed
checkpointspoints a replay could resume from, as AgentCheckpoint

The system declares that it records the trajectory, on the decorator and in oloproof.yaml:

system:
  name: support-agent
  version: slice-e-example
  callable: app:run
  records: [agent_trajectory/v1]

Without records:, the agent evaluators refuse to run rather than count every case as missing.

What a case declares

A case that should call particular tools, in a particular order, says so under expected.tools:

{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}

expected.tools: [] says the case expects no tool calls at all, and is judged on that. Leaving the key out says tool choice does not apply to the case: agent_tool_sequence and agent_no_undeclared_tool exclude it with no_declared_tools. It leaves the denominator rather than counting as a pass, because a rate inflated with cases that were never measured is not a rate. A value that is not a list of tool names is refused before the run.

The evaluators

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_tool_called, tool_name: lookup_order}
  - {type: agent_no_tool_loop, max_repeats: 2}
  - {type: agent_tool_sequence}
  - {type: agent_constraints_satisfied, constraints: [no_deletion]}
  - {type: agent_max_steps, max_steps: 10}
metrics:
  - {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
  - {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3
TypePasses whenTakes
agent_tool_calledthe tool was called at least min_calls timestool_name, min_calls (default 1)
agent_no_tool_loopno identical call, same tool and arguments, repeats more than max_repeats times in a rowmax_repeats (default 2)
agent_tool_sequencethe tools called match expected.toolsordered (default true)
agent_no_undeclared_toolno tool outside expected.tools was callednothing
agent_constraints_satisfiedevery named constraint the environment recorded passedconstraints
agent_max_stepsthe trace took at most max_steps stepsmax_steps

agent_steps and agent_tool_calls are the distributions behind the last one, as quantile metrics.

oloproof run
│ answer_correct                 │ 82.5%    │ [67.2%, 92.7%]       │ 33 / 40 observed · 0 missing · 0 excluded   │
│ agent_tool_lookup_order_called │ 100.0%   │ [86.8%, 100.0%]      │ 39 / 39 observed · 1 missing · 0 excluded   │
│ agent_no_tool_loop             │ 74.4%    │ [56.1%, 87.4%]       │ 29 / 39 observed · 1 missing · 0 excluded   │
│ agent_tool_sequence            │ 60.0%    │ [43.3%, 75.2%]       │ 24 / 40 observed · 0 missing · 0 excluded   │
│ agent_constraints_satisfied    │ 92.5%    │ [79.6%, 98.5%]       │ 37 / 40 observed · 0 missing · 0 excluded   │
│ agent_steps_le_10              │ 97.5%    │ [86.8%, 100.0%]      │ 39 / 40 observed · 0 missing · 0 excluded   │
│ steps_p95                      │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50                 │ 2 calls  │ [2, 3] calls         │ p50 of 39 observed · 1 missing · 0 excluded │
│ no-deletion         │ agent_constraints_satisfied │ FAIL                  │ observed_failures_exceed_limit │

A constraint gates like any other criterion. no-deletion is an observed-count rule with max_failures: 0: three cases called delete_customer, and "this must not happen in the suite we ran" needs no interval to decide.

A trace that stops short

One case hits the agent's step limit, so its trace is recorded truncated. Every count over a truncated trace is a lower bound, and a lower bound settles some questions and not others:

CriterionThe truncated caseWhy
agent_steps_le_10failsthe steps recorded already prove a limit of ten was broken
agent_tool_lookup_order_calledmissingthe call may be in the part that was not recorded
agent_no_tool_loopmissinga repeat reaching the end of the trace may continue past it
steps_p95, tool_calls_p50missinga quantile over lower bounds is not a quantile

Missing rather than excluded: the case stays in the denominator unobserved, and the interval allows it to have gone either way. That is why agent_tool_lookup_order_called reads [86.8%, 100.0%] over 39 observed cases rather than the tighter interval 39 cases alone would give.

Slices over trajectories

first_tool groups cases by the tool the agent reached for first, whatever tools the run called. repeated_action splits cases that repeated an identical call from those that did not. trajectory_length:4,8 buckets cases at 1-4, 5-8 and 9 or more steps, at bounds you declare because a bucket boundary changes what a slice says.

│ first_tool=search          │ agent_no_tool_loop             │ 0.0%     │ [0.0%, 57.9%]        │ 0 / 6 observed · 1 missing · 0 excluded ·          │
│ first_tool=search          │ agent_tool_sequence            │ 0.0%     │ [0.0%, 41.0%]        │ 0 / 7 observed · 0 missing · 0 excluded ·          │
│ repeated_action=false      │ answer_correct                 │ 93.1%    │ [77.2%, 99.2%]       │ 27 / 29 observed · 0 missing · 0 excluded ·        │
│ repeated_action=true       │ answer_correct                 │ 60.0%    │ [26.2%, 87.9%]       │ 6 / 10 observed · 0 missing · 0 excluded ·         │

A loop is a signal, not an explanation. Nothing in the output says a repeated call is why a case failed; only a replay that removed it could.

Several agents

examples/triage_agents/ runs a team of three: triage hands each request to billing or tech, and a refund billing may not issue is handed to a person. In a system of several agents, each step names the agent that took it, and a transfer of control is a handoff step:

steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))

A trajectory names the agent of every step or of none. A recording that names only some is refused, and one that names none is missing for every criterion below rather than passing it. A hand-off step is optional, since control also changes hands when the next step is taken by another agent. It is how a trace records a transfer to an agent that takes no step, such as a person.

A case declares the agents it should pass through under expected.route:

{"id": "case_012", "input": {"topic": "billing", "request": "I was charged twice for order ord-012", "order_id": "ord-012", "behaviour": "escalated"}, "expected": {"answer": "escalated", "route": ["triage", "billing", "person"]}, "metadata": {"topic": "billing", "behaviour": "escalated"}}
evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_route}
  - type: agent_tool_permissions
    permissions:
      triage: []
      billing: [lookup_order, issue_refund]
      tech: [search_kb]
  - {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
TypePasses whenTakes
agent_routethe agents that held control, repeats collapsed, are the case's expected.routenothing
agent_tool_permissionsevery tool call was made by an agent the map lets call that toolpermissions
agent_max_handoffscontrol changed hands at most max_handoffs timesmax_handoffs

The route includes a hand-off's receiver, so a trace that ends by escalating to a person routes to the person. A case with no expected.route is not counted by agent_route. The permission map is closed: an agent it does not list may call no tool.

│ answer_correct         │ 96.7%    │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route            │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1%    │ [73.4%, 99.2%]  │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2    │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │

In two requests tech issued a refund. Both pass answer_correct and fail agent_tool_permissions: the customer got the right answer from a system that broke its own permissions. One request is handed back and forth between billing and tech until the loop's bound. Its trace is truncated, so it fails the route and the hand-off bound, which its prefix already settles, and is missing for permissions, which it does not.

route groups cases by the route their agents took, written triage>billing. A truncated trace joins no route group, because its route is a prefix of wherever its agents went next.

Nothing in the output says which agent is to blame. A route divergence says where two routes part. That an agent caused a failure is a claim about what would have happened had it acted otherwise, which only a replay substituting that agent could show, and Oloproof does not run one.

Where to go next