Guides
Agents and tools
Oloproof does not drive an agent. Your agent runs its own loop, calls its own tools, and records what happened as an agent_trajectory/v1 artifact. Every agent metric is read out of that record, so the record is the whole integration.
examples/support_agent/ is the project this page runs: forty refund requests, a deterministic agent, and no provider credentials.
Recording a trajectory
Inside the system, build the steps as they happen and hand them to the case recorder:
from oloproof import AgentConstraintCheck, AgentStep, AgentTrajectory, current_case, system
@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
steps = []
for name, arguments in plan(case):
steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
result = getattr(tools, name)(**arguments)
steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))
current_case().agent_trajectory(
AgentTrajectory(
steps=tuple(steps),
terminal_status="success" if refunded else "failure",
truncated=truncated,
step_limit=STEP_LIMIT if truncated else None,
constraints=(
AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),
),
)
)
return {"answer": "refunded" if refunded else "unresolved"}What each part is for:
| Field | What it records |
|---|---|
| AgentStep.kind | message, tool_call, tool_result, observation, decision or final |
| AgentStep.tool_name, arguments, result | the call and what came back |
| terminal_status | success, failure or unknown |
| truncated, step_limit | that the agent hit its bound and the trace stops short |
| constraints | checks your environment made, each with the step it observed |
| checkpoints | points a replay could resume from, as AgentCheckpoint |
The system declares that it records the trajectory, on the decorator and in oloproof.yaml:
system:
name: support-agent
version: slice-e-example
callable: app:run
records: [agent_trajectory/v1]Without records:, the agent evaluators refuse to run rather than count every case as missing.
What a case declares
A case that should call particular tools, in a particular order, says so under expected.tools:
{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}expected.tools: [] says the case expects no tool calls at all, and is judged on that. Leaving the key out says tool choice does not apply to the case: agent_tool_sequence and agent_no_undeclared_tool exclude it with no_declared_tools. It leaves the denominator rather than counting as a pass, because a rate inflated with cases that were never measured is not a rate. A value that is not a list of tool names is refused before the run.
The evaluators
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_tool_called, tool_name: lookup_order}
- {type: agent_no_tool_loop, max_repeats: 2}
- {type: agent_tool_sequence}
- {type: agent_constraints_satisfied, constraints: [no_deletion]}
- {type: agent_max_steps, max_steps: 10}
metrics:
- {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
- {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3| Type | Passes when | Takes |
|---|---|---|
| agent_tool_called | the tool was called at least min_calls times | tool_name, min_calls (default 1) |
| agent_no_tool_loop | no identical call, same tool and arguments, repeats more than max_repeats times in a row | max_repeats (default 2) |
| agent_tool_sequence | the tools called match expected.tools | ordered (default true) |
| agent_no_undeclared_tool | no tool outside expected.tools was called | nothing |
| agent_constraints_satisfied | every named constraint the environment recorded passed | constraints |
| agent_max_steps | the trace took at most max_steps steps | max_steps |
agent_steps and agent_tool_calls are the distributions behind the last one, as quantile metrics.
oloproof run│ answer_correct │ 82.5% │ [67.2%, 92.7%] │ 33 / 40 observed · 0 missing · 0 excluded │
│ agent_tool_lookup_order_called │ 100.0% │ [86.8%, 100.0%] │ 39 / 39 observed · 1 missing · 0 excluded │
│ agent_no_tool_loop │ 74.4% │ [56.1%, 87.4%] │ 29 / 39 observed · 1 missing · 0 excluded │
│ agent_tool_sequence │ 60.0% │ [43.3%, 75.2%] │ 24 / 40 observed · 0 missing · 0 excluded │
│ agent_constraints_satisfied │ 92.5% │ [79.6%, 98.5%] │ 37 / 40 observed · 0 missing · 0 excluded │
│ agent_steps_le_10 │ 97.5% │ [86.8%, 100.0%] │ 39 / 40 observed · 0 missing · 0 excluded │
│ steps_p95 │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50 │ 2 calls │ [2, 3] calls │ p50 of 39 observed · 1 missing · 0 excluded ││ no-deletion │ agent_constraints_satisfied │ FAIL │ observed_failures_exceed_limit │A constraint gates like any other criterion. no-deletion is an observed-count rule with max_failures: 0: three cases called delete_customer, and "this must not happen in the suite we ran" needs no interval to decide.
A trace that stops short
One case hits the agent's step limit, so its trace is recorded truncated. Every count over a truncated trace is a lower bound, and a lower bound settles some questions and not others:
| Criterion | The truncated case | Why |
|---|---|---|
| agent_steps_le_10 | fails | the steps recorded already prove a limit of ten was broken |
| agent_tool_lookup_order_called | missing | the call may be in the part that was not recorded |
| agent_no_tool_loop | missing | a repeat reaching the end of the trace may continue past it |
| steps_p95, tool_calls_p50 | missing | a quantile over lower bounds is not a quantile |
Missing rather than excluded: the case stays in the denominator unobserved, and the interval allows it to have gone either way. That is why agent_tool_lookup_order_called reads [86.8%, 100.0%] over 39 observed cases rather than the tighter interval 39 cases alone would give.
Slices over trajectories
first_tool groups cases by the tool the agent reached for first, whatever tools the run called. repeated_action splits cases that repeated an identical call from those that did not. trajectory_length:4,8 buckets cases at 1-4, 5-8 and 9 or more steps, at bounds you declare because a bucket boundary changes what a slice says.
│ first_tool=search │ agent_no_tool_loop │ 0.0% │ [0.0%, 57.9%] │ 0 / 6 observed · 1 missing · 0 excluded · │
│ first_tool=search │ agent_tool_sequence │ 0.0% │ [0.0%, 41.0%] │ 0 / 7 observed · 0 missing · 0 excluded · │
│ repeated_action=false │ answer_correct │ 93.1% │ [77.2%, 99.2%] │ 27 / 29 observed · 0 missing · 0 excluded · │
│ repeated_action=true │ answer_correct │ 60.0% │ [26.2%, 87.9%] │ 6 / 10 observed · 0 missing · 0 excluded · │A loop is a signal, not an explanation. Nothing in the output says a repeated call is why a case failed; only a replay that removed it could.
Several agents
examples/triage_agents/ runs a team of three: triage hands each request to billing or tech, and a refund billing may not issue is handed to a person. In a system of several agents, each step names the agent that took it, and a transfer of control is a handoff step:
steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))A trajectory names the agent of every step or of none. A recording that names only some is refused, and one that names none is missing for every criterion below rather than passing it. A hand-off step is optional, since control also changes hands when the next step is taken by another agent. It is how a trace records a transfer to an agent that takes no step, such as a person.
A case declares the agents it should pass through under expected.route:
{"id": "case_012", "input": {"topic": "billing", "request": "I was charged twice for order ord-012", "order_id": "ord-012", "behaviour": "escalated"}, "expected": {"answer": "escalated", "route": ["triage", "billing", "person"]}, "metadata": {"topic": "billing", "behaviour": "escalated"}}evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_route}
- type: agent_tool_permissions
permissions:
triage: []
billing: [lookup_order, issue_refund]
tech: [search_kb]
- {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]| Type | Passes when | Takes |
|---|---|---|
| agent_route | the agents that held control, repeats collapsed, are the case's expected.route | nothing |
| agent_tool_permissions | every tool call was made by an agent the map lets call that tool | permissions |
| agent_max_handoffs | control changed hands at most max_handoffs times | max_handoffs |
The route includes a hand-off's receiver, so a trace that ends by escalating to a person routes to the person. A case with no expected.route is not counted by agent_route. The permission map is closed: an agent it does not list may call no tool.
│ answer_correct │ 96.7% │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1% │ [73.4%, 99.2%] │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2 │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │In two requests tech issued a refund. Both pass answer_correct and fail agent_tool_permissions: the customer got the right answer from a system that broke its own permissions. One request is handed back and forth between billing and tech until the loop's bound. Its trace is truncated, so it fails the route and the hand-off bound, which its prefix already settles, and is missing for permissions, which it does not.
route groups cases by the route their agents took, written triage>billing. A truncated trace joins no route group, because its route is a prefix of wherever its agents went next.
Nothing in the output says which agent is to blame. A route divergence says where two routes part. That an agent caused a failure is a claim about what would have happened had it acted otherwise, which only a replay substituting that agent could show, and Oloproof does not run one.
Where to go next
- Recording what a system did covers the other typed artifacts.
- Slices covers slice support and why slices never gate.