Skip to content

الأدلة الإرشادية

Recording what a system did

ترجمة هذه الصفحة قديمة، لذا تُعرض بالإنجليزية.

A system's output is what an evaluator reads by default. Everything else it did — the prompt it built, the passages it retrieved, the tools it called, the tokens it spent — is recorded alongside the output as artifacts and usage, and every artifact is stored content-addressed with the execution that produced it.

From a Python system

Inside a system, current_case() is the recorder for the case being executed:

import time

from oloproof import current_case, system


@system(name="example-support-bot", version="2")
def answer(case):
    started = time.perf_counter()
    text = case["question"].lower()
    label = "refund" if "refund" in text else "other"
    recorder = current_case()
    recorder.artifact("trace", {"steps": [{"name": "classify", "label": label}], "prompt": text})
    recorder.usage(input_tokens=42, output_tokens=7, cost_usd=0.0001)
    recorder.artifact("stage_timings/v1", {"generation_ms": (time.perf_counter() - started) * 1000})
    return {"answer": "Refunds are available within 30 days." if label == "refund" else "A support agent will follow up.", "label": label}

artifact(kind, data) records any JSON under a kind you name. A kind is a lowercase name with an optional version: trace, prompt, tool_log/v2. usage(...) records what the call consumed; a value left out stays unrecorded rather than becoming zero.

Typed kinds

Six kinds have a schema, and a payload that does not match it stops the run with exit 2 before it can be stored:

KindHoldsRecorder method
retrieval/v1the ranked candidates a retriever returnedcurrent_case().retrieval(...)
context/v1the passages assembled for generation, and what was droppedcurrent_case().context(...)
citations/v1the ids an answer citescurrent_case().citations(...)
stage_timings/v1retrieval_ms, assembly_ms, generation_mscurrent_case().artifact(...)
agent_trajectory/v1an agent's steps, constraint checks and checkpointscurrent_case().agent_trajectory(...)
conversation/v1a multi-turn conversationcurrent_case().artifact(...)

A misnamed field is refused with the field it did not accept. The system above, recording generation where the schema says generation_ms:

Configuration error: malformed stage_timings/v1 artifact: generation: Extra inputs are not permitted

The retrieval, agent and conversation evaluators read these kinds. A system that records one for them declares it on the system, so an evaluator that needs it can refuse to run rather than count every case as missing:

system:
  name: support-agent
  version: slice-e-example
  callable: app:run
  records: [agent_trajectory/v1]

A staged RAG system records its own and takes no records:.

From an HTTP system

An HTTP system cannot call a recorder, so it names the response fields to keep, as a kind and a dotted path into the response:

system:
  name: example-support-bot
  version: "1"
  http:
    url: http://127.0.0.1:8766/answer
    output_path: result
    artifacts:
      trace: debug

output_path picks the output out of the response, and each entry under artifacts records another part of it. Against a server answering {"result": {"answer": ..., "label": ...}, "debug": {"model": "stub"}}, a case's output is {"answer":"Refunds take 30 days.","label":"refund"} and its artifacts are {"trace":[{"model":"stub"}]}.

Where it ends up

oloproof export RUN_ID writes a bundle whose cases.jsonl carries each case's artifacts, decoded:

jq -c '.artifacts' .oloproof/bundles/RUN_ID/cases.jsonl
{"trace":[{"prompt":"can i get a refund? #0","steps":[{"label":"refund","name":"classify"}]}],"stage_timings/v1":[{"assembly_ms":null,"generation_ms":0.323874999594409,"retrieval_ms":null}]}

and the usage on the execution, with jq -c '.execution.usage':

{"input_tokens":42,"output_tokens":7,"cost_usd":0.0001}

oloproof inspect RUN_ID --case CASE_ID shows the input, output and judgments, and not the artifacts; the bundle is where they are read.

oloproof usage totals what was recorded:

Usage for 2026-09-01 to 2026-10-01
  billable  30 cases evaluated
  runs      1 (reported, never charged)
  work      30 executions, 0 judgments
            30 uncacheable judgments (reported, never charged)
  cost      $0.0030
  tokens    1K in, 210 out
  stored    7.6 KiB held now

This period has not closed, so every figure is a total so far.

Artifacts stay on the machine that produced them. oloproof push sends metrics, intervals and decisions to a workspace, and sends raw content — inputs, outputs, artifacts — only for the categories egress: in oloproof.yaml names.

Reading an artifact in an evaluator

A custom evaluator reads an artifact by declaring it in reads, which is part of what keys its cached judgments:

from oloproof import evaluate
from oloproof.evaluators import evaluator

from app import answer


@evaluator(criterion="one_step", reads=("output", "artifacts.trace"))
def one_step(case):
    [trace] = case.artifacts("trace")
    return len(trace["steps"]) == 1


result = evaluate(system=answer, dataset="data/example.jsonl", evaluators=[one_step])
for metric in result.metrics:
    print(metric.metric, metric.estimate, metric.n_observed, metric.n_missing)
one_step 1.0 30 0

case.artifacts(name) takes the kind without its version, so case.artifacts("context") reads context/v1, and returns every artifact of that kind the case recorded, in order.

Latency, tokens and cost as metrics

The execution record already holds latency and whatever usage was recorded, so a distribution over them needs only a metrics: entry:

metrics:
  - {id: latency_p95, type: quantile, source: latency_ms, quantile: 0.95}
  - {id: tokens_in_p50, type: quantile, source: input_tokens, quantile: 0.5}
│ latency_p95   │ 10.84 ms  │ [8.23, no bound] ms │ p95 of 30 observed · 0 missing · 0 excluded │
│ tokens_in_p50 │ 42 tokens │ [42, 42] tokens     │ p50 of 30 observed · 0 missing · 0 excluded │

The sources are latency_ms, input_tokens, output_tokens and cost_usd, and for agents agent_steps and agent_tool_calls. A cached execution keeps the latency measured when it ran. no bound is an honest answer: thirty cases cannot put an upper bound on a 95th percentile.

Production traces

Everything above records what a system did while Oloproof ran it. Traffic in production is recorded by your own OpenTelemetry instrumentation, and oloproof collect receives it. Install the extra and start the collector beside your project:

pip install 'oloproof[collect]'
oloproof collect

It listens for OTLP/HTTP on http://127.0.0.1:4318/v1/traces, the standard address, so a stock exporter needs only its endpoint:

from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter

exporter = OTLPSpanExporter(endpoint="http://127.0.0.1:4318/v1/traces")

Spans that follow the OpenTelemetry GenAI conventions or OpenInference become one trace per request: its input and output, the model, token counts and latency, its tool calls as agent_trajectory/v1 and its retrievals as retrieval/v1, the same kinds a run records. Spans following neither convention are kept as timings. A trace is sealed once its spans have stopped arriving for ten seconds (--settle); a span arriving after that is not kept.

Traces are stored in the project's own store. When this terminal is logged in, sealed traces are also pushed to the workspace through the project's egress policy, so under the default only digests, timings, models and token counts leave your machine, and the names your instrumentation gave traces, spans, tools and agents leave as digests (add span_names under egress: to send them as written); --no-push keeps them local. Without a login, nothing is pushed anywhere. Set OLOPROOF_COLLECT_TOKEN to require Authorization: Bearer on every export when the collector listens beyond this machine.

A failure seen in production is worth more as a case the next evaluation runs. List the traces this project holds and promote one to a dataset:

oloproof traces list
oloproof traces promote TRACE_ID --dataset data/regressions.jsonl --expected '{"answer": "..."}'

The case takes the trace's input, the reference answer you give (or the trace's own output with --expected-output, when a person judged it right), and in metadata the trace id and digests it came from. The same input is not added twice. The dataset is then a new suite: run it, and an Owner pins the new run in the workspace settings to make it the evaluation the workspace verifies.

Evaluating production

A sample of production is only evidence if nobody chose it. Your collector decides which traces the workspace ever sees, so Oloproof has the collector commit to everything it received before anyone draws from it, and the workspace draws with its own randomness.

Give the collector a runner key an Owner or Admin registered for the project (oloproof runner keygen makes one; register its public key in the workspace settings), and start it logged in:

export OLOPROOF_RUNNER_KEY_FILE=~/.oloproof/runner.key OLOPROOF_RUNNER_KEY_ID=rk_...
oloproof collect

At the end of each UTC hour it signs a count and a Merkle root of every trace it sealed in that hour, an empty hour included, and pushes that commitment. Content stays where it is.

An Owner or Admin then draws a sample of whole hours on the project's Production page. The workspace picks positions at random among the committed traces. Each collector answers its own positions with the traces it committed there and a proof, signed with its key, and the workspace checks every proof. Every live collector that has committed must have committed every hour of the range since its first commitment: otherwise the sample is refused, naming the collector and the hour, so a collector cannot keep an hour out by staying quiet. A collector is not required for hours before its first commitment, so start it, logged in and with its key, before its traffic starts. When you retire a collector, revoke its key: from then on its hours are left out of new samples, which name it, since only its key could answer for them. A sample left unanswered for a day is refused.

Evaluate the sample where the content is:

oloproof traces evaluate SAMPLE_ID

Each drawn trace becomes a case with its input and no reference answer, and its output is the one production recorded, with its token counts and its recorded retrievals and tool calls. Nothing is executed again. Your project's evaluators judge them and the run is pushed; oloproof collect --evaluate does this for every sample it proves. Use evaluators that need no reference answer. A latency metric measures the replay, not production; read production latency from the traces themselves.

Pin a run of a traffic sample as the project's evaluation in the workspace settings. Its evaluators and metrics are pinned, and its cases are not: the workspace then decides every sample on its first run under the pinned policy, after checking that the run is the sample's replay and judged exactly the drawn traces' recorded outputs. Judging a sample again changes nothing. Judged decisions need PPI as everywhere else: open a measurement queue on the sample's run, and reviewers label a random sample of it with content fetched from your collector. Comparison rules are not decided for traffic, since no two samples hold the same cases.

The Production page also watches each binary metric across samples, with a 95% confidence sequence over every sample asked for since the pinned one, in the order they were asked for. A sample that was refused, expired or not evaluated within a day counts as missing, so leaving one out only widens the sequence. Each point is written once and never revised, so the sequence holds at every look. If it crosses a pinned threshold, the page says at which sample, and the alarm stays while the same evaluation, level and rule are pinned (pinning another starts a new sequence): it is evidence that at least one sampled hour's judged rate was past the threshold. It decides nothing; each sample's decision is the workspace gate's.

What this cannot catch: traffic your application never sends to the collector, and a collector whose key is used to commit to a curated set. Those are the same limits as a runner's key; see "Verified evidence" in the gating guide. And whoever holds the content can leave a sample undecided by not answering or not evaluating it: the page shows that, and the monitor counts it as missing.

Where to go next