Skip to content
AI systems evaluation

Build with evidence.

Oloproof helps teams evaluate AI systems with repeatable tests, diagnostics, and decision-grade evidence.

Runs in your CI. Your evidence stays on your machine.

oloproof / compare · olotalk-docs-assistant
Summary Compare Diagnostics Evidence
BASELINE
olotalk docs assistant · rerank off
CANDIDATE
olotalk docs assistant · rerank on
DATASET
olotalk golden panel · 58 cases
METRIC BASELINE CANDIDATE PAIRED Δ 95% CI STATUS
page_hit_1 0.618 0.745 +6.2 [-36.3, +45.5] — Inconclusive
page_hit_3 0.683 0.742 +3.7 [-73.9, +76.2] — Inconclusive
answer_correct 0.685 0.760 +6.5 [-39.4, +48.4] — Inconclusive
page_hit_1 observed 55 / 58 51 / 58 −4 in interval ! Bounded
page_hit_3 observed 41 / 58 31 / 58 −10 in interval ! Bounded
answer_correct observed 54 / 58 50 / 58 −4 in interval ! Bounded
WHY EVALUATION

Prompt testing by hand does not survive production.

Teams ship model and prompt changes weekly. Spot-checking a handful of outputs in a playground tells you almost nothing about what changed, for whom, or whether it is safe to release.

Oloproof replaces that with a measured record: the same cases, the same judges, the same seeds, run on every change.

01
Silent regressions

A prompt edit fixes one failure class and quietly breaks another. Nobody notices until a customer does.

02
Unrepeatable results

On a live RAG system, about one observed case in nine changed its verdict between identical measurements.

03
Scores without meaning

A live harness scored its 13.8% of HTTP 500s as zeros, charging outages to the model. An average cannot carry a decision.

04
No record to show

When risk, legal, or a customer asks what you tested, screenshots in a thread are not an answer.

CORE CAPABILITIES

Four things every release decision needs.

MEASURE
Repeatable evaluations

Versioned datasets, pinned seeds, and declarative judges. Every run reproduces exactly, months later.

COMPARE
Baseline vs candidate

Deltas with confidence intervals and thresholds, so a two-point move is never mistaken for a win.

DIAGNOSE
Failure modes, not averages

Clustered failures by segment, intent, and tool path, each linked to the exact traces behind it.

DECIDE
Decision-grade records

A release recommendation with the evidence attached — exportable for review, audit, and customers.

Diagnostics23 failed cases, gold context
Unresolved: gold context did not settle it11 cases
Generation failure: right context, wrong answer7 cases
Retrieval miss: answer reachable, not retrieved5 cases
Rerank on minus off95% interval, points
hit@1
hit@3
correct

Every interval crosses the dashed zero line. Read off the point estimates alone, the reranker would have shipped.

PLATFORM

Built for teams that have to justify the release.

Evaluation suites as code

Define datasets, judges, and thresholds in version control. Review evaluation changes the way you review code.

Judges you can audit

Programmatic checks, LLM judges, trained classifiers that score text on your own machine, and human labels in one pipeline. Each judge's agreement with people is measured with an interval, and one that has not cleared its bar cannot decide a release.

Regression gates in CI

Fail a pull request when a metric crosses its threshold. The gate reports which cases moved and by how much.

Trace-level evidence

Every score links to inputs, outputs, tool calls, retrieved context, and judge rationale. Nothing is a black box.

Cost and latency budgets

Quality is never the only axis. Track spend and tail latency alongside correctness, under the same thresholds.

Evidence for people who need the decision

Share the run, diagnostics, and recommendation without changing the measured decision it contains.

HOW IT WORKS

Three steps, then it runs itself.

1
Define the suite

Bring your own cases as a JSONL file. Declare judges and the thresholds a release must clear.

2
Run on every change

One command locally, the same command in CI. Prompt, model, retrieval, and tool changes all get the same treatment.

3
Decide with the evidence

Read the deltas, open the failing traces, and keep the evidence attached to the call you made.

~/olodemo · after oloproof init
 oloproof run
Run run_01M3B0CCPA0JFQQ7HWRKYNWFJC [DECIDED/COMPLETE]
Gate: ALLOW (exit 0)

  Rule         Metric       State  Reasons
  label-floor  exact_label  PASS   lower_bound_meets_minimum

  Metric       Estimate  Interval          N
  exact_label  100.0%    [88.4%, 100.0%]   30 / 30 observed · 0 missing

Cache: execution 0 hit/30 miss; judgment 0 hit/30 miss

 oloproof run
Run run_01M3B0CXTTVPTBXAFK3WNGFX42 [DECIDED/COMPLETE]
Gate: ALLOW (exit 0)
Cache: execution 30 hit/0 miss; judgment 0 hit/30 miss
EVIDENCE YOU CAN HAND OVER

Every number traces back to a case.

Runs are immutable and dated. Datasets, prompts, judges, and configuration are versioned together, so any result can be reproduced or challenged.

Immutable
runs, with full lineage
Self-hosted
on your infrastructure, your data stays yours
Bounded
missing cases priced, never dropped

Read off the point estimates, the reranker delivers +12.7 points of hit@1. Read honestly, this panel cannot tell whether it helps or hurts.

Oloproof against a live RAG system, 22 September 2026
Oloproof
Evidence bundle
oloproof export RUN_ID
run.jsonthe run, its suite, system and evaluators
cases.jsonlevery case: input, output, judgments, digests
diagnoses.jsonlinterventions and what they recovered
signoffs.jsonlwho shipped against a blocked gate, and why
An exported evidence bundle: the run and every case behind its numbers, as files you keep.
THE SAME DISCIPLINE, POINTED HERE

What this is not, yet.

Everything below is true on the day you are reading it.

  1. 01

    There is no hosted deployment you can sign up for. Access is by invitation.

  2. 02

    There is no pricing. The unit is a case evaluated; the rates are not set.

  3. 03

    There is no security audit, no SOC 2 report and no DPA. There are no customers and no case studies, so there are none quoted.

  4. 04

    Notifications are email only: no Slack, pager or webhook. There is no report document. A run is read in the workbench or exported as a bundle.

  5. 05

    Production evaluation is not built: traces are received on your own machine, and nothing scores production traffic yet.

  6. 06

    It has been run end to end against exactly one live third-party system. That is the evidence above, and it is one system.

Measure before you ship.

Bring one system and one dataset. Start with one local run and keep the evidence on your machine.