Build with evidence.
Oloproof helps teams evaluate AI systems with repeatable tests, diagnostics, and decision-grade evidence.
Runs in your CI. Your evidence stays on your machine.
Prompt testing by hand does not survive production.
Teams ship model and prompt changes weekly. Spot-checking a handful of outputs in a playground tells you almost nothing about what changed, for whom, or whether it is safe to release.
Oloproof replaces that with a measured record: the same cases, the same judges, the same seeds, run on every change.
A prompt edit fixes one failure class and quietly breaks another. Nobody notices until a customer does.
On a live RAG system, about one observed case in nine changed its verdict between identical measurements.
A live harness scored its 13.8% of HTTP 500s as zeros, charging outages to the model. An average cannot carry a decision.
When risk, legal, or a customer asks what you tested, screenshots in a thread are not an answer.
Four things every release decision needs.
Versioned datasets, pinned seeds, and declarative judges. Every run reproduces exactly, months later.
Deltas with confidence intervals and thresholds, so a two-point move is never mistaken for a win.
Clustered failures by segment, intent, and tool path, each linked to the exact traces behind it.
A release recommendation with the evidence attached — exportable for review, audit, and customers.
Every interval crosses the dashed zero line. Read off the point estimates alone, the reranker would have shipped.
Built for teams that have to justify the release.
Define datasets, judges, and thresholds in version control. Review evaluation changes the way you review code.
Programmatic checks, LLM judges, trained classifiers that score text on your own machine, and human labels in one pipeline. Each judge's agreement with people is measured with an interval, and one that has not cleared its bar cannot decide a release.
Fail a pull request when a metric crosses its threshold. The gate reports which cases moved and by how much.
Every score links to inputs, outputs, tool calls, retrieved context, and judge rationale. Nothing is a black box.
Quality is never the only axis. Track spend and tail latency alongside correctness, under the same thresholds.
Share the run, diagnostics, and recommendation without changing the measured decision it contains.
Three steps, then it runs itself.
Bring your own cases as a JSONL file. Declare judges and the thresholds a release must clear.
One command locally, the same command in CI. Prompt, model, retrieval, and tool changes all get the same treatment.
Read the deltas, open the failing traces, and keep the evidence attached to the call you made.
Every number traces back to a case.
Runs are immutable and dated. Datasets, prompts, judges, and configuration are versioned together, so any result can be reproduced or challenged.
Read off the point estimates, the reranker delivers +12.7 points of hit@1. Read honestly, this panel cannot tell whether it helps or hurts.
What this is not, yet.
Everything below is true on the day you are reading it.
- 01
There is no hosted deployment you can sign up for. Access is by invitation.
- 02
There is no pricing. The unit is a case evaluated; the rates are not set.
- 03
There is no security audit, no SOC 2 report and no DPA. There are no customers and no case studies, so there are none quoted.
- 04
Notifications are email only: no Slack, pager or webhook. There is no report document. A run is read in the workbench or exported as a bundle.
- 05
Production evaluation is not built: traces are received on your own machine, and nothing scores production traffic yet.
- 06
It has been run end to end against exactly one live third-party system. That is the evidence above, and it is one system.
Measure before you ship.
Bring one system and one dataset. Start with one local run and keep the evidence on your machine.