Guides
Classifiers and regressors
A predictive model is evaluated like any other system: a callable returns a prediction for each case, and evaluators read it. What changes is the denominators. Accuracy, recall and precision are three rates over three different sets of rows, and a model can look good on one while failing the question the business is asking.
examples/churn_model/ is the project this page runs: two hundred accounts, a deterministic churn model, and no provider credentials.
The prediction
The system returns a label and, if the model has one, the score behind it:
from oloproof import system
@system(name="churn-model", version="slice-f-example")
def run(account):
score = churn_score(account)
return {"label": score >= 0.5, "score": round(score, 4)}A case declares the truth under expected:
{"expected": {"label": false}, "id": "account_000", "input": {"recent_upgrade": true, "support_contacts": 0, "tenure_months": 0}, "metadata": {"plan": "enterprise"}}The `predictive:` block
Where the prediction, its score and the truth live is declared once, for the project:
predictive:
label_field: label
score_field: score
expected_field: label
positive: true
calibration_bins: 10
thresholds: [0.3, 0.4, 0.5, 0.6, 0.7]| Field | Default | Meaning |
|---|---|---|
| label_field | label | the output field holding the predicted label |
| score_field | score | the output field holding the score behind it |
| expected_field | label | the field under expected holding the truth |
| positive | true | which label value counts as positive; refund, 1 or true |
| calibration_bins | 10 | how many score bands the calibration table uses |
| thresholds | none | cut-offs to sweep; each is evidence, never a recommendation |
| average | none | macro or micro, for an aggregate over several classes |
Every predictive evaluator that does not state field, expected_field or positive itself takes them from this block, so one suite has one positive class. The block is also what produces the confusion counts, the calibration table and the threshold sweep; a project without it gets the metrics and nothing beside them.
A positive class that no case's label matches is refused before anything runs, because recall over it would be a rate over nothing. The same project with positive: churned:
Configuration error: evaluator 'recall' counts 'churned' as the positive class, and no case's 'label' is 'churned' (labels: False, True); declare `positive:` on the evaluator or in the `predictive:` blockThe evaluators
evaluators:
- {type: predictive_correct, criterion: accuracy}
- {type: predictive_recall, criterion: recall}
- {type: predictive_precision, criterion: precision}
- {type: predictive_brier, criterion: brier}
- {type: predictive_log_loss, criterion: log_loss, clip: 0.02}
- {type: predictive_ranking, criterion: rank}
metrics:
- {id: roc_auc, type: ranking, criterion: rank, statistic: roc_auc}
- {id: pr_auc, type: ranking, criterion: rank, statistic: average_precision}
slices: [metadata.plan, "confidence:0.5"]
min_slice_support: 20predictive_log_loss requires clip, because one confident mistake is otherwise infinite. predictive_ranking is the criterion a ranking metric orders rows by, and it needs a metrics: entry naming the statistic: ROC-AUC and average precision answer different questions, and the engine will not pick one for you.
oloproof run│ accuracy │ 88.5% │ [83.2%, 92.6%] │ 177 / 200 observed · 0 missing · 0 excluded │
│ recall │ 81.2% │ [69.5%, 90.0%] │ 52 / 64 observed · 0 missing · 136 excluded │
│ precision │ 82.5% │ [70.9%, 91.0%] │ 52 / 63 observed · 0 missing · 137 excluded │
│ brier │ 0.120 │ [0.094, 0.154] │ mean of 200 observed · 0 missing · 0 excluded │
│ log_loss │ 0.389 │ [0.323, 0.499] │ mean of 200 observed · 0 missing · 0 excluded │
│ roc_auc │ 92.3% │ [69.3%, 100.0%] │ roc_auc over 64 positive · 136 negative · 0 missing · 0 excluded │
│ pr_auc │ 86.5% │ │ average_precision over 64 positive · 136 negative · 0 missing · 0 excluded │Read the excluded column. Recall is measured over the 64 accounts that churned, so the other 136 are excluded from it; precision over the 63 the model flagged. Thirty-two percent of these accounts churn, so a model that predicts nobody churns is 68% accurate and finds no one. A floor on accuracy alone would pass it, which is why the example's policy puts a floor on each rate.
pr_auc has an estimate and no interval. At two hundred rows its interval is genuinely weaker than ROC-AUC's, and the engine withholds a bound it cannot support rather than show one. A rule on it reads:
pr: INSUFFICIENT_EVIDENCE (interval_unavailable)Beside the metrics
The confusion counts are counts, not rates:
│ actually positive │ 52 │ 12 │
│ actually negative │ 11 │ 125 │A release rule naming one is a configuration error, because a count is not a metric:
Configuration error: release rule 'fp' refers to unknown metric 'false_positives'The calibration table sets what the model claimed against what happened, by score band:
│ 0.2-0.3 │ 26.7% │ 0.0% │ 34 rows │
│ 0.5-0.6 │ 53.4% │ 81.0% │ 21 rows │And the threshold sweep shows what each declared cut-off would have measured:
│ 0.3 │ 57.4% │ 96.9% │ 62/108 predicted positive · 62/64 actual positive │
│ 0.5 │ 82.5% │ 81.2% │ 52/63 predicted positive · 52/64 actual positive │
│ 0.7 │ 100.0% │ 37.5% │ 24/24 predicted positive · 24/64 actual positive │The sweep is titled Thresholds (exploratory; recommends nothing). Which cut-off is right depends on what a false positive costs against a false negative, and that is not something the engine can know.
Regression
A regressor is scored by absolute error, which needs the range its targets lie in:
evaluators:
- type: predictive_absolute_error
criterion: days_error
field: days
expected_field: days
target_range: [0, 20]rules:
- id: error-budget
metric: days_error
max: 1.5│ days_error │ 1.02 │ [0.78, 1.88] │ mean of 120 observed · 0 missing · 0 excluded │error-budget: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)target_range is required, not defaulted. An absolute error is a bounded mean, and its interval holds only within a range every value lies in. A wider range is a wider interval, so declare the range the targets can actually take. The rule cannot decide: the estimate is inside the budget, and 120 orders cannot yet show that the true mean error is.
Where to go next
- Slices covers confidence: bands and slice support.
- Comparison rules covers comparing two model versions.