Skip to content

Guides

Classifiers and regressors

A predictive model is evaluated like any other system: a callable returns a prediction for each case, and evaluators read it. What changes is the denominators. Accuracy, recall and precision are three rates over three different sets of rows, and a model can look good on one while failing the question the business is asking.

examples/churn_model/ is the project this page runs: two hundred accounts, a deterministic churn model, and no provider credentials.

The prediction

The system returns a label and, if the model has one, the score behind it:

from oloproof import system


@system(name="churn-model", version="slice-f-example")
def run(account):
    score = churn_score(account)
    return {"label": score >= 0.5, "score": round(score, 4)}

A case declares the truth under expected:

{"expected": {"label": false}, "id": "account_000", "input": {"recent_upgrade": true, "support_contacts": 0, "tenure_months": 0}, "metadata": {"plan": "enterprise"}}

The `predictive:` block

Where the prediction, its score and the truth live is declared once, for the project:

predictive:
  label_field: label
  score_field: score
  expected_field: label
  positive: true
  calibration_bins: 10
  thresholds: [0.3, 0.4, 0.5, 0.6, 0.7]
FieldDefaultMeaning
label_fieldlabelthe output field holding the predicted label
score_fieldscorethe output field holding the score behind it
expected_fieldlabelthe field under expected holding the truth
positivetruewhich label value counts as positive; refund, 1 or true
calibration_bins10how many score bands the calibration table uses
thresholdsnonecut-offs to sweep; each is evidence, never a recommendation
averagenonemacro or micro, for an aggregate over several classes

Every predictive evaluator that does not state field, expected_field or positive itself takes them from this block, so one suite has one positive class. The block is also what produces the confusion counts, the calibration table and the threshold sweep; a project without it gets the metrics and nothing beside them.

A positive class that no case's label matches is refused before anything runs, because recall over it would be a rate over nothing. The same project with positive: churned:

Configuration error: evaluator 'recall' counts 'churned' as the positive class, and no case's 'label' is 'churned' (labels: False, True); declare `positive:` on the evaluator or in the `predictive:` block

The evaluators

evaluators:
  - {type: predictive_correct, criterion: accuracy}
  - {type: predictive_recall, criterion: recall}
  - {type: predictive_precision, criterion: precision}
  - {type: predictive_brier, criterion: brier}
  - {type: predictive_log_loss, criterion: log_loss, clip: 0.02}
  - {type: predictive_ranking, criterion: rank}
metrics:
  - {id: roc_auc, type: ranking, criterion: rank, statistic: roc_auc}
  - {id: pr_auc, type: ranking, criterion: rank, statistic: average_precision}
slices: [metadata.plan, "confidence:0.5"]
min_slice_support: 20

predictive_log_loss requires clip, because one confident mistake is otherwise infinite. predictive_ranking is the criterion a ranking metric orders rows by, and it needs a metrics: entry naming the statistic: ROC-AUC and average precision answer different questions, and the engine will not pick one for you.

oloproof run
│ accuracy  │ 88.5%    │ [83.2%, 92.6%]  │ 177 / 200 observed · 0 missing · 0 excluded                                │
│ recall    │ 81.2%    │ [69.5%, 90.0%]  │ 52 / 64 observed · 0 missing · 136 excluded                                │
│ precision │ 82.5%    │ [70.9%, 91.0%]  │ 52 / 63 observed · 0 missing · 137 excluded                                │
│ brier     │ 0.120    │ [0.094, 0.154]  │ mean of 200 observed · 0 missing · 0 excluded                              │
│ log_loss  │ 0.389    │ [0.323, 0.499]  │ mean of 200 observed · 0 missing · 0 excluded                              │
│ roc_auc   │ 92.3%    │ [69.3%, 100.0%] │ roc_auc over 64 positive · 136 negative · 0 missing · 0 excluded           │
│ pr_auc    │ 86.5%    │                 │ average_precision over 64 positive · 136 negative · 0 missing · 0 excluded │

Read the excluded column. Recall is measured over the 64 accounts that churned, so the other 136 are excluded from it; precision over the 63 the model flagged. Thirty-two percent of these accounts churn, so a model that predicts nobody churns is 68% accurate and finds no one. A floor on accuracy alone would pass it, which is why the example's policy puts a floor on each rate.

pr_auc has an estimate and no interval. At two hundred rows its interval is genuinely weaker than ROC-AUC's, and the engine withholds a bound it cannot support rather than show one. A rule on it reads:

pr: INSUFFICIENT_EVIDENCE (interval_unavailable)

Beside the metrics

The confusion counts are counts, not rates:

│ actually positive │ 52                 │ 12                 │
│ actually negative │ 11                 │ 125                │

A release rule naming one is a configuration error, because a count is not a metric:

Configuration error: release rule 'fp' refers to unknown metric 'false_positives'

The calibration table sets what the model claimed against what happened, by score band:

│ 0.2-0.3 │ 26.7%   │ 0.0%     │ 34 rows │
│ 0.5-0.6 │ 53.4%   │ 81.0%    │ 21 rows │

And the threshold sweep shows what each declared cut-off would have measured:

│ 0.3     │ 57.4%     │ 96.9%  │ 62/108 predicted positive · 62/64 actual positive │
│ 0.5     │ 82.5%     │ 81.2%  │ 52/63 predicted positive · 52/64 actual positive  │
│ 0.7     │ 100.0%    │ 37.5%  │ 24/24 predicted positive · 24/64 actual positive  │

The sweep is titled Thresholds (exploratory; recommends nothing). Which cut-off is right depends on what a false positive costs against a false negative, and that is not something the engine can know.

Regression

A regressor is scored by absolute error, which needs the range its targets lie in:

evaluators:
  - type: predictive_absolute_error
    criterion: days_error
    field: days
    expected_field: days
    target_range: [0, 20]
rules:
  - id: error-budget
    metric: days_error
    max: 1.5
│ days_error │ 1.02     │ [0.78, 1.88] │ mean of 120 observed · 0 missing · 0 excluded │
error-budget: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)

target_range is required, not defaulted. An absolute error is a bounded mean, and its interval holds only within a range every value lies in. A wider range is a wider interval, so declare the range the targets can actually take. The rule cannot decide: the estimate is inside the budget, and 120 orders cannot yet show that the true mean error is.

Where to go next