Skip to content

指南

分类器与回归器

预测模型的评估方式与其他系统相同:一个可调用对象为每个用例返回一个预测,评估器读取该预测。变化的是分母。准确率、召回率和精确率是三个不同行集合上的三个比率,一个模型可能在其中一个上表现良好,却无法回答业务真正关心的问题。

examples/churn_model/ 是本页所运行的项目:两百个账户、一个确定性的流失模型,无需任何提供方凭据。

预测

系统返回一个标签,以及(如果模型有的话)该标签背后的分数:

from oloproof import system


@system(name="churn-model", version="slice-f-example")
def run(account):
    score = churn_score(account)
    return {"label": score >= 0.5, "score": round(score, 4)}

用例在 expected 下声明真实值:

{"expected": {"label": false}, "id": "account_000", "input": {"recent_upgrade": true, "support_contacts": 0, "tenure_months": 0}, "metadata": {"plan": "enterprise"}}

`predictive:` 块

预测、其分数以及真实值位于何处,在项目级别声明一次:

predictive:
  label_field: label
  score_field: score
  expected_field: label
  positive: true
  calibration_bins: 10
  thresholds: [0.3, 0.4, 0.5, 0.6, 0.7]
字段默认值含义
label_fieldlabel存放预测标签的输出字段
score_fieldscore存放标签背后分数的输出字段
expected_fieldlabelexpected 下存放真实值的字段
positivetrue哪个标签值算作正类;refund、1 或 true
calibration_bins10校准表使用多少个分数区间
thresholds无要扫描的截断值;每一个都是证据,绝不是推荐
average无macro 或 micro,用于多个类别的汇总

每个未自行声明 field、expected_field 或 positive 的预测评估器都会从此块中获取它们,因此一个测试套件只有一个正类。混淆计数、校准表和阈值扫描也由此块产生;没有此块的项目只会得到指标,除此之外什么都没有。

如果没有任何用例的标签与正类匹配,该正类会在任何内容运行之前被拒绝,因为基于它的召回率将是一个空集合上的比率。同一个项目使用 positive: churned 时:

Configuration error: evaluator 'recall' counts 'churned' as the positive class, and no case's 'label' is 'churned' (labels: False, True); declare `positive:` on the evaluator or in the `predictive:` block

评估器

evaluators:
  - {type: predictive_correct, criterion: accuracy}
  - {type: predictive_recall, criterion: recall}
  - {type: predictive_precision, criterion: precision}
  - {type: predictive_brier, criterion: brier}
  - {type: predictive_log_loss, criterion: log_loss, clip: 0.02}
  - {type: predictive_ranking, criterion: rank}
metrics:
  - {id: roc_auc, type: ranking, criterion: rank, statistic: roc_auc}
  - {id: pr_auc, type: ranking, criterion: rank, statistic: average_precision}
slices: [metadata.plan, "confidence:0.5"]
min_slice_support: 20

predictive_log_loss 要求提供 clip,否则一个高置信度的错误就会导致无穷大。predictive_ranking 是排序指标用来对行排序的判据,它需要一个指明统计量的 metrics: 条目:ROC-AUC 和平均精确率回答的是不同的问题,引擎不会替你选择。

oloproof run
│ accuracy  │ 88.5%    │ [83.2%, 92.6%]  │ 177 / 200 observed · 0 missing · 0 excluded                                │
│ recall    │ 81.2%    │ [69.5%, 90.0%]  │ 52 / 64 observed · 0 missing · 136 excluded                                │
│ precision │ 82.5%    │ [70.9%, 91.0%]  │ 52 / 63 observed · 0 missing · 137 excluded                                │
│ brier     │ 0.120    │ [0.094, 0.154]  │ mean of 200 observed · 0 missing · 0 excluded                              │
│ log_loss  │ 0.389    │ [0.323, 0.499]  │ mean of 200 observed · 0 missing · 0 excluded                              │
│ roc_auc   │ 92.3%    │ [69.3%, 100.0%] │ roc_auc over 64 positive · 136 negative · 0 missing · 0 excluded           │
│ pr_auc    │ 86.5%    │                 │ average_precision over 64 positive · 136 negative · 0 missing · 0 excluded │

请看 excluded 列。召回率是在已流失的 64 个账户上度量的,因此其余 136 个被排除在外;精确率则是在模型标记的 63 个账户上度量的。这些账户中有 32% 会流失,因此一个预测没有人流失的模型准确率为 68%,却一个也找不到。仅对准确率设置下限会让它通过,这就是示例策略为每个比率都设置下限的原因。

pr_auc 有估计值,但没有区间。在两百行的规模下,它的区间确实比 ROC-AUC 的更弱,引擎宁可不给出无法支撑的界限,也不会展示一个。针对它的规则显示为:

pr: INSUFFICIENT_EVIDENCE (interval_unavailable)

指标之外

混淆计数是计数,而不是比率:

│ actually positive │ 52                 │ 12                 │
│ actually negative │ 11                 │ 125                │

在发布规则中引用某个计数属于配置错误,因为计数不是指标:

Configuration error: release rule 'fp' refers to unknown metric 'false_positives'

校准表按分数区间将模型声称的概率与实际发生的情况进行对照:

│ 0.2-0.3 │ 26.7%   │ 0.0%     │ 34 rows │
│ 0.5-0.6 │ 53.4%   │ 81.0%    │ 21 rows │

阈值扫描则展示每个已声明的截断值会度量出什么:

│ 0.3     │ 57.4%     │ 96.9%  │ 62/108 predicted positive · 62/64 actual positive │
│ 0.5     │ 82.5%     │ 81.2%  │ 52/63 predicted positive · 52/64 actual positive  │
│ 0.7     │ 100.0%    │ 37.5%  │ 24/24 predicted positive · 24/64 actual positive  │

该扫描的标题是 Thresholds (exploratory; recommends nothing)。哪个截断值合适取决于假阳性与假阴性各自的代价,而这是引擎无从得知的。

回归

回归器按绝对误差评分,这需要知道其目标值所处的范围:

evaluators:
  - type: predictive_absolute_error
    criterion: days_error
    field: days
    expected_field: days
    target_range: [0, 20]
rules:
  - id: error-budget
    metric: days_error
    max: 1.5
│ days_error │ 1.02     │ [0.78, 1.88] │ mean of 120 observed · 0 missing · 0 excluded │
error-budget: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)

target_range 是必填项,没有默认值。绝对误差是一个有界均值,其区间只有在所有值都位于某个范围内时才成立。范围越宽,区间越宽,因此请声明目标值实际可能取到的范围。该规则无法作出判定:估计值在预算之内,但 120 个订单还不足以表明真实的平均误差也在预算之内。

下一步

  • 切片 介绍 confidence: 区间与切片支持度。
  • 比较规则 介绍如何比较两个模型版本。