指南
分类器与回归器
预测模型的评估方式与其他系统相同:一个可调用对象为每个用例返回一个预测,评估器读取该预测。变化的是分母。准确率、召回率和精确率是三个不同行集合上的三个比率,一个模型可能在其中一个上表现良好,却无法回答业务真正关心的问题。
examples/churn_model/ 是本页所运行的项目:两百个账户、一个确定性的流失模型,无需任何提供方凭据。
预测
系统返回一个标签,以及(如果模型有的话)该标签背后的分数:
from oloproof import system
@system(name="churn-model", version="slice-f-example")
def run(account):
score = churn_score(account)
return {"label": score >= 0.5, "score": round(score, 4)}用例在 expected 下声明真实值:
{"expected": {"label": false}, "id": "account_000", "input": {"recent_upgrade": true, "support_contacts": 0, "tenure_months": 0}, "metadata": {"plan": "enterprise"}}`predictive:` 块
预测、其分数以及真实值位于何处,在项目级别声明一次:
predictive:
label_field: label
score_field: score
expected_field: label
positive: true
calibration_bins: 10
thresholds: [0.3, 0.4, 0.5, 0.6, 0.7]| 字段 | 默认值 | 含义 |
|---|---|---|
| label_field | label | 存放预测标签的输出字段 |
| score_field | score | 存放标签背后分数的输出字段 |
| expected_field | label | expected 下存放真实值的字段 |
| positive | true | 哪个标签值算作正类;refund、1 或 true |
| calibration_bins | 10 | 校准表使用多少个分数区间 |
| thresholds | 无 | 要扫描的截断值;每一个都是证据,绝不是推荐 |
| average | 无 | macro 或 micro,用于多个类别的汇总 |
每个未自行声明 field、expected_field 或 positive 的预测评估器都会从此块中获取它们,因此一个测试套件只有一个正类。混淆计数、校准表和阈值扫描也由此块产生;没有此块的项目只会得到指标,除此之外什么都没有。
如果没有任何用例的标签与正类匹配,该正类会在任何内容运行之前被拒绝,因为基于它的召回率将是一个空集合上的比率。同一个项目使用 positive: churned 时:
Configuration error: evaluator 'recall' counts 'churned' as the positive class, and no case's 'label' is 'churned' (labels: False, True); declare `positive:` on the evaluator or in the `predictive:` block评估器
evaluators:
- {type: predictive_correct, criterion: accuracy}
- {type: predictive_recall, criterion: recall}
- {type: predictive_precision, criterion: precision}
- {type: predictive_brier, criterion: brier}
- {type: predictive_log_loss, criterion: log_loss, clip: 0.02}
- {type: predictive_ranking, criterion: rank}
metrics:
- {id: roc_auc, type: ranking, criterion: rank, statistic: roc_auc}
- {id: pr_auc, type: ranking, criterion: rank, statistic: average_precision}
slices: [metadata.plan, "confidence:0.5"]
min_slice_support: 20predictive_log_loss 要求提供 clip,否则一个高置信度的错误就会导致无穷大。predictive_ranking 是排序指标用来对行排序的判据,它需要一个指明统计量的 metrics: 条目:ROC-AUC 和平均精确率回答的是不同的问题,引擎不会替你选择。
oloproof run│ accuracy │ 88.5% │ [83.2%, 92.6%] │ 177 / 200 observed · 0 missing · 0 excluded │
│ recall │ 81.2% │ [69.5%, 90.0%] │ 52 / 64 observed · 0 missing · 136 excluded │
│ precision │ 82.5% │ [70.9%, 91.0%] │ 52 / 63 observed · 0 missing · 137 excluded │
│ brier │ 0.120 │ [0.094, 0.154] │ mean of 200 observed · 0 missing · 0 excluded │
│ log_loss │ 0.389 │ [0.323, 0.499] │ mean of 200 observed · 0 missing · 0 excluded │
│ roc_auc │ 92.3% │ [69.3%, 100.0%] │ roc_auc over 64 positive · 136 negative · 0 missing · 0 excluded │
│ pr_auc │ 86.5% │ │ average_precision over 64 positive · 136 negative · 0 missing · 0 excluded │请看 excluded 列。召回率是在已流失的 64 个账户上度量的,因此其余 136 个被排除在外;精确率则是在模型标记的 63 个账户上度量的。这些账户中有 32% 会流失,因此一个预测没有人流失的模型准确率为 68%,却一个也找不到。仅对准确率设置下限会让它通过,这就是示例策略为每个比率都设置下限的原因。
pr_auc 有估计值,但没有区间。在两百行的规模下,它的区间确实比 ROC-AUC 的更弱,引擎宁可不给出无法支撑的界限,也不会展示一个。针对它的规则显示为:
pr: INSUFFICIENT_EVIDENCE (interval_unavailable)指标之外
混淆计数是计数,而不是比率:
│ actually positive │ 52 │ 12 │
│ actually negative │ 11 │ 125 │在发布规则中引用某个计数属于配置错误,因为计数不是指标:
Configuration error: release rule 'fp' refers to unknown metric 'false_positives'校准表按分数区间将模型声称的概率与实际发生的情况进行对照:
│ 0.2-0.3 │ 26.7% │ 0.0% │ 34 rows │
│ 0.5-0.6 │ 53.4% │ 81.0% │ 21 rows │阈值扫描则展示每个已声明的截断值会度量出什么:
│ 0.3 │ 57.4% │ 96.9% │ 62/108 predicted positive · 62/64 actual positive │
│ 0.5 │ 82.5% │ 81.2% │ 52/63 predicted positive · 52/64 actual positive │
│ 0.7 │ 100.0% │ 37.5% │ 24/24 predicted positive · 24/64 actual positive │该扫描的标题是 Thresholds (exploratory; recommends nothing)。哪个截断值合适取决于假阳性与假阴性各自的代价,而这是引擎无从得知的。
回归
回归器按绝对误差评分,这需要知道其目标值所处的范围:
evaluators:
- type: predictive_absolute_error
criterion: days_error
field: days
expected_field: days
target_range: [0, 20]rules:
- id: error-budget
metric: days_error
max: 1.5│ days_error │ 1.02 │ [0.78, 1.88] │ mean of 120 observed · 0 missing · 0 excluded │error-budget: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)target_range 是必填项,没有默认值。绝对误差是一个有界均值,其区间只有在所有值都位于某个范围内时才成立。范围越宽,区间越宽,因此请声明目标值实际可能取到的范围。该规则无法作出判定:估计值在预算之内,但 120 个订单还不足以表明真实的平均误差也在预算之内。