Skip to content

指南

分類器與迴歸模型

預測模型的評估方式與任何其他系統相同:一個可呼叫物件為每個案例回傳一個預測,再由評估器讀取。不同之處在於分母。準確率、召回率與精確率是針對三組不同資料列的三種比率,一個模型可能在其中一項上看起來很好,卻答不出業務真正在問的問題。

examples/churn_model/ 是本頁所執行的專案:兩百個帳戶、一個確定性的流失模型,而且不需要任何提供者憑證。

預測

系統會回傳一個標籤,若模型有分數,也會一併回傳其背後的分數:

from oloproof import system


@system(name="churn-model", version="slice-f-example")
def run(account):
    score = churn_score(account)
    return {"label": score >= 0.5, "score": round(score, 4)}

案例在 expected 下宣告真實值:

{"expected": {"label": false}, "id": "account_000", "input": {"recent_upgrade": true, "support_contacts": 0, "tenure_months": 0}, "metadata": {"plan": "enterprise"}}

`predictive:` 區塊

預測、其分數與真實值位於何處,只需為整個專案宣告一次:

predictive:
  label_field: label
  score_field: score
  expected_field: label
  positive: true
  calibration_bins: 10
  thresholds: [0.3, 0.4, 0.5, 0.6, 0.7]
欄位預設值意義
label_fieldlabel存放預測標籤的輸出欄位
score_fieldscore存放其背後分數的輸出欄位
expected_fieldlabelexpected 下存放真實值的欄位
positivetrue哪個標籤值算作正類;refund、1 或 true
calibration_bins10校準表使用多少個分數區間
thresholds無要掃描的截斷點;每一個都是證據,絕非建議
average無macro 或 micro,用於跨多個類別的彙總

凡是未自行指定 field、expected_field 或 positive 的預測評估器,都會從這個區塊取得這些設定,因此一個套件只有一個正類。這個區塊也負責產生混淆計數、校準表與門檻掃描;沒有這個區塊的專案只會得到指標,此外什麼都沒有。

若沒有任何案例的標籤符合所指定的正類,該設定會在任何東西執行前就被拒絕,因為針對它的召回率將是對空集合求比率。同一個專案改用 positive: churned 時:

Configuration error: evaluator 'recall' counts 'churned' as the positive class, and no case's 'label' is 'churned' (labels: False, True); declare `positive:` on the evaluator or in the `predictive:` block

評估器

evaluators:
  - {type: predictive_correct, criterion: accuracy}
  - {type: predictive_recall, criterion: recall}
  - {type: predictive_precision, criterion: precision}
  - {type: predictive_brier, criterion: brier}
  - {type: predictive_log_loss, criterion: log_loss, clip: 0.02}
  - {type: predictive_ranking, criterion: rank}
metrics:
  - {id: roc_auc, type: ranking, criterion: rank, statistic: roc_auc}
  - {id: pr_auc, type: ranking, criterion: rank, statistic: average_precision}
slices: [metadata.plan, "confidence:0.5"]
min_slice_support: 20

predictive_log_loss 需要 clip,否則一個自信的錯誤就會變成無限大。predictive_ranking 是排名指標用來排序資料列的準則,它需要一個指明統計量的 metrics: 項目:ROC-AUC 與平均精確率回答的是不同的問題,引擎不會替你挑選其中一個。

oloproof run
│ accuracy  │ 88.5%    │ [83.2%, 92.6%]  │ 177 / 200 observed · 0 missing · 0 excluded                                │
│ recall    │ 81.2%    │ [69.5%, 90.0%]  │ 52 / 64 observed · 0 missing · 136 excluded                                │
│ precision │ 82.5%    │ [70.9%, 91.0%]  │ 52 / 63 observed · 0 missing · 137 excluded                                │
│ brier     │ 0.120    │ [0.094, 0.154]  │ mean of 200 observed · 0 missing · 0 excluded                              │
│ log_loss  │ 0.389    │ [0.323, 0.499]  │ mean of 200 observed · 0 missing · 0 excluded                              │
│ roc_auc   │ 92.3%    │ [69.3%, 100.0%] │ roc_auc over 64 positive · 136 negative · 0 missing · 0 excluded           │
│ pr_auc    │ 86.5%    │                 │ average_precision over 64 positive · 136 negative · 0 missing · 0 excluded │

請看 excluded 欄。召回率是針對 64 個已流失的帳戶量測的,因此其他 136 個帳戶被排除在外;精確率則針對模型標記的 63 個帳戶。這些帳戶中有百分之三十二會流失,因此一個預測沒有人會流失的模型,準確率為 68%,卻一個也找不到。只對準確率設下限會讓它通過,這就是為什麼範例的政策對每一種比率都設了下限。

pr_auc 有估計值但沒有區間。在兩百列資料下,它的區間確實比 ROC-AUC 的更弱,引擎會扣下它無法支持的界限,而不是勉強顯示一個。針對它的規則顯示為:

pr: INSUFFICIENT_EVIDENCE (interval_unavailable)

指標之外

混淆計數是計數,而不是比率:

│ actually positive │ 52                 │ 12                 │
│ actually negative │ 11                 │ 125                │

指名其中一項的發布規則屬於設定錯誤,因為計數不是指標:

Configuration error: release rule 'fp' refers to unknown metric 'false_positives'

校準表依分數區間,將模型宣稱的結果與實際發生的結果對照:

│ 0.2-0.3 │ 26.7%   │ 0.0%     │ 34 rows │
│ 0.5-0.6 │ 53.4%   │ 81.0%    │ 21 rows │

而門檻掃描則顯示每個宣告的截斷點會量測到什麼:

│ 0.3     │ 57.4%     │ 96.9%  │ 62/108 predicted positive · 62/64 actual positive │
│ 0.5     │ 82.5%     │ 81.2%  │ 52/63 predicted positive · 52/64 actual positive  │
│ 0.7     │ 100.0%    │ 37.5%  │ 24/24 predicted positive · 24/64 actual positive  │

這項掃描的標題是 Thresholds (exploratory; recommends nothing)。哪個截斷點才正確,取決於偽陽性相對於偽陰性的代價,而這不是引擎能夠知道的。

迴歸

迴歸模型以絕對誤差評分,這需要知道其目標值所在的範圍:

evaluators:
  - type: predictive_absolute_error
    criterion: days_error
    field: days
    expected_field: days
    target_range: [0, 20]
rules:
  - id: error-budget
    metric: days_error
    max: 1.5
│ days_error │ 1.02     │ [0.78, 1.88] │ mean of 120 observed · 0 missing · 0 excluded │
error-budget: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)

target_range 是必填的,沒有預設值。絕對誤差是一個有界平均數,其區間只在每個值都落入的範圍內才成立。範圍越寬,區間就越寬,因此請宣告目標值實際可能取到的範圍。這條規則無法做出判定:估計值落在預算之內,但 120 筆訂單還無法證明真實的平均誤差也是如此。

下一步

  • 切片 介紹 confidence: 區間與切片支持度。
  • 比較規則 介紹如何比較兩個模型版本。