Skip to content

开始使用

Python API

本页的英文版本在翻译之后已有更改。英文页面为最新版本。 阅读英文版

CLI 能做的一切,库也都能做。当评估应当放在脚本、notebook 或测试套件中,而不是放在配置文件旁边时,就使用它。

from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, Contains, JsonSchema, Regex, RubricJudge, evaluator

oloproof_core 下的所有内容都属于引擎内部实现,不在这个公开接口之内。

系统接收的是用例的输入,而不是用例本身

这是最值得先弄对的一件事,因为弄错了它会悄无声息地失败。

@system(name="support-bot", version="1")
def answer(case):
    return {"label": "refund" if "refund" in case["question"].lower() else "other"}

数据集中的一行如下所示:

{"id":"refund_00","input":{"question":"Can I get a refund? #0"},"expected":{"label":"refund"}}

函数接收的是 input 对象,因此 case["question"] 就是问题,而不存在 case["expected"]。它返回的内容就是评估器读取的输出,因此 ExactMatch(field="label") 会将返回的 label 与该用例预期的 label 进行比较。

运行一次评估

result = evaluate(
    system=answer,
    dataset="data/example.jsonl",
    evaluators=[ExactMatch(criterion="exact_label", field="label")],
)

for metric in result.metrics:
    print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)
exact_label 1.0 lower=0.8842966917779722 upper=1.0 30 0

三十个用例全部正确,而区间下界仍然低至 88.4%。无论估计值是多少,三十个用例能确立的也就只有这么多。

查看失败的用例

如果系统抛出异常,该用例会被记为缺失而不是错误,指标也会如实报告:

exact_label None lower=0.0 upper=1.0 0 30

估计值为 None、区间覆盖整个取值范围,意味着没有观测到任何东西。在信任一个估计值之前,请先检查 n_missing:一个每个用例都抛出异常的系统,产生的结果对象在结构上完全正常,只有这个计数能告诉你它是空的。

这是分母约定在正常发挥作用,而不是它的失效——出错的用例会被纳入界限,而不是被丢弃——但没有任何东西会替你抛出异常,所以这项检查需要你自己来做。

编写你自己的评估器

@evaluator 会把一个函数变成评估器。该函数接收一个参数,即用例,并从中读取所需的内容:case.output 是系统返回的内容,case.expected 是该行的 expected 对象。

from oloproof.evaluators import evaluator

@evaluator(criterion="known_label", reads=("output",))
def known_label(case):
    return case.output["label"] in {"refund", "other"}

result = evaluate(system=answer, dataset="data/example.jsonl", evaluators=[known_label])
known_label 1.0 lower=0.8842966917779722 upper=1.0 30 0

reads 声明它读取用例的哪些部分,用于溯源;默认值是 ("output", "expected")。它返回 True 或 False;如果它声明了 value_type="score" 以及分数所在的 score_range,则返回一个分数。它的版本包含定义它的文件的摘要,因此编辑该文件会使它已缓存的评判结果失效。

写成 def f(output, expected) 形式的函数会在声明时就被拒绝,并附上可用的写法,而不是在每个用例上失败。

应用策略

evaluate 接受一个策略并对运行做出决策,所用的规则与 CLI 从 release.yaml 中读取的相同:

result = evaluate(
    system=answer,
    dataset="data/example.jsonl",
    evaluators=[ExactMatch(criterion="exact_label", field="label")],
    policy="release.yaml",
)
print(result.gate)

如果没有策略,result.gate 为 None:没有可以不达标的规则,因此也就没有需要决定的东西。

下一步