开始使用
Python API
本页的英文版本在翻译之后已有更改。英文页面为最新版本。 阅读英文版
CLI 能做的一切,库也都能做。当评估应当放在脚本、notebook 或测试套件中,而不是放在配置文件旁边时,就使用它。
from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, Contains, JsonSchema, Regex, RubricJudge, evaluatoroloproof_core 下的所有内容都属于引擎内部实现,不在这个公开接口之内。
系统接收的是用例的输入,而不是用例本身
这是最值得先弄对的一件事,因为弄错了它会悄无声息地失败。
@system(name="support-bot", version="1")
def answer(case):
return {"label": "refund" if "refund" in case["question"].lower() else "other"}数据集中的一行如下所示:
{"id":"refund_00","input":{"question":"Can I get a refund? #0"},"expected":{"label":"refund"}}函数接收的是 input 对象,因此 case["question"] 就是问题,而不存在 case["expected"]。它返回的内容就是评估器读取的输出,因此 ExactMatch(field="label") 会将返回的 label 与该用例预期的 label 进行比较。
运行一次评估
result = evaluate(
system=answer,
dataset="data/example.jsonl",
evaluators=[ExactMatch(criterion="exact_label", field="label")],
)
for metric in result.metrics:
print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)exact_label 1.0 lower=0.8842966917779722 upper=1.0 30 0三十个用例全部正确,而区间下界仍然低至 88.4%。无论估计值是多少,三十个用例能确立的也就只有这么多。
查看失败的用例
如果系统抛出异常,该用例会被记为缺失而不是错误,指标也会如实报告:
exact_label None lower=0.0 upper=1.0 0 30估计值为 None、区间覆盖整个取值范围,意味着没有观测到任何东西。在信任一个估计值之前,请先检查 n_missing:一个每个用例都抛出异常的系统,产生的结果对象在结构上完全正常,只有这个计数能告诉你它是空的。
这是分母约定在正常发挥作用,而不是它的失效——出错的用例会被纳入界限,而不是被丢弃——但没有任何东西会替你抛出异常,所以这项检查需要你自己来做。
编写你自己的评估器
@evaluator 会把一个函数变成评估器。该函数接收一个参数,即用例,并从中读取所需的内容:case.output 是系统返回的内容,case.expected 是该行的 expected 对象。
from oloproof.evaluators import evaluator
@evaluator(criterion="known_label", reads=("output",))
def known_label(case):
return case.output["label"] in {"refund", "other"}
result = evaluate(system=answer, dataset="data/example.jsonl", evaluators=[known_label])known_label 1.0 lower=0.8842966917779722 upper=1.0 30 0reads 声明它读取用例的哪些部分,用于溯源;默认值是 ("output", "expected")。它返回 True 或 False;如果它声明了 value_type="score" 以及分数所在的 score_range,则返回一个分数。它的版本包含定义它的文件的摘要,因此编辑该文件会使它已缓存的评判结果失效。
写成 def f(output, expected) 形式的函数会在声明时就被拒绝,并附上可用的写法,而不是在每个用例上失败。
应用策略
evaluate 接受一个策略并对运行做出决策,所用的规则与 CLI 从 release.yaml 中读取的相同:
result = evaluate(
system=answer,
dataset="data/example.jsonl",
evaluators=[ExactMatch(criterion="exact_label", field="label")],
policy="release.yaml",
)
print(result.gate)如果没有策略,result.gate 为 None:没有可以不达标的规则,因此也就没有需要决定的东西。
下一步
- 编写测试套件 以配置的形式介绍同样的评估器。
- 记录系统做了什么 介绍 current_case() 以及如何在评估器中读取产物。
- 将候选与基线进行比较 介绍区间所服务的工作流程。