Skip to content

시작하기

빠른 시작

엔진을 설치하고, 프로젝트 골격을 만들고, 여러분의 머신에서 첫 실행과 게이트 결정을 얻습니다. 여러분이 push하지 않는 한 아무것도 머신 밖으로 나가지 않습니다.

이 단계들은 다시 타이핑한 것이 아니라 README의 내용을 여기에 렌더링한 것입니다. README.md는 개발자가 GitHub에서 만나는 사본이자 온보딩 테스트가 코드와 대조하는 사본이므로, 그것이 원본으로 남고 이 페이지는 scripts/generate_quickstart.py로 그로부터 생성됩니다.

개발용 설치

저장소 루트에서 다음을 실행합니다.

python3 -m venv .venv
.venv/bin/python -m pip install -e '.[dev]'
pnpm install --frozen-lockfile
make check
pnpm --filter @oloproof/web check

공개 Python 인터페이스는 다음과 같습니다.

from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, Contains, JsonSchema, Regex, RubricJudge, evaluator

oloproof_core 아래의 모든 것은 엔진 내부 구현입니다.

빠른 시작

작은 프로젝트를 생성합니다.

. .venv/bin/activate
oloproof init /tmp/oloproof-demo
cd /tmp/oloproof-demo
oloproof run

평가기

oloproof.yaml이 받는 type: 값은 다음과 같습니다.

결정론적exact_match, contains, regex, json_schema
LLM 심사 모델rubric_judge, groundedness_judge, citation_support_judge; probability_judge: 한 토큰의 확률로 답하는 예/아니요, 선택, 점수 질문; cascade: 먼저 확률 심사 모델, 그것이 확신하지 못하는 곳에서만 두 번째 심사 모델
모델model_classifier: TEI 서버에서 동작하는 학습된 분류기 또는 NLI 모델로, 임계값에 대해 채점합니다
검색hit_rate, recall, mrr, ndcg, citation_validity
에이전트agent_max_steps, agent_tool_called, agent_no_tool_loop, agent_tool_sequence, agent_no_undeclared_tool, agent_constraints_satisfied
멀티 에이전트agent_route, agent_tool_permissions, agent_max_handoffs
예측predictive_correct, predictive_precision, predictive_recall, predictive_ranking, predictive_brier, predictive_log_loss, predictive_absolute_error

심사는 자신의 provider(anthropic, openai 또는 openai_compatible), model, 그리고 rubric_text 또는 rubric_file로 된 루브릭을 명시합니다. OpenAI API를 사용하지만 OpenAI가 아닌 제공자(Groq, Together, 로컬 서버)에 있는 심사는 base_url과 api_key_env도 필요합니다.

evaluators:
  - type: rubric_judge
    criterion: answer_correct
    provider: openai_compatible
    model: openai/gpt-oss-20b
    base_url: https://api.groq.com/openai/v1
    api_key_env: GROQ_API_KEY
    rubric_text: "PASS if the answer conveys the same fact as the reference."

평가기가 받지 않는 필드를 지정하면, 그 평가기가 받는 필드 목록으로 응답합니다.

시스템이나 평가기를 다시 실행하지 않고 기존 실행을 다시 결정합니다.

oloproof gate RUN_ID --policy release.yaml

실패 목록과 케이스 하나를 살펴봅니다.

oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case refund_00

번들을 내보내고 웹 워크벤치에서 엽니다.

oloproof export RUN_ID
OLOPROOF_BUNDLE_DIR="$PWD/.oloproof/bundles" pnpm --dir /path/to/Oloproof --filter @oloproof/web dev

pnpm이 apps/web에서 앱을 시작하므로 번들 경로는 절대 경로여야 합니다. 워크벤치는 워크스페이스 인덱스에서 열리며, OLOPROOF_BUNDLE_DIR 하나만 있으면 그곳에 local/bundles 프로젝트로 나타납니다. 한 머신에서 여러 프로젝트를 다루려면 대신 workbench.json이 들어 있는 디렉터리를 OLOPROOF_WORKBENCH_DIR로 지정하십시오(docs/API_CONTRACTS.md).