LLM evaluation and regression harness
Baseline blocks releases when quality regresses.
Run a fixed golden test set on every prompt, model, or retrieval change. Metrics publish on the pull request and the merge fails when a threshold is missed.
baseline gate · pull request #418suite: payments-support (124 cases)model: anthropic.claude-3-5-sonnet (us-east-1)metric value threshold status────────────────────────────────────────────────accuracy 0.91 ≥ 0.88 ✓ passgroundedness 0.84 ≥ 0.90 ✗ failretrieval_hit_rate 0.96 ≥ 0.92 ✓ passp95_latency_ms 820 ≤ 1200 ✓ passcost_per_query_usd 0.014 ≤ 0.020 ✓ pass→ merge blocked until groundedness recoversThe problem
Without a gate, every prompt edit is a production experiment.
Baseline turns evaluation into a CI check: same cases, same thresholds, every change.
01 / fixed set
Golden cases you own
Define expected behavior for the questions that matter. The suite runs unchanged until you deliberately update it.
02 / thresholds
Fail the build, not the user
Set minimum accuracy, groundedness, hit rate, latency, and cost. A regression blocks merge before it reaches production.
03 / history
Compare across changes
Each run is stored with git metadata so you can see which commit moved a metric and whether it was intentional.
Metrics
What each run measures
Answer quality
Accuracy and groundedness against labeled expected answers and citation requirements.Retrieval hit rate
Whether the right documents appear in the candidate set before generation, which matters when ACL filters shrink the pool.Latency
p50 and p95 end-to-end time per case so performance regressions fail the build.Cost per query
Token and model spend estimated per case using your configured pricing tables.Custom assertions
Regex, JSON schema, or script hooks for domain rules that numeric scores alone cannot capture.
CI integration
A check on every relevant pull request
The CLI runs your suite, writes JUnit for your CI dashboard, and exits non-zero when a gate fails.
Pair with Clearance to include permission-aware retrieval traces in failure reports.
name: baseline-evalon: pull_request: paths: - "prompts/**" - "retrieval/**" - "models/**"jobs: eval: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Run golden set run: baseline run --suite payments-support --format junit env: BASELINE_API: ${{ secrets.BASELINE_API_URL }} BASELINE_TOKEN: ${{ secrets.BASELINE_TOKEN }} - name: Enforce thresholds run: baseline gate --min-accuracy 0.88 --min-groundedness 0.90FAQ
Common questions about Baseline
Run your golden set against our stack in a call.
Bring five to ten real questions. We wire a CI gate and show how failures surface on the pull request.