Skip to content

LLM evaluation and regression harness

Baseline blocks releases when quality regresses.

Run a fixed golden test set on every prompt, model, or retrieval change. Metrics publish on the pull request and the merge fails when a threshold is missed.

PR check outputExample
baseline gate · pull request #418suite: payments-support (124 cases)model: anthropic.claude-3-5-sonnet (us-east-1)metric              value    threshold   status────────────────────────────────────────────────accuracy            0.91     ≥ 0.88      ✓ passgroundedness        0.84     ≥ 0.90      ✗ failretrieval_hit_rate  0.96     ≥ 0.92      ✓ passp95_latency_ms      820      ≤ 1200      ✓ passcost_per_query_usd  0.014    ≤ 0.020     ✓ pass→ merge blocked until groundedness recovers

The problem

Without a gate, every prompt edit is a production experiment.

Baseline turns evaluation into a CI check: same cases, same thresholds, every change.

  1. 01 / fixed set

    Golden cases you own

    Define expected behavior for the questions that matter. The suite runs unchanged until you deliberately update it.

  2. 02 / thresholds

    Fail the build, not the user

    Set minimum accuracy, groundedness, hit rate, latency, and cost. A regression blocks merge before it reaches production.

  3. 03 / history

    Compare across changes

    Each run is stored with git metadata so you can see which commit moved a metric and whether it was intentional.

Metrics

What each run measures

  • Answer quality

    Accuracy and groundedness against labeled expected answers and citation requirements.
  • Retrieval hit rate

    Whether the right documents appear in the candidate set before generation, which matters when ACL filters shrink the pool.
  • Latency

    p50 and p95 end-to-end time per case so performance regressions fail the build.
  • Cost per query

    Token and model spend estimated per case using your configured pricing tables.
  • Custom assertions

    Regex, JSON schema, or script hooks for domain rules that numeric scores alone cannot capture.

CI integration

A check on every relevant pull request

The CLI runs your suite, writes JUnit for your CI dashboard, and exits non-zero when a gate fails.

Pair with Clearance to include permission-aware retrieval traces in failure reports.

.github/workflows/baseline-eval.ymlExample
name: baseline-evalon:  pull_request:    paths:      - "prompts/**"      - "retrieval/**"      - "models/**"jobs:  eval:    runs-on: ubuntu-latest    steps:      - uses: actions/checkout@v4      - name: Run golden set        run: baseline run --suite payments-support --format junit        env:          BASELINE_API: ${{ secrets.BASELINE_API_URL }}          BASELINE_TOKEN: ${{ secrets.BASELINE_TOKEN }}      - name: Enforce thresholds        run: baseline gate --min-accuracy 0.88 --min-groundedness 0.90

FAQ

Common questions about Baseline

Run your golden set against our stack in a call.

Bring five to ten real questions. We wire a CI gate and show how failures surface on the pull request.