Spec: Blind A/B Comparison for Skill Version Testing
Issue: #558
Depends on: #562 (skill regression testing — merged via #775)
Pattern reference: docs/specs/562-skill-regression-testing.md
Problem
Comparing two versions of a skill requires controlled, unbiased evaluation. Without a blind comparison mechanism, skill authors rely on gut feel or sequential evaluation (which introduces order bias). There is no way to determine whether version B is genuinely better than version A, or whether observed differences are within noise.
Proposed Solution
Add a blind A/B comparison mode to the eval command. Two members run against the same suite; each task's outputs are scored independently by an LLM judge that does not know which member produced which output. Scores are aggregated across tasks and a paired t-test determines whether the observed difference is statistically significant.
Data Model
Metrics
Three new judge metrics for A/B comparison:
| Metric | Rubric |
|---|---|
| clarity | How clear, well-structured, and easy to understand the response is. |
| completeness | How thoroughly the response addresses all aspects of the task. |
| accuracy | How factually correct and precise the response is. |
These complement the existing metrics (faithfulness, relevance, context_recall, answer_correctness) and are optimized for direct output-to-output comparison.
Pure Functions (src/evals/abComparison.ts)
pairedTTest: computes mean difference, sample standard error, t = mean / SE. Uses normal approximation with small-sample correction for the two-tailed p-value.cohensD: mean difference / standard deviation of differences.effectLabel: |d| < 0.2 negligible, < 0.5 small, < 0.8 medium, else large.
Blind Judge (src/evals/BlindJudge.ts)
- For each metric, scores outputA and outputB independently via two separate
judge.score()calls. - No randomization needed since each output is scored in isolation (no order bias).
- Computes deltas (B - A) per metric.
CLI Integration (src/commands/eval.ts)
New flag: --ab <memberB>
agenthood eval <memberA> --ab <memberB> --suite <path>
When --ab is set:
- Run memberA against the suite → reportA
- Run memberB against the suite → reportB
- For each task, call
BlindJudge.compareTask()with both outputs - Aggregate scores and run significance tests
- Print comparison table and result
Exit codes: 0 = A wins or tie, 1 = B wins (useful for CI gating on regression).
Output Format
A/B Comparison — the-scribe (A) vs the-scribe-v2 (B)
Suite: the-scribe.json | Tasks: 5 | Metrics: clarity, completeness, accuracy
Per-task:
Task 1: "Write a commit message..."
A: cl 0.85, co 0.90, ac 0.88 | avg 0.877
B: cl 0.92, co 0.87, ac 0.91 | avg 0.900
Δ: +0.023 → B
...
Aggregate:
A B Δ
clarity: 0.83 0.89 +0.06
comp.: 0.88 0.85 -0.03
acc.: 0.86 0.90 +0.04
Winner: B (2/3 metrics, significant, p=0.032)
Effect: medium (Cohen's d = 0.45)
Out of Scope (YAGNI)
- Multiple judges (one LLM judge is sufficient for now)
- Visualization / dashboards
- Automatic CI gating on A/B results (exit code is enough)
- Cross-suite A/B comparison
- History tracking for A/B results
- Per-metric significance testing (only overall)
Testing Strategy
- Unit (
tests/unit/evals/abComparison.test.ts): Pure functions with synthetic data. t-test against known values, Cohen's effect size, winner determination, edge cases (empty, single task, all ties). - Unit (
tests/unit/evals/BlindJudge.test.ts): Mock judge returns fixed scores. Verify correct mapping of scores to A/B, delta computation, metric coverage.