“The agent feels better after the prompt change” is not evidence. **evalspec** turns agent quality into gold labels: for every question in your eval set, you define what the correct behavior *is* — answer, ABSTAIN, CLARIFY, or a specific tool call — and the harness measures what the model actually did. ## What it does - `evalspec-run` — the harness. Pluggable agents: OpenAI, HTTP endpoints, mocks. - `evalspec-split` — 80/20 stratified holdout, so you tune on one set and report on another. - `evalspec-compare` — side-by-side model comparison with an HTML report. - `evalspec-regression` — a CI gate with 3% tolerance and dataset-hash version-freezing. A prompt change that quietly breaks abstention fails the build instead of reaching production. - `evalspec-leakage` — a parametric filter that catches eval questions leaking into training or prompt context. ## Why I built it I kept rebuilding the same bespoke eval harness for every agent project. Generalizing the philosophy into a library paid off immediately: replacing one bespoke harness with evalspec cut it from 700 lines to 120. The deeper opinion baked into the tool: knowing when an agent should *refuse to answer* is as important as knowing when it answers correctly — so abstention is a first-class gold label, not an afterthought. ## Get it `pip install evalspec` — MIT license. Source on [GitHub](https://github.com/Ref34t/evalspec), package on [PyPI](https://pypi.org/project/evalspec/).