Part IV · Production
Chapter 14
Evaluation and Regression Testing
Correctness here is statistical, so testing is sampling — the pipeline is easy; the case set is the work, and it is the thing that makes every earlier decision defensible.
Deliverable: A gated eval suite: a real case set, deterministic scorers first, a calibrated judge, and canary with cheap rollback.
What's inside
11 topics
- 14.1Why Testing Is Sampling
- 14.2Building the Case Set
- 14.3Deterministic Scorers First
- 14.4Model Judges and Their Calibration
- 14.5Continuous Regression Detection
- 14.6The Release Gate
- 14.7Canary and Rollback
- 14.8Cost and Latency as Gated Metrics
- 14.9Adversarial and Safety Suites
- 14.10Human Review That Scales
- 14.11Three Evaluation Architectures Compared
Preparing PDF viewer…
A note on this content
The book and its chapters are my personal learning notes — compiled from online research and hands-on practice, with most of the content AI-generated from that research and learning. It is not a peer-reviewed publication, and I make no claim that it is 100% error-free. If you spot a mistake, I'd genuinely appreciate hearing about it — contact me.