LLM Regression Report Analyzer
Use paired case-level results when available. Do not infer significance from rounded aggregate percentages alone.
Workflow
- Verify dataset, rubric, judge, sampling, model settings, and traffic conditions are comparable.
- Compute absolute and relative deltas by metric and critical slice.
- Inspect paired wins/losses and cluster failures by task, language, length, tool, and risk.
- Separate random variance, evaluator drift, infrastructure errors, and genuine model behavior changes.
- Apply hard gates to safety and critical-task metrics before weighted averages.
- Include latency, retry rate, token use, and cost per successful task.
- Recommend release, canary, fix-and-retest, or rollback with explicit evidence.
Output
Return data-quality warnings, metric deltas, worst slices, representative failure IDs, confidence limitations, and a decision with required follow-up tests. Never hide a critical regression inside an improved overall mean.
Scan to join WeChat group