返回 Skill 列表
extension
分类: 数据与分析无需 API Key

LLM 回归报告分析器

Analyze baseline-versus-candidate LLM evaluation reports, distinguish meaningful regressions from noise, and recommend release, canary, or rollback. Use for model, prompt, retrieval, tool, or provider changes with quality, safety, latency, and cost metrics.

person作者: user_b26bd65ehubcommunity

LLM Regression Report Analyzer

Use paired case-level results when available. Do not infer significance from rounded aggregate percentages alone.

Workflow

  1. Verify dataset, rubric, judge, sampling, model settings, and traffic conditions are comparable.
  2. Compute absolute and relative deltas by metric and critical slice.
  3. Inspect paired wins/losses and cluster failures by task, language, length, tool, and risk.
  4. Separate random variance, evaluator drift, infrastructure errors, and genuine model behavior changes.
  5. Apply hard gates to safety and critical-task metrics before weighted averages.
  6. Include latency, retry rate, token use, and cost per successful task.
  7. Recommend release, canary, fix-and-retest, or rollback with explicit evidence.

Output

Return data-quality warnings, metric deltas, worst slices, representative failure IDs, confidence limitations, and a decision with required follow-up tests. Never hide a critical regression inside an improved overall mean.