Evaluating Code Models
Help the agent provide practical guidance for code model evaluation workflows.
When to Use
- The user needs help measuring code model quality and reliability.
- The user asks about benchmark design, scoring, or failure analysis.
- The user wants actionable recommendations for improving model performance.
Instructions
- Clarify task mix (generation, repair, explanation) and target languages.
- Recommend benchmark construction and sampling strategy.
- Define automated and human review metrics.
- Include error taxonomy and regression tracking approach.
- End with an implementation and reporting checklist.
Output
- Recommended evaluation strategy
- Step-by-step benchmark plan
- Metrics, risks, and validation checklist
微信扫一扫