返回 Skill 列表
extension
分类: 内容与媒体无需 API Key

Evaluating Code Models

关于评估代码生成和代码理解模型的实际指导。

person作者: jakexiaohubgithub

Evaluating Code Models

Help the agent provide practical guidance for code model evaluation workflows.

When to Use

  • The user needs help measuring code model quality and reliability.
  • The user asks about benchmark design, scoring, or failure analysis.
  • The user wants actionable recommendations for improving model performance.

Instructions

  1. Clarify task mix (generation, repair, explanation) and target languages.
  2. Recommend benchmark construction and sampling strategy.
  3. Define automated and human review metrics.
  4. Include error taxonomy and regression tracking approach.
  5. End with an implementation and reporting checklist.

Output

  • Recommended evaluation strategy
  • Step-by-step benchmark plan
  • Metrics, risks, and validation checklist