返回 Skill 列表
extension
分类: 开发与工程无需 API Key

generalization-evaluator

跨领域评估以估计通用性和发现盲点。在需要评估广泛能力、跨领域比较模型或识别缺失技能时使用。

person作者: jakexiaohubgithub

Generalization Evaluator

Use this skill to measure generality across domains and identify weak coverage.

Workflow

  1. Load a task set (use references/task_set.example.json).
  2. Run the task set with a consistent runner.
  3. Score pass/fail per task and summarize by domain.
  4. Rank gaps by impact.

Scripts

  • Run: python scripts/run_eval.py --tasks references/task_set.example.json --runner ollama --model qwen3:latest

Output Expectations

  • Provide a domain score table and a short summary of weaknesses.
  • List the top 3 skill gaps with suggested skill actions.