D9: Eval Validation (20 points)
D9: Eval Validation (20 points)
Purpose: Verify the skill has been validated at runtime through eval scenarios, proving agents actually follow its instructions.
Scoring:
| Points | Signal |
|---|---|
| 17–20 | Complete evals with ≥80% instruction coverage, ≥3 valid scenarios |
| 13–16 | Evals present with partial coverage or incomplete scenarios |
| 7–12 | Evals directory exists but missing key files |
| 1–6 | Minimal eval structure, no coverage data |
| 0 | No evals directory |
Core principle: Static quality (D1–D8) is necessary but not sufficient. Runtime validation proves the skill actually changes agent behaviour.
Components
1. Eval Directory Structure (4 points)
evals/directory exists with proper layout- Follows the skill eval harness conventions (scenario-N/{task.md,criteria.json,capability.txt}; criteria sum to 100)
2. Instruction Inventory (3 points)
instructions.jsonpresent and non-empty- Every instruction extracted from
SKILL.md - Each instruction classified by
why_given:reminder,new knowledge,preference
3. Coverage Statistics (6 points)
summary.jsonwithinstructions_coveragedata (3 points)- Coverage percentage ≥ 80% (3 points)
4. Valid Scenarios (4 points)
- ≥ 3 scenarios with complete structure (
task.md+criteria.json+capability.txt) - Each
criteria.jsonsums to exactly 100
5. Criteria Quality (3 points)
- 10+ checklist items per scenario
- Binary yes/no criteria traceable to specific instructions
- No instruction leakage in
task.md
Relationship to D1 and D3
When instructions.json exists, its data enriches other dimensions:
- D1 (Knowledge Delta): The
why_givendistribution (new knowledge+preferencevsreminder) provides a more accurate expert content ratio than heuristics alone. - D3 (Anti-Pattern Quality): Instructions containing NEVER/ALWAYS/anti-pattern keywords are cross-referenced with scenario coverage for a stronger signal.
Creating Evals
Author scenarios in cmd/assets/evals/scenario-N/ (one directory per
scenario; each contains task.md, criteria.json, capability.txt).
Run the native eval runner:
# Structural gate — deterministic, no key needed, required every PR.
./dist/skill-auditor eval ./cmd/assets --fail-below 0
# LLM-judge advisory — bring your own key. See .env.example for variables.
ANTHROPIC_API_KEY=... ./dist/skill-auditor eval ./cmd/assets --json --samples 3 --cost-log
Examples
High Eval Validation (19/20):
skill-name/evals/
instructions.json # 28 instructions extracted
summary.json # 100% coverage, 5 scenarios
summary_infeasible.json
scenario-0/ # task.md + criteria.json (sum=100) + capability.txt
scenario-1/
scenario-2/
scenario-3/
scenario-4/
Low Eval Validation (4/20):
skill-name/evals/
instructions.json # present but only 5 instructions
# no summary.json, no scenarios
Zero Eval Validation (0/20):
skill-name/
SKILL.md # no evals/ directory at all
Academic References
@article{rehan2026tdad,
title = {Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications},
author = {T. Rehan},
year = {2026},
journal = {arXiv preprint arXiv:2603.08806},
eprint = {2603.08806},
archivePrefix = {arXiv},
url = {<https://arxiv.org/abs/2603.08806}>
}
@article{alami2026camouflage,
title = {Cognitive Camouflage: Specification Gaming in LLM-Generated Code Evades Holistic Evaluation but Not Adversarial Execution},
author = {Alami},
year = {2026},
journal = {SSRN},
url = {<https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6512960}>
}
@article{wangmutationtesting,
title = {A Comprehensive Study on Large Language Models for Mutation Testing},
author = {Wang and Chen and Deng and Lin and Harman and others},
journal = {ACM Transactions on Software Engineering and Methodology},
publisher = {ACM},
url = {<https://dl.acm.org/doi/abs/10.1145/3805038}>
}
@article{pan2026benchmarks,
title = {Re-Evaluating Code LLM Benchmarks Under Semantic Mutation},
author = {Pan and Hu and Xia and Yang},
journal = {arXiv preprint arXiv:2506.17369},
eprint = {2506.17369},
archivePrefix = {arXiv},
url = {<https://arxiv.org/abs/2506.17369}>
}
@article{bouafifprimg,
title = {PrimG: Efficient LLM-Driven Test Generation Using Mutant Prioritization},
author = {Bouafif and Hamdaqa and Zulkoski},
journal = {Proceedings of the 2025 ACM Conference on Automated Software Engineering},
publisher = {ACM},
url = {<https://dl.acm.org/doi/abs/10.1145/3756681.3756991}>
}