Agentic Eval
Agentic Eval is an agent skill (a SKILL.md file) from github/awesome-copilot. Patterns and techniques for evaluating and improving AI agent outputs. It works with Claude Code and GitHub Copilot and has 39,749 GitHub stars across a repository of 7 listed skills.
github.com/github/awesome-copilot/skills/agentic-eval (opens in a new tab)
- Official
- Multi-skill repo
- Testing & QA
- Actively maintained
Add this skill
Claude
Claude Code loads skills from ~/.claude/skills/ (all projects) or .claude/skills/ (one project):
git clone --depth 1 https://github.com/github/awesome-copilot.git
cp -r awesome-copilot/skills/agentic-eval ~/.claude/skills/agentic-eval # personal, or .claude/skills in a projectIn the Claude apps, zip the agentic-eval folder and upload it under Customize > Skills > + > Upload a skill (code execution must be on).
ChatGPT / Codex
Codex reads skills from .agents/skills/ in a repo or ~/.agents/skills/ for every project:
git clone --depth 1 https://github.com/github/awesome-copilot.git
cp -r awesome-copilot/skills/agentic-eval .agents/skills/agentic-eval # repo; ~/.agents/skills for all projectsStandalone skills also load in the ChatGPT desktop app.
Cursor
Cursor loads skills from .cursor/skills/ (or ~/.cursor/skills/) and also reads .claude/skills/:
git clone --depth 1 https://github.com/github/awesome-copilot.git
cp -r awesome-copilot/skills/agentic-eval .cursor/skills/agentic-eval # project; ~/.cursor/skills for all projectsSource (checked Oct 7, 2026): code.claude.com/docs/en/skills (opens in a new tab), support.claude.com/en/articles/12512180-using-skills-in-claude (opens in a new tab), learn.chatgpt.com/docs/build-skills (opens in a new tab), cursor.com/docs/context/skills (opens in a new tab)
What this skill does
Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimizer pipelines for quality-critical generation - Creating test-driven code refinement workflows - Designing rubric-based or LLM-as-judge evaluation systems - Adding iterative improvement to agent outputs (code, reports, analysis) - Measuring and improving agent response quality Patterns for self-improvement through iterative…
More skills in github/awesome-copilot
7 skills are listed from this repository.
Similar skills
More testing & qa skills
Systematic Debugging
obra/superpowers
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
Testing & QAShellTest Driven Development
obra/superpowers
Use when implementing any feature or bugfix, before writing implementation code.
Testing & QAShellVerification Before Completion
obra/superpowers
Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims…
Testing & QAShellFinishing A Development Branch
obra/superpowers
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work.
Testing & QAShellDiagnosing Superpowers
obra/superpowers
Use when a superpowers session went wrong and your human partner wants to know why — repeated work, ignored plans, stumbles, poor results, a skill that didn't fire, "it took too long", "why is it so…
Testing & QAShellTDD
mattpocock/skills
Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.
Testing & QAShell
What is the Agentic Eval skill?
Agentic Eval is an agent skill (a SKILL.md file) from github/awesome-copilot. Patterns and techniques for evaluating and improving AI agent outputs. It works with Claude Code and GitHub Copilot and has 39,749 GitHub stars across a repository of 7 listed skills. Its SKILL.md lives at github.com/github/awesome-copilot/skills/agentic-eval.
How do I install the Agentic Eval skill?
Copy the agentic-eval folder (the one containing SKILL.md) into ~/.claude/skills/ for Claude Code, .agents/skills/ for Codex or .cursor/skills/ for Cursor. The agent picks it up automatically when a task matches its description.
Is the Agentic Eval skill free?
Yes. The repository is open source under the MIT license.
Is Agentic Eval maintained?
The repository's most recent commit was on Oct 7, 2026. appsgit only lists skills from repositories with a commit in the last six months.