Google Agents CLI Eval
Google Agents CLI Eval is an agent skill (a SKILL.md file) from google/agents-cli. It should be used when the user wants to "run an evaluation", "evaluate my agent", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures"… It works with Claude Code and Codex and has 6,058 GitHub stars across a repository of 6 listed skills.
github.com/google/agents-cli/skills/google-agents-cli-eval (opens in a new tab)
- Official
- Multi-skill repo
- Needs API key
- DevOps & cloud
- Actively maintained
Add this skill
Claude
Claude Code loads skills from ~/.claude/skills/ (all projects) or .claude/skills/ (one project):
git clone --depth 1 https://github.com/google/agents-cli.git
cp -r agents-cli/skills/google-agents-cli-eval ~/.claude/skills/google-agents-cli-eval # personal, or .claude/skills in a projectIn the Claude apps, zip the google-agents-cli-eval folder and upload it under Customize > Skills > + > Upload a skill (code execution must be on).
ChatGPT / Codex
Codex reads skills from .agents/skills/ in a repo or ~/.agents/skills/ for every project:
git clone --depth 1 https://github.com/google/agents-cli.git
cp -r agents-cli/skills/google-agents-cli-eval .agents/skills/google-agents-cli-eval # repo; ~/.agents/skills for all projectsStandalone skills also load in the ChatGPT desktop app.
Cursor
Cursor loads skills from .cursor/skills/ (or ~/.cursor/skills/) and also reads .claude/skills/:
git clone --depth 1 https://github.com/google/agents-cli.git
cp -r agents-cli/skills/google-agents-cli-eval .cursor/skills/google-agents-cli-eval # project; ~/.cursor/skills for all projectsSource (checked Oct 7, 2026): code.claude.com/docs/en/skills (opens in a new tab), support.claude.com/en/articles/12512180-using-skills-in-claude (opens in a new tab), learn.chatgpt.com/docs/build-skills (opens in a new tab), cursor.com/docs/context/skills (opens in a new tab)
What this skill does
This skill should be used when the user wants to "run an evaluation", "evaluate my agent", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodology and the Quality Flywheel. Covers eval metrics, dataset schema, LLM-as-judge scoring, and common failure causes. Applies to any agents-cli project, whatever framework the agent is written in.
When it triggers
- used when the user wants to "run an evaluation", "evaluate my agent", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodol
- use for agent API code patterns (ADK: use google-agents-cli-adk-code), deployment (use google-agents-cli-deploy), or project scaffolding (use google-agents-cli-scaffold).
- "run an evaluation"
- "evaluate my agent"
- "evaluate my ADK agent"
- "write an eval dataset"
- "analyze eval failures"
- "compare eval results"
More skills in google/agents-cli
6 skills are listed from this repository.
Similar skills
More devops & cloud skills
Caveman Setup
JuliusBrussee/caveman
Wire a repository through the Caveman Cloud gateway so every LLM request is measured, with no behavior change.
DevOps & cloudGoObservability And Instrumentation
addyosmani/agent-skills
Instruments code so production behavior is visible and diagnosable.
DevOps & cloudJavaScriptShipping And Launch
addyosmani/agent-skills
Prepares production launches. Use when preparing to deploy to production, or when asking what needs to be in place before shipping.
DevOps & cloudJavaScriptRuview Applications
ruvnet/RuView
Run RuView sensing applications — presence/occupancy, breathing & heart rate, activity & fall detection, 17-keypoint pose estimation (WiFlow), sleep monitoring & apnea screening, environment…
DevOps & cloudRustRuview Mmwave
ruvnet/RuView
Set up and run RuView mmWave / FMCW radar sensing — ESP32-C6 + Seeed MR60BHA2 (60 GHz, heart rate / breathing rate / presence) and HLK-LD2410 (24 GHz, presence + distance), plus mmWave↔WiFi-CSI…
DevOps & cloudRustAgentcore
vercel-labs/agent-browser
Run agent-browser on AWS Bedrock AgentCore cloud browsers.
OfficialDevOps & cloudRust
What is the Google Agents CLI Eval skill?
Google Agents CLI Eval is an agent skill (a SKILL.md file) from google/agents-cli. It should be used when the user wants to "run an evaluation", "evaluate my agent", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures"… It works with Claude Code and Codex and has 6,058 GitHub stars across a repository of 6 listed skills. Its SKILL.md lives at github.com/google/agents-cli/skills/google-agents-cli-eval.
How do I install the Google Agents CLI Eval skill?
Copy the google-agents-cli-eval folder (the one containing SKILL.md) into ~/.claude/skills/ for Claude Code, .agents/skills/ for Codex or .cursor/skills/ for Cursor. The agent picks it up automatically when a task matches its description.
Is the Google Agents CLI Eval skill free?
Yes. The repository is open source under the Apache-2.0 license. The skill mentions an API key or token for an external service, which may need its own account.
Is Google Agents CLI Eval maintained?
The repository's most recent commit was on Oct 6, 2026. Its latest release is v1.9.0. appsgit only lists skills from repositories with a commit in the last six months.