Managed Deep Agents
langchain-ai/langchain-skills
INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.
This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs…
$ npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pifferologo/cloud-agents-cli google-agents-cli-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pifferologo/cloud-agents-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/google-agents-cli-eval .claude/skills/google-agents-cli-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "google-agents-cli-eval" agent skill from https://github.com/pifferologo/cloud-agents-cli/tree/main/skills/google-agents-cli-eval into .claude/skills/google-agents-cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-agents-cli-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pifferologo/cloud-agents-cli/tree/main/skills/google-agents-cli-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pifferologo/cloud-agents-cli google-agents-cli-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pifferologo/cloud-agents-cli.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/google-agents-cli-eval .agents/skills/google-agents-cli-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "google-agents-cli-eval" agent skill from https://github.com/pifferologo/cloud-agents-cli/tree/main/skills/google-agents-cli-eval into .agents/skills/google-agents-cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-agents-cli-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pifferologo/cloud-agents-cli google-agents-cli-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pifferologo/cloud-agents-cli.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/google-agents-cli-eval .cursor/skills/google-agents-cli-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "google-agents-cli-eval" agent skill from https://github.com/pifferologo/cloud-agents-cli/tree/main/skills/google-agents-cli-eval into .cursor/skills/google-agents-cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-agents-cli-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pifferologo/cloud-agents-cli.git --path skills/google-agents-cli-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pifferologo/cloud-agents-cli google-agents-cli-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pifferologo/cloud-agents-cli.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/google-agents-cli-eval .gemini/skills/google-agents-cli-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "google-agents-cli-eval" agent skill from https://github.com/pifferologo/cloud-agents-cli/tree/main/skills/google-agents-cli-eval into .gemini/skills/google-agents-cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-agents-cli-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pifferologo/cloud-agents-cli google-agents-cli-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pifferologo/cloud-agents-cli.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/google-agents-cli-eval .github/skills/google-agents-cli-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "google-agents-cli-eval" agent skill from https://github.com/pifferologo/cloud-agents-cli/tree/main/skills/google-agents-cli-eval into .github/skills/google-agents-cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-agents-cli-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pifferologo/cloud-agents-cli google-agents-cli-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pifferologo/cloud-agents-cli.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/google-agents-cli-eval .opencode/skills/google-agents-cli-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "google-agents-cli-eval" agent skill from https://github.com/pifferologo/cloud-agents-cli/tree/main/skills/google-agents-cli-eval into .opencode/skills/google-agents-cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-agents-cli-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
google-agents-cli-evalThis skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs…
Google Agents CLI Eval is an agent skill from pifferologo/cloud-agents-cli. This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodology and the Quality Flywheel. Covers eval metrics, dataset schema, LLM-as-judge scoring, and common failure causes. Do NOT use for API code patterns (use google-agents-cli-adk-code), deployment (use google-agents-cli-deploy), or project scaffolding (use…
Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/builtin-tools-eval.md`, `references/dataset_schema.md` and `references/metrics-guide.md`).
It sits in Development, covering Project scaffolding, LLM evaluation and Deployment. The repository describes itself as: google cloud agent cli for Drive, Gmail, Calendar, Sheets, Docs, Chat, Admin, and more. Dynamically built from piffer labs. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5957f5a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.astral.shcloud.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Google Agents CLI Eval loads about 6.8k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 136 tokens; SKILL.md has 2,604 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from pifferologo/cloud-agents-cli at commit 5957f5a, republished under its Apache-2.0 licence (© pifferologo). 2,604 words, ~6,762 tokens.
.claude/skills/google-agents-cli-eval/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Requires:
agents-cli(uv tool install google-agents-cli) — install uv first if needed.
Scaffolded project? If you used
/google-agents-cli-scaffold, you already haveagents-cli eval run(chainsgenerate+grade),tests/eval/datasets/, andtests/eval/eval_config.yaml. Start with executingeval runand iterate from there.
| File | Contents |
|---|---|
references/dataset_schema.md | Canonical EvaluationDataset schema — all field types, JSON examples for single-turn / multi-turn / multi-agent, common mistakes |
references/metrics-guide.md | Complete metrics reference — all built-in metrics, match types, custom metrics, judge model config |
references/user-simulation.md | Dynamic conversation testing — eval dataset synthesize flags, what scenarios are, compatible metrics |
references/builtin-tools-eval.md | google_search and model-internal tools — trajectory behavior, metric compatibility |
references/multimodal-eval.md | Multimodal inputs — eval dataset schema, built-in metric limitations, custom evaluator pattern |
Improving agent quality is iterative. The 5 stages below describe the loop. Each stage has a Default path (you, the coding agent, do the work directly) and an Opt-in CLI command that delegates to the Agent Platform Eval Service for better quality and scale.
Default: Use or edit the scaffolded tests/eval/datasets/basic-dataset.json to define single-turn eval inputs. Start with 1–2 cases.
Opt-in: agents-cli eval dataset synthesize — runs e2e user simulation against your live agent to synthesize multi-turn eval datasets. Prefer when testing multi-turn conversations but lacking data. Output includes traces, so you can skip Stage 2 and go directly to eval grade.
agents-cli eval generate — executes the agent over the dataset and writes traces to artifacts/traces/. Run this when you wrote the dataset by hand in Stage 1 (default path). Skip this stage if you used eval dataset synthesize — that command already produced traces.
agents-cli eval grade — scores the traces and writes results_<ts>.{json,html} to artifacts/grade_results/. No opt-in alternative; this is the core. Always run, regardless of how Stages 1 and 2 produced the traces.
Shortcut:
agents-cli eval runchains Stages 2 + 3 in one command using the defaultartifacts/traces/directory between them. Use it for the common path; drop back to the two-step form when you need a custom traces location or want to grade an existing traces file.
Default: Open the latest artifacts/grade_results/results_<ts>.html (or .json) and identify failed metrics — see What to fix when scores fail below for the fix table.
Opt-in: agents-cli eval analyze — runs LLM-based failure clustering and root-cause analysis over the grade results. Prefer when you have 10+ failing cases and want categorized failure modes instead of case-by-case reading.
Default: Edit the agent — adjust prompts, tool descriptions, instructions, or eval dataset based on the failure analysis. See What to fix when scores fail below for the failure → fix mapping.
Opt-in: agents-cli eval optimize — runs ADK GEPA prompt optimization against a target metric. Suitable for prompt-only failures. The optimized prompt appears in the command output; capture it and apply it to the agent. For the full per-iteration trace, set print_detailed_results: true in your optimization config file.
Long-running and expensive. GEPA optimization makes many LLM calls and can take a long time. Do not run it unless the user explicitly asks for prompt optimization. When you do run it, iterate as far as possible with manual fixes first, then run a single final
eval optimize— never loop on this command.
Iterate stages 2 → 3 → 4 → 5 → 2 (or 1 → 3 → 4 → 5 → 1 if using synthesize). After each fix, run agents-cli eval compare <prev_results>.json <new_results>.json to confirm the target metric improved without regressing others. Expect 5–10+ iterations per case before it passes — this is normal. Only after a case passes should you expand coverage with more eval cases.
When doing 5+ iterations, maintain a task list of which cases are fixed, which are still failing, and what fixes you've tried. Prevents re-attempting the same fix.
Recognize these rationalizations and push back — they always cost more time than they save:
| Shortcut | Why it fails |
|---|---|
| "I'll tune the eval thresholds down to make it pass" | Lowering thresholds hides real failures. If the agent can't meet the bar, fix the agent — don't move the bar. |
| "This eval case is flaky, I'll skip it" | Flaky evals reveal non-determinism in your agent. Fix with temperature=0, rubric-based metrics, or more specific instructions — don't delete the signal. |
| "I just need to fix the eval dataset, not the agent" | If you're always adjusting expected outputs, your agent has a behavior problem. Fix the instructions or tool logic first. |
Pick built-in metrics by what you want to measure. Multi-turn metrics evaluate the full conversation; single-turn metrics evaluate one prompt-response pair (with intermediate tool calls). When no built-in fits, write a custom metric (see Evaluation Configuration Schema below).
| Goal | Recommended built-in metrics |
|---|---|
| Did the agent achieve the user's goal? (catch-all for multi-turn agents) | multi_turn_task_success |
| Was the agent's reasoning path logical and efficient? | multi_turn_trajectory_quality |
| Quality of tool / function calling across turns | multi_turn_tool_use_quality |
| Final response quality (no ground-truth reference needed) | final_response_quality |
| Factual grounding (catch hallucinated claims, e.g., RAG agents) | hallucination |
| Safety policy compliance | safety |
| Domain-specific check no built-in covers | Write a custom LLMMetric (LLM-judge) or CodeExecutionMetric (deterministic Python). See Evaluation Configuration Schema below. |
Run agents-cli eval metric list to see all available built-ins. For full metric definitions and rubric details, see the Agent Platform metric docs and references/metrics-guide.md.
After agents-cli eval grade completes, inspect the latest artifacts/grade_results/results_<timestamp>.json (or open the .html file) for per-case scores and judge rationales — that's the input to every fix decision below.
| Failure | What to change |
|---|---|
multi_turn_task_success low | The agent isn't completing the user's goal — fix orchestration, missing tool calls, premature termination, or wrong tool selection |
multi_turn_trajectory_quality low | The agent reaches the goal inefficiently or takes wrong steps — refine planning prompts, tighten instruction order, or remove redundant tool calls |
multi_turn_tool_use_quality low | Fix tool descriptions, parameter docstrings, or agent instructions for tool selection |
final_response_quality low | Read the auto-generated rubric verdicts; refine agent instructions to address the worst-scoring criterion (often clarity, completeness, or instruction-following) |
hallucination low | Tighten agent instructions to stay grounded in tool output; verify the tool actually returned the data the agent claimed |
safety low | Add safety guardrails to instructions; review the violating content category in the rubric verdict |
| Agent calls wrong tools | Fix tool descriptions, agent instructions, or tool_config |
| Agent calls extra tools | Add strict stop instructions, or switch to multi_turn_tool_use_quality |
After applying a fix, rerun agents-cli eval generate && agents-cli eval grade and use agents-cli eval compare <prev_results>.json <new_results>.json to confirm the fix improved the target metric without regressing others.
All agents-cli eval subcommands support --help for the authoritative flag list and defaults — run agents-cli eval <subcommand> --help (or agents-cli eval dataset <subcommand> --help) when in doubt. The examples below show the most common invocations; flags can change between releases.
eval generateRuns an agent over an evaluation dataset and writes traces to disk.
# Basic — uses tests/eval/datasets/, writes to artifacts/traces/
agents-cli eval generate
# Advanced — custom dataset and output dir
agents-cli eval generate --dataset tests/eval/datasets/custom.json -o ./custom_traces/eval gradeScores generated traces against built-in or custom metrics. Writes timestamped results_<YYYYMMDD_HHMMSS>.json (consumed by eval compare) and .html (open in a browser) into the output dir, and prints a summary table to the console.
# Basic — defaults: traces from artifacts/traces/, results to artifacts/grade_results/,
# metrics from tests/eval/eval_config.yaml's metrics_to_run
agents-cli eval grade
# Advanced 1 — grade traces from a non-default location (the canonical
# pairing for `eval generate --output custom_traces/`)
agents-cli eval grade --traces custom_traces/
# Advanced 2 — pick built-in metrics, custom output dir
agents-cli eval grade --metrics tool_use_quality,safety --output ./out/
# Advanced 3 — load metrics to run from a config file (YAML or JSON) on a specified trace file.
agents-cli eval grade --traces ./artifacts/traces/trace_1.json --config tests/eval/eval_config.yamlSee Evaluation Configuration Schema below for the config file format.
eval compareDiffs two results_*.json files produced by eval grade. Run it after a fix to confirm the target metric improved without regressing others.
agents-cli eval compare baseline.json candidate.jsoneval metric listLists the built-in metric names usable with eval grade --metrics.
agents-cli eval metric listeval analyzeRuns LLM-based failure clustering and root-cause analysis over a results_*.json produced by eval grade. Use when you have 10+ failing cases and want categorized failure modes instead of reading the HTML case-by-case. Supported --metric values: multi_turn_task_success, multi_turn_tool_use_quality.
# Basic — analyze a results file with default settings
agents-cli eval analyze --eval-result artifacts/grade_results/results_<ts>.json
# Advanced — restrict to a specific metric and cap loss clusters
agents-cli eval analyze \
--eval-result artifacts/grade_results/results_<ts>.json \
--metric multi_turn_tool_use_quality \
--top-k 5 \
--output artifacts/analysis_<ts>.jsoneval dataset synthesizeGenerates user scenarios server-side from your agent's tools and instructions, then plays each scenario against an LLM-backed user simulator. The output is a graded-ready trace file with full agent_data.turns populated — feed it directly to eval grade (skip eval generate).
# Basic — generate 3 default scenarios (up to 5 turns each) into artifacts/traces/
# (where eval grade reads from by default, so synthesize → grade works without flags)
agents-cli eval dataset synthesize
# Advanced — guide scenario generation with optional instruction and environment context
agents-cli eval dataset synthesize \
-n 5 \
--instruction "Customer asking about refunds" \
--environment-context "E-commerce support" \
--max-turns 8 \
-o tests/eval/datasets/refund_scenarios.jsonFor scenario semantics, the full eval dataset synthesize flag table, and which simulator internals are not user-configurable, see references/user-simulation.md.
eval optimizeRuns ADK GEPA prompt optimization against a target metric. Suitable after eval grade identifies prompt-only failures (wording, not tool/orchestration logic). --dataset and --target-metric override values in --config when both are passed. Long-running and expensive — see Stage 5 of the Quality Flywheel for usage guidance.
# Basic — optimize against a single metric on a dataset
agents-cli eval optimize --dataset tests/eval/datasets/basic-dataset.json --target-metric final_response_quality
# Advanced — drive multi-metric / multi-dataset optimization from a config file
agents-cli eval optimize --config tests/eval/optimization_config.jsoneval submit / eval results (cloud-side)The managed, asynchronous counterpart to the local path, for large or CI-driven runs: eval submit hands the dataset and metrics to the Agent Platform Eval Service, and eval results polls and downloads the scores. Pass --resource-name <agent> to also run inference server-side (managed generate + grade); omit it to grade an existing trace (managed grade).
# Grade an existing trace server-side; returns a run resource name to poll
agents-cli eval submit --dataset tests/eval/datasets/basic-dataset.json --dest gs://my-bucket
# Add --resource-name projects/<p>/locations/<l>/reasoningEngines/<id> to run inference too
agents-cli eval results --run-id <run-resource-name>An EvaluationDataset is a JSON file with an eval_cases array. Cases come in two shapes depending on how they're used:
eval generate) — a user prompt or a partial conversation ending in a user prompt. The agent runs and produces traces.eval grade) — a complete trace including the agent's responses and tool calls. Normally produced by eval generate or eval dataset synthesize; you don't write these by hand.See references/dataset_schema.md for the full canonical schema, all field types, and common mistakes.
Two shapes are supported.
(a) Simple single-turn prompt — what the scaffolded tests/eval/datasets/basic-dataset.json uses. The agent runs from scratch.
{
"eval_cases": [
{
"eval_case_id": "greeting",
"prompt": {
"role": "user",
"parts": [{"text": "Hello, what can you help me with?"}]
}
},
{
"eval_case_id": "weather_query",
"prompt": {
"role": "user",
"parts": [{"text": "What's the weather like in San Francisco?"}]
}
}
]
}(b) Multi-turn continuation via agent_data — partial conversation, last turn ends with a user message. Use to continue an existing conversation; the agent's next response is what gets evaluated.
{
"eval_cases": [
{
"eval_case_id": "booking_followup",
"agent_data": {
"agents": {
"flight_booking_agent": {
"agent_id": "flight_booking_agent",
"instruction": "You are a helpful flight booking assistant."
}
},
"turns": [
{
"turn_index": 0,
"events": [
{"author": "user", "content": {"parts": [{"text": "I want to book a flight to Paris."}]}},
{"author": "flight_booking_agent", "content": {"parts": [{"text": "I found a flight for $800. Do you want to book it?"}]}}
]
},
{
"turn_index": 1,
"events": [
{"author": "user", "content": {"parts": [{"text": "Yes, please book it."}]}}
]
}
]
}
}
]
}Complete trace — agent responses, tool calls, and tool responses all present. Normally produced by eval generate or eval dataset synthesize; shown here so you can recognize the shape when debugging.
{
"eval_cases": [
{
"eval_case_id": "weather_query",
"agent_data": {
"agents": {
"weather_agent": {
"agent_id": "weather_agent",
"instruction": "You are a helpful weather assistant."
}
},
"turns": [
{
"turn_index": 0,
"events": [
{"author": "user", "content": {"parts": [{"text": "What's the weather in San Francisco?"}]}},
{"author": "weather_agent", "content": {"parts": [{"function_call": {"name": "get_weather", "args": {"city": "San Francisco"}}}]}},
{"author": "weather_agent", "content": {"parts": [{"function_response": {"name": "get_weather", "response": {"temp_f": 62, "conditions": "foggy"}}}]}},
{"author": "weather_agent", "content": {"parts": [{"text": "It's currently 62°F and foggy in San Francisco."}]}}
]
}
]
}
}
]
}Key conventions: authors are "user", agent IDs from the agents map, or "tool"; tool calls use function_call parts and tool results use function_response parts. See references/dataset_schema.md for multi-agent examples and the full type reference.
agents-cli eval grade --config <path> accepts a single configuration file in either YAML (.yaml / .yml) or JSON (.json). The file declares two parts:
metrics_to_run — the selection list of metric names to execute on this run. Names resolve to built-in metrics first, then to entries in custom_metrics.custom_metrics — a definition pool of custom metrics available to this project. Defining a metric here does not run it; it must also appear in metrics_to_run (or be passed via --metrics name1,name2 on the CLI, which is equivalent to overriding metrics_to_run for that invocation).Minimal example (YAML preferred — human-readable, no JSON escaping for prompts and Python):
metrics_to_run:
- multi_turn_task_success # built-in
- example_llm_metric # selected from custom_metrics pool below
- agent_turn_count # selected from custom_metrics pool below
custom_metrics:
- name: example_llm_metric
prompt_template: |
Rate the agent's response 1-5 for helpfulness and accuracy.
Prompt: {prompt}
Final response: {response}
Full trace (for tool-call and reasoning context): {agent_data}
Return JSON: {"score": <1|2|3|4|5>, "explanation": "<reason>"}
- name: agent_turn_count
custom_function: |
def evaluate(instance):
turns = (instance.get("agent_data") or {}).get("turns", [])
return {'score': len(turns)}JSON is also accepted (same field names, with prompt_template and custom_function as escaped strings) — but always prefer YAML for human-readable configs.
Each entry in custom_metrics is dispatched by field: presence of custom_function makes it a CodeExecutionMetric (deterministic Python); otherwise it's an LLMMetric (LLM-as-judge with prompt_template). Run agents-cli eval metric list to see available built-ins. For full custom-metric field reference (judge model options, sampling counts), see references/metrics-guide.md.
Agent trace field model. For datasets produced by agents-cli eval generate (or eval dataset synthesize), each eval case exposes three standard fields to a metric:
{prompt} — the user message (or first user turn).{response} — the agent's final text response, extracted from the last text-bearing event. In custom_function callbacks this is instance['response'] with shape {"role": "model", "parts": [{"text": "..."}]}.{agent_data} — the full structured turns/events trace, useful when the judge needs to reason about tool calls or intermediate reasoning.{reference} and {context} resolve only when the eval case has reference / context fields populated (e.g., golden-answer datasets); they are not populated by eval generate / eval dataset synthesize.
Code-based metrics default to local in-process execution (no GCP project or region required, but the evaluate(instance) function runs with the CLI's privileges). Set execution: "remote" on the metric to run it server-side in Vertex AI's CodeExecutionMetric sandbox instead — that path requires a configured GCP project + region.
Evaluating agent tool usage using strict sequence matching is fragile because agents may call helper tools (like searches or geocoding) in different orders or perform extra proactive steps.
Instead, use multi_turn_tool_use_quality / multi_turn_trajectory_quality. These metrics automatically generate content-based and intent-based adaptive rubrics, assessing technical correctness and technical sequence logic semantically using an LLM judge rather than forcing a rigid match.
The App object's name parameter MUST match the directory containing your agent:
# CORRECT - matches the "app" directory
app = App(root_agent=root_agent, name="app")
# WRONG - causes "Session not found" errors
app = App(root_agent=root_agent, name="flight_booking_assistant")Each eval case runs in its own fresh in-memory session (eval generate creates a new InMemorySessionService and session id per case). Multi-turn within a case works via agent_data.turns, but behavior that depends on a separate prior session — e.g. Memory Bank recall across sessions — can't be exercised by eval. Validate cross-session continuity with pytest integration tests instead.
eval grade, eval submit, and eval dataset synthesize default to the global endpoint — they don't inherit the manifest region (the eval services support only a subset of regions). eval analyze is global-only; eval generate runs locally and follows the project region. So you normally don't configure anything for eval.
Override per run with --region <REGION> (e.g. data residency); the service rejects an unsupported one:
400 FAILED_PRECONDITION: Unsupported region for Vertex Evaluation Service: <region>No eval region fits your data-residency rules? Fall back to local custom metrics — a custom_metrics entry with a custom_function (execution: local, the default) grades in-process with no GCP region required. You lose the managed built-in metrics, but your custom_function can still call an LLM judge in a compliant region itself — so LLM-as-judge grading stays available anywhere.
before_agent_callback Pattern (State Initialization)Always use a callback to initialize session state variables used in your instruction template. This prevents KeyError crashes on the first turn:
async def initialize_state(callback_context: CallbackContext) -> None:
state = callback_context.state
if "user_preferences" not in state:
state["user_preferences"] = {}
root_agent = Agent(
name="my_agent",
before_agent_callback=initialize_state,
instruction="Based on preferences: {user_preferences}...",
)Models with "thinking" enabled may skip tool calls. Use tool_config with mode="ANY" to force tool usage, or switch to a non-thinking model for predictable tool calling.
| Symptom | Cause | Fix |
|---|---|---|
| Agent mentions data not in tool output | Hallucination | Tighten agent instructions; add hallucination metric |
| "Session not found" error | App name mismatch | Ensure App name matches directory name |
| Score fluctuates between runs | Non-deterministic model | Set temperature=0 or use rubric-based eval with multiple samples |
tool_use_quality score low | Wrong tool selected or invalid arguments passed | Refine tool descriptions, instructions, or parameter documentation |
| LLM judge ignores image/audio in eval | get_text_from_content() skips non-text parts | Use custom metric with vision-capable judge (see references/multimodal-eval.md) |
User says: "tool_use_quality is low, what's wrong?"
artifacts/grade_results/results_<timestamp>.html (or read the .json) and find the rubric verdicts the adaptive metric generated for the failing case.artifacts/traces/.agents-cli eval generate && agents-cli eval grade.agents-cli eval compare <prev>.json <new>.json to confirm the score improved.Don't assert that eval passes — show the evidence. Concrete output prevents false confidence and catches issues early.
agents-cli eval generate and agents-cli eval grade one final time.agents-cli eval grade output with all cases above threshold. This is the gate — no exceptions./google-agents-cli-workflow — Development workflow and the spec-driven build-evaluate-deploy lifecycle/google-agents-cli-adk-code — ADK Python API quick reference for writing agent code/google-agents-cli-scaffold — Project creation and enhancement with agents-cli scaffold create / scaffold enhance/google-agents-cli-deploy — Deployment targets, CI/CD pipelines, and production workflows/google-agents-cli-observability — Cloud Trace, logging, and monitoring for debugging agent behavior© pifferologo, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/google-agents-cli-eval of pifferologo/cloud-agents-cli.
Open the folder on GitHubat commit 5957f5a
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in pifferologo/cloud-agents-cli, which our catalogue first saw on October 7, 2026.
Google Agents CLI Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Google Agents CLI Eval this skillpifferologo/cloud-agents-cli | 129 | 1 repos | ~6.8k | Automated safety check: Pass | Apache-2.0 | |
| Managed Deep Agentslangchain-ai/langchain-skills | 1.3k | — | ~8.7k | Automated safety check: Notes | MIT | |
| Adk Agent Builderjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~960 | Automated safety check: Pass | MIT | |
| Create Vechain Dappvechain/x-app-template | 450 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Deployjustrach/merjs | 357 | — | ~171 | Automated safety check: Pass | MIT | |
| Spec Optimizeleo-kuang-ai/spec-first | 107 | — | ~13k | Automated safety check: Pass | MIT |
langchain-ai/langchain-skills
INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.
jeremylongshore/tons-of-skills-marketplace
Scaffold production-ready AI agents on Google's Agent Development Kit (ADK): ReAct-style single agents, multi-agent orchestration (Sequential/Parallel/Loop), tool wiring, evaluation, and optional…
vechain/x-app-template
Scaffold a VeChain dApp with Next.js, VeChain Kit, Chakra UI v3, and GitHub Pages deployment.
justrach/merjs
Full production build — codegen, compile, prerender, and prepare for deployment.
leo-kuang-ai/spec-first
Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.
matlab/simulink-agentic-toolkit
Configure Simulink models for Embedded Coder (ERT), Simulink Coder (GRT rapid-prototyping), or AUTOSAR code generation.
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "write agent code", "build an agent with ADK", "add a tool", "create a callback", "define an agent", "use state management", or needs ADK (Agent…
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "set up tracing", "monitor my ADK agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring…
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "create an agent project", "start a new ADK project", "build me a new agent", "add CI/CD to my project", "add deployment", "enhance my project", or…
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "deploy an agent", "deploy my ADK agent", "set up CI/CD", "configure secrets", "troubleshoot a deployment", or needs guidance on Agent Runtime, Cloud…
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "publish an agent", "publish my ADK agent", "register an agent with Gemini Enterprise", "publish to Gemini Enterprise", or needs guidance on the…
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "develop an agent", "build an agent using ADK", "run the agent locally", "debug agent code", "test an agent", "deploy an agent", "publish an agent"…
Categories
This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs…. Google Agents CLI Eval is an agent skill from pifferologo/cloud-agents-cli. This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodology and the Quality Flywheel.
Google Agents CLI Eval fits situations like: wants to run an evaluation; evaluate my ADK agent; write an eval dataset; analyze eval failures.
Run `npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a claude-code`. Or copy the skill folder (skills/google-agents-cli-eval in pifferologo/cloud-agents-cli) into .claude/skills/google-agents-cli-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a codex`. Or copy the skill folder (skills/google-agents-cli-eval in pifferologo/cloud-agents-cli) into .agents/skills/google-agents-cli-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pifferologo/cloud-agents-cli --skill google-agents-cli-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/google-agents-cli-eval, .gemini/skills/google-agents-cli-eval, .github/skills/google-agents-cli-eval and .opencode/skills/google-agents-cli-eval in your project.
Going by SKILL.md and its folder, Google Agents CLI Eval needs the command-line tools its instructions call (uv).
SKILL.md names 2 domains. As links in the text: docs.astral.sh and cloud.google.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Google Agents CLI Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Google Agents CLI Eval: Managed Deep Agents (langchain-ai/langchain-skills, 1.3k stars), Adk Agent Builder (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Create Vechain Dapp (vechain/x-app-template, 450 stars) and Deploy (justrach/merjs, 357 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pifferologo (a GitHub user) maintains it in pifferologo/cloud-agents-cli, which has 129 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 3, 2026.
Source: pifferologo/cloud-agents-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.