Official agent skill

Agent Platform Eval Flywheel

by google in google/skills

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Agent Platform Eval Flywheel

skills CLI
$ npx skills add google/skills --skill agent-platform-eval-flywheel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills agent-platform-eval-flywheel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/agent-platform-eval-flywheel .claude/skills/agent-platform-eval-flywheel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-platform-eval-flywheel
GitHub stars
21k
Token cost
~6.7k tokens
SKILL.md length
2,136 words
Files
13 (incl. scripts, references)
Skills in repo
145
Repo updated
First seen
Licence
Apache-2.0

At a glance

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology.

  • Works in 5 steps: Prepare Data → Run Inference → Grade (always run) → …
  • Generating synthetic user scenarios
  • SKILL.md covers When to use this skill, Safety & Confirmation Tiers…, Setup and The Quality Flywheel, plus 4 more sections
  • Runs Python scripts from its folder; calls python3 and pip

What it does

Agent Platform Eval Flywheel is an agent skill from google/skills, published by the product's own GitHub organization. Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results before and after a fix, or when guidance is needed on Agent Platform eval methodology — including dataset schema, LLM-as-judge scoring, and common failure causes. For fine-tuning, use agent-platform-tuning. For general…

Its SKILL.md is about 6.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts and reference files (for example `references/dataset_schema.md`, `references/deployment.md` and `references/failure_patterns.md`).

It sits in AI & LLM Engineering, covering LLM evaluation, Fine-tuning and Deployment. It works with Google Cloud. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Generating synthetic user scenarios
  • Evaluating an agent
  • Building an eval dataset
  • Writing evaluation metrics

Example prompts

  • “Use the agent-platform-eval-flywheel skill to measure and improves the quality of AI models and agents on Google Cloud using the Eval Quality…”
  • “/agent-platform-eval-flywheel”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Prepare Data
  2. Run Inference
  3. Grade (always run)
  4. Analyze Failures
  5. Optimize & Iterate

What it can do on your machine

Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.cloud.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Platform Eval Flywheel loads about 6.7k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 149 tokens; SKILL.md has 2,136 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~149
When it runs · the whole SKILL.md, loaded when a task matches
~6.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 2,136 words, ~6,739 tokens.

Download SKILL.mdSave it as .claude/skills/agent-platform-eval-flywheel/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
agent-platform-eval-flywheel
description
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results before and after a fix, or when guidance is needed on Agent Platform eval methodology — including dataset schema, LLM-as-judge scoring, and common failure causes. For fine-tuning, use agent-platform-tuning. For general production deployment, use agent-platform-deploy.
metadata.version
1.0.2
metadata.category
AiAndMachineLearning

Agent Platform Eval Flywheel Skill

Help users evaluate and iteratively improve GenAI models and agents using the Agent Platform GenAI Evaluation SDK (google.genai / agentplatform).

When to use this skill

  • Evaluating GenAI agents or models with the Agent Platform GenAI Evaluation SDK (client.evals.evaluate()).
  • Creating evaluation datasets from session traces, pandas DataFrames, or synthetic generation.
  • Selecting, configuring, or writing custom evaluation metrics.
  • Analyzing rubric verdicts, loss patterns, and clustering failures.
  • Suggesting concrete code/prompt improvements based on eval results.
  • Evaluating a model served on an Agent Platform endpoint (BYOM) or a Model-as-a-Service (MaaS) model by ID — including deploying the model first if needed. For this case, follow references/deployment.md and use the endpoint_evaluation.py / maas_evaluation.py scripts.

Safety & Confirmation Tiers (CRITICAL)

Before executing any commands or scripts on behalf of the user, you MUST adhere to the following safety tiers based on the action requested:

  1. Tier R: Read-only (inspect_results.py, compare_results.py, validate_dataset.py, parse_adk_traces.py, render_html_report.py)
    • Rule: No confirmation needed. You may execute these helper scripts immediately to inspect data, validate schemas, parse traces, or compare evaluation results.
  2. Tier M: Read-only with Compute Costs (client.evals.run_inference, client.evals.evaluate, client.evals.generate_conversation_scenarios, client.evals.generate_loss_clusters)
    • Rule: These operations invoke LLMs or remote evaluation services that consume compute resources and incur costs. This requires interactive confirmation with 'Yes'/'No' options.
    • Confirmation for EVERY evaluation run: Every evaluation, re-evaluation, metric update, parameter change, or synthetic scenario generation requires its own dry-run preview and interactive confirmation. Never execute a second evaluation, comparison pass, or modified evaluation without presenting a new confirmation preview and obtaining user approval.
    • Same-turn restriction: Do not run the evaluation in the same turn as presenting the confirmation prompt. End your turn after asking and wait for the user's reply; only execute after explicit 'Yes' / approval. Printing a preview and then calling the tool before the user can answer does not count as obtaining confirmation.
    • No Pre-Execution of Remote Evaluation: NEVER execute client.evals.evaluate(), client.evals.run_inference(), client.evals.generate_conversation_scenarios(), or run any script invoking these remote operations before user confirmation. In the initial turn, you may prepare local data structures and compose the script, but you MUST present the dry-run preview card and obtain explicit user confirmation before running any remote evaluation or scenario generation call.
    • Immediate Execution Upon Approval: Once the user explicitly approves (e.g., 'Yes', 'Approved', 'Go ahead', 'Proceed'), proceed directly to executing the previewed evaluation script via run_command and report the results. Do not conclude the turn without executing the approved action.

Setup

The scripts need vertexai (from google-cloud-aiplatform[evaluation]), google-genai, pandas, and requests. Do not create a virtual environment — it starts empty and hides packages the environment already provides, forcing a redundant install. Probe, and install only what is missing:

bash
python3 -c "import vertexai, google.genai, pandas, requests" \
  || pip install 'google-cloud-aiplatform[evaluation]>=1.163.0' 'google-genai>=1.0.0'

The version specifiers must stay quoted: unquoted, bash reads >=1.154.0 as a redirect and silently writes an empty file instead of constraining the install.

Need GOOGLE_CLOUD_PROJECT and GOOGLE_CLOUD_LOCATION.

  • Preserve User Project and Location: Always prioritize the user's explicitly provided project and location (e.g. project='<PROJECT_NUMBER>', location='us-central1'). Never change or override the user's requested location to 'global' unless the user explicitly requested 'global'.
  • Missing Parameters: If the user's request omits the project or location, you MUST pause in your response and ask the user for the missing location/project before preparing or running the evaluation.
Correct SDK entrypoints
python
import agentplatform
client = agentplatform.Client(project=PROJECT, location=LOCATION)

client.evals.run_inference(model=..., src=...)
client.evals.evaluate(dataset=..., metrics=...)
client.evals.generate_conversation_scenarios(...)

Two imports that look plausible and are not:

  • from agentplatform.types import evals -- ModuleNotFoundError. types is a module, not a package; use from agentplatform import types.
  • from vertexai.evaluation import PointwiseMetric, EvalTask -- the superseded SDK. Its classes take different arguments (PointwiseMetric has no system_instruction), so code written against it fails with TypeError rather than an import error. Use agentplatform throughout.

The Quality Flywheel

Five stages, run in order on the first pass, then loop 2 → 5 until quality targets are met.

Shortcuts that waste time
ShortcutWhy it fails
"I'll tune the metric threshold downHides real failures. Fix the agent,
: so it passes." : not the bar. :
"This case is flaky, I'll skip it."Flakiness reveals non-determinism in
: : the agent. Fix with temperature=0 :
: : or stricter instructions. :
"I just need to fix the evalIf expected outputs keep moving, the
: dataset, not the agent." : agent has a behavior problem. :
"I can tell from the trace it worksSelf-grading doesn't generalize.
: — skip Stage 3." : Always run evaluate() and read :
: : scores. :
"One iteration is enough."Expect 5–10+ iterations. Stopping
: : early leaves regressions on other :
: : metrics undetected. :
1. Prepare Data

Produce an EvaluationDataset. There are three input shapes, pick the one that matches the data the user already has:

  • EvalCase list (single-turn or multi-turn):

    python
    from agentplatform import types
    from google.genai import types as genai_types
    
    # prompt/reference/response values are Content, not str. UserContent and
    # ModelContent wrap a plain string and set the right role.
    dataset = types.EvaluationDataset(eval_cases=[
        types.EvalCase(
            prompt=genai_types.UserContent("What is 2+2?"),
            responses=[types.ResponseCandidate(
                response=genai_types.ModelContent("4"))],
            reference=types.ResponseCandidate(
                response=genai_types.ModelContent("4")),
        ),
        # For multi-turn agent traces, set agent_data instead of prompt/responses.
    ])

    Multi-turn agent traces wrap each conversation in AgentData → ConversationTurn → AgentEvent. See references/dataset_schema.md for the full type hierarchy.

  • Pandas DataFrame (tabular sources — CSV, BigQuery, Sheets):

    python
    import pandas as pd
    from agentplatform import types
    
    df = pd.DataFrame({
        "prompt":    ["What is 2+2?", "Capital of France?"],
        "response":  ["4",            "Paris"],
        "reference": ["4",            "Paris"],
    })
    dataset = types.EvaluationDataset(eval_dataset_df=df)

    Column names must match the fields the chosen metrics expect (see references/dataset_schema.md for the per-metric requirements table).

  • Cold start (no data at all): synthesize scenarios server-side with client.evals.generate_conversation_scenarios(agent=..., config=...) -- the parameter is agent or agent_info, not agents, and config is required. The config class is types.evals.UserScenarioGenerationConfig, not types.UserScenarioGenerationConfig. Set its user_scenario_count (1-100): it defaults to None, the client accepts that, and the server rejects the call with 400 INVALID_ARGUMENT. count is a separate field and does not substitute for it. Stage 2 plays the scenarios out.

    • CRITICAL - Underspecified Requests: When asked to synthesize scenarios, if the request omits required parameters (such as location, environment_data, simulation_instruction, or model_name), do NOT assume defaults or guess values. You MUST pause in your first turn and explicitly ask the user for the missing information (e.g., "Please provide the missing simulation instructions, environment data, model name, and location"). Only proceed with the dry-run preview after the user provides them.
    • Friction & Parameter Changes: When asked to generate synthetic user scenarios, if the user modifies requested parameters (such as scenario count, model, or instructions) or pushes back, you MUST present a revised dry-run confirmation card with the updated parameters and wait for explicit user approval before executing generation code via run_command. Do NOT generate scenarios directly in plain text.
  • Managed Agents (Gemini Agents API): evaluate agents created with the Managed Agents API. Use generate_conversation_scenarios to create test scenarios from the agent's configuration, run_inference to execute the agent, and evaluate to score the traces. These functions now accept managed agents and interaction ids as input. You can also evaluate existing interactions recorded via the Interactions API using InteractionsDataSource. See references/sdk_patterns.md Pattern 8 for the full code pattern.

For ADK session dumps, use scripts/parse_adk_traces.py instead of writing the conversion by hand.

2. Run Inference

Populate responses/traces on the dataset. Skip this stage if traces are already complete (e.g., production logs or replay).

python
# Agent eval — pass a callable wrapping the user's ADK Agent/App.
client.evals.run_inference(model=agent_callable, src=dataset)

# Model eval — pass a model ID directly.
client.evals.run_inference(model="gemini-2.5-flash", src=dataset)

# Synthesized scenarios — let the simulator drive.
client.evals.run_inference(
    model=agent_callable,
    src=dataset,
    user_simulator_config=UserSimulatorConfig(max_turn=10),
)

# DataFrame also works as src= — no EvalCase wrapping needed.
client.evals.run_inference(model="gemini-2.5-flash", src=df)

# Managed Agent — pass an agent resource name.
AGENT_RESOURCE = f"projects/{PROJECT_ID}/locations/global/agents/{AGENT_ID}"
client.evals.run_inference(
    agent=AGENT_RESOURCE,
    src=scenarios,
    config={"user_simulator_config": {"max_turn": 3}},
)
Show full SKILL.md (988 more words)Show less
3. Grade (always run)
python
result = client.evals.evaluate(dataset=dataset, metrics=[...])
result.show()  # Interactive HTML report with scores, rubrics, and traces.

Pick metrics by what you want to measure. Full catalog in references/metric_registry.md.

If the user names a metric, use it directly. Every identifier in the tables below (general_quality, text_quality, instruction_following, hallucination, grounding, safety, multi_turn_*, final_response_*, tool_use_quality) is a types.RubricMetric.<UPPERCASE_NAME> accessor — pass it straight into metrics=[types.RubricMetric.GENERAL_QUALITY, ...]. Do not scaffold a custom LLMMetric for a name that appears here, and do not reach for vertexai.evaluation.EvalTask / PointwiseMetric / MetricPromptTemplateExamples — that SDK is superseded (see Setup).

Agent metrics (multi-turn, adaptive rubrics) — start here for agent eval.

GoalMetric
Did the agent achieve the user's goal?multi_turn_task_success
Was the reasoning path logical and efficient?multi_turn_trajectory_quality
Tool/function calling quality across turnsmulti_turn_tool_use_quality
Overall conversational qualitymulti_turn_general_quality
Final response quality (no reference needed)final_response_quality
Final response vs. a golden referencefinal_response_match
Single-turn tool usetool_use_quality

General quality metrics (single-turn, adaptive rubrics) — for model eval.

GoalMetric
Overall response quality (recommended starting point)general_quality
Linguistic quality (fluency, coherence, grammar)text_quality
Adherence to specific constraints / instructionsinstruction_following

Static rubric metrics (fixed criteria) — apply alongside the above.

GoalMetric
Catch hallucinated claims (RAG, factual answers)hallucination
Factuality / consistency against provided contextgrounding
Safety policy compliancesafety

Domain-specific check no built-in covers: write a custom metric.

  • Predefined: types.RubricMetric.<NAME> — server-side AutoRater, no judge model needed.
  • Custom LLM-as-a-judge: types.LLMMetric with prompt_template or types.MetricPromptBuilder for structured rubrics. Always set judge_model; it defaults to None and every case then fails with 400 INVALID_ARGUMENT: Error parsing JSON.
    • Judge Model Selection: If the user specifies a judge model (e.g. gemini-2.5-pro), use it. If the user omits the judge model or states they do not have information / preference for one, default to gemini-2.5-flash as the judge model in the dry-run preview card and ask for confirmation to run the evaluation. Do NOT halt or refuse to evaluate when the user does not specify a judge model.
  • Custom code: types.CodeExecutionMetric with a custom_function string containing def evaluate(instance: dict) for remote sandboxed execution; or types.Metric with custom_function=<callable> for local execution.

Always persist the result so Stage 4 and 5 can read it. Save both JSON (machine-readable, diffable) and HTML (human-readable, linkable):

python
import datetime
from pathlib import Path

from agentplatform._genai import _evals_visualization

out_dir = Path("artifacts/grade_results")
out_dir.mkdir(parents=True, exist_ok=True)
ts = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")

# fallback=str, or a DataFrame-backed dataset raises PydanticSerializationError.
result_json = result.model_dump_json(fallback=str)
(out_dir / f"results_{ts}.json").write_text(result_json)

html = _evals_visualization.get_evaluation_html(result_json)
(out_dir / f"results_{ts}.html").write_text(str(html))

Or after the fact: scripts/render_html_report.py --type evaluation or scripts/inspect_results.py --save-html.

4. Analyze Failures

Read summary_metrics and eval_case_results — never fabricate scores. Use scripts/inspect_results.py --failing-only to filter to failures.

For each failed metric, see references/failure_patterns.md for deeper diagnoses. The compact mapping:

Failing metricWhat to change
multi_turn_task_success lowThe agent isn't completing the goal —
: : fix orchestration, missing tool calls, :
: : premature termination, wrong tool :
: : selection. :
multi_turn_trajectory_quality lowThe agent reaches the goal
: : inefficiently — refine planning :
: : prompts, remove redundant tool calls. :
multi_turn_tool_use_quality lowFix tool descriptions, parameter
: : docstrings, or agent instructions for :
: : tool selection. :
final_response_quality lowRead auto-generated rubric verdicts;
: : refine instructions to address the :
: : worst-scoring criterion. :
final_response_match lowThe agent's final answer doesn't match
: : the golden reference — adjust response :
: : format or update the reference. :
hallucination lowTighten instructions to stay grounded
: : in tool output; verify the tool :
: : actually returned the claimed data. :
grounding lowThe response contradicts the provided
: : context — add explicit "cite only from :
: : context" instructions. :
safety lowAdd safety guardrails; review the
: : violating content category in the :
: : rubric verdict. :
general_quality / text_qualityAdjust system instruction wording; the
: low : model's default phrasing is too :
: : generic for the task. :
instruction_following lowThe agent is ignoring constraints —
: : restate them in the system instruction :
: : or use stricter wording. :
Agent calls wrong toolsFix tool descriptions, agent
: : instructions, or tool_config. :
Agent calls extra toolsAdd explicit stop instructions, or
: : switch to :
: : multi_turn_tool_use_quality to :
: : surface the extra calls in the rubric. :

For 10+ failures on the same metric, use the Error Analysis service to cluster failures into themes (L1/L2 taxonomy categories) instead of reading every trace:

python
# Only supports multi_turn_task_success and multi_turn_tool_use_quality.
# Service runs in the global region.
analysis_client = agentplatform.Client(project="PROJECT_ID", location="global")
response = analysis_client.evals.generate_loss_clusters(
    eval_result=result,
    metric="multi_turn_task_success",
    config={"max_top_cluster_count": 5},
)
for r in response.results:
    for cluster in r.clusters:
        print(
            f"[{cluster.taxonomy_entry.l1_category}/"
            f"{cluster.taxonomy_entry.l2_category}] "
            f"{cluster.item_count} cases — {cluster.taxonomy_entry.description}"
        )

Save response.model_dump_json() and render with scripts/render_html_report.py --type loss-analysis.

5. Optimize & Iterate

Apply a fix targeting the failing metric. Re-run Stage 3. Compare with scripts/compare_results.py --baseline <prev> --candidate <new> to confirm the target improved AND no other metric regressed.

Track progress across iterations:

IterationMetric AMetric BChange made
Baseline0.620.55—
v20.780.68Added grounding prompt
v30.810.72Fixed tool selection

Expect 5–10+ iterations per failing case. Only after a case passes should you expand coverage with more eval cases.

Proving your work

Never claim eval results you didn't read from an actual result object.

  • After running eval, print the summary_metrics table (scripts/inspect_results.py).
  • After a fix, show before/after via scripts/compare_results.py.
  • Before declaring success, confirm ALL cases pass — not just the one you were working on.

If you can't produce the evidence (SDK call failed, result truncated, metric unsupported), say so explicitly. Don't paper over gaps.

Rules of Engagement

  1. Always Plan First: Before writing a script, output a <plan> block detailing the steps you are about to take.
  2. Step-by-Step Execution: Prepare the data and evaluation script, present the dry-run confirmation card with full parameters, wait for user approval, execute only after explicit confirmation, then inspect and analyze results. Do NOT run evaluation calls before user confirmation.
  3. Standard Python: Use standard Python imports (import agentplatform, from google.genai import types). Don't use internal import paths.
  4. Verify Before Guessing: When unsure about SDK types or metrics, check the SDK source code rather than guessing or hallucinating.
  5. Never End a Turn Silently: Every turn must end with a non-empty, informative text reply to the user summarizing the actions taken or presenting the next steps. Returning nothing reads as a failure no matter what the tools did.

SDK Quick Reference

python
import agentplatform
from agentplatform import types
from google.genai import types as genai_types
import pandas as pd

# Initialize client
client = agentplatform.Client(project="PROJECT_ID", location="LOCATION")

# --- SINGLE-TURN EVAL (pandas DataFrame) -- RECOMMENDED ---
# The converter wraps plain strings for you.
df = pd.DataFrame({
    "prompt":   ["Q1", "Q2"],
    "response": ["A1", "A2"],
})
dataset = types.EvaluationDataset(eval_dataset_df=df)

# --- SINGLE-TURN EVAL (direct EvalCase) ---
# Verbose and easy to get wrong; see references/dataset_schema.md for the
# exact types before using this form.
dataset = types.EvaluationDataset(eval_cases=[
    types.EvalCase(
        prompt=genai_types.UserContent("Query here"),
        responses=[types.ResponseCandidate(
            response=genai_types.ModelContent("Model response here"))],
        reference=types.ResponseCandidate(
            response=genai_types.ModelContent("Ground truth here")),
    ),
])

# --- MULTI-TURN AGENT EVAL ---
agent_data = types.evals.AgentData(
    agents={"my_agent": types.evals.AgentConfig(
        agent_id="my_agent", instruction="You are helpful.")},
    turns=[types.evals.ConversationTurn(turn_index=0, events=[
        types.evals.AgentEvent(author="user",
            content=genai_types.Content(role="user",
                parts=[genai_types.Part(text="Hello")])),
        types.evals.AgentEvent(author="my_agent",
            content=genai_types.Content(role="model",
                parts=[genai_types.Part(text="Hi! How can I help?")])),
    ])],
)
dataset = types.EvaluationDataset(
    eval_cases=[types.EvalCase(agent_data=agent_data)])

# --- METRICS ---
predefined = types.RubricMetric.MULTI_TURN_TRAJECTORY_QUALITY
custom_llm = types.LLMMetric(name="tone",
    prompt_template="Is this polite? Response: {response}")
custom_code = types.CodeExecutionMetric(name="check",
    custom_function='def evaluate(instance): return {"score": 1.0}')

# --- EVALUATE ---
result = client.evals.evaluate(dataset=dataset, metrics=[predefined])

# --- RESULTS ---
for s in result.summary_metrics:
    print(f"{s.metric_name}: mean={s.mean_score}, pass_rate={s.pass_rate}")
for case in result.eval_case_results:
    for cand in case.response_candidate_results:
        for name, r in cand.metric_results.items():
            print(f"  {name}: score={r.score}, explanation={r.explanation}")

See references/sdk_patterns.md for advanced patterns: synthetic data generation, pairwise comparison, MetricPromptBuilder, multi-agent evaluation.

Bundled scripts

ScriptWhen to use
validate_dataset.pyBefore Stage 3 — catch malformed EvaluationDataset JSON.
parse_adk_traces.pyStage 1 — convert ADK session dumps to the canonical dataset shape.
inspect_results.pyStages 3/4 — render summary + per-case scores. --save-html for a browsable report.
compare_results.pyStage 5 — diff baseline vs. candidate, detect regressions.
render_html_report.pyRender HTML from a saved result JSON or loss-clusters JSON.
endpoint_evaluation.pyStages 2/3 against a deployed Agent Platform endpoint (BYOM). See references/deployment.md.
maas_evaluation.pyStages 2/3 against a Model-as-a-Service model by ID. See references/deployment.md.

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references) in skills/cloud/agent-platform-eval-flywheel of google/skills.

  • SKILL.md
  • references/dataset_schema.md
  • references/deployment.md
  • references/failure_patterns.md
  • references/metric_registry.md
  • references/sdk_patterns.md
  • scripts/compare_results.py
  • scripts/endpoint_evaluation.py
  • scripts/inspect_results.py
  • scripts/maas_evaluation.py
  • scripts/parse_adk_traces.py
  • scripts/render_html_report.py
  • scripts/validate_dataset.py

Open the folder on GitHubat commit 8a1ac05

Compare with similar skills

Agent Platform Eval Flywheel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Platform Eval Flywheel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Platform Eval Flywheel this skillgoogle/skills21k—~6.7kAutomated safety check: PassApache-2.0
Vertex AI Geminimajiayu000/claude-skill-registry6661 repos~2.1kAutomated safety check: PassApache-2.0
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples791—~2kAutomated safety check: PassApache-2.0
Aqua CLIoracle/accelerated-data-science125—~2.1kAutomated safety check: PassUPL-1.0
Jd Gap Analysisstarkyru/learn-ai105—~1.9kAutomated safety check: PassMIT

Similar skills

  • Vertex AI Gemini

    majiayu000/claude-skill-registry

    Google Cloud Vertex AI for enterprise Gemini deployments — production scaling, fine-tuning, and MLOps.

    666 GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    791 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Aqua CLI

    oracle/accelerated-data-science

    Official

    Complete CLI reference for the ADS AQUA command-line interface (ads aqua).

    125 GitHub stars~2.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Jd Gap Analysis

    starkyru/learn-ai

    Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.

    105 GitHub stars~1.9k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • ML Training Run Verifier

    Leeroo-AI/superml

    Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.

    195 GitHub stars~3.8k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from google/skills

All 145 skills in this repo
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated today
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Official

    Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.

    21k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Agent Platform Eval Flywheel

What does Agent Platform Eval Flywheel do?

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Agent Platform Eval Flywheel is an agent skill from google/skills, published by the product's own GitHub organization. Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology.

When should I use Agent Platform Eval Flywheel?

Agent Platform Eval Flywheel fits situations like: generating synthetic user scenarios; evaluating an agent; building an eval dataset; writing evaluation metrics.

How do I install Agent Platform Eval Flywheel in Claude Code?

Run `npx skills add google/skills --skill agent-platform-eval-flywheel -a claude-code`. Or copy the skill folder (skills/cloud/agent-platform-eval-flywheel in google/skills) into .claude/skills/agent-platform-eval-flywheel in your project. Claude Code loads it when a task matches its description.

How do I install Agent Platform Eval Flywheel in Codex?

Run `npx skills add google/skills --skill agent-platform-eval-flywheel -a codex`. Or copy the skill folder (skills/cloud/agent-platform-eval-flywheel in google/skills) into .agents/skills/agent-platform-eval-flywheel in your project. Codex loads it when a task matches its description.

Can I use Agent Platform Eval Flywheel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill agent-platform-eval-flywheel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-platform-eval-flywheel, .gemini/skills/agent-platform-eval-flywheel, .github/skills/agent-platform-eval-flywheel and .opencode/skills/agent-platform-eval-flywheel in your project.

What does Agent Platform Eval Flywheel need to run?

Going by SKILL.md and its folder, Agent Platform Eval Flywheel needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and pip). Our summary lists: Python 3.

Does Agent Platform Eval Flywheel access the network?

SKILL.md names 1 domain. As links in the text: docs.cloud.google.com. This is read from the text; nothing was executed.

Is Agent Platform Eval Flywheel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Platform Eval Flywheel use?

Agent Platform Eval Flywheel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Platform Eval Flywheel use?

About 6.7k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to Agent Platform Eval Flywheel?

Skills that share tags, products or a category with Agent Platform Eval Flywheel: Vertex AI Gemini (majiayu000/claude-skill-registry, 666 stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars), Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 791 stars) and Aqua CLI (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Platform Eval Flywheel?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.