Vertex AI Gemini
majiayu000/claude-skill-registry
Google Cloud Vertex AI for enterprise Gemini deployments — production scaling, fine-tuning, and MLOps.
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology.
$ npx skills add google/skills --skill agent-platform-eval-flywheel -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills agent-platform-eval-flywheel --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/agent-platform-eval-flywheel .claude/skills/agent-platform-eval-flywheel && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-platform-eval-flywheel" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-eval-flywheel into .claude/skills/agent-platform-eval-flywheel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-eval-flywheel", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/agent-platform-eval-flywheelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill agent-platform-eval-flywheel -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills agent-platform-eval-flywheel --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/agent-platform-eval-flywheel .agents/skills/agent-platform-eval-flywheel && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-platform-eval-flywheel" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-eval-flywheel into .agents/skills/agent-platform-eval-flywheel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-eval-flywheel", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill agent-platform-eval-flywheel -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills agent-platform-eval-flywheel --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/agent-platform-eval-flywheel .cursor/skills/agent-platform-eval-flywheel && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-platform-eval-flywheel" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-eval-flywheel into .cursor/skills/agent-platform-eval-flywheel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-eval-flywheel", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/agent-platform-eval-flywheel--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill agent-platform-eval-flywheel -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills agent-platform-eval-flywheel --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/agent-platform-eval-flywheel .gemini/skills/agent-platform-eval-flywheel && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-platform-eval-flywheel" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-eval-flywheel into .gemini/skills/agent-platform-eval-flywheel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-eval-flywheel", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills agent-platform-eval-flywheelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill agent-platform-eval-flywheel -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/agent-platform-eval-flywheel .github/skills/agent-platform-eval-flywheel && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-platform-eval-flywheel" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-eval-flywheel into .github/skills/agent-platform-eval-flywheel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-eval-flywheel", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill agent-platform-eval-flywheel -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills agent-platform-eval-flywheel --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/agent-platform-eval-flywheel .opencode/skills/agent-platform-eval-flywheel && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-platform-eval-flywheel" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-eval-flywheel into .opencode/skills/agent-platform-eval-flywheel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-eval-flywheel", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-platform-eval-flywheelMeasures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology.
Agent Platform Eval Flywheel is an agent skill from google/skills, published by the product's own GitHub organization. Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results before and after a fix, or when guidance is needed on Agent Platform eval methodology — including dataset schema, LLM-as-judge scoring, and common failure causes. For fine-tuning, use agent-platform-tuning. For general…
Its SKILL.md is about 6.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts and reference files (for example `references/dataset_schema.md`, `references/deployment.md` and `references/failure_patterns.md`).
It sits in AI & LLM Engineering, covering LLM evaluation, Fine-tuning and Deployment. It works with Google Cloud. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 7 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.cloud.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Platform Eval Flywheel loads about 6.7k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 149 tokens; SKILL.md has 2,136 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 2,136 words, ~6,739 tokens.
.claude/skills/agent-platform-eval-flywheel/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Help users evaluate and iteratively improve GenAI models and agents using the
Agent Platform GenAI Evaluation SDK (google.genai / agentplatform).
client.evals.evaluate()).endpoint_evaluation.py / maas_evaluation.py scripts.Before executing any commands or scripts on behalf of the user, you MUST adhere to the following safety tiers based on the action requested:
inspect_results.py, compare_results.py,
validate_dataset.py, parse_adk_traces.py, render_html_report.py)client.evals.run_inference,
client.evals.evaluate, client.evals.generate_conversation_scenarios,
client.evals.generate_loss_clusters)client.evals.evaluate(), client.evals.run_inference(),
client.evals.generate_conversation_scenarios(), or run any script
invoking these remote operations before user confirmation. In the
initial turn, you may prepare local data structures and compose the
script, but you MUST present the dry-run preview card and obtain
explicit user confirmation before running any remote evaluation or
scenario generation call.run_command and report
the results. Do not conclude the turn without executing the approved
action.The scripts need vertexai (from google-cloud-aiplatform[evaluation]),
google-genai, pandas, and requests. Do not create a virtual
environment — it starts empty and hides packages the environment already
provides, forcing a redundant install. Probe, and install only what is missing:
python3 -c "import vertexai, google.genai, pandas, requests" \
|| pip install 'google-cloud-aiplatform[evaluation]>=1.163.0' 'google-genai>=1.0.0'The version specifiers must stay quoted: unquoted, bash reads >=1.154.0 as a
redirect and silently writes an empty file instead of constraining the install.
Need GOOGLE_CLOUD_PROJECT and GOOGLE_CLOUD_LOCATION.
project='<PROJECT_NUMBER>',
location='us-central1'). Never change or override the user's requested
location to 'global' unless the user explicitly requested 'global'.import agentplatform
client = agentplatform.Client(project=PROJECT, location=LOCATION)
client.evals.run_inference(model=..., src=...)
client.evals.evaluate(dataset=..., metrics=...)
client.evals.generate_conversation_scenarios(...)Two imports that look plausible and are not:
from agentplatform.types import evals -- ModuleNotFoundError. types is
a module, not a package; use from agentplatform import types.from vertexai.evaluation import PointwiseMetric, EvalTask -- the
superseded SDK. Its classes take different arguments (PointwiseMetric has
no system_instruction), so code written against it fails with TypeError
rather than an import error. Use agentplatform throughout.Five stages, run in order on the first pass, then loop 2 → 5 until quality targets are met.
| Shortcut | Why it fails |
|---|---|
| "I'll tune the metric threshold down | Hides real failures. Fix the agent, |
| : so it passes." : not the bar. : | |
| "This case is flaky, I'll skip it." | Flakiness reveals non-determinism in |
: : the agent. Fix with temperature=0 : | |
| : : or stricter instructions. : | |
| "I just need to fix the eval | If expected outputs keep moving, the |
| : dataset, not the agent." : agent has a behavior problem. : | |
| "I can tell from the trace it works | Self-grading doesn't generalize. |
: — skip Stage 3." : Always run evaluate() and read : | |
| : : scores. : | |
| "One iteration is enough." | Expect 5–10+ iterations. Stopping |
| : : early leaves regressions on other : | |
| : : metrics undetected. : |
Produce an EvaluationDataset. There are three input shapes, pick the one that
matches the data the user already has:
EvalCase list (single-turn or multi-turn):
from agentplatform import types
from google.genai import types as genai_types
# prompt/reference/response values are Content, not str. UserContent and
# ModelContent wrap a plain string and set the right role.
dataset = types.EvaluationDataset(eval_cases=[
types.EvalCase(
prompt=genai_types.UserContent("What is 2+2?"),
responses=[types.ResponseCandidate(
response=genai_types.ModelContent("4"))],
reference=types.ResponseCandidate(
response=genai_types.ModelContent("4")),
),
# For multi-turn agent traces, set agent_data instead of prompt/responses.
])Multi-turn agent traces wrap each conversation in AgentData →
ConversationTurn → AgentEvent. See
references/dataset_schema.md for the full
type hierarchy.
Pandas DataFrame (tabular sources — CSV, BigQuery, Sheets):
import pandas as pd
from agentplatform import types
df = pd.DataFrame({
"prompt": ["What is 2+2?", "Capital of France?"],
"response": ["4", "Paris"],
"reference": ["4", "Paris"],
})
dataset = types.EvaluationDataset(eval_dataset_df=df)Column names must match the fields the chosen metrics expect (see references/dataset_schema.md for the per-metric requirements table).
Cold start (no data at all): synthesize scenarios server-side with
client.evals.generate_conversation_scenarios(agent=..., config=...) -- the
parameter is agent or agent_info, not agents, and config is
required. The config class is types.evals.UserScenarioGenerationConfig,
not types.UserScenarioGenerationConfig. Set its user_scenario_count
(1-100): it defaults to None, the client accepts that, and the server
rejects the call with 400 INVALID_ARGUMENT. count is a separate field
and does not substitute for it. Stage 2 plays the scenarios out.
location,
environment_data, simulation_instruction, or model_name), do NOT
assume defaults or guess values. You MUST pause in your first turn and
explicitly ask the user for the missing information (e.g., "Please
provide the missing simulation instructions, environment data, model
name, and location"). Only proceed with the dry-run preview after the
user provides them.run_command. Do NOT generate scenarios directly in plain text.Managed Agents (Gemini Agents API): evaluate agents created with the
Managed Agents API.
Use generate_conversation_scenarios to create test scenarios from the
agent's configuration, run_inference to execute the agent, and evaluate
to score the traces. These functions now accept managed agents and
interaction ids as input. You can also evaluate existing interactions
recorded via the Interactions API using InteractionsDataSource. See
references/sdk_patterns.md Pattern 8 for the
full code pattern.
For ADK session dumps, use scripts/parse_adk_traces.py instead of writing the
conversion by hand.
Populate responses/traces on the dataset. Skip this stage if traces are already complete (e.g., production logs or replay).
# Agent eval — pass a callable wrapping the user's ADK Agent/App.
client.evals.run_inference(model=agent_callable, src=dataset)
# Model eval — pass a model ID directly.
client.evals.run_inference(model="gemini-2.5-flash", src=dataset)
# Synthesized scenarios — let the simulator drive.
client.evals.run_inference(
model=agent_callable,
src=dataset,
user_simulator_config=UserSimulatorConfig(max_turn=10),
)
# DataFrame also works as src= — no EvalCase wrapping needed.
client.evals.run_inference(model="gemini-2.5-flash", src=df)
# Managed Agent — pass an agent resource name.
AGENT_RESOURCE = f"projects/{PROJECT_ID}/locations/global/agents/{AGENT_ID}"
client.evals.run_inference(
agent=AGENT_RESOURCE,
src=scenarios,
config={"user_simulator_config": {"max_turn": 3}},
)result = client.evals.evaluate(dataset=dataset, metrics=[...])
result.show() # Interactive HTML report with scores, rubrics, and traces.Pick metrics by what you want to measure. Full catalog in references/metric_registry.md.
If the user names a metric, use it directly. Every identifier in the tables
below (general_quality, text_quality, instruction_following,
hallucination, grounding, safety, multi_turn_*, final_response_*,
tool_use_quality) is a types.RubricMetric.<UPPERCASE_NAME> accessor — pass
it straight into metrics=[types.RubricMetric.GENERAL_QUALITY, ...]. Do not
scaffold a custom LLMMetric for a name that appears here, and do not reach for
vertexai.evaluation.EvalTask / PointwiseMetric /
MetricPromptTemplateExamples — that SDK is superseded (see Setup).
Agent metrics (multi-turn, adaptive rubrics) — start here for agent eval.
| Goal | Metric |
|---|---|
| Did the agent achieve the user's goal? | multi_turn_task_success |
| Was the reasoning path logical and efficient? | multi_turn_trajectory_quality |
| Tool/function calling quality across turns | multi_turn_tool_use_quality |
| Overall conversational quality | multi_turn_general_quality |
| Final response quality (no reference needed) | final_response_quality |
| Final response vs. a golden reference | final_response_match |
| Single-turn tool use | tool_use_quality |
General quality metrics (single-turn, adaptive rubrics) — for model eval.
| Goal | Metric |
|---|---|
| Overall response quality (recommended starting point) | general_quality |
| Linguistic quality (fluency, coherence, grammar) | text_quality |
| Adherence to specific constraints / instructions | instruction_following |
Static rubric metrics (fixed criteria) — apply alongside the above.
| Goal | Metric |
|---|---|
| Catch hallucinated claims (RAG, factual answers) | hallucination |
| Factuality / consistency against provided context | grounding |
| Safety policy compliance | safety |
Domain-specific check no built-in covers: write a custom metric.
types.RubricMetric.<NAME> — server-side AutoRater, no
judge model needed.types.LLMMetric with prompt_template or
types.MetricPromptBuilder for structured rubrics. Always set
judge_model; it defaults to None and every case then fails with 400 INVALID_ARGUMENT: Error parsing JSON.gemini-2.5-pro), use it. If the user omits the judge model or states
they do not have information / preference for one, default to
gemini-2.5-flash as the judge model in the dry-run preview card and
ask for confirmation to run the evaluation. Do NOT halt or refuse to
evaluate when the user does not specify a judge model.types.CodeExecutionMetric with a custom_function string
containing def evaluate(instance: dict) for remote sandboxed execution; or
types.Metric with custom_function=<callable> for local execution.Always persist the result so Stage 4 and 5 can read it. Save both JSON (machine-readable, diffable) and HTML (human-readable, linkable):
import datetime
from pathlib import Path
from agentplatform._genai import _evals_visualization
out_dir = Path("artifacts/grade_results")
out_dir.mkdir(parents=True, exist_ok=True)
ts = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
# fallback=str, or a DataFrame-backed dataset raises PydanticSerializationError.
result_json = result.model_dump_json(fallback=str)
(out_dir / f"results_{ts}.json").write_text(result_json)
html = _evals_visualization.get_evaluation_html(result_json)
(out_dir / f"results_{ts}.html").write_text(str(html))Or after the fact: scripts/render_html_report.py --type evaluation or
scripts/inspect_results.py --save-html.
Read summary_metrics and eval_case_results — never fabricate scores. Use
scripts/inspect_results.py --failing-only to filter to failures.
For each failed metric, see references/failure_patterns.md for deeper diagnoses. The compact mapping:
| Failing metric | What to change |
|---|---|
multi_turn_task_success low | The agent isn't completing the goal — |
| : : fix orchestration, missing tool calls, : | |
| : : premature termination, wrong tool : | |
| : : selection. : | |
multi_turn_trajectory_quality low | The agent reaches the goal |
| : : inefficiently — refine planning : | |
| : : prompts, remove redundant tool calls. : | |
multi_turn_tool_use_quality low | Fix tool descriptions, parameter |
| : : docstrings, or agent instructions for : | |
| : : tool selection. : | |
final_response_quality low | Read auto-generated rubric verdicts; |
| : : refine instructions to address the : | |
| : : worst-scoring criterion. : | |
final_response_match low | The agent's final answer doesn't match |
| : : the golden reference — adjust response : | |
| : : format or update the reference. : | |
hallucination low | Tighten instructions to stay grounded |
| : : in tool output; verify the tool : | |
| : : actually returned the claimed data. : | |
grounding low | The response contradicts the provided |
| : : context — add explicit "cite only from : | |
| : : context" instructions. : | |
safety low | Add safety guardrails; review the |
| : : violating content category in the : | |
| : : rubric verdict. : | |
general_quality / text_quality | Adjust system instruction wording; the |
| : low : model's default phrasing is too : | |
| : : generic for the task. : | |
instruction_following low | The agent is ignoring constraints — |
| : : restate them in the system instruction : | |
| : : or use stricter wording. : | |
| Agent calls wrong tools | Fix tool descriptions, agent |
: : instructions, or tool_config. : | |
| Agent calls extra tools | Add explicit stop instructions, or |
| : : switch to : | |
: : multi_turn_tool_use_quality to : | |
| : : surface the extra calls in the rubric. : |
For 10+ failures on the same metric, use the Error Analysis service to cluster failures into themes (L1/L2 taxonomy categories) instead of reading every trace:
# Only supports multi_turn_task_success and multi_turn_tool_use_quality.
# Service runs in the global region.
analysis_client = agentplatform.Client(project="PROJECT_ID", location="global")
response = analysis_client.evals.generate_loss_clusters(
eval_result=result,
metric="multi_turn_task_success",
config={"max_top_cluster_count": 5},
)
for r in response.results:
for cluster in r.clusters:
print(
f"[{cluster.taxonomy_entry.l1_category}/"
f"{cluster.taxonomy_entry.l2_category}] "
f"{cluster.item_count} cases — {cluster.taxonomy_entry.description}"
)Save response.model_dump_json() and render with scripts/render_html_report.py --type loss-analysis.
Apply a fix targeting the failing metric. Re-run Stage 3. Compare with
scripts/compare_results.py --baseline <prev> --candidate <new> to confirm the
target improved AND no other metric regressed.
Track progress across iterations:
| Iteration | Metric A | Metric B | Change made |
|---|---|---|---|
| Baseline | 0.62 | 0.55 | — |
| v2 | 0.78 | 0.68 | Added grounding prompt |
| v3 | 0.81 | 0.72 | Fixed tool selection |
Expect 5–10+ iterations per failing case. Only after a case passes should you expand coverage with more eval cases.
Never claim eval results you didn't read from an actual result object.
summary_metrics table
(scripts/inspect_results.py).scripts/compare_results.py.If you can't produce the evidence (SDK call failed, result truncated, metric unsupported), say so explicitly. Don't paper over gaps.
<plan> block
detailing the steps you are about to take.import agentplatform,
from google.genai import types). Don't use internal import paths.import agentplatform
from agentplatform import types
from google.genai import types as genai_types
import pandas as pd
# Initialize client
client = agentplatform.Client(project="PROJECT_ID", location="LOCATION")
# --- SINGLE-TURN EVAL (pandas DataFrame) -- RECOMMENDED ---
# The converter wraps plain strings for you.
df = pd.DataFrame({
"prompt": ["Q1", "Q2"],
"response": ["A1", "A2"],
})
dataset = types.EvaluationDataset(eval_dataset_df=df)
# --- SINGLE-TURN EVAL (direct EvalCase) ---
# Verbose and easy to get wrong; see references/dataset_schema.md for the
# exact types before using this form.
dataset = types.EvaluationDataset(eval_cases=[
types.EvalCase(
prompt=genai_types.UserContent("Query here"),
responses=[types.ResponseCandidate(
response=genai_types.ModelContent("Model response here"))],
reference=types.ResponseCandidate(
response=genai_types.ModelContent("Ground truth here")),
),
])
# --- MULTI-TURN AGENT EVAL ---
agent_data = types.evals.AgentData(
agents={"my_agent": types.evals.AgentConfig(
agent_id="my_agent", instruction="You are helpful.")},
turns=[types.evals.ConversationTurn(turn_index=0, events=[
types.evals.AgentEvent(author="user",
content=genai_types.Content(role="user",
parts=[genai_types.Part(text="Hello")])),
types.evals.AgentEvent(author="my_agent",
content=genai_types.Content(role="model",
parts=[genai_types.Part(text="Hi! How can I help?")])),
])],
)
dataset = types.EvaluationDataset(
eval_cases=[types.EvalCase(agent_data=agent_data)])
# --- METRICS ---
predefined = types.RubricMetric.MULTI_TURN_TRAJECTORY_QUALITY
custom_llm = types.LLMMetric(name="tone",
prompt_template="Is this polite? Response: {response}")
custom_code = types.CodeExecutionMetric(name="check",
custom_function='def evaluate(instance): return {"score": 1.0}')
# --- EVALUATE ---
result = client.evals.evaluate(dataset=dataset, metrics=[predefined])
# --- RESULTS ---
for s in result.summary_metrics:
print(f"{s.metric_name}: mean={s.mean_score}, pass_rate={s.pass_rate}")
for case in result.eval_case_results:
for cand in case.response_candidate_results:
for name, r in cand.metric_results.items():
print(f" {name}: score={r.score}, explanation={r.explanation}")See references/sdk_patterns.md for advanced
patterns: synthetic data generation, pairwise comparison, MetricPromptBuilder,
multi-agent evaluation.
| Script | When to use |
|---|---|
validate_dataset.py | Before Stage 3 — catch malformed EvaluationDataset JSON. |
parse_adk_traces.py | Stage 1 — convert ADK session dumps to the canonical dataset shape. |
inspect_results.py | Stages 3/4 — render summary + per-case scores. --save-html for a browsable report. |
compare_results.py | Stage 5 — diff baseline vs. candidate, detect regressions. |
render_html_report.py | Render HTML from a saved result JSON or loss-clusters JSON. |
endpoint_evaluation.py | Stages 2/3 against a deployed Agent Platform endpoint (BYOM). See references/deployment.md. |
maas_evaluation.py | Stages 2/3 against a Model-as-a-Service model by ID. See references/deployment.md. |
© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts, references) in skills/cloud/agent-platform-eval-flywheel of google/skills.
Open the folder on GitHubat commit 8a1ac05
Agent Platform Eval Flywheel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Platform Eval Flywheel this skillgoogle/skills | 21k | — | ~6.7k | Automated safety check: Pass | Apache-2.0 | |
| Vertex AI Geminimajiayu000/claude-skill-registry | 666 | 1 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 791 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Aqua CLIoracle/accelerated-data-science | 125 | — | ~2.1k | Automated safety check: Pass | UPL-1.0 | |
| Jd Gap Analysisstarkyru/learn-ai | 105 | — | ~1.9k | Automated safety check: Pass | MIT |
majiayu000/claude-skill-registry
Google Cloud Vertex AI for enterprise Gemini deployments — production scaling, fine-tuning, and MLOps.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
oracle/accelerated-data-science
Complete CLI reference for the ADS AQUA command-line interface (ads aqua).
starkyru/learn-ai
Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.
Leeroo-AI/superml
Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
google/skills
Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.
Works with
Categories
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Agent Platform Eval Flywheel is an agent skill from google/skills, published by the product's own GitHub organization. Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology.
Agent Platform Eval Flywheel fits situations like: generating synthetic user scenarios; evaluating an agent; building an eval dataset; writing evaluation metrics.
Run `npx skills add google/skills --skill agent-platform-eval-flywheel -a claude-code`. Or copy the skill folder (skills/cloud/agent-platform-eval-flywheel in google/skills) into .claude/skills/agent-platform-eval-flywheel in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill agent-platform-eval-flywheel -a codex`. Or copy the skill folder (skills/cloud/agent-platform-eval-flywheel in google/skills) into .agents/skills/agent-platform-eval-flywheel in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill agent-platform-eval-flywheel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-platform-eval-flywheel, .gemini/skills/agent-platform-eval-flywheel, .github/skills/agent-platform-eval-flywheel and .opencode/skills/agent-platform-eval-flywheel in your project.
Going by SKILL.md and its folder, Agent Platform Eval Flywheel needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and pip). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: docs.cloud.google.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Agent Platform Eval Flywheel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.7k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agent Platform Eval Flywheel: Vertex AI Gemini (majiayu000/claude-skill-registry, 666 stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars), Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 791 stars) and Aqua CLI (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.