Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Quality Flywheel

skills CLI
$ npx skills add GoogleCloudPlatform/vertex-ai-samples --skill quality-flywheel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GoogleCloudPlatform/vertex-ai-samples quality-flywheel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GoogleCloudPlatform/vertex-ai-samples.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/quality-flywheel .claude/skills/quality-flywheel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quality-flywheel
GitHub stars
791
Token cost
~2k tokens
SKILL.md length
607 words
Files
10 (incl. scripts, references)
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

  • Works in 6 steps: Setup & Project Initialization → Dataset Creation & Formatting → Metric Selection & Customization → …
  • Asked to evaluate my agent
  • SKILL.md covers When to use this skill and Workflow
  • Runs Python scripts from its folder

What it does

Quality Flywheel is an agent skill from GoogleCloudPlatform/vertex-ai-samples. Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK. Creates eval datasets (from session traces or synthetic generation), selects and configures metrics (RubricMetric, LLMMetric, CodeExecutionMetric), executes evals via client.evals.evaluate(), and analyzes results to suggest concrete fixes. Supports both single-turn model evaluation and multi-turn agent trajectory evaluation. Use when asked to "evaluate my agent", "evaluate my model", "create eval dataset", "run evals", "analyze…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `EVAL.yaml`, `TEST.md` and `references/dataset_schema.md`).

It sits in AI & LLM Engineering, covering LLM evaluation, Machine learning and Test generation. It works with Vertex AI, Google Cloud and Google Gemini. The repository describes itself as: Notebooks, code samples, sample apps, and other resources that demonstrate how to use, develop and manage machine learning and generative AI workflows using Google Cloud Vertex AI. The licence is Apache-2.0.

When your agent uses it

  • Asked to evaluate my agent
  • Evaluate my model
  • Create eval dataset
  • Analyze eval results

Example prompts

  • “evaluate my agent”
  • “evaluate my model”
  • “create eval dataset”
  • “/quality-flywheel”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Setup & Project Initialization
  2. Dataset Creation & Formatting
  3. Metric Selection & Customization
  4. Automated Execution
  5. Result Analysis & Auto-Optimization
  6. Iterate (The Flywheel)

What it can do on your machine

Read from SKILL.md and the folder at commit d0aed81. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quality Flywheel loads about 2k tokens when it runs, and up to ~9.9k if it reads all its reference files. Until then it costs about 155 tokens; SKILL.md has 607 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~155
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from GoogleCloudPlatform/vertex-ai-samples at commit d0aed81, republished under its Apache-2.0 licence (© GoogleCloudPlatform). 607 words, ~1,979 tokens.

Download SKILL.mdSave it as .claude/skills/quality-flywheel/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
quality-flywheel
description
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK. Creates eval datasets (from session traces or synthetic generation), selects and configures metrics (RubricMetric, LLMMetric, CodeExecutionMetric), executes evals via client.evals.evaluate(), and analyzes results to suggest concrete fixes. Supports both single-turn model evaluation and multi-turn agent trajectory evaluation. Use when asked to "evaluate my agent", "evaluate my model", "create eval dataset", "run evals", "analyze eval results", "which metrics should I use", "generate test data", or "improve quality".

Quality Flywheel Skill

You are the Quality Flywheel — an expert in GenAI evaluation. Your mission is to help users evaluate and iteratively improve their GenAI models and agents using the Google GenAI Evaluation SDK (google.genai / vertexai).

When to use this skill

  • Evaluating GenAI agents or models using client.evals.evaluate()
  • Creating synthetic datasets or ingesting session traces
  • Selecting, configuring, or writing custom evaluation metrics
  • Analyzing rubric verdicts and loss patterns
  • Suggesting concrete code/prompt improvements based on eval results

Workflow

Follow this workflow sequentially when assisting users:

Step 0. Setup & Project Initialization
  • CRITICAL: Before generating or executing any scripts, obtain the GCP Project ID and Location (e.g., global, us-central1). Check environment variables first (GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION). If not found, ask the user.
  • Newer Gemini models may only be available in the global region — use location="global" if the user wants to use them.
Step 1. Dataset Creation & Formatting
  • Parse Inputs: Convert user-provided descriptions into the SDK formats (EvalCase, AgentData, ConversationTurn, EvaluationDataset). See references/dataset_schema.md for the full type hierarchy and examples.

  • Single-Turn (Model Eval): Create EvalCase objects with prompt strings. Use client.evals.run_inference(model=..., src=dataset) to populate model responses if needed.

  • Multi-Turn (Agent Eval): If the user wants to test a multi-turn agent but lacks data:

    1. Generate Scenarios: Use client.evals.generate_user_scenarios with a UserScenarioGenerationConfig specifying user_scenario_count, simulation_instruction, and environment_data.
    2. Run Inference: Use client.evals.run_inference with a user_simulator_config to simulate interactions up to max_turn.
Step 2. Metric Selection & Customization

Use the quick-reference table to pick metrics. For the full catalog, see references/metric_registry.md.

Use CaseRecommended Metrics
RAG / QAhallucination_v1, grounding_v1, general_quality_v1
Tool-use agenttool_use_quality_v1, multi_turn_task_success_v1, tool_call_valid, tool_name_match
Multi-turn conversationmulti_turn_general_quality_v1, multi_turn_text_quality_v1, safety_v1
Code generationCodeExecutionMetric (custom), exact_match, instruction_following_v1
SummarizationRubricMetric.SUMMARIZATION_QUALITY, rouge_l_sum
Single-turn model evalgeneral_quality_v1, text_quality_v1, instruction_following_v1
  • Predefined: Access via types.RubricMetric.<NAME>. Server-side AutoRater — no judge model needed.
  • Custom LLM-as-a-judge: types.LLMMetric with prompt_template or types.MetricPromptBuilder for structured rubrics.
  • Custom Code: types.CodeExecutionMetric with a custom_function string containing def evaluate(instance: dict) for remote sandboxed execution. Or types.Metric with custom_function=<callable> for local execution.
Step 3. Automated Execution
  • Generate a complete Python evaluation script using client.evals.evaluate(dataset=..., metrics=...).
  • Save the script to a file and execute it to get real results.
  • Ensure the script prints results in a parseable format (JSON).
Show full SKILL.md (251 more words)Show less
Step 4. Result Analysis & Auto-Optimization
  • Read the stdout/stderr from the evaluation run.

  • CRITICAL — DO NOT HALLUCINATE: Only analyze the exact summary_metrics and eval_case_results returned by the executed script. Never fabricate scores or results.

  • Perform loss pattern analysis: Identify why a model or agent failed based on the returned explanations and rubric verdicts. See references/failure_patterns.md for common failure modes and their fixes.

  • Suggest concrete improvements to the user's prompt, system instruction, or agent code based on the failed examples.

Step 5. Iterate (The Flywheel)

After applying fixes, re-run evaluation (Step 3) and compare results. Repeat until quality targets are met. Track progress across iterations:

IterationMetric AMetric BChange Made
Baseline0.620.55—
v20.780.68Added grounding prompt
v30.810.72Fixed tool selection
Rules of Engagement
  1. Always Plan First: Before writing a script, output a <plan> block detailing the steps you are about to take.
  2. Step-by-Step Execution: Write the script, execute it, wait for output, then analyze. Don't do everything in one response.
  3. Standard Python: Use standard Python imports (import vertexai, from google.genai import types). Don't use internal import paths.
  4. Verify Before Guessing: When unsure about SDK types or metrics, check the SDK source code rather than guessing or hallucinating.
Error Handling

If execution returns a traceback:

  1. Analyze the error immediately.
  2. Fix the script.
  3. Run again.
  4. Keep iterating until success or user input is needed.
SDK Quick Reference
python
import vertexai
from vertexai import Client, types
from google.genai import types as genai_types

# Initialize client
client = vertexai.Client(project="PROJECT_ID", location="LOCATION")

# --- SINGLE-TURN EVAL ---
dataset = types.EvaluationDataset(eval_cases=[
    types.EvalCase(prompt="Query here", response="Model response here"),
])

# --- MULTI-TURN AGENT EVAL ---
agent_data = types.evals.AgentData(
    agents={"my_agent": types.evals.AgentConfig(
        agent_id="my_agent", instruction="You are helpful.")},
    turns=[types.evals.ConversationTurn(turn_index=0, events=[
        types.evals.AgentEvent(author="user",
            content=genai_types.Content(role="user",
                parts=[genai_types.Part(text="Hello")])),
        types.evals.AgentEvent(author="my_agent",
            content=genai_types.Content(role="model",
                parts=[genai_types.Part(text="Hi! How can I help?")])),
    ])],
)
dataset = types.EvaluationDataset(
    eval_cases=[types.EvalCase(agent_data=agent_data)])

# --- METRICS ---
predefined = types.RubricMetric.MULTI_TURN_TRAJECTORY_QUALITY
custom_llm = types.LLMMetric(name="tone",
    prompt_template="Is this polite? Response: {response}")
custom_code = types.CodeExecutionMetric(name="check",
    custom_function='def evaluate(instance): return 1.0')

# --- EVALUATE ---
result = client.evals.evaluate(dataset=dataset, metrics=[predefined])

# --- RESULTS ---
for s in result.summary_metrics:
    print(f"{s.metric_name}: mean={s.mean_score}, pass_rate={s.pass_rate}")
for case in result.eval_case_results:
    for cand in case.response_candidate_results:
        for name, r in cand.metric_results.items():
            print(f"  {name}: score={r.score}, explanation={r.explanation}")

See references/sdk_patterns.md for advanced patterns: synthetic data generation, pairwise comparison, MetricPromptBuilder, multi-agent evaluation.

© GoogleCloudPlatform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references) in skills/quality-flywheel of GoogleCloudPlatform/vertex-ai-samples.

  • SKILL.md
  • EVAL.yaml
  • TEST.md
  • references/dataset_schema.md
  • references/failure_patterns.md
  • references/metric_registry.md
  • references/sdk_patterns.md
  • scripts/generate_eval_code.py
  • scripts/parse_adk_traces.py
  • scripts/validate_dataset.py

Open the folder on GitHubat commit d0aed81

Compare with similar skills

Quality Flywheel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quality Flywheel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quality Flywheel this skillGoogleCloudPlatform/vertex-ai-samples791—~2kAutomated safety check: PassApache-2.0
Antigravityyuting0624/antigravity-for-claude-code374—~9.1kAutomated safety check: PassMIT
Gemini APIgoogle/skills21k3 repos~2.6kAutomated safety check: PassApache-2.0
Vertex AI API DevJetBrains/skills3631 repos~2.4kAutomated safety check: PassNone
Vertex AI Geminimajiayu000/claude-skill-registry6661 repos~2.1kAutomated safety check: PassApache-2.0
Gemini Sttmajiayu000/claude-skill-registry6661 repos~893Automated safety check: NotesMIT

Similar skills

  • Antigravity

    yuting0624/antigravity-for-claude-code

    Run the Antigravity CLI (Gemini) as a collaborating AI inside Claude Code, with intelligent model routing across the software development lifecycle.

    374 GitHub stars~9.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Gemini API

    google/skills

    Official

    A skill your agent uses when the user asks about using Gemini in an enterprise environment or explicitly mentions Vertex AI, Google Cloud, or Agent Platform.

    21k GitHub starsUsed in 3 repos~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Vertex AI API Dev

    JetBrains/skills

    Official

    Guides the usage of Gemini API on Google Cloud Vertex AI with the Gen AI SDK.

    363 GitHub starsUsed in 1 repo~2.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Vertex AI Gemini

    majiayu000/claude-skill-registry

    Google Cloud Vertex AI for enterprise Gemini deployments — production scaling, fine-tuning, and MLOps.

    666 GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Gemini Stt

    majiayu000/claude-skill-registry

    Transcribe audio files using Google's Gemini API or Vertex AI

    666 GitHub starsUsed in 1 repo~893 tokens
    Media & CreativeAuto-check: notes
  • Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing.

    97k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from GoogleCloudPlatform/vertex-ai-samples

  • Liveapi Service

    GoogleCloudPlatform/vertex-ai-samples

    Generates a LiveAPI client service class in the user's chosen programming language.

    791 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Vertex AI

    GoogleCloudPlatform/vertex-ai-samples

    Primary Router for Vertex AI skills. An agent skill from GoogleCloudPlatform/vertex-ai-samples.

    791 GitHub stars~522 tokensUpdated yesterday
    Auto-check passed

Questions about Quality Flywheel

What does Quality Flywheel do?

Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK. Quality Flywheel is an agent skill from GoogleCloudPlatform/vertex-ai-samples. Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

When should I use Quality Flywheel?

Quality Flywheel fits situations like: asked to evaluate my agent; evaluate my model; create eval dataset; analyze eval results.

How do I install Quality Flywheel in Claude Code?

Run `npx skills add GoogleCloudPlatform/vertex-ai-samples --skill quality-flywheel -a claude-code`. Or copy the skill folder (skills/quality-flywheel in GoogleCloudPlatform/vertex-ai-samples) into .claude/skills/quality-flywheel in your project. Claude Code loads it when a task matches its description.

How do I install Quality Flywheel in Codex?

Run `npx skills add GoogleCloudPlatform/vertex-ai-samples --skill quality-flywheel -a codex`. Or copy the skill folder (skills/quality-flywheel in GoogleCloudPlatform/vertex-ai-samples) into .agents/skills/quality-flywheel in your project. Codex loads it when a task matches its description.

Can I use Quality Flywheel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GoogleCloudPlatform/vertex-ai-samples --skill quality-flywheel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quality-flywheel, .gemini/skills/quality-flywheel, .github/skills/quality-flywheel and .opencode/skills/quality-flywheel in your project.

What does Quality Flywheel need to run?

Going by SKILL.md and its folder, Quality Flywheel needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Quality Flywheel access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quality Flywheel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Quality Flywheel use?

Quality Flywheel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quality Flywheel use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.9k tokens, read only when the agent opens those files.

What are the alternatives to Quality Flywheel?

Skills that share tags, products or a category with Quality Flywheel: Antigravity (yuting0624/antigravity-for-claude-code, 374 stars), Gemini API (google/skills, 21k stars), Vertex AI API Dev (JetBrains/skills, 363 stars) and Vertex AI Gemini (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quality Flywheel?

GoogleCloudPlatform (a GitHub organization) maintains it in GoogleCloudPlatform/vertex-ai-samples, which has 791 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 6, 2026.

Source: GoogleCloudPlatform/vertex-ai-samples on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.