Converts CXAS golden evaluations to SCRAPI SimulationEvals test cases.

Apache-2.0Auto-check passedTesting & QA

Install Cxas Sim Eval

skills CLI
$ npx skills add GoogleCloudPlatform/cxas-scrapi --skill cxas-sim-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GoogleCloudPlatform/cxas-scrapi cxas-sim-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GoogleCloudPlatform/cxas-scrapi.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/cxas-sim-eval .claude/skills/cxas-sim-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cxas-sim-eval
GitHub stars
106
Token cost
~1.2k tokens
SKILL.md length
466 words
Files
5 (incl. scripts)
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Converts CXAS golden evaluations to SCRAPI SimulationEvals test cases.

  • Works in 10 steps: Check Environment → Get App Name and Output Directory → Fetch Evaluations → …
  • Generating high-level
  • SKILL.md covers Steps, Automation Scripts and Interpreting Cognitive…
  • Runs Python scripts from its folder; calls python and gcloud

What it does

Cxas Sim Eval is an agent skill from GoogleCloudPlatform/cxas-scrapi. Converts CXAS golden evaluations to SCRAPI SimulationEvals test cases. Use when generating high-level, goal-oriented test cases from turn-by-turn evaluation JSONs, and when enriching test expectations with inferred tool calls.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/convert_eval.py`, `scripts/fetch_app_data.py` and `scripts/fetch_tool_schemas.py`).

It sits in Testing & QA, covering Test generation. It works with Google Cloud. The repository describes itself as: A powerful Python API, CLI, and set of Agent Skills for CX Agent Studio to automate, evaluate, and scale your agents with ease. The licence is Apache-2.0.

When your agent uses it

  • Generating high-level
  • Goal-oriented test cases from turn-by-turn evaluation JSONs
  • When enriching test expectations with inferred tool calls

Example prompts

  • “Use the cxas-sim-eval skill to convert CXAS golden evaluations to SCRAPI SimulationEvals test cases”
  • “/cxas-sim-eval”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Check Environment
  2. Get App Name and Output Directory
  3. Fetch Evaluations
  4. Fetch Tool Schemas
  5. Fetch Agent Tools Configuration
  6. Convert Evaluations
  7. Fetch Evaluations and Agent Config
  8. Fetch Tool Schemas
  9. Convert Evaluations
  10. Run Evaluations

What it can do on your machine

Read from SKILL.md and the folder at commit ffba639. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gcloud, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cxas Sim Eval loads about 1.2k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 466 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from GoogleCloudPlatform/cxas-scrapi at commit ffba639, republished under its Apache-2.0 licence (© GoogleCloudPlatform). 466 words, ~1,184 tokens.

Download SKILL.mdSave it as .claude/skills/cxas-sim-eval/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
cxas-sim-eval
description
Converts CXAS golden evaluations to SCRAPI SimulationEvals test cases. Use when generating high-level, goal-oriented test cases from turn-by-turn evaluation JSONs, and when enriching test expectations with inferred tool calls.

CXAS Evaluation to Simulation Converter

This skill helps convert turn-by-turn CXAS golden evaluations into high-level, goal-oriented test cases for the SCRAPI SimulationEvals framework. It analyzes the agent's tools to enrich expectations with specific tool calls.


Steps

1. Check Environment

Ensure cxas_scrapi is installed as a python package. You can check this by running:

bash
python -c "import cxas_scrapi"

Ensure gcloud is authenticated properly:

bash
gcloud auth list

If needed, login with:

bash
gcloud auth login
2. Get App Name and Output Directory

[!IMPORTANT] You MUST ask the user for the full resource name of the app/agent (e.g., projects/.../locations/.../apps/...) and the base output directory before proceeding with any execution steps.

Ask the user for these values.

3. Fetch Evaluations

Fetch the list of evaluations using the CES API. Save each evaluation as a JSON file named after its display name under [output_dir]/golden_evals/.

4. Fetch Tool Schemas

Fetch the full schemas for all tools available in the app and save them under [output_dir]/tools/.

5. Fetch Agent Tools Configuration

Fetch the list of tools and toolsets used by the agent and save the configuration (e.g., to [output_dir]/agent_tools.json).

6. Convert Evaluations

Run the conversion script (convert_eval.py) to process the fetched evaluations and save the converted test cases under [output_dir]/sim_evals/.

Automation Scripts

Three scripts are available to automate the process:

1. Fetch Evaluations and Agent Config

scripts/fetch_app_data.py

Fetches evaluations and the list of tools used by the agent from the CES API.

Usage:

bash
python .agents/skills/cxas-sim-eval/scripts/fetch_app_data.py \
  --app-name "projects/.../locations/.../apps/..." \
  --output-dir /path/to/output_directory
2. Fetch Tool Schemas

scripts/fetch_tool_schemas.py

Fetches the full schemas for all tools available in the app.

Usage:

bash
python .agents/skills/cxas-sim-eval/scripts/fetch_tool_schemas.py \
  --app-name "projects/.../locations/.../apps/..." \
  --output-dir /path/to/output_directory
3. Convert Evaluations

scripts/convert_eval.py

Converts the fetched evaluations to simulation test cases, using the fetched tool schemas to infer expectations.

Usage:

bash
python .agents/skills/cxas-sim-eval/scripts/convert_eval.py \
  --output-dir /path/to/output_directory \
  --parallelism 5
Show full SKILL.md (203 more words)Show less
4. Run Evaluations

scripts/run_evals.py

Runs the simulation evaluations, logs raw results, and generates a combined HTML report.

Cognitive Diagnostics Analysis: If the agent has the intercept_and_score_reasoning tool enabled, this script will automatically extract and analyze the agent's internal monologue for failed evaluations. It detects issues like overthinking, hesitation, and backtracking. Furthermore, it correlates these diagnostics with the agent's instructions to generate actionable suggestions for improvement directly in the HTML report.

Usage:

bash
python .agents/skills/cxas-sim-eval/scripts/run_evals.py \
  --app-name "projects/.../locations/.../apps/..." \
  --output-dir /path/to/output_directory \
  --parallelism 5 \
  --start-index 0 \
  --end-index 10

Interpreting Cognitive Diagnostics

When running evaluations with the intercept_and_score_reasoning tool enabled, the system extracts diagnostics to help you identify issues in agent reasoning.

Key Signals
  1. Overthinking (Verbosity)

    • Symptom: Internal monologue exceeds 350 or 600 characters.
    • Meaning: The agent is struggling to process complex or circular instructions.
    • Fix: Simplify instructions. Break down complex tasks into smaller, linear steps.
  2. Hedging

    • Symptom: Use of words like "might be", "guess", "unsure", "assume".
    • Meaning: The agent is uncertain about its next action, often due to missing edge case handling.
    • Fix: Add explicit instructions for the scenario the agent is unsure about.
  3. Backtracking

    • Symptom: Use of words like "wait", "actually", "on second thought".
    • Meaning: The agent is abandoning a plan mid-turn or correcting itself, indicating unclear triggers.
    • Fix: Clarify triggers and state transitions in instructions.

© GoogleCloudPlatform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in .agents/skills/cxas-sim-eval of GoogleCloudPlatform/cxas-scrapi.

  • SKILL.md
  • scripts/convert_eval.py
  • scripts/fetch_app_data.py
  • scripts/fetch_tool_schemas.py
  • scripts/run_evals.py

Open the folder on GitHubat commit ffba639

Compare with similar skills

Cxas Sim Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cxas Sim Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cxas Sim Eval this skillGoogleCloudPlatform/cxas-scrapi106—~1.2kAutomated safety check: PassApache-2.0
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples791—~2kAutomated safety check: PassApache-2.0
Ak Dev New Secret Provideryaalalabs/agent-kernel191—~2.6kAutomated safety check: PassApache-2.0
Write and Verify Playwright Testsappsmithorg/appsmith41k—~2.9kAutomated safety check: NotesApache-2.0
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Engine E2Ewix/react-native-navigation13k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    791 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ak Dev New Secret Provider

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new built-in secret provider to Agent Kernel's secret-resolution capability (beyond env and awsssm).

    191 GitHub stars~2.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check: notes
  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Engine E2E

    wix/react-native-navigation

    Official

    Run Wix Engine (mobile-apps-engine) iOS E2E tests locally to validate RNN changes.

    13k GitHub stars~1.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed

More from GoogleCloudPlatform/cxas-scrapi

All 14 skills in this repo
  • Cxas Agent Foundry

    GoogleCloudPlatform/cxas-scrapi

    End-to-end GECX/CXAS/CES conversational agent lifecycle -- build agents from requirements (PRD-to-agent), create and run evals (goldens, simulations, tool tests, callback tests), debug failures, and…

    106 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Cxas Autolabel Rules

    GoogleCloudPlatform/cxas-scrapi

    Author, validate, and manage Contact Center AI (CCAI) Insights Autolabeling Rules.

    106 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Cxas Configurable Dashboards

    GoogleCloudPlatform/cxas-scrapi

    Author, validate, and manage Contact Center AI (CCAI) Insights Configurable Dashboards.

    106 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Cxas Dfcx Migration

    GoogleCloudPlatform/cxas-scrapi

    Migrate Dialogflow CX (DFCX) agents to CXAS (Customer Experience Agent Studio) agents.

    106 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Cxas Loss Analysis

    GoogleCloudPlatform/cxas-scrapi

    Retrieves non-contained CCAI Insights conversations (losses), uses agent intelligence to cluster them into common failure patterns, and generates a markdown report.

    106 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Cxas Composite Voice Agent Optimizer

    GoogleCloudPlatform/cxas-scrapi

    Audits, optimizes, and remediates CXAS agent configurations for Gemini Composite V1 voice naturalness, persona styling, and multi-language coverage directly in local workspaces with cxas-scrapi.

    106 GitHub stars~10k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Cxas Sim Eval

What does Cxas Sim Eval do?

Converts CXAS golden evaluations to SCRAPI SimulationEvals test cases. Cxas Sim Eval is an agent skill from GoogleCloudPlatform/cxas-scrapi. Converts CXAS golden evaluations to SCRAPI SimulationEvals test cases.

When should I use Cxas Sim Eval?

Cxas Sim Eval fits situations like: generating high-level; goal-oriented test cases from turn-by-turn evaluation JSONs; when enriching test expectations with inferred tool calls.

How do I install Cxas Sim Eval in Claude Code?

Run `npx skills add GoogleCloudPlatform/cxas-scrapi --skill cxas-sim-eval -a claude-code`. Or copy the skill folder (.agents/skills/cxas-sim-eval in GoogleCloudPlatform/cxas-scrapi) into .claude/skills/cxas-sim-eval in your project. Claude Code loads it when a task matches its description.

How do I install Cxas Sim Eval in Codex?

Run `npx skills add GoogleCloudPlatform/cxas-scrapi --skill cxas-sim-eval -a codex`. Or copy the skill folder (.agents/skills/cxas-sim-eval in GoogleCloudPlatform/cxas-scrapi) into .agents/skills/cxas-sim-eval in your project. Codex loads it when a task matches its description.

Can I use Cxas Sim Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GoogleCloudPlatform/cxas-scrapi --skill cxas-sim-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cxas-sim-eval, .gemini/skills/cxas-sim-eval, .github/skills/cxas-sim-eval and .opencode/skills/cxas-sim-eval in your project.

What does Cxas Sim Eval need to run?

Going by SKILL.md and its folder, Cxas Sim Eval needs Python for the scripts in its folder and the command-line tools its instructions call (python and gcloud). Our summary lists: Python 3.

Does Cxas Sim Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cxas Sim Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Cxas Sim Eval use?

Cxas Sim Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cxas Sim Eval use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cxas Sim Eval?

Skills that share tags, products or a category with Cxas Sim Eval: Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 791 stars), Ak Dev New Secret Provider (yaalalabs/agent-kernel, 191 stars), Write and Verify Playwright Tests (appsmithorg/appsmith, 41k stars) and Adk Verify Snippets (google/adk-python, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cxas Sim Eval?

GoogleCloudPlatform (a GitHub organization) maintains it in GoogleCloudPlatform/cxas-scrapi, which has 106 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 6, 2026.

Source: GoogleCloudPlatform/cxas-scrapi on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.