Official agent skill

Testing Skill

by pydantic in pydantic/pydantic-ai

Record, rewrite, and debug VCR cassettes for HTTP recordings.

OfficialMITAuto-check: notesTesting & QA

Install Testing Skill

skills CLI
$ npx skills add pydantic/pydantic-ai --skill testing-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pydantic/pydantic-ai testing-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pydantic/pydantic-ai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/testing-skill .claude/skills/testing-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-skill
GitHub stars
20k
Token cost
~839 tokens
SKILL.md length
198 words
Files
2
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

Record, rewrite, and debug VCR cassettes for HTTP recordings.

  • Works in 3 steps: Record cassettes → Verify recordings → Review snapshots
  • Running tests with --record-mode
  • SKILL.md covers Prerequisites, Important flags, Recording Cassettes and Parsing Cassettes, plus 1 more section
  • Runs Python scripts from its folder; calls uv and git

What it does

Testing Skill is an agent skill from pydantic/pydantic-ai, published by the product's own GitHub organization. Record, rewrite, and debug VCR cassettes for HTTP recordings. Use when running tests with --record-mode, verifying cassette playback, or inspecting request/response bodies in YAML cassettes.

Its SKILL.md is about 840 tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `parse_cassette.py`).

It sits in Testing & QA. The repository describes itself as: How Python does AI. Agents, realtime voice, image generation, embeddings. Every model, every interface, typed end to end. The licence is MIT.

When your agent uses it

  • Running tests with --record-mode
  • Verifying cassette playback
  • Inspecting request/response bodies in YAML cassettes

Example prompts

  • “/testing-skill”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash(uv run pytest *), Bash(uv run python .claude/skills/testing-skill/parse_cassette.py *), Bash(source .env && uv run pytest *), Bash(git diff *)

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Record cassettes
  2. Verify recordings
  3. Review snapshots

What it can do on your machine

Read from SKILL.md and the folder at commit 402a2ed. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(uv run pytest *)
    • Bash(uv run python .claude/skills/testing-skill/parse_cassette.py *)
    • Bash(source .env && uv run pytest *)
    • Bash(git diff *)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing Skill loads about 839 tokens when it runs. Until then it costs about 51 tokens; SKILL.md has 198 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~839

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:4
    -skill/parse_cassette.py *), Bash(source .env && uv run pytest *), Bash(git diff *)
  • NoteMentions a .env fileSKILL.md:13
    - Verify `.env` exists: `test -f .env && echo 'ok' || echo 'missing'`
  • NoteMentions a .env fileSKILL.md:28
    source .env && uv run pytest path/to/test.py::test_function_name -v --tb=line --record-mode=rewrite
  • NoteMentions a .env fileSKILL.md:33
    source .env && uv run pytest path/to/test.py::test_one path/to/test.py::test_two -v --tb=line --record-mode=rewrite
  • NoteMentions a .env fileSKILL.md:40
    source .env && uv run pytest path/to/test.py::test_function_name -vv --tb=line
  • NoteMentions a .env fileSKILL.md:84
    source .env && uv run pytest tests/models/test_openai.py::test_chat_completion -v --tb=line --record-mode=rewrite
  • NoteMentions a .env fileSKILL.md:87
    source .env && uv run pytest tests/models/test_openai.py::test_chat_completion -vv --tb=line

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pydantic/pydantic-ai at commit 402a2ed, republished under its MIT licence (© pydantic). 198 words, ~839 tokens.

Download SKILL.mdSave it as .claude/skills/testing-skill/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
testing-skill
description
Record, rewrite, and debug VCR cassettes for HTTP recordings. Use when running tests with --record-mode, verifying cassette playback, or inspecting request/response bodies in YAML cassettes.
allowed-tools
Bash(uv run pytest *), Bash(uv run python .claude/skills/testing-skill/parse_cassette.py *), Bash(source .env && uv run pytest *), Bash(git diff *)

Pytest VCR Workflow

Use this skill when recording or re-recording VCR cassettes for tests, or when debugging cassette contents.

Prerequisites

  • Verify .env exists: test -f .env && echo 'ok' || echo 'missing'
  • Missing API keys will cause clear test errors at runtime

Important flags

  • --record-mode=rewrite : Record cassettes (works for both new and existing)
  • --lf : Run only the last failed tests
  • -vv : Verbose output
  • --tb=line : Short traceback output
  • -k="" : Run tests matching the given substring expression

Recording Cassettes

Step 1: Record cassettes
bash
source .env && uv run pytest path/to/test.py::test_function_name -v --tb=line --record-mode=rewrite

Multiple tests can be specified:

bash
source .env && uv run pytest path/to/test.py::test_one path/to/test.py::test_two -v --tb=line --record-mode=rewrite
Step 2: Verify recordings

Run the same tests WITHOUT --record-mode to verify cassettes play back correctly:

bash
source .env && uv run pytest path/to/test.py::test_function_name -vv --tb=line
Step 3: Review snapshots

If tests use snapshot() assertions:

  • The test run in Step 2 auto-fills snapshot content
  • Review the generated snapshot files to ensure they match expected output
  • You only review - don't manually write snapshot contents
  • Snapshots capture what the test actually produced, additional to explicit assertions

Parsing Cassettes

Parse VCR cassette YAML files to inspect request/response bodies without dealing with raw YAML.

Usage
bash
uv run python .claude/skills/testing-skill/parse_cassette.py <cassette_path> [--interaction N]
Examples
bash
# Parse all interactions in a cassette
uv run python .claude/skills/testing-skill/parse_cassette.py tests/models/cassettes/test_foo/test_bar.yaml

# Parse only interaction 1 (0-indexed)
uv run python .claude/skills/testing-skill/parse_cassette.py tests/models/cassettes/test_foo/test_bar.yaml --interaction 1
Output

For each interaction, shows:

  • Request: method, URI, parsed body (truncated base64)
  • Response: status code, parsed body (truncated base64)

Base64 strings longer than 100 chars are truncated for readability.

Full Workflow Example

bash
# 1. Record cassette
source .env && uv run pytest tests/models/test_openai.py::test_chat_completion -v --tb=line --record-mode=rewrite

# 2. Verify playback and fill snapshots
source .env && uv run pytest tests/models/test_openai.py::test_chat_completion -vv --tb=line

# 3. Review test code diffs (excludes cassettes)
git diff tests/ -- ':!**/cassettes/**'

# 4. List new/changed cassettes (name only - use parse_cassette.py to inspect)
git diff --name-only tests/ -- '**/cassettes/**'

# 5. Inspect cassette contents if needed
uv run python .claude/skills/testing-skill/parse_cassette.py tests/models/cassettes/test_openai/test_chat_completion.yaml

© pydantic, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/testing-skill of pydantic/pydantic-ai.

  • SKILL.md
  • parse_cassette.py

Open the folder on GitHubat commit 402a2ed

Compare with similar skills

Testing Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing Skill this skillpydantic/pydantic-ai20k—~839Automated safety check: NotesMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Diagnosing Bugsfossasia/eventyay-interpretation1.6k31 repos~2.1kAutomated safety check: PassApache-2.0
TDDfossasia/eventyay-interpretation1.6k28 repos~1.1kAutomated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Diagnosing Bugs

    fossasia/eventyay-interpretation

    Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 31 repos~2.1k tokens
    Testing & QAAuto-check passed
  • TDD

    fossasia/eventyay-interpretation

    Test-driven development. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 28 repos~1.1k tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • Context Driven Development

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…

    380 GitHub starsUsed in 8 repos~2.9k tokens
    Testing & QAAuto-check passed

More from pydantic/pydantic-ai

All 20 skills in this repo
  • Pydantic AI Harness

    pydantic/pydantic-ai

    Official

    Adds optional capabilities to Pydantic AI agents from pydantic-ai-harness, led by Code Mode, which runs many tool calls as one sandboxed Python script.

    20k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Building Pydantic AI Agents

    pydantic/pydantic-ai

    Official

    Build AI agents with Pydantic AI — tools, capabilities (including on-demand loading), workspaces, structured output, streaming, testing, and multi-agent patterns.

    20k GitHub stars~8k tokensUpdated today
    Auto-check passed
  • Complete Partial PR

    pydantic/pydantic-ai

    Official

    Evaluate and complete an issue or PR where the submitted patch fixes only a narrow symptom of the reported pain point.

    20k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Migrating Agno To Pydantic AI

    pydantic/pydantic-ai

    Official

    Migrate Python Agno applications to Pydantic AI and, only when needed, Pydantic AI Harness.

    20k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Official

    Migrate Python applications from the Claude Agent SDK to Pydantic AI and, only when needed, Pydantic AI Harness.

    20k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Official

    Migrates Python LangChain Deep Agents applications to Pydantic AI and Pydantic AI Harness while preserving the application's observed behavior.

    20k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Testing Skill

What does Testing Skill do?

Record, rewrite, and debug VCR cassettes for HTTP recordings. Testing Skill is an agent skill from pydantic/pydantic-ai, published by the product's own GitHub organization. Record, rewrite, and debug VCR cassettes for HTTP recordings.

When should I use Testing Skill?

Testing Skill fits situations like: running tests with --record-mode; verifying cassette playback; inspecting request/response bodies in YAML cassettes.

How do I install Testing Skill in Claude Code?

Run `npx skills add pydantic/pydantic-ai --skill testing-skill -a claude-code`. Or copy the skill folder (.claude/skills/testing-skill in pydantic/pydantic-ai) into .claude/skills/testing-skill in your project. Claude Code loads it when a task matches its description.

How do I install Testing Skill in Codex?

Run `npx skills add pydantic/pydantic-ai --skill testing-skill -a codex`. Or copy the skill folder (.claude/skills/testing-skill in pydantic/pydantic-ai) into .agents/skills/testing-skill in your project. Codex loads it when a task matches its description.

Can I use Testing Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pydantic/pydantic-ai --skill testing-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-skill, .gemini/skills/testing-skill, .github/skills/testing-skill and .opencode/skills/testing-skill in your project.

What does Testing Skill need to run?

Going by SKILL.md and its folder, Testing Skill needs Python for the scripts in its folder and the command-line tools its instructions call (uv and git). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(uv run pytest *), Bash(uv run python .claude/skills/testing-skill/parse_cassette.py *), Bash(source .env && uv run pytest *), Bash(git diff *).

Does Testing Skill access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Testing Skill safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Testing Skill use?

Testing Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing Skill use?

About 839 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing Skill?

Skills that share tags, products or a category with Testing Skill: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (fossasia/eventyay-interpretation, 1.6k stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing Skill?

pydantic (a GitHub organization, an official publisher) maintains it in pydantic/pydantic-ai, which has 20,465 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 7, 2026.

Source: pydantic/pydantic-ai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.