Official agent skill

Test Pyramid

by kubernetes-sigs in kubernetes-sigs/agent-sandbox

Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need…

OfficialApache-2.0Auto-check passedTesting & QA

Install Test Pyramid

skills CLI
$ npx skills add kubernetes-sigs/agent-sandbox --skill test-pyramid -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kubernetes-sigs/agent-sandbox test-pyramid --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kubernetes-sigs/agent-sandbox.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/test-pyramid .claude/skills/test-pyramid && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-pyramid
GitHub stars
4.2k
Token cost
~1.6k tokens
SKILL.md length
834 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need…

  • Works in 5 steps: Inventory → Characterize every E2E test → Classify → …
  • The user asks about test pyramid
  • SKILL.md covers Phase 1 — Inventory, Phase 2 — Characterize every…, Phase 3 — Classify and Phase 4 — Verify before…, plus 1 more section
  • Calls git

What it does

Test Pyramid is an agent skill from kubernetes-sigs/agent-sandbox, published by the product's own GitHub organization. Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need E2E coverage, and where coverage is duplicated. Use when the user asks about test pyramid, test rebalancing, "should this be an e2e or unit test", E2E-to-unit migration, or slow/flaky E2E suites that might shrink.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test strategy, Unit testing and End-to-end testing. It works with Kubernetes, Python and Git. The repository describes itself as: agent-sandbox enables easy management of isolated, stateful, singleton workloads, ideal for use cases like AI agent runtimes and reinforcement learning (RL). The licence is Apache-2.0.

When your agent uses it

  • The user asks about test pyramid
  • Test rebalancing
  • Should this be an e2e
  • E2E-to-unit migration

Example prompts

  • “should this be an e2e or unit test”
  • “/test-pyramid”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Inventory
  2. Characterize every E2E test
  3. Classify
  4. Verify before reporting
  5. Report (inline only; write no files)

What it can do on your machine

Read from SKILL.md and the folder at commit b668082. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Pyramid loads about 1.6k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 834 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kubernetes-sigs/agent-sandbox at commit b668082, republished under its Apache-2.0 licence (© kubernetes-sigs). 834 words, ~1,642 tokens.

Download SKILL.mdSave it as .claude/skills/test-pyramid/SKILL.md (or your agent's skills folder).
name
test-pyramid
description
Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need E2E coverage, and where coverage is duplicated. Use when the user asks about test pyramid, test rebalancing, "should this be an e2e or unit test", E2E-to-unit migration, or slow/flaky E2E suites that might shrink.

Test Pyramid Analysis

Goal: anything that can be tested as a unit test should be a unit test; only what genuinely needs a real cluster should be E2E. The end state is a pyramid — many fast unit tests, few E2E tests. This skill produces an evidence-backed rebalancing report; it does NOT move tests itself unless the user asks afterward.

Phase 1 — Inventory

Enumerate both layers and count actual test functions (not just files), so the report can show the pyramid shape numerically:

  • Unit tests: git ls-files '*_test.go' | grep -v '^test/e2e/' plus git ls-files 'test/e2e/framework/*_test.go' and Python unit tests across all client packages (discover with git ls-files 'clients/**' | grep '/test[s]*/unit/' — covers clients/python/, clients/integrations/deepagents/, and clients/integrations/mcp-server/). Count func Test... per package (grep -c '^func Test') and def test_ for Python.
  • E2E / system tests: test/e2e/*_test.go (excluding test/e2e/framework/), plus test/e2e/extensions/ (density, python-runtime, rollout, and other cluster-physics E2E), test/e2e/clients/python/ (SDK E2E), dev/tools/test-migration.py (upgrade/rollback), and test/stress/ (load). Count func Test... (and def test_ for Python / migration tests), and for table-driven E2E, the sub-scenarios.
  • Note per-layer runtime cost if discoverable (CI job durations from dev/ci/, TestGrid tab names) — the payoff argument for each migration is time and flake surface removed from presubmit.

Phase 2 — Characterize every E2E test

Read each E2E test body (fan out parallel subagents over batches of 3-5 files for speed; each returns structured notes). For every test, record:

  1. What it arranges (objects applied, cluster preconditions).
  2. What it asserts — split assertions into:
    • Cluster-physics assertions: pod actually scheduled/running, kubelet behavior, image pulls, real networking/routing (sandbox-router paths), LoadBalancer/Gateway, RBAC enforcement, webhook admission via real API server, CRD conversion via real storage, controller<->controller timing, upgrade/rollback state survival.
    • Logic assertions: field values on objects after a reconcile, label/annotation stamping, status conditions, owner references, name hashing, defaulting, spec conversion, error classification, requeue decisions — anything a reconciler computes deterministically from inputs.
  3. The seam: which function/reconciler produces each logic assertion's value (e.g. isAdoptable, computeAndSetStatus, merge_flaky_by_test). If you cannot name the seam, you cannot claim a unit test can cover it.

Phase 3 — Classify

  • E2E → unit candidate: every logic assertion whose seam is reachable with the repo's existing unit patterns — table-driven tests with controller-runtime's fake client and newScheme(t) (see extensions/controllers/sandboxclaim_controller_test.go for the house style), direct function calls, or Python unittest.mock. A whole E2E test is a removal candidate only if ALL its assertions are logic assertions; otherwise propose extracting the logic assertions to unit tests and thinning the E2E to its cluster-physics core.
  • Keep as E2E: tests dominated by cluster-physics assertions. Do not propose unit-testing scheduling, pod readiness, real router traffic, webhook round-trips, or migration/upgrade behavior — a fake client proves nothing there.
  • Unit → E2E promotion (rare, keep the pyramid in mind): unit tests that mock so much they only test the mock (assert the seam is meaningfully exercised), or critical user journeys (create claim → adopt → route traffic → shutdown) with no E2E smoke path at all. Prefer ONE thin journey test over per-feature E2E.
  • Redundant coverage: the same seam asserted at both layers with the same inputs — keep the unit test, list the E2E assertion as prunable.
Show full SKILL.md (327 more words)Show less

Phase 4 — Verify before reporting

Every suggestion must survive these checks (drop or downgrade to "uncertain" if not):

  • Name the exact seam (file:line of the function) and confirm it is callable without a cluster — no unexported entanglement with live clients that fake clients can't satisfy.
  • Check a unit test for that seam doesn't already exist (use package-qualified symbol search across *_test.go and verify the test invokes the function); if it exists, the finding is "redundant E2E assertion", not "missing unit test".
  • Confirm the E2E test would still have a reason to exist after extraction, or explicitly state it can be deleted and what residual smoke coverage (if any) replaces it.
  • Beware behaviors that LOOK like logic but are cluster-coupled: informer cache timing, conversion-webhook storage effects, owner-reference garbage collection, Prow retest semantics, anything the migration test covers. When unsure, classify as keep-E2E and say why.

Phase 5 — Report (inline only; write no files)

  1. Pyramid snapshot: test-function counts per layer (unit / E2E / migration / stress) now vs. after adopting all suggestions, plus CI-time estimate if available.
  2. E2E → unit table: E2E test (file:line) | assertions to extract | seam (file:line) | proposed unit test location + shape (table-driven case to add vs. new test) | what remains of the E2E test (thinned / deleted).
  3. Unit → E2E table (expect this to be short — the pyramid demands it): gap | why unit level cannot cover it | proposed E2E home (existing file to extend before new file).
  4. Redundant coverage list.
  5. Top 5 quick wins ranked by (CI time + flake history removed) vs. effort — cross-reference flake evidence where available: open kind/flake issues, TestGrid history for the presubmit tabs, or dev/tools/flake-report output if that tool exists in your checkout. Migrating a flaky E2E assertion to a unit test is the highest-value move.

Rank suggestions by confidence; separate "verified" (seam confirmed, patterns exist) from "needs maintainer judgment" (cluster-coupling unclear). Do not inflate the E2E→unit list — a wrong migration that deletes real coverage is worse than a kept E2E test.

© kubernetes-sigs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/test-pyramid of kubernetes-sigs/agent-sandbox.

Open the folder on GitHubat commit b668082

Compare with similar skills

Test Pyramid next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Pyramid compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Pyramid this skillkubernetes-sigs/agent-sandbox4.2k—~1.6kAutomated safety check: PassApache-2.0
Testing Patternssoftspark/ai-toolkit179—~1.6kAutomated safety check: PassApache-2.0
Testing Test Strategydzhalaevd/Donatello135—~1.7kAutomated safety check: PassApache-2.0
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
Designing TestsCloudAI-X/claude-workflow-v21.4k1 repos~1.5kAutomated safety check: PassMIT
Studio Testingsupabase/supabase111k—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Testing Patterns

    softspark/ai-toolkit

    Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.

    179 GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Testing Test Strategy

    dzhalaevd/Donatello

    Use as the lead skill when choosing what backend tests to write, reviewing test plans, or balancing unit, integration, API, concurrency, and E2E coverage

    135 GitHub stars~1.7k tokensUpdated 5 days ago
    Testing & QAAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed
  • Studio Testing

    supabase/supabase

    Official

    Testing strategy for Supabase Studio. An agent skill from supabase/supabase.

    111k GitHub stars~2.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Python Testing

    affaan-m/ECC

    Python testing strategies using pytest, TDD methodology, fixtures, mocking, parametrization, and coverage requirements.

    276k GitHub starsUsed in 6 repos~4.7k tokens
    Testing & QAAuto-check passed

More from kubernetes-sigs/agent-sandbox

  • Bump Go Version

    kubernetes-sigs/agent-sandbox

    Official

    Bumps the Go version to the latest release across go.mod, tools.mod, and Dockerfiles.

    4.2k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Fix Flakes

    kubernetes-sigs/agent-sandbox

    Official

    Diagnose and fix flaky tests tracked as open kind/flake issues in kubernetes-sigs/agent-sandbox — reproduce the flake, apply a minimal fix, and open a PR linking the issue.

    4.2k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • K8s Agent Sandbox MCP

    kubernetes-sigs/agent-sandbox

    Official

    An MCP server skill for managing Kubernetes sandboxes. An agent skill from kubernetes-sigs/agent-sandbox.

    4.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • K8s API Conventions

    kubernetes-sigs/agent-sandbox

    Official

    Guides the agent to follow Kubernetes API conventions for OSS standards.

    4.2k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Dev Rules

    kubernetes-sigs/agent-sandbox

    Official

    Enforces project-specific development rules and conventions.

    4.2k GitHub stars~447 tokensUpdated today
    Auto-check passed
  • Triage Issues

    kubernetes-sigs/agent-sandbox

    Official

    Triage open GitHub issues for kubernetes-sigs/agent-sandbox by mapping them to roadmap.md and assigning k8s priority labels + Kanban Priority (P0–P4) on Project

    4.2k GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Categories

Questions about Test Pyramid

What does Test Pyramid do?

Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need…. Test Pyramid is an agent skill from kubernetes-sigs/agent-sandbox, published by the product's own GitHub organization. Analyze the repo's unit and E2E tests and propose rebalancing toward a test pyramid — which E2E tests (or assertions inside them) can be covered by unit tests, which unit-level gaps genuinely need E2E coverage, and where coverage is duplicated.

When should I use Test Pyramid?

Test Pyramid fits situations like: the user asks about test pyramid; test rebalancing; should this be an e2e; E2E-to-unit migration.

How do I install Test Pyramid in Claude Code?

Run `npx skills add kubernetes-sigs/agent-sandbox --skill test-pyramid -a claude-code`. Or copy the skill folder (.agents/skills/test-pyramid in kubernetes-sigs/agent-sandbox) into .claude/skills/test-pyramid in your project. Claude Code loads it when a task matches its description.

How do I install Test Pyramid in Codex?

Run `npx skills add kubernetes-sigs/agent-sandbox --skill test-pyramid -a codex`. Or copy the skill folder (.agents/skills/test-pyramid in kubernetes-sigs/agent-sandbox) into .agents/skills/test-pyramid in your project. Codex loads it when a task matches its description.

Can I use Test Pyramid in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kubernetes-sigs/agent-sandbox --skill test-pyramid -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-pyramid, .gemini/skills/test-pyramid, .github/skills/test-pyramid and .opencode/skills/test-pyramid in your project.

What does Test Pyramid need to run?

Going by SKILL.md and its folder, Test Pyramid needs the command-line tools its instructions call (git). Our summary lists: Python 3.

Does Test Pyramid access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Pyramid safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Pyramid use?

Test Pyramid is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Pyramid use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Pyramid?

Skills that share tags, products or a category with Test Pyramid: Testing Patterns (softspark/ai-toolkit, 179 stars), Testing Test Strategy (dzhalaevd/Donatello, 135 stars), OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars) and Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Pyramid?

kubernetes-sigs (a GitHub organization, an official publisher) maintains it in kubernetes-sigs/agent-sandbox, which has 4,198 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: kubernetes-sigs/agent-sandbox on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.