Agent skill

Ma Sandbox Test Plan

by michelangelo-ai in michelangelo-ai/michelangelo

Build, test, and verify a sandbox change across Go, JS, and Python.

Apache-2.0Auto-check passedTesting & QA

Install Ma Sandbox Test Plan

skills CLI
$ npx skills add michelangelo-ai/michelangelo --skill ma-sandbox-test-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install michelangelo-ai/michelangelo ma-sandbox-test-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/michelangelo-ai/michelangelo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ma-sandbox-test-plan .claude/skills/ma-sandbox-test-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ma-sandbox-test-plan
GitHub stars
118
Token cost
~2k tokens
SKILL.md length
933 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Build, test, and verify a sandbox change across Go, JS, and Python.

  • Works in 6 steps: Detect stack → Build + lint (gate) → Unit tests (gate) → …
  • The user asks to verify a change
  • SKILL.md covers Prerequisites, Stage 0: Detect stack, Stage 1: Build + lint (gate) and Stage 2: Unit tests (gate), plus 4 more sections
  • Calls git, bash and kubectl

What it does

Ma Sandbox Test Plan is an agent skill from michelangelo-ai/michelangelo. Build, test, and verify a sandbox change across Go, JS, and Python. Use when the user asks to verify a change or generate evidence that a feature works.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test generation. It works with Python, JavaScript and Model Context Protocol. The repository describes itself as: Michelangelo AI: Uber's end-to-end machine learning platform. The licence is Apache-2.0.

When your agent uses it

  • The user asks to verify a change
  • Generate evidence that a feature works

Example prompts

  • “/ma-sandbox-test-plan”

Requirements

  • Python 3
  • Node.js

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Detect stack
  2. Build + lint (gate)
  3. Unit tests (gate)
  4. Integration setup
  5. Integration scenarios
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 491a9b2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • bash
    • kubectl
    • go
    • yarn
    • poetry
    • claude
    • vitest
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, kubectl and yarn, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ma Sandbox Test Plan loads about 2k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 933 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from michelangelo-ai/michelangelo at commit 491a9b2, republished under its Apache-2.0 licence (© michelangelo-ai). 933 words, ~2,046 tokens.

Download SKILL.mdSave it as .claude/skills/ma-sandbox-test-plan/SKILL.md (or your agent's skills folder).
name
ma-sandbox-test-plan
description
Build, test, and verify a sandbox change across Go, JS, and Python. Use when the user asks to verify a change or generate evidence that a feature works.
argument-hint
<description of what changed>
user-invocable
true

A pipeline that gates on failure at every stage. The output of each stage IS the evidence — don't summarize it away. See .claude/skills/ma-sandbox-references/evidence-quality.md for what makes evidence good vs. hollow.

Prerequisites

Stage 4's JavaScript scenarios need a Playwright MCP server connected in this session. This repo doesn't ship one in .mcp.json — it's a per-user setup, not shared repo infra. If you don't have one connected yet, add it once with:

bash
claude mcp add playwright -- npx -y @playwright/mcp@latest

MCP servers only load at session startup, so this doesn't take effect in the session that's currently running the skill — end this session and start a new one, then re-invoke /ma-sandbox-test-plan. Check /mcp to confirm the connection before running this skill.

Stage 0: Detect stack

bash
CHANGED_FILES=$(git diff --name-only origin/main...HEAD)

Categorize by path prefix: go/ → Go, javascript/ → JS, python/ → Python. Mixed changes run every applicable stage below, once per stack, each gated independently — a Go build failure doesn't block the JS pipeline from also reporting its own result.

Stage 1: Build + lint (gate)

StackCommands
Gogo build ./cmd/<service>/... for each changed go/cmd/<service>/; go vet ./go/components/...
JSyarn typecheck, yarn lint
Pythonpoetry check

On any failure: stop that stack's pipeline, report the command and full error output, and do not proceed to Stage 2 for that stack. Other stacks (in a mixed change) still run independently.

Stage 2: Unit tests (gate)

Scope to the packages/dirs actually touched — derive this from Stage 0's $CHANGED_FILES directly, there's no fixed mapping table to maintain:

StackCommand
Gogo test ./<dir>/... for each unique directory containing a changed .go file under go/components/ or go/cmd/
JSyarn test <changed-dir-glob> (this repo runs vitest --run, which accepts path filters)
Pythonpytest <changed-dir>

On failure: stop, report the failing test output, do not proceed to Stage 3 for that stack.

Stage 3: Integration setup

Only reached for stacks whose Stage 1–2 passed.

Go backend:

bash
bash $(git rev-parse --show-toplevel)/.claude/skills/ma-sandbox-scripts/check-backend.sh
OutputAction
cleanNo local Go changes deployed — skip Go integration scenarios (Stage 4 Go section), note this in "Known gaps"
deployed <service>Enable debug logging before capturing any evidence: bash $(git rev-parse --show-toplevel)/.claude/skills/ma-sandbox-scripts/set_log_level.sh <service> debug
needs-deploy <service>Stop and respond: "You have local changes to <service> that aren't deployed. Run /ma-sandbox-deploy first, then re-run /ma-sandbox-test-plan." Do not proceed for this stack.

JavaScript UI:

bash
VITE_INFO=$(bash $(git rev-parse --show-toplevel)/.claude/skills/ma-sandbox-scripts/find-vite-port.sh)
VITE_PORT=$(echo "$VITE_INFO" | awk '{print $1}')
VITE_PID=$(echo "$VITE_INFO" | awk '{print $2}')
VITE_MODE=$(echo "$VITE_INFO" | awk '{print $3}')  # "existing" or "started"
  • Exit 0 → use $VITE_PORT in place of 5173 in all URLs below. Record $VITE_MODE/$VITE_PID for cleanup.
  • Non-zero exit → stop immediately and report the failure. Do not fall back to guessing a port via curl/lsof — a successful response doesn't confirm which branch is being served.
  • If the Go backend is also deployed (Stage 3 said deployed <service>), use the full sandbox at http://localhost:8090 instead of Vite, and skip straight to Stage 4.

Python:

bash
poetry run ma sandbox health

Confirm the services relevant to the change are running.

Stage 4: Integration scenarios

Before running scenarios, list the code paths identified from the diff and note any paths skipped — this list becomes the "Known gaps" input for Stage 5.

Go

For each controller/reconcile path touched by the diff:

  1. Apply a named test CR (test-<feature>-<scenario> — see .claude/skills/ma-sandbox-references/sandbox-data.md for demo-data conventions; prefer creating a new entity over disturbing shared demo state, per evidence principle 7).
  2. Capture logs scoped to that resource:
    bash
    bash $(git rev-parse --show-toplevel)/.claude/skills/ma-sandbox-scripts/capture_service_logs.sh <service> --resource <test-name>
  3. Verify state independently via kubectl get/kubectl describe — don't rely on logs alone.
  4. Show the full lifecycle: creation, intermediate phase transitions, terminal state, and (if the diff touches them) immutability enforcement or garbage collection. See .claude/skills/ma-sandbox-references/evidence-quality.md #8–9.
Show full SKILL.md (375 more words)Show less
JavaScript
  1. Playwright setup. Use whichever Playwright MCP server is connected in this session — check /mcp if unsure which name it registered under. None connected → stop and report the command from Prerequisites above; don't guess or fall back to a non-MCP browser tool. Clean up stale screenshots: rm -rf tmp/ma-sandbox-test-plan && mkdir -p tmp/ma-sandbox-test-plan (repo-root tmp/ is already gitignored — screenshots are run evidence, not .claude/ config, and shouldn't sit in a dot-prefixed directory Finder hides by default)

  2. Derive the route. Read .claude/skills/ma-sandbox-references/routes.md for the phase/entity keyword tables and resolution algorithm. Resolve phase and entity as independent signals, then combine. No matching row → stop and report, don't guess.

  3. Navigate and check preconditions.

    • browser_snapshot to confirm the page loaded (no error boundary, no blank state) — stop and report if it didn't.
    • Read .claude/skills/ma-sandbox-references/sandbox-data.md for what state each feature needs. If the required entity/state is missing, create it (prefer the UI over a full re-seed) and document every setup step taken.
    • Do not use yab/curl/direct API calls for setup — kubectl for inspection, UI or ma CLI for creation.
  4. Exercise the feature — click, fill, trigger the interactions the description calls for.

  5. Screenshot after each meaningful state: browser_take_screenshot → tmp/ma-sandbox-test-plan/{state-name}.png.

  6. Check console errors: browser_console_messages(level: "error"). Classify:

    ClassificationMatches
    BlockingUncaught, TypeError, ReferenceError, Failed to fetch, Cannot read properties, React render error
    Non-blockingdeprecated, 404 /favicon, CORS preflight, OPTIONS, sourceMap, net::ERR_ABORTED for known-missing mock data

    Anything matching neither list: treat as Blocking (unknown = assume blocking).

Python
  1. Run the relevant CLI commands or kubectl apply the relevant resources.
  2. Capture service logs (capture_service_logs.sh) and CLI stdout/stderr.
  3. Verify state via kubectl get.

Stage 5: Report

Print the integration evidence following .claude/skills/ma-sandbox-references/evidence-quality.md's guidance. Build/lint/unit-test results from stages 1–2 are gates, not evidence — don't include them (CI covers that).

Cleanup

  • If Vite was started by this run ($VITE_MODE = started): ask "I started a Vite dev server (PID $VITE_PID) for this session. Kill it now?" If yes: kill "$VITE_PID". If $VITE_MODE = existing, leave it alone.
  • If log level was changed in Stage 3: ask whether to revert — bash $(git rev-parse --show-toplevel)/.claude/skills/ma-sandbox-scripts/set_log_level.sh <service> info.
  • Test CRs left in the cluster: list them in the report. Don't auto-delete — the user may want to inspect them further or reuse them for a follow-up run.

© michelangelo-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/ma-sandbox-test-plan of michelangelo-ai/michelangelo.

Open the folder on GitHubat commit 491a9b2

Compare with similar skills

Ma Sandbox Test Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ma Sandbox Test Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ma Sandbox Test Plan this skillmichelangelo-ai/michelangelo118—~2kAutomated safety check: PassApache-2.0
Crap Analyzerswingerman/engineer154—~1.2kAutomated safety check: PassMIT
Zizkadb TestZIZKA-AI-SL/ZizkaDB129—~358Automated safety check: PassCustom licence
Approval Testing Toolkitlexler/skill-factory239—~1.2kAutomated safety check: PassApache-2.0
Anti Detect Browserantibrow/anti-detect-browser-skills17—~9.8kAutomated safety check: WarnMIT
Property Based Testingforyourhealth111-pixel/Vibe-Skills3.6k—~4.8kAutomated safety check: NotesApache-2.0

Similar skills

  • Crap Analyzer

    swingerman/engineer

    A skill your agent uses to produce a risk-based refactor + test plan for recently-changed code on a diff/branch/PR by computing CRAP (complexity × untested) on changed methods.

    154 GitHub stars~1.2k tokensUpdated 18 days ago
    Testing & QAAuto-check passed
  • Zizkadb Test

    ZIZKA-AI-SL/ZizkaDB

    Run the full ZizkaDB test suite across all layers — lint, Python unit tests, SDK tests, MCP tests, TypeScript tests, and dashboard build verification.

    129 GitHub stars~358 tokensUpdated today
    Testing & QAAuto-check passed
  • Approval Testing Toolkit

    lexler/skill-factory

    Writes snapshot-style approval tests in Python, JavaScript, TypeScript or Java, comparing output against an approved file instead of writing individual assertions.

    239 GitHub stars~1.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Anti Detect Browser

    antibrow/anti-detect-browser-skills

    Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…

    17 GitHub stars~9.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: warnings
  • Property Based Testing

    foryourhealth111-pixel/Vibe-Skills

    Property-based testing with fast-check (TypeScript/JavaScript) and Hypothesis (Python).

    3.6k GitHub stars~4.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes
  • Test Case Reducer

    ArabelaTso/Skills-4-SE

    Automatically reduces bug-triggering test cases to minimal form while preserving the failure.

    253 GitHub stars~2.5k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from michelangelo-ai/michelangelo

All 9 skills in this repo
  • Ma Sandbox Debug

    michelangelo-ai/michelangelo

    Tail logs, inspect pods, and diagnose unhealthy services in a running Michelangelo sandbox.

    118 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Ma Sandbox Deploy

    michelangelo-ai/michelangelo

    Build a Go service binary, package it into a Docker image, import into k3d, and deploy via helm sync.

    118 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Ma Sandbox Setup

    michelangelo-ai/michelangelo

    Canonical setup sequence for the Michelangelo local sandbox.

    118 GitHub stars~936 tokensUpdated yesterday
    Auto-check passed
  • Update Docs

    michelangelo-ai/michelangelo

    Update Michelangelo documentation. An agent skill from michelangelo-ai/michelangelo.

    118 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Ma Design Interview

    michelangelo-ai/michelangelo

    Structured interview for designing and implementing changes to the Michelangelo platform.

    118 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Ma Sandbox Reset

    michelangelo-ai/michelangelo

    Tear down the Michelangelo sandbox cluster and recreate it from scratch.

    118 GitHub stars~312 tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Ma Sandbox Test Plan

What does Ma Sandbox Test Plan do?

Build, test, and verify a sandbox change across Go, JS, and Python. Ma Sandbox Test Plan is an agent skill from michelangelo-ai/michelangelo. Build, test, and verify a sandbox change across Go, JS, and Python.

When should I use Ma Sandbox Test Plan?

Ma Sandbox Test Plan fits situations like: the user asks to verify a change; generate evidence that a feature works.

How do I install Ma Sandbox Test Plan in Claude Code?

Run `npx skills add michelangelo-ai/michelangelo --skill ma-sandbox-test-plan -a claude-code`. Or copy the skill folder (.claude/skills/ma-sandbox-test-plan in michelangelo-ai/michelangelo) into .claude/skills/ma-sandbox-test-plan in your project. Claude Code loads it when a task matches its description.

How do I install Ma Sandbox Test Plan in Codex?

Run `npx skills add michelangelo-ai/michelangelo --skill ma-sandbox-test-plan -a codex`. Or copy the skill folder (.claude/skills/ma-sandbox-test-plan in michelangelo-ai/michelangelo) into .agents/skills/ma-sandbox-test-plan in your project. Codex loads it when a task matches its description.

Can I use Ma Sandbox Test Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add michelangelo-ai/michelangelo --skill ma-sandbox-test-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ma-sandbox-test-plan, .gemini/skills/ma-sandbox-test-plan, .github/skills/ma-sandbox-test-plan and .opencode/skills/ma-sandbox-test-plan in your project.

What does Ma Sandbox Test Plan need to run?

Going by SKILL.md and its folder, Ma Sandbox Test Plan needs the command-line tools its instructions call (git, bash, kubectl, go, yarn and poetry). Our summary lists: Python 3; Node.js.

Does Ma Sandbox Test Plan access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ma Sandbox Test Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ma Sandbox Test Plan use?

Ma Sandbox Test Plan is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ma Sandbox Test Plan use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ma Sandbox Test Plan?

Skills that share tags, products or a category with Ma Sandbox Test Plan: Crap Analyzer (swingerman/engineer, 154 stars), Zizkadb Test (ZIZKA-AI-SL/ZizkaDB, 129 stars), Approval Testing Toolkit (lexler/skill-factory, 239 stars) and Anti Detect Browser (antibrow/anti-detect-browser-skills, 17 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ma Sandbox Test Plan?

michelangelo-ai (a GitHub organization) maintains it in michelangelo-ai/michelangelo, which has 118 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 10, 2026.

Source: michelangelo-ai/michelangelo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.