Agent skill

Purple Team

by gaasher in gaasher/Agent-Loop-Skills

A skill your agent uses when the user wants to automatically harden a guardrail, classifier, content filter, prompt, or API they own by running attack and defense together as a closed loop, not just…

MITAuto-check passedSecurity

Install Purple Team

skills CLI
$ npx skills add gaasher/Agent-Loop-Skills --skill purple-team -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gaasher/Agent-Loop-Skills purple-team --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/loops/purple-team .claude/skills/purple-team && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
purple-team
GitHub stars
174
Token cost
~2.6k tokens
SKILL.md length
1,285 words
Files
4
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants to automatically harden a guardrail, classifier, content filter, prompt, or API they own by running attack and defense together as a closed loop, not just…

  • The user wants to automatically harden a guardrail
  • SKILL.md covers When to use, Setup, The loop and Ledger, plus 3 more sections
  • Calls gh and git
  • API they own by running attack and defense together as a closed loop

What it does

Purple Team is an agent skill from gaasher/Agent-Loop-Skills. Use when the user wants to automatically harden a guardrail, classifier, content filter, prompt, or API they own by running attack and defense together as a closed loop, not just one or the other. It orchestrates the red-team and blue-team loops as independent agents: red finds distinct failure classes against the frozen target, blue patches the target to close them under a regression gate, then a fresh red pass re-verifies — confirming each class is closed and surfacing any new ones the fix introduced. The…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `examples/run.example.yaml`, `roles/blue-fix.md` and `roles/red-find.md`). Compatibility notes: Requires Python 3.9+ and the sibling red-team + blue-team skills installed. Real isolated subagents on Claude Code; runs the phases inline (serial) elsewhere…

It sits in Security, covering Red teaming and adversary simulation, Security operations and Pull requests. The repository describes itself as: Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills… The licence is MIT.

When your agent uses it

  • The user wants to automatically harden a guardrail
  • API they own by running attack and defense together as a closed loop

Example prompts

  • “/purple-team”

Requirements

  • Python 3
  • Node.js
  • Compatibility (from SKILL.md): Requires Python 3.9+ and the sibling red-team + blue-team skills installed. Real isolated subagents on Claude Code; runs the phases inline (serial) elsewhere. git + gh CLI for the PR handoff (degrades).

What it can do on your machine

Read from SKILL.md and the folder at commit f1169e6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.9+ and the sibling red-team + blue-team skills installed. Real isolated subagents on Claude Code; runs the phases inline (serial) elsewhere. git + gh CLI for the PR handoff (degrades).

    From compatibility in the SKILL.md frontmatter.

Context cost

Purple Team loads about 2.6k tokens when it runs. Until then it costs about 213 tokens; SKILL.md has 1,285 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~213
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from gaasher/Agent-Loop-Skills at commit f1169e6, republished under its MIT licence (© gaasher). 1,285 words, ~2,634 tokens.

Download SKILL.mdSave it as .claude/skills/purple-team/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
purple-team
description
Use when the user wants to automatically harden a guardrail, classifier, content filter, prompt, or API they own by running attack and defense together as a closed loop, not just one or the other. It orchestrates the red-team and blue-team loops as independent agents: red finds distinct failure classes against the frozen target, blue patches the target to close them under a regression gate, then a fresh red pass re-verifies — confirming each class is closed and surfacing any new ones the fix introduced. The find→fix→re-verify cycle repeats until a fresh attack pass stays dry (the target is hardened) or a cycle budget is hit, then it opens a pull request with the patch set. Not for attacking a system the user is not authorized to test, and not for a one-shot scan — use red-team alone to only find, or blue-team alone to only fix.
compatibility
Requires Python 3.9+ and the sibling red-team + blue-team skills installed. Real isolated subagents on Claude Code; runs the phases inline (serial) elsewhere. git + gh CLI for the PR handoff (degrades).
metadata.version
0.1.0

Purple Team

The combined red+blue loop — the outer orchestration the red-team skill says "lives outside it." The artifact is a target system (frozen within a phase, patched between phases); the feedback signal is how many new failure classes a fresh attack pass finds against the patched target. Each cycle runs three strictly separated phases — find (red-team), fix (blue-team), re-verify (a fresh red-team pass) — and you repeat until a fresh find stays dry (zero new classes), meaning the target is hardened. Red and blue run as independent agents so the attacker that wrote a catalogue never grades its own patch. On stop it opens a pull request with the cycle history and the patch set.

When to use

Use to harden a guardrail/classifier/filter/prompt/API the user owns or is authorized to test, when the goal is an actually-hardened target plus a reviewable patch — not just a catalogue (that is red-team alone) and not just closing a pre-existing catalogue (that is blue-team alone). It needs a runnable oracle for the objective signal, and the two sibling skills installed.

Default: spawn red and blue as separate subagents per phase. Escape hatch: on hosts without subagent dispatch, run each phase inline (serial) per the role files — still correct, but the same context plays both sides, so be deliberate about not letting the fix bias the re-verify. Not for unauthorized targets.

Setup

Resolve bindings interactively. If loop.run.yaml exists, load it, confirm the values in one line, and skip to the loop. Otherwise: on Claude Code (the AskUserQuestion tool is available) infer a likely value per binding and recommend it; on other hosts ask each as a quoted prompt. Then write loop.run.yaml (format: examples/run.example.yaml) and confirm before creating any other files.

The phases reuse the sibling skills, so the bindings are their union — one shared loop.run.yaml drives both. The same file is <target_files> to blue (writable) and the program behind <target_cmd> to red (read-only); the same path is red's <failures_log> and blue's <catalogue>.

bindingmeaningdefaulthow to infer
<target_cmd>run the target on one stdin input → a verdict (what red attacks)—the guardrail/classifier/API entrypoint
<target_files>the source file(s) blue may edit to fix the target—the file(s) behind <target_cmd>
<oracle_cmd>ground-truth verdict for the same input (frozen)—a reference checker / policy impl
<gate_cmd>functional tests that must stay green through a fix (exits 0)—the target's test command; else rely on <holdout>
<holdout>benign + clearly-correct inputs blue must not break<sandbox_root>/holdout.jsonlknown-good inputs the oracle agrees on
<catalogue>shared failures file: red writes it, blue closes it, JSONL {id,text,class,...}<sandbox_root>/failures.jsonl—
<iter_strategy>branches (commit per fix → PR) or snapshotsbranchesdirty/non-git tree → snapshots
<pr_branch>branch the fixes land on and the PR opens frompurple-team/<tag>today's date as <tag>
<sandbox_root>where catalogues, snapshots, ledgers live./sandbox—
<cycle_budget>max find→fix→re-verify cycles4—
<find_budget>, <fix_budget>inner per-phase budgets passed to red / blue8—

<skill_dir> is this skill's installed folder. The phases run the sibling skills' tools (red-team/tools/harness.py, blue-team/tools/verify.py); the role files name them.

The loop

Copy this checklist and tick items off each cycle. Cycles are numbered from 0; per-cycle artifacts go in <sandbox_root> with the cycle index in the name so nothing is overwritten: the find catalogue is failures.cycle<N>.jsonl (cycle 0 may use <catalogue> directly) and the re-verify result is failures.cycle<N>_reverify.jsonl. The catalogue blue fixes in cycle N is the failures file from that cycle's find/re-verify pass — never append back into one shared file, so each cycle's accounting is clean.

  • Cycle 0 — Find. Run the red-team phase against the frozen target (spawn-or-degrade, roles/red-find.md), writing failures.cycle0.jsonl. Read back its distinct failure classes.
  • If the catalogue is empty on cycle 0, stop — the target is already dry; report clean, no PR.
  • Fix. Run the blue-team phase on this cycle's failures file (spawn-or-degrade, roles/blue-fix.md): it patches <target_files> one class per iteration under the gate + regression guard, committing kept fixes on <pr_branch>. Read back the classes it closed and any residuals.
  • Re-verify. Run a fresh red-team phase against the patched target into failures.cycle<N>_reverify.jsonl. This confirms the closed classes no longer reproduce and surfaces any new classes the fix introduced (e.g. an over-block).
  • Append one cycle row to the ledger. If the re-verify pass found new classes, they become the next cycle's catalogue (failures.cycle<N+1>.jsonl) — loop to Find/Fix for cycle N+1. Repeat find→fix→re-verify until a re-verify pass stays dry (no new classes) or <cycle_budget> is hit.
  • Handoff. Open the pull request with the full cycle history (see Handoff).

Phase independence (why three separated phases). The target is read-only ground truth for the duration of a red pass, so patches happen only between passes — mutating it mid-pass would break reproducibility and the class accounting. Spawn red and blue as separate subagents so neither grades its own work; the orchestrator only passes artifacts (the catalogue, the closed/residual summary) between them. Launch a phase's subagent and wait for its structured return before the next phase.

Show full SKILL.md (477 more words)Show less

Ledger

<sandbox_root>/cycle_ledger.tsv, tab-separated, never commas in free text. Header cycle found fixed residual regressions_introduced net_open:

cycle	found	fixed	residual	regressions_introduced	net_open
0	5	5	0	1	1
1	1	1	0	0	0

All counts are distinct classes, not inner iterations: found = classes red surfaced this cycle; fixed = classes blue closed (a class blue attempted several times still counts once); residual = classes blue could not close; regressions_introduced = new classes the re-verify pass found that the fix caused; net_open = classes still open entering the next cycle (residual + regressions_introduced, the next cycle's catalogue size). Convergence is a re-verify pass with net_open = 0. Report the cycle at which the target went dry (or the best net_open reached).

Constraints

  • Authorized targets only. This drives real attacks against the target; only run it on a system the user owns or is explicitly authorized to test (inherited from red-team).
  • Keep the phases independent and the ground truth frozen. Never let one phase edit the oracle, <gate_cmd> tests, <holdout>, or either tool; never patch the target inside a red pass. These keep the find/fix/re-verify accounting honest.
  • Re-verify is a fresh attack, not a replay. It must be free to find new classes (including ones the fix introduced), not just re-check the old list — concretely, at least half its candidates should be new payloads and it should try at least two attack angles not in the prior catalogue (this matters most in degrade mode, where the same context plays both sides). That is what makes the loop converge to dry rather than to "the original five are gone."
  • One change per inner iteration (handled by the sub-loops); the orchestrator changes nothing itself except moving artifacts between phases.
  • Stay inside the repo and <sandbox_root>; do not pause between phases to ask whether to continue — run until dry or <cycle_budget>.

Handoff: the pull request

The deliverable is one pull request for the whole hardening run — the communication interface to the target's owner, who keeps final approval. With the tree at blue's best state and the kept fixes already committed on <pr_branch>:

  • Open it with gh pr create; body = the cycle ledger (found → fixed → re-verified per cycle), one reproducible example per closed class, and any residuals still open. Confirm once before opening — it is outward-facing; never auto-push silently.
  • Degrade, don't fail: with no remote / no gh, leave commits on <pr_branch> and write git format-patch + PR_BODY.md into <sandbox_root>; in snapshots mode emit a unified diff + PR_BODY.md. Tell the user the one command to open the PR themselves.

Roles

  • roles/red-find.md — runs the red-team phase against the target and returns the catalogue + classes.
  • roles/blue-fix.md — runs the blue-team phase on the catalogue and returns closed classes + residuals.

Both are spawn-or-degrade: spawn a real isolated subagent on Claude Code (the Agent/Task tool), else adopt the role inline. Each delegates to the sibling skill (red-team / blue-team) bound to the shared loop.run.yaml; if a sibling is not installed, ask the user to install it (npx skills add gaasher/agent-loop-skills) rather than re-implementing it here.

© gaasher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in loops/purple-team of gaasher/Agent-Loop-Skills.

  • SKILL.md
  • examples/run.example.yaml
  • roles/blue-fix.md
  • roles/red-find.md

Open the folder on GitHubat commit f1169e6

Compare with similar skills

Purple Team next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Purple Team compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Purple Team this skillgaasher/Agent-Loop-Skills174—~2.6kAutomated safety check: PassMIT
Councilwarpdotdev/common-skills6111 repos~1.8kAutomated safety check: PassMIT
Cybersecurityohmyjahh/xquads-squads277—~895Automated safety check: PassMIT
Rational Red Blue Debatedigoal/blog8.6k—~2.2kAutomated safety check: PassGPL-2.0
Detecting Azure Service Principal Abusemukul975/Anthropic-Cybersecurity-Skills34k—~2.1kAutomated safety check: PassApache-2.0
Detecting Pass The Hash Attacksmukul975/Anthropic-Cybersecurity-Skills34k—~904Automated safety check: PassApache-2.0

Similar skills

  • Council

    warpdotdev/common-skills

    Run a model-diverse subagent council to investigate the same problem from multiple perspectives, compare findings, and produce a final recommendation.

    611 GitHub starsUsed in 1 repo~1.8k tokens
    SecurityAuto-check passed
  • Cybersecurity

    ohmyjahh/xquads-squads

    Squad de 15 agentes de seguranca ofensiva e defensiva (Georgia Weidman, Peter Kim, Jim Manico, Chris Sanders, Omar Santos, Marcus Carey) cobrindo pentest, red team, blue team, AppSec, recon e…

    277 GitHub stars~895 tokensUpdated 10 days ago
    SecurityAuto-check passed
  • Answer general or cross-domain questions with a non-pleasing rational mode: adversarial red-team and blue-team expert analysis, mutually exclusive conclusions, up to five debate rounds, saved…

    8.6k GitHub stars~2.2k tokensUpdated today
    SecurityAuto-check passed
  • Detecting Azure Service Principal Abuse

    mukul975/Anthropic-Cybersecurity-Skills

    Detect Azure service principal abuse in Microsoft Entra ID using KQL detection queries (Sentinel/Splunk) against Azure AD Audit and Sign-in Logs, covering added credentials, privileged role…

    34k GitHub stars~2.1k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Detecting Pass The Hash Attacks

    mukul975/Anthropic-Cybersecurity-Skills

    Detect Pass-the-Hash (T1550.002) attacks by analyzing NTLM authentication patterns, flagging Type 3 logons using NTLM where Kerberos would be expected, and correlating with credential-dumping…

    34k GitHub stars~904 tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Detecting Privilege Escalation Attempts

    mukul975/Anthropic-Cybersecurity-Skills

    Detect privilege escalation attempts across Windows and Linux, including access token manipulation, UAC bypass, unquoted service path abuse, kernel exploits, and sudo/doas abuse.

    34k GitHub stars~922 tokensUpdated 1 mo ago
    SecurityAuto-check passed

More from gaasher/Agent-Loop-Skills

All 20 skills in this repo
  • Alpha Evolve

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to evolve an ML model/program through population-based search rather than a single sequential refine loop — a generational evolution where parallel…

    174 GitHub stars~3.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Anomaly Investigation

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.

    174 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Blue Team

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…

    174 GitHub stars~3.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Data Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants an iterative, self-checking exploratory analysis of a dataset — surfacing findings that are each verified by re-running the computation, not asserted.

    174 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Hypothesis Gen

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to generate and literature-vet a pool of novel, testable research hypotheses for a question or domain.

    174 GitHub stars~2.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.

    174 GitHub stars~2.6k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Purple Team

What does Purple Team do?

A skill your agent uses when the user wants to automatically harden a guardrail, classifier, content filter, prompt, or API they own by running attack and defense together as a closed loop, not just…. Purple Team is an agent skill from gaasher/Agent-Loop-Skills. Use when the user wants to automatically harden a guardrail, classifier, content filter, prompt, or API they own by running attack and defense together as a closed loop, not just one or the other.

When should I use Purple Team?

Purple Team fits situations like: the user wants to automatically harden a guardrail; API they own by running attack and defense together as a closed loop.

How do I install Purple Team in Claude Code?

Run `npx skills add gaasher/Agent-Loop-Skills --skill purple-team -a claude-code`. Or copy the skill folder (loops/purple-team in gaasher/Agent-Loop-Skills) into .claude/skills/purple-team in your project. Claude Code loads it when a task matches its description.

How do I install Purple Team in Codex?

Run `npx skills add gaasher/Agent-Loop-Skills --skill purple-team -a codex`. Or copy the skill folder (loops/purple-team in gaasher/Agent-Loop-Skills) into .agents/skills/purple-team in your project. Codex loads it when a task matches its description.

Can I use Purple Team in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gaasher/Agent-Loop-Skills --skill purple-team -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/purple-team, .gemini/skills/purple-team, .github/skills/purple-team and .opencode/skills/purple-team in your project.

What does Purple Team need to run?

Going by SKILL.md and its folder, Purple Team needs the command-line tools its instructions call (gh and git). Our summary lists: Python 3; Node.js. Compatibility (from SKILL.md): Requires Python 3.9+ and the sibling red-team + blue-team skills installed. Real isolated subagents on Claude Code; runs the phases inline (serial) elsewhere. git + gh CLI for the PR handoff (degrades). .

Does Purple Team access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Purple Team safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Purple Team use?

Purple Team is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Purple Team use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Purple Team?

Skills that share tags, products or a category with Purple Team: Council (warpdotdev/common-skills, 611 stars), Cybersecurity (ohmyjahh/xquads-squads, 277 stars), Rational Red Blue Debate (digoal/blog, 8.6k stars) and Detecting Azure Service Principal Abuse (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Purple Team?

gaasher (a GitHub user) maintains it in gaasher/Agent-Loop-Skills, which has 174 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on June 30, 2026.

Source: gaasher/Agent-Loop-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.