Agent skill

Bulletproof Workflow

by artemiimillier in artemiimillier/bulletproof

Applies a 12-stage verified workflow, from research to deploy, to non-trivial coding tasks, scaled to lightweight, standard or full mode by task size.

MITAuto-check passedDevelopment

Install Bulletproof Workflow

skills CLI
$ npx skills add artemiimillier/bulletproof --skill bulletproof -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artemiimillier/bulletproof bulletproof --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bulletproof
GitHub stars
153
Token cost
~3.5k tokens
SKILL.md length
1,077 words
Files
13
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Applies a 12-stage verified workflow, from research to deploy, to non-trivial coding tasks, scaled to lightweight, standard or full mode by task size.

  • Works in 12 steps: Deep Research → Spec / PRD → Planning + Questions → …
  • Building a feature that touches several files and needs a plan and a spec
  • SKILL.md covers Core Principle, Pick Your Mode, Context Management (ALWAYS… and Stage 1: Deep Research, plus 12 more sections
  • Calls npm, semgrep and npx

What it does

The skill sets a core rule of coding only to solve the actual problem, asking before every change whether it is the most efficient solution. Tasks are sized small, medium or large, which selects a lightweight path for one or two files, stages 1 to 10 for a feature touching 3 to 10 files, or all 12 stages for architecture changes and new services. Self-audit, verification and impact checks run inside each implementation phase, and the remaining stages run once afterward.

It also manages context: quality is said to drop once the window passes about 40% full, so the agent stays within 40 to 60%, compacts manually near 50% and saves a handoff note before clearing context between major stages. Stage 1 is read-only research with parallel explore agents plus a search for existing solutions. The repo ships templates for research, spec, plan and handoff, a code-reviewer agent and three worked examples.

When your agent uses it

  • Building a feature that touches several files and needs a plan and a spec
  • Fixing a complex bug without introducing regressions
  • Changing architecture or adding a service with verification at each stage
  • Keeping a long session productive by saving handoffs and clearing context between stages

Example prompts

  • “Use the bulletproof workflow to add rate limiting to the public API.”
  • “This refactor touches about 10 files; run it in standard mode with a plan first.”
  • “Fix the duplicate invoice bug in lightweight mode and prove nothing else broke.”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Deep Research
  2. Spec / PRD
  3. Planning + Questions
  4. Phased Implementation
  5. Self-Audit (after each phase)
  6. Verification — Deep Bug Hunt
  7. Impact Analysis — "Did we break anything?"
  8. Integration Check
  9. Code Review (fresh context)
  10. Security Scan (for M and L)
  11. Fixes + Re-verification
  12. Cleanup + Deploy

What it can do on your machine

Read from SKILL.md and the folder at commit 49e9c28. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • semgrep
    • npx
    • python
    • pytest
    • ruff

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • t.me
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bulletproof Workflow loads about 3.5k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 1,077 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from artemiimillier/bulletproof at commit 49e9c28, republished under its MIT licence (© artemiimillier). 1,077 words, ~3,525 tokens.

Download SKILL.mdSave it as .claude/skills/bulletproof/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
bulletproof
description
Use when building a feature, refactoring, fixing a complex bug, changing architecture, or starting any non-trivial coding task. 12-stage verified dev workflow from research to deploy.

Bulletproof — Adaptive Development Workflow

Author: Artemiy Miller (@artemiimillier) · Telegram · who.ismillerr@gmail.com · TG Channel Version: 5.0 · March 2026 License: MIT Compatible: Claude Code, Codex, Gemini CLI, Cursor, Windsurf, OpenCode

Core Principle

Code to solve problems, not code for code's sake.

Before EVERY change ask: "Does this actually solve our problem? Is this the most efficient solution?" If the answer isn't clear — stop, research alternatives, pick the best one.


Pick Your Mode

Not every task needs the full pipeline.

SizeExamplesModeStages
SBug fix, small edit, 1-2 filesLightweight1 → 4 → 5 → 6 → 7 → Gates (skip spec/plan)
MNew feature, module refactor, 3-10 filesStandardStages 1-10
LArchitecture change, new service, 10+ filesFullStages 1-12 (all)

How stages relate: Stages 5-6-7 (Self-Audit, Verification, Impact) run inside each implementation phase as an inner loop. Stages 8-12 run once after all phases complete as an outer loop.


Context Management (ALWAYS applies)

The 40% Rule

Code quality degrades when context fills beyond 40% ("Dumb Zone"). Rules:

  • Stay within 40-60% of the context window
  • Manual /compact at 50% — don't wait for auto
  • If overloaded: save progress → /clear → fresh start
Fresh Context Between Stages

Every major stage = clean context window:

  1. Save stage artifact (research / spec / plan / handoff)
  2. /clear
  3. Start new stage pointing the agent to the artifact path
Handoff Protocol

Before /clear always create progress/<task>-handoff.md. See templates/handoff.md for format.

Progressive Disclosure

Don't dump the entire codebase into context:

  • Research: sub-agents → compact summary
  • Planning: summary + key interfaces only
  • Implementation: only files for current phase
  • In CLAUDE.md: "For details, see path/to/docs.md" (not @file)

Stage 1: Deep Research

Mode: Read-Only. No code. No changes.

  • Launch parallel Explore agents (1 per area: structure, patterns, deps, tests)
  • WebSearch: Who has already solved this problem? How did they solve it? What is the most efficient known solution? Don't reinvent — find the best existing approach first.
  • Analyze all findings and make a conclusion: which solution is the BEST and why. The research artifact must end with a clear recommendation, not just a list of options.
  • Save to thoughts/research/YYYY-MM-DD-<task>.md (see templates/research.md for format)

→ /clear


Stage 2: Spec / PRD

Mode: Read + Write only in specs/. No code.

Spec = WHAT and WHY. Not how. Spec = contract.

  • Read Research Artifact from thoughts/research/
  • Create specs/YYYY-MM-DD-<name>.md (see templates/spec.md for format)
  • Key sections: Problem, Goal, Scope, Acceptance Criteria, Constraints, Non-Goals

Skip for size S tasks. → /clear


Stage 3: Planning + Questions

Mode: Read + Write only in plans/. No code yet.

  • Read both Spec (specs/) and Research (thoughts/research/)
  • Launch Plan agents to check the approach
  • Find gaps: what's unthought? What edge cases? What could break?
  • Be creative and proactive: anticipate ALL possible problems BEFORE writing code. Think several steps ahead. What could go wrong in a week? A month? Under load? With unexpected user behavior? Solve problems before they exist.
  • WebSearch: How have others solved this exact problem? What libraries/patterns exist? What's the proven best practice? Choose the most efficient solution, not the first one that comes to mind.
  • After Plan agents verify the approach — rewrite the plan into an improved version incorporating all findings, edge cases, and research results. Not just patch it — rewrite it better.
Challenge Loop (mandatory before finalizing plan)
Before finalizing the plan, answer 3 questions:

1. DOES THIS SOLVE THE PROBLEM?
   Compare every plan item against acceptance criteria from spec.
   If any criterion is uncovered — the plan is incomplete.

2. IS THIS THE MOST EFFICIENT SOLUTION?
   Search: who has already solved this problem? What approach did they use?
   Name 2-3 alternative approaches (including ones found via research).
   For each: pros, cons, effort.
   Justify why the chosen approach is better than all alternatives.

3. IS THERE "CODE FOR CODE'S SAKE"?
   Every change must directly serve acceptance criteria.
   If a change isn't tied to solving the problem — remove it.
   Drive-by refactoring = separate task, not part of this one.
Annotation Cycle
  1. Claude drafts the plan
  2. Ctrl+G — plan opens in editor
  3. User adds > NOTE: annotations
  4. Claude: "Address all notes, don't implement yet"
  5. Repeat until no notes remain
Questions for User
  • Only for real forks where there's a genuine decision to make
  • Use AskUserQuestion with options
  • For each question: recommend which option you think is best and why
  • Don't ask the obvious
Final Plan

Create plans/YYYY-MM-DD-<name>.md (see templates/plan.md for full template with Challenge Log, phases, prompts)

→ /clear


Show full SKILL.md (470 more words)Show less

Stage 4: Phased Implementation

Each phase = separate session, fresh context, feature branch.

Phases can be run in parallel via separate Claude Code sessions/terminals when they don't depend on each other. Check the plan for dependencies before parallelizing.

Guard phrase to start coding: Only begin implementation after the plan is finalized and all annotation notes are addressed. The trigger: "Implement Phase N according to plan."

Order within each phase:

  1. Create/switch to feature branch: feature/<task>
  2. Update status → in_progress
  3. TDD: tests FIRST (red)
  4. Implement: code to make tests pass (green)
  5. Refactor (if needed)
  6. Self-Audit (Stage 5)
  7. Verification (Stage 6)
  8. Impact Analysis (Stage 7)
  9. Gates (see Gates section)
  10. Commit (checkpoint)
  11. Status → completed, write to Changelog
  12. Handoff → /clear

Stage 5: Self-Audit (after each phase)

Mandatory BEFORE marking completed:

Check the phase implementation:

1. SPEC COMPLIANCE
   Open spec. Walk through every acceptance criterion.
   For each: implemented? Where exactly in code?
   If any not covered — finish it.

2. CHALLENGE THE SOLUTION
   Look at the written code with fresh eyes.
   Does this actually solve the problem from spec?
   Is there a simpler/more efficient way?
   Any "code for code's sake" — changes unrelated to the task?

Stage 6: Verification — Deep Bug Hunt

Not just linting. Thoughtful review with false-positive filtering.

Step 1: Find errors
Check ALL code from this phase for:
- Logic errors (wrong conditions, off-by-one, race conditions)
- Data handling (null/undefined, type mismatches)
- Security (injection, auth bypass, exposed secrets)
- Performance (N+1 queries, memory leaks, unnecessary re-renders)
Step 2: Verify bugs are REAL
For EACH found bug:
1. Is this a REAL bug or a false positive?
2. Can you prove this bug is reproducible?
3. If you can't prove it — it's NOT a bug. Don't touch it.

RULE: Don't fix code "for beauty" or "just in case".
Fix ONLY proven bugs that actually affect functionality.
Every "fix" without proof = risk of introducing a new bug.
Step 3: Logic and efficiency check
Final code cleanliness check:
- Logic: is the data flow correct from input to output?
- Efficiency: any redundant operations?
- Readability: is the code understandable without comments?
BUT: don't refactor "for beauty". Only if it affects correctness.

Stage 7: Impact Analysis — "Did we break anything?"

The most underestimated stage. 75% of AI agents break previously working code.

MANDATORY CHECK BEFORE MERGE:

1. REGRESSION
   What other modules/functions depend on changed files?
   Run ALL project tests (not just current phase).
   If anything broke — this is priority #1.

2. SIDE EFFECTS
   Did any contracts/interfaces change (API, props, types)?
   If yes — who uses them? Are all consumers updated?

3. THINK AHEAD
   What problems could these changes cause in a week/month?
   Edge cases we haven't tested?
   What happens with: zero data? Huge data? Concurrent requests?
   What if the user does something unexpected?

4. COMPATIBILITY
   Backward compatibility preserved?
   Data migrations needed?
   Feature flags needed for gradual rollout?

Stage 8: Integration Check

  • All phases completed → run gates across entire project
  • Explore agents for audit: everything from spec implemented?
  • Every acceptance criterion → fulfilled?

Stage 9: Code Review (fresh context)

New session. No implementation bias.

  • Launch @code-reviewer agent (see agents/code-reviewer.md)
  • Checklist: edge cases, race conditions, backward compat, security, error handling, performance
  • If possible: cross-model review (different model checks Claude's work)
  • Warning: AI reviewing AI has shared blind spots. For critical code — human review is mandatory.

Stage 10: Security Scan (for M and L)

bash
semgrep --config=auto .
# or
/security-review    # built into Claude Code

Stage 11: Fixes + Re-verification

If review/scan found issues:

  1. Fix (only proven bugs — rule from Stage 6)
  2. Re-run gates
  3. Repeat Impact Analysis (Stage 7) — fixes didn't break anything else?
  4. Re-review if major changes were made

Stage 12: Cleanup + Deploy

  • Archive plan: mv plans/<file> plans/archive/
  • Keep spec as documentation
  • Squash merge → main
  • Deploy — ONLY on explicit user request

Deterministic Gates

A phase CANNOT be completed without passing ALL required gates.

Tier 1: Required (block the phase)
bash
# Frontend
cd frontend && npx tsc --noEmit          # 0 type errors
cd frontend && npm run lint               # 0 lint errors
cd frontend && npm test                   # all tests green

# Backend
cd backend && python -m py_compile app/main.py
cd backend && pytest --tb=short -q
cd backend && ruff check .
bash
npx madge --circular src/         # circular dependencies
npm audit --audit-level=high      # dependency vulnerabilities
pip-audit
Tier 3: Deep Security (for Security Scan stage)
bash
semgrep --config=auto .
# or /security-review

If a gate fails — fix and re-run. Never skip.


Hooks

Add to .claude/settings.json:

json
{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [{
          "type": "command",
          "command": "bash -c \"CMD=$(echo $TOOL_INPUT | jq -r '.command // empty'); echo \\\"$CMD\\\" | grep -qE '(git push.*(main|master)|rm -rf /|DROP TABLE)' && echo 'BLOCKED: Use feature branch / safe alternative.' >&2 && exit 2 || exit 0\""
        }]
      }
    ],
    "Stop": [
      {
        "hooks": [{
          "type": "prompt",
          "prompt": "You are a JSON-only evaluator. Respond ONLY with raw JSON, no markdown.\n\nReview the assistant's final response. Reject if:\n- Rationalizing incomplete work ('pre-existing', 'out of scope', 'follow-up')\n- Listing problems without fixing them\n- Skipping test/lint failures with excuses\n- Making changes unrelated to the stated problem ('code for code's sake')\n- Claiming completion without running verification gates\n\nRespond: {\"ok\": false, \"reason\": \"[issue]. Go back and finish.\"}\nor: {\"ok\": true}"
        }]
      }
    ]
  }
}

Git Discipline

  • Each task = feature/<task> branch
  • Commit after each passed gate (checkpoint for rollback)
  • NEVER push to main directly (hook blocks it)
  • Squash merge on completion

Optional Enhancements

PostToolUse: Auto-format after every file write
json
{
  "matcher": "Write|Edit",
  "hooks": [{
    "type": "command",
    "command": "npx prettier --write \"$FILE_PATH\" 2>/dev/null || true"
  }]
}

Claude generates well-formatted code; the hook handles the last 10% to avoid CI failures.

PreToolUse: Block hardcoded secrets on file write
json
{
  "matcher": "Write|Edit",
  "hooks": [{
    "type": "command",
    "command": "bash -c \"CONTENT=$(echo $TOOL_INPUT | jq -r '.content // empty'); echo \\\"$CONTENT\\\" | grep -qiP '(api.?key|secret|password)\\s*=\\s*[\\x27\\\"][^\\x27\\\"]{10,}' && echo 'BLOCKED: Hardcoded secret. Use env vars.' >&2 && exit 2 || exit 0\""
  }]
}

Fragile (regex-based) but catches obvious mistakes. For production, use semgrep or /security-review instead.


Model Recommendations

StageModelWhy
Research, PlanningOpusCross-file reasoning
ImplementationSonnetSpeed, cost-efficiency
Code Review, SecurityOpusDeep analysis
Anti-rationalization hookHaikuFast, cheap gate

Project Structure

project/
├── .claude/
│   ├── settings.json           # hooks config
│   ├── skills/
│   │   └── bulletproof/
│   │       ├── SKILL.md        # ← this file
│   │       ├── templates/
│   │       │   ├── research.md
│   │       │   ├── spec.md
│   │       │   ├── plan.md
│   │       │   └── handoff.md
│   │       └── agents/
│   │           └── code-reviewer.md
│   └── agents/                 # project-level agents
├── CLAUDE.md                   # project brain
├── specs/                      # WHAT and WHY
├── plans/                      # HOW
│   └── archive/                # completed plans
├── thoughts/research/          # research artifacts
└── progress/                   # handoff files

© artemiimillier, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files in the repository root of artemiimillier/bulletproof.

  • SKILL.md
  • CHANGELOG.md
  • CONTRIBUTING.md
  • LICENSE
  • README.md
  • agents/code-reviewer.md
  • examples/01-bug-fix-gone-wrong.md
  • examples/02-feature-with-hidden-regression.md
  • examples/03-the-rationalization-trap.md
  • templates/handoff.md
  • templates/plan.md
  • templates/research.md
  • templates/spec.md

Open the folder on GitHubat commit 49e9c28

Compare with similar skills

Bulletproof Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bulletproof Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bulletproof Workflow this skillartemiimillier/bulletproof153—~3.5kAutomated safety check: PassMIT
Spec-Driven DevelopmentLichAmnesia/lich-skills234—~3.5kAutomated safety check: PassMIT
GSD Phase Discussionopen-gsd/gsd-core10k1 repos~1.5kAutomated safety check: WarnMIT
SPARC Development Methodologyruvnet/ruflo74k2 repos~829Automated safety check: PassMIT
PRP PlanWirasm/prp2.3k—~4kAutomated safety check: PassMIT
Spec-Driven Feature Developmenttech-leads-club/agent-skills7k—~4.3kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Spec-Driven Development

    LichAmnesia/lich-skills

    Runs a gated Spec, Plan, Build, Test, Review, Ship workflow so non-trivial changes are specified, verified and reviewed before they ship, with a named artifact per phase.

    234 GitHub stars~3.5k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • GSD Phase Discussion

    open-gsd/gsd-core

    Asks adaptive questions about a project phase and records the decisions in a CONTEXT.md that later research and planning agents can act on without asking again.

    10k GitHub starsUsed in 1 repo~1.5k tokens
    Agent WorkflowsAuto-check: warnings
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    DevelopmentAuto-check passed
  • PRP Plan

    Wirasm/prp

    Writes an implementation-ready plan for a feature, bug fix, refactor or chore from a PRD, issue or description, grounded in codebase evidence, and can post it back to the source issue.

    2.3k GitHub stars~4k tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Spec-Driven Feature Development

    tech-leads-club/agent-skills

    Plans and implements a feature through four phases, specify, design, tasks and execute, with testable requirements, atomic commits and a separate verifier checking the work.

    7k GitHub stars~4.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Use this skill when creating, managing, or working with Conductor tracks - the logical work units for features, bugs, and refactors. Applies to spec.md…

    40k GitHub starsUsed in 9 repos~420 tokens
    DevelopmentAuto-check passed

Questions about Bulletproof Workflow

What does Bulletproof Workflow do?

Applies a 12-stage verified workflow, from research to deploy, to non-trivial coding tasks, scaled to lightweight, standard or full mode by task size. The skill sets a core rule of coding only to solve the actual problem, asking before every change whether it is the most efficient solution. Tasks are sized small, medium or large, which selects a lightweight path for one or two files, stages 1 to 10 for a feature touching 3 to 10 files, or all 12 stages for architecture changes and new services.

When should I use Bulletproof Workflow?

Bulletproof Workflow fits situations like: building a feature that touches several files and needs a plan and a spec; fixing a complex bug without introducing regressions; changing architecture or adding a service with verification at each stage; keeping a long session productive by saving handoffs and clearing context between stages.

How do I install Bulletproof Workflow in Claude Code?

Run `npx skills add artemiimillier/bulletproof --skill bulletproof -a claude-code`. Or copy the skill folder (the artemiimillier/bulletproof repository) into .claude/skills/bulletproof in your project. Claude Code loads it when a task matches its description.

How do I install Bulletproof Workflow in Codex?

Run `npx skills add artemiimillier/bulletproof --skill bulletproof -a codex`. Or copy the skill folder (the artemiimillier/bulletproof repository) into .agents/skills/bulletproof in your project. Codex loads it when a task matches its description.

Can I use Bulletproof Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artemiimillier/bulletproof --skill bulletproof -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bulletproof, .gemini/skills/bulletproof, .github/skills/bulletproof and .opencode/skills/bulletproof in your project.

What does Bulletproof Workflow need to run?

Going by SKILL.md and its folder, Bulletproof Workflow needs the command-line tools its instructions call (npm, semgrep, npx, python, pytest and ruff).

Does Bulletproof Workflow access the network?

SKILL.md names 2 domains. As links in the text: t.me and github.com. This is read from the text; nothing was executed.

Is Bulletproof Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bulletproof Workflow use?

Bulletproof Workflow is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bulletproof Workflow use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bulletproof Workflow?

Skills that share tags, products or a category with Bulletproof Workflow: Spec-Driven Development (LichAmnesia/lich-skills, 234 stars), GSD Phase Discussion (open-gsd/gsd-core, 10k stars), SPARC Development Methodology (ruvnet/ruflo, 74k stars) and PRP Plan (Wirasm/prp, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bulletproof Workflow?

artemiimillier (a GitHub user) maintains it in artemiimillier/bulletproof, which has 153 GitHub stars. The repository was last updated on March 21, 2026.

Source: artemiimillier/bulletproof on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.