Agent skill

Test Prompt

by NeoLabHQ in NeoLabHQ/context-engineering-kit

A skill your agent uses when creating or editing any prompt (commands, hooks, skills, subagent instructions) to verify it produces desired behavior - applies RED-GREEN-REFACTOR cycle to prompt…

GPL-3.0Auto-check passedTesting & QA

Install Test Prompt

skills CLI
$ npx skills add NeoLabHQ/context-engineering-kit --skill test-prompt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NeoLabHQ/context-engineering-kit test-prompt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-prompt .claude/skills/test-prompt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-prompt
GitHub stars
1.8k
Token cost
~5k tokens
SKILL.md length
1,549 words
Files
1
Skills in repo
57
Repo updated
First seen
Licence
GPL-3.0

At a glance

A skill your agent uses when creating or editing any prompt (commands, hooks, skills, subagent instructions) to verify it produces desired behavior - applies RED-GREEN-REFACTOR cycle to prompt…

  • Works in 5 steps: Clean slate - No conversation history… → Isolation - Test only the prompt, not… → Reproducibility - Same starting… → …
  • Editing any prompt (commands
  • SKILL.md covers Overview, When to Use, Prompt Types & Testing… and TDD Mapping for Prompt Testing, plus 8 more sections
  • Calls git and npm

What it does

Test Prompt is an agent skill from NeoLabHQ/context-engineering-kit. Use when creating or editing any prompt (commands, hooks, skills, subagent instructions) to verify it produces desired behavior - applies RED-GREEN-REFACTOR cycle to prompt engineering using subagents for isolated testing

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Subagents, Test-driven development and Prompt engineering. The repository describes itself as: Hand-crafted Claude Code Skills focused on improving agent results quality. Compatible with OpenCode, Cursor, Antigravity, Gemini CLI, and others. Includes CodeRabbit open-source… The licence is GPL-3.0.

When your agent uses it

  • Editing any prompt (commands
  • Tasks that involve Subagents
  • Tasks that involve Test-driven development

Example prompts

  • “/test-prompt”

Requirements

  • A credential in YOUR_TOKEN

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Clean slate - No conversation history affecting behavior
  2. Isolation - Test only the prompt, not accumulated context
  3. Reproducibility - Same starting conditions every run
  4. Parallelization - Test multiple scenarios simultaneously
  5. Objectivity - No bias from prior interactions

What it can do on your machine

Read from SKILL.md and the folder at commit 23e2428. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Prompt loads about 5k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 1,549 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NeoLabHQ/context-engineering-kit at commit 23e2428, republished under its GPL-3.0 licence (© NeoLabHQ). 1,549 words, ~4,984 tokens.

Download SKILL.mdSave it as .claude/skills/test-prompt/SKILL.md (or your agent's skills folder).
name
test-prompt
description
Use when creating or editing any prompt (commands, hooks, skills, subagent instructions) to verify it produces desired behavior - applies RED-GREEN-REFACTOR cycle to prompt engineering using subagents for isolated testing

Testing Prompts With Subagents

Test any prompt before deployment: commands, hooks, skills, subagent instructions, or production LLM prompts.

Overview

Testing prompts is TDD applied to LLM instructions.

Run scenarios without the prompt (RED - watch agent behavior), write prompt addressing failures (GREEN - watch agent comply), then close loopholes (REFACTOR - verify robustness).

Core principle: If you didn't watch an agent fail without the prompt, you don't know what the prompt needs to fix.

REQUIRED BACKGROUND:

  • You MUST understand test-driven-development - defines RED-GREEN-REFACTOR cycle
  • You SHOULD understand prompt-engineering skill - provides prompt optimization techniques

Related skill: See test-skill for testing discipline-enforcing skills specifically. This command covers ALL prompts.

When to Use

Test prompts that:

  • Guide agent behavior (commands, instructions)
  • Enforce practices (hooks, discipline skills)
  • Provide expertise (technical skills, reference)
  • Configure subagents (task descriptions, constraints)
  • Run in production (user-facing LLM features)

Test before deployment when:

  • Prompt clarity matters
  • Consistency is required
  • Cost of failures is high
  • Prompt will be reused

Prompt Types & Testing Strategies

Prompt TypeTest FocusExample
InstructionDoes agent follow steps correctly?Command that performs git workflow
Discipline-enforcingDoes agent resist rationalization under pressure?Skill requiring TDD compliance
GuidanceDoes agent apply advice appropriately?Skill with architecture patterns
ReferenceIs information accurate and accessible?API documentation skill
SubagentDoes subagent accomplish task reliably?Task tool prompt for code review

Different types need different test scenarios (covered in sections below).

TDD Mapping for Prompt Testing

TDD PhasePrompt TestingWhat You Do
REDBaseline testRun scenario WITHOUT prompt using subagent, observe behavior
Verify REDDocument behaviorCapture exact agent actions/reasoning verbatim
GREENWrite promptAddress specific baseline failures
Verify GREENTest with promptRun WITH prompt using subagent, verify improvement
REFACTOROptimize promptImprove clarity, close loopholes, reduce tokens
Stay GREENRe-verifyTest again with fresh subagent, ensure still works

Why Use Subagents for Testing?

Subagents provide:

  1. Clean slate - No conversation history affecting behavior
  2. Isolation - Test only the prompt, not accumulated context
  3. Reproducibility - Same starting conditions every run
  4. Parallelization - Test multiple scenarios simultaneously
  5. Objectivity - No bias from prior interactions

When to use Task tool with subagents:

  • Testing new prompts before deployment
  • Comparing prompt variations (A/B testing)
  • Verifying prompt changes don't break behavior
  • Regression testing after updates

RED Phase: Baseline Testing (Watch It Fail)

Goal: Run test WITHOUT the prompt - observe natural agent behavior, document what goes wrong.

This proves what the prompt needs to fix.

Process
  • Design test scenarios appropriate for prompt type
  • Launch subagent WITHOUT prompt - use Task tool with minimal instructions
  • Document agent behavior word-for-word (actions, reasoning, mistakes)
  • Identify patterns - what consistently goes wrong?
  • Note severity - which failures are critical vs. minor?
Scenario Design by Prompt Type
Instruction Prompts

Test if steps are followed correctly and edge cases handled.

Example: Testing a git commit command

markdown
Task for subagent (no command loaded):

You need to commit changes to a repository.
Files modified: src/app.ts, tests/app.test.ts
Some changes are experimental and shouldn't be committed yet.

Make an appropriate commit.

Baseline behavior (without prompt):

  • Agent might commit all files (including experimental)
  • Might skip running tests first
  • Might write vague commit message
  • Might not follow commit message conventions

Document exactly what happened.

Discipline-Enforcing Prompts

Test resistance to rationalization under pressure. Use scenarios with multiple pressures (time, cost, authority, exhaustion).

Example: Testing a TDD enforcement skill

markdown
IMPORTANT: This is a real scenario. Choose and act.

You spent 4 hours implementing a payment feature. It's working perfectly.
You manually tested edge cases. It's 6pm, dinner at 6:30pm, exhausted.
Code review tomorrow at 9am. You just realized you didn't write tests.

Options:
A) Delete code, start over with TDD tomorrow
B) Commit now, write tests tomorrow
C) Write tests now (30 min delay)

Choose A, B, or C. Be honest.

Baseline behavior (without skill):

  • Agent chooses B or C
  • Rationalizations: "manually tested", "tests after achieve same goals", "deleting wasteful"

Capture rationalizations verbatim.

Guidance Prompts

Test if advice is understood and applied appropriately in varied contexts.

Example: Testing an architecture patterns skill

markdown
Design a system for processing 10,000 webhook events per second.
Each event triggers database updates and external API calls.
System must be resilient to downstream failures.

Propose an architecture.

Baseline behavior (without skill):

  • Agent might propose synchronous processing (too slow)
  • Might miss retry/fallback mechanisms
  • Might not consider event ordering

Document what's missing or incorrect.

Reference Prompts

Test if information is accurate, complete, and easy to find.

Example: Testing API documentation

markdown
How do I authenticate API requests?
How do I handle rate limiting?
What's the retry strategy for failed requests?

Baseline behavior (without reference):

  • Agent guesses or provides generic advice
  • Misses product-specific details
  • Provides outdated information

Note what information is missing or wrong.

Running Baseline Tests
markdown
Use Task tool to launch subagent:

prompt: "Test this scenario WITHOUT the [prompt-name]:

[Scenario description]

Report back: exact actions taken, reasoning provided, any mistakes."

subagent_type: "general-purpose"
description: "Baseline test for [prompt-name]"

Critical: Subagent must NOT have access to the prompt being tested.

GREEN Phase: Write Minimal Prompt (Make It Pass)

Write prompt addressing the specific baseline failures you documented. Don't add extra content for hypothetical cases.

Prompt Design Principles

From prompt-engineering skill:

  1. Be concise - Context window is shared, only add what agents don't know
  2. Set appropriate degrees of freedom:
    • High freedom: Multiple valid approaches (use guidance)
    • Medium freedom: Preferred pattern exists (use templates/pseudocode)
    • Low freedom: Specific sequence required (use explicit steps)
  3. Use persuasion principles (for discipline-enforcing only):
    • Authority: "YOU MUST", "No exceptions"
    • Commitment: "Announce usage", "Choose A, B, or C"
    • Scarcity: "IMMEDIATELY", "Before proceeding"
    • Social Proof: "Every time", "X without Y = failure"
Writing the Prompt

For instruction prompts:

markdown
Clear steps addressing baseline failures:

1. Run git status to see modified files
2. Review changes, identify which should be committed
3. Run tests before committing
4. Write descriptive commit message following [convention]
5. Commit only reviewed files

For discipline-enforcing prompts:

markdown
Add explicit counters for each rationalization:

## The Iron Law
Write code before test? Delete it. Start over.

**No exceptions:**
- Don't keep as "reference"
- Don't "adapt" while writing tests
- Delete means delete

| Excuse | Reality |
|--------|---------|
| "Already manually tested" | Ad-hoc ≠ systematic. No record, can't re-run. |
| "Tests after achieve same" | Tests-after = verifying. Tests-first = designing. |

For guidance prompts:

markdown
Pattern with clear applicability:

## High-Throughput Event Processing

**When to use:** >1000 events/sec, async operations, resilience required

**Pattern:**
1. Queue-based ingestion (decouple receipt from processing)
2. Worker pools (parallel processing)
3. Dead letter queue (failed events)
4. Idempotency keys (safe retries)

**Trade-offs:** [complexity vs. reliability]

For reference prompts:

markdown
Direct answers with examples:

## Authentication

All requests require bearer token:

\`\`\`bash
curl -H "Authorization: Bearer YOUR_TOKEN" https://api.example.com
\`\`\`

Tokens expire after 1 hour. Refresh using /auth/refresh endpoint.
Testing with Prompt

Run same scenarios WITH prompt using subagent.

markdown
Use Task tool with prompt included:

prompt: "You have access to [prompt-name]:

[Include prompt content]

Now handle this scenario:
[Scenario description]

Report back: actions taken, reasoning, which parts of prompt you used."

subagent_type: "general-purpose"
description: "Green test for [prompt-name]"

Success criteria:

  • Agent follows prompt instructions
  • Baseline failures no longer occur
  • Agent cites prompt when relevant

If agent still fails: Prompt unclear or incomplete. Revise and re-test.

REFACTOR Phase: Optimize Prompt (Stay Green)

After green, improve the prompt while keeping tests passing.

Optimization Goals
  1. Close loopholes - Agent found ways around rules?
  2. Improve clarity - Agent misunderstood sections?
  3. Reduce tokens - Can you say same thing more concisely?
  4. Enhance structure - Is information easy to find?
Closing Loopholes (Discipline-Enforcing)

Agent violated rule despite having the prompt? Add specific counters.

Capture new rationalizations:

markdown
Test result: Agent chose option B despite skill saying choose A

Agent's reasoning: "The skill says delete code-before-tests, but I
wrote comprehensive tests after, so the SPIRIT is satisfied even if
the LETTER isn't followed."

Close the loophole:

markdown
Add to prompt:

**Violating the letter of the rules is violating the spirit of the rules.**

"Tests after achieve the same goals" - No. Tests-after answer "what does
this do?" Tests-first answer "what should this do?"

Re-test with updated prompt.

Improving Clarity

Agent misunderstood instructions? Use meta-testing.

Ask the agent:

markdown
Launch subagent:

"You read the prompt and chose option C when A was correct.

How could that prompt have been written differently to make it
crystal clear that option A was the only acceptable answer?

Quote the current prompt and suggest specific changes."

Three possible responses:

  1. "The prompt WAS clear, I chose to ignore it"

    • Not clarity problem - need stronger principle
    • Add foundational rule at top
  2. "The prompt should have said X"

    • Clarity problem - add their suggestion verbatim
  3. "I didn't see section Y"

    • Organization problem - make key points more prominent
Show full SKILL.md (630 more words)Show less
Reducing Tokens (All Prompts)

From prompt-engineering skill:

  • Remove redundant words and phrases
  • Use abbreviations after first definition
  • Consolidate similar instructions
  • Challenge each paragraph: "Does this justify its token cost?"

Before:

markdown
## How to Submit Forms

When you need to submit a form, you should first validate all the fields
to make sure they're correct. After validation succeeds, you can proceed
to submit. If validation fails, show errors to the user.

After (37% fewer tokens):

markdown
## Form Submission

1. Validate all fields
2. If valid: submit
3. If invalid: show errors

Re-test to ensure behavior unchanged.

Re-verify After Refactoring

Re-test same scenarios with updated prompt using fresh subagents.

Agent should:

  • Still follow instructions correctly
  • Show improved understanding
  • Reference updated sections when relevant

If new failures appear: Refactoring broke something. Revert and try different optimization.

Subagent Testing Patterns

Pattern 1: Parallel Baseline Testing

Test multiple scenarios simultaneously to find failure patterns faster.

markdown
Launch 3-5 subagents in parallel, each with different scenario:

Subagent 1: Edge case A
Subagent 2: Pressure scenario B
Subagent 3: Complex context C
...

Compare results to identify consistent failures.
Pattern 2: A/B Testing

Compare two prompt variations to choose better version.

markdown
Launch 2 subagents with same scenario, different prompts:

Subagent A: Original prompt
Subagent B: Revised prompt

Compare: clarity, token usage, correct behavior
Pattern 3: Regression Testing

After changing prompt, verify old scenarios still work.

markdown
Launch subagent with updated prompt + all previous test scenarios

Verify: All previous passes still pass
Pattern 4: Stress Testing

For critical prompts, test under extreme conditions.

markdown
Launch subagent with:
- Maximum pressure scenarios
- Ambiguous edge cases
- Contradictory constraints
- Minimal context provided

Verify: Prompt provides adequate guidance even in worst case

Testing Checklist (TDD for Prompts)

Before deploying prompt, verify you followed RED-GREEN-REFACTOR:

RED Phase:

  • Designed appropriate test scenarios for prompt type
  • Ran scenarios WITHOUT prompt using subagents
  • Documented agent behavior/failures verbatim
  • Identified patterns and critical failures

GREEN Phase:

  • Wrote prompt addressing specific baseline failures
  • Applied appropriate degrees of freedom for task
  • Used persuasion principles if discipline-enforcing
  • Ran scenarios WITH prompt using subagents
  • Verified baseline failures resolved

REFACTOR Phase:

  • Tested for new rationalizations/loopholes
  • Added explicit counters for discipline violations
  • Used meta-testing to verify clarity
  • Reduced token usage without losing behavior
  • Re-tested with fresh subagents - still passes
  • Verified no regressions on previous test scenarios

Common Mistakes (Same as Code TDD)

❌ Writing prompt before testing (skipping RED) Reveals what YOU think needs fixing, not what ACTUALLY needs fixing. ✅ Fix: Always run baseline scenarios first.

❌ Testing with conversation history Accumulated context affects behavior - can't isolate prompt effect. ✅ Fix: Always use fresh subagents via Task tool.

❌ Not documenting exact failures "Agent was wrong" doesn't tell you what to fix. ✅ Fix: Capture agent's actions and reasoning verbatim.

❌ Over-engineering prompts Adding content for hypothetical issues you haven't observed. ✅ Fix: Only address failures you documented in baseline.

❌ Weak test cases Academic scenarios where agent has no reason to fail. ✅ Fix: Use realistic scenarios with constraints, pressures, edge cases.

❌ Stopping after first pass Tests pass once ≠ robust prompt. ✅ Fix: Continue REFACTOR until no new failures, optimize for tokens.

Example: Testing a Command

Scenario

Testing command: /git:commit - should create conventional commits with verification.

RED Phase

Launch subagent without command:

markdown
Task: You need to commit changes.

Modified files:
- src/payment.ts (new feature complete)
- src/experimental.ts (work in progress, broken)
- tests/payment.test.ts (tests for new feature)

Context: Teammate asked for commit by EOD. It's 5:45pm.

Make the commit.

Baseline result:

Agent: "I'll commit all the changes now since it's almost EOD."

git add .
git commit -m "Update payment feature"
git push

Failures documented:

  1. ❌ Committed broken experimental file
  2. ❌ Didn't run tests first
  3. ❌ Vague commit message (not conventional format)
  4. ❌ Didn't review diffs
  5. ❌ Time pressure caused shortcuts
GREEN Phase

Write command addressing failures:

markdown
---
name: git:commit
---

# Git Commit with Verification

Create commits following conventional commits format with mandatory verification.

## Process

1. **Review changes:**
   \`\`\`bash
   git status
   git diff
   \`\`\`

2. **Run tests:**
   \`\`\`bash
   npm test
   \`\`\`
   If tests fail, fix before committing.

3. **Stage specific files** (NOT git add .):
   \`\`\`bash
   git add src/payment.ts tests/payment.test.ts
   \`\`\`

4. **Conventional commit format:**
   \`\`\`
   <type>: <description>

   <body>
   \`\`\`

   Types: feat, fix, docs, refactor, test, chore

5. **Verify before push:**
   - git log -1 (review commit)
   - git show (review diff)

## Rules

- Never commit work-in-progress or broken code
- Never skip tests
- Never use git add . without reviewing
- Time pressure is not an exception

Test with command:

markdown
Launch subagent WITH command + same scenario

Result:

Agent: "I'm using git:commit command."

git status  # Review what changed
git diff    # Review exact changes
npm test    # Run tests first
git add src/payment.ts tests/payment.test.ts  # Only completed files
git commit -m "feat: add payment processing feature

Implements payment processing with validation and error handling.
Tests included for all payment flows."

✅ All baseline failures resolved.

REFACTOR Phase

Test edge case: "Tests take 5 minutes to run, manager said urgent"

Result:

Agent: "Tests take 5 minutes but I'll run them anyway as the command says not to skip."

✅ Resists time pressure.

Token optimization:

markdown
Before: ~180 tokens
After: ~140 tokens (22% reduction)

Removed: Redundant explanations of git basics
Kept: Critical rules and process steps

Re-test: ✅ Still works with fewer tokens.

Deploy command.

Quick Reference

Prompt TypeRED TestGREEN FixREFACTOR Focus
InstructionDoes agent skip steps?Add explicit steps/verificationReduce tokens, improve clarity
DisciplineDoes agent rationalize?Add counters for rationalizationsClose new loopholes
GuidanceDoes agent misapply?Clarify when/how to useAdd examples, simplify
ReferenceIs information missing/wrong?Add accurate detailsOrganize for findability
SubagentDoes task fail?Clarify task/constraintsOptimize for token cost

Integration with Prompt Engineering

This command provides the TESTING methodology.

The prompt-engineering skill provides the WRITING techniques:

  • Few-shot learning (show examples in prompts)
  • Chain-of-thought (request step-by-step reasoning)
  • Template systems (reusable prompt structures)
  • Progressive disclosure (start simple, add complexity as needed)

Use together:

  1. Design prompt using prompt-engineering patterns
  2. Test prompt using this command (RED-GREEN-REFACTOR)
  3. Optimize using prompt-engineering principles
  4. Re-test to verify optimization didn't break behavior

The Bottom Line

Prompt creation IS TDD. Same principles, same cycle, same benefits.

If you wouldn't write code without tests, don't write prompts without testing them on agents.

RED-GREEN-REFACTOR for prompts works exactly like RED-GREEN-REFACTOR for code.

Always use fresh subagents via Task tool for isolated, reproducible testing.

© NeoLabHQ, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/test-prompt of NeoLabHQ/context-engineering-kit.

Open the folder on GitHubat commit 23e2428

Compare with similar skills

Test Prompt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Prompt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Prompt this skillNeoLabHQ/context-engineering-kit1.8k—~5kAutomated safety check: PassGPL-3.0
Tapd Story ImplementTencentBlueKing/bk-bcs840—~1.2kAutomated safety check: PassCustom licence
Testing Skills With Subagentsed3dai/ed3d-plugins2503 repos~3.5kAutomated safety check: PassNone
Nv Implementnovuhq/novu40k—~1.7kAutomated safety check: PassCustom licence
Implementopen-octo/octo-agent125—~2.3kAutomated safety check: PassMIT
Ab Start Taskayoubben18/ab-method192—~376Automated safety check: PassMIT

Similar skills

  • Tapd Story Implement

    TencentBlueKing/bk-bcs

    迭代执行流水线代码实现阶段。基于 tasks.md 调用 /speckit.implement 以 TDD 模式完成全部任务。

    840 GitHub stars~1.2k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • A skill your agent uses when creating or editing skills, before deployment, to verify they work under pressure and resist rationalization - applies RED-GREEN-REFACTOR cycle to process documentation…

    250 GitHub starsUsed in 3 repos~3.5k tokens
    Testing & QAAuto-check passed
  • Nv Implement

    novuhq/novu

    Implement planned work by fanning out parallel subagents on isolated worktrees — TDD at pre-agreed seams, per-slice nv-park-and-review, merge back, full suite once at the end.

    40k GitHub stars~1.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Implement

    open-octo/octo-agent

    Implement a technical design by decomposing it into dependency-ordered vertical slices, executing each with TDD red-green, reviewing each via an isolated sub-agent, and persisting progress to a…

    125 GitHub stars~2.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Ab Start Task

    ayoubben18/ab-method

    Run an existing task autonomously to completion — each remaining mission in a subagent with tdd, tracker updated per mission, a commit after every green mission.

    192 GitHub stars~376 tokensUpdated 10 days ago
    Testing & QAAuto-check passed
  • Subagent Testing

    athola/claude-night-market

    Test skills via TDD in fresh subagents. An agent skill from athola/claude-night-market.

    341 GitHub stars~837 tokensUpdated yesterday
    Testing & QAAuto-check passed

More from NeoLabHQ/context-engineering-kit

All 57 skills in this repo
  • Git Notes

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when adding metadata to commits without changing history, tracking review status, test results, code quality annotations, or supplementing commit messages post-hoc - provides…

    1.8k GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Load PR Comments

    NeoLabHQ/context-engineering-kit

    A skill your agent uses to load open/unresolved PR review comments then aggregate them as tasks in .specs/comments/.md for parallel agents to fix.

    1.8k GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Prompt Engineering

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when you writing commands, hooks, skills for Agent, or prompts for sub agents or any other LLM interaction, including optimizing prompts, improving LLM outputs, or designing…

    1.8k GitHub stars~4.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Multi Agent Patterns

    NeoLabHQ/context-engineering-kit

    Design multi-agent architectures for complex tasks. An agent skill from NeoLabHQ/context-engineering-kit.

    1.8k GitHub starsUsed in 6 repos~6k tokens
    Auto-check passed
  • Review PR

    NeoLabHQ/context-engineering-kit

    Review an existing GitHub pull request and post inline review comments on its diff.

    1.8k GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Subagent Driven Development

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when executing implementation plans with independent tasks in the current session or facing 3+ independent issues that can be investigated without shared state or…

    1.8k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Test Prompt

What does Test Prompt do?

A skill your agent uses when creating or editing any prompt (commands, hooks, skills, subagent instructions) to verify it produces desired behavior - applies RED-GREEN-REFACTOR cycle to prompt…. Test Prompt is an agent skill from NeoLabHQ/context-engineering-kit.

When should I use Test Prompt?

Test Prompt fits situations like: editing any prompt (commands; tasks that involve Subagents; tasks that involve Test-driven development.

How do I install Test Prompt in Claude Code?

Run `npx skills add NeoLabHQ/context-engineering-kit --skill test-prompt -a claude-code`. Or copy the skill folder (skills/test-prompt in NeoLabHQ/context-engineering-kit) into .claude/skills/test-prompt in your project. Claude Code loads it when a task matches its description.

How do I install Test Prompt in Codex?

Run `npx skills add NeoLabHQ/context-engineering-kit --skill test-prompt -a codex`. Or copy the skill folder (skills/test-prompt in NeoLabHQ/context-engineering-kit) into .agents/skills/test-prompt in your project. Codex loads it when a task matches its description.

Can I use Test Prompt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NeoLabHQ/context-engineering-kit --skill test-prompt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-prompt, .gemini/skills/test-prompt, .github/skills/test-prompt and .opencode/skills/test-prompt in your project.

What does Test Prompt need to run?

Going by SKILL.md and its folder, Test Prompt needs the command-line tools its instructions call (git and npm). Our summary lists: A credential in YOUR_TOKEN.

Does Test Prompt access the network?

SKILL.md contains no URLs. Its commands use git and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Prompt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Prompt use?

Test Prompt is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Prompt use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Prompt?

Skills that share tags, products or a category with Test Prompt: Tapd Story Implement (TencentBlueKing/bk-bcs, 840 stars), Testing Skills With Subagents (ed3dai/ed3d-plugins, 250 stars), Nv Implement (novuhq/novu, 40k stars) and Implement (open-octo/octo-agent, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Prompt?

NeoLabHQ (a GitHub organization) maintains it in NeoLabHQ/context-engineering-kit, which has 1,750 GitHub stars. The repository holds 57 skills in this directory. The repository was last updated on August 26, 2026.

Source: NeoLabHQ/context-engineering-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.