Agent skill

Prompt Lab

by Mathews-Tom in Mathews-Tom/armory

LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites.

MITAuto-check passedAI & LLM Engineering

Install Prompt Lab

skills CLI
$ npx skills add Mathews-Tom/armory --skill prompt-lab -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory prompt-lab --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/prompt-lab .claude/skills/prompt-lab && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prompt-lab
GitHub stars
329
Token cost
~2.1k tokens
SKILL.md length
685 words
Files
6 (incl. references)
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites.

  • Works in 5 steps: Define Objective → Analyze Current Prompt → Generate Variants → …
  • : prompt engineering
  • SKILL.md covers Reference Files, Prerequisites, Workflow and Output Format, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Prompt Lab is an agent skill from Mathews-Tom/armory. LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites. Triggers on: "prompt engineering", "generate prompt variants", "A/B test prompts", "optimize prompt", "improve this prompt". NOT for SKILL.md files, use skill-evaluator.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `evals/cases.yaml`, `references/evaluation-metrics.md` and `references/failure-modes.md`).

It sits in AI & LLM Engineering, covering Prompt engineering, Quizzes and assessments and Test generation. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • : prompt engineering
  • Generate prompt variants
  • A/B test prompts
  • Optimize prompt

Example prompts

  • “prompt engineering”
  • “generate prompt variants”
  • “A/B test prompts”
  • “/prompt-lab”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Define Objective
  2. Analyze Current Prompt
  3. Generate Variants
  4. Design Evaluation
  5. Output

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prompt Lab loads about 2.1k tokens when it runs, and up to ~7.3k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 685 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 685 words, ~2,147 tokens.

Download SKILL.mdSave it as .claude/skills/prompt-lab/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
prompt-lab
description
LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites. Triggers on: "prompt engineering", "generate prompt variants", "A/B test prompts", "optimize prompt", "improve this prompt". NOT for SKILL.md files, use skill-evaluator.
metadata.version
1.1.1
metadata.category
development
metadata.tags
prompt-engineering, evaluation, few-shot, chain-of-thought
metadata.difficulty
intermediate
metadata.phase
build

Prompt Lab

Replaces trial-and-error prompt engineering with structured methodology: objective definition, current prompt analysis, variant generation (instruction clarity, example strategies, output format specification), evaluation rubric design, test case creation, and failure mode identification.

Reference Files

FileContentsLoad When
references/prompt-patterns.mdPrompt structure catalog: zero-shot, few-shot, CoT, persona, structured outputAlways
references/evaluation-metrics.mdQuality metrics (accuracy, format compliance, completeness), rubric designEvaluation needed
references/failure-modes.mdCommon prompt failure taxonomy, detection strategies, mitigationsFailure analysis requested
references/output-constraints.mdTechniques for constraining LLM output format, JSON mode, schema enforcementFormat control needed

Prerequisites

  • Clear objective: what should the prompt accomplish?
  • Target model (GPT-4, Claude, open-source) — prompting techniques vary by model
  • Current prompt (if improving) or task description (if creating)

Workflow

Phase 1: Define Objective
  1. Task specification — What should the LLM produce? Be specific: "Classify customer support tickets into 5 categories" not "Handle support tickets."
  2. Success criteria — How do you know the output is correct? Define measurable criteria before writing any prompt.
  3. Failure modes — What does a bad output look like? Missing information? Wrong format? Hallucinated content? Refusal to answer?
Phase 2: Analyze Current Prompt

If an existing prompt is provided:

  1. Structure assessment — Is the instruction clear? Are examples provided? Is the output format specified?
  2. Ambiguity detection — Where could the model misinterpret the instruction?
  3. Missing components — What's not specified that should be? (output format, tone, length constraints, edge case handling)
  4. Failure mode mapping — Which known failure patterns (see references/failure-modes.md) apply to this prompt?
Phase 3: Generate Variants

Create 2-4 prompt variants, each testing a different hypothesis:

Variant TypeHypothesisWhen to Use
Direct instructionClear instruction is sufficientSimple tasks, capable models
Few-shotExamples improve output consistencyPattern-following tasks
Chain-of-thoughtReasoning improves accuracyMulti-step logic, math, analysis
Persona/roleRole framing improves tone/expertiseDomain-specific tasks
Structured outputFormat specification prevents errorsJSON, CSV, specific templates

For each variant:

  • State the hypothesis (why this variant might work)
  • Identify the risk (what could go wrong)
  • Provide the complete prompt text
Phase 4: Design Evaluation
  1. Rubric — Define weighted criteria:

    CriterionWhat It MeasuresTypical Weight
    CorrectnessOutput matches expected answer30-50%
    Format complianceFollows specified structure15-25%
    CompletenessAll required elements present15-25%
    ConcisenessNo unnecessary content5-15%
    Tone/styleMatches requested voice5-10%
  2. Test cases — Minimum 5 cases covering:

    • Happy path (standard input)
    • Edge cases (unusual but valid input)
    • Adversarial cases (inputs designed to confuse)
    • Boundary cases (minimum/maximum input)
Phase 5: Output

Present variants, rubric, and test cases in a structured format ready for execution.

Show full SKILL.md (272 more words)Show less

Output Format

text
## Prompt Lab: {Task Name}

### Objective
{What the prompt should achieve — specific and measurable}

### Success Criteria
- [ ] {Criterion 1 — measurable}
- [ ] {Criterion 2 — measurable}

### Current Prompt Analysis
{If existing prompt provided}
- **Strengths:** {what works}
- **Weaknesses:** {what fails or is ambiguous}
- **Missing:** {what's not specified}

### Variants

#### Variant A: {Strategy Name}

{Complete prompt text}

text
**Hypothesis:** {Why this approach might work}
**Risk:** {What could go wrong}

#### Variant B: {Strategy Name}

{Complete prompt text}

text
**Hypothesis:** {Why this approach might work}
**Risk:** {What could go wrong}

#### Variant C: {Strategy Name}

{Complete prompt text}

text
**Hypothesis:** {Why this approach might work}
**Risk:** {What could go wrong}

### Evaluation Rubric

| Criterion | Weight | Scoring |
|-----------|--------|---------|
| {criterion} | {%} | {how to score: 0-3 scale or pass/fail} |

### Test Cases

| # | Input | Expected Output | Tests Criteria |
|---|-------|-----------------|---------------|
| 1 | {standard input} | {expected} | Correctness, Format |
| 2 | {edge case} | {expected} | Completeness |
| 3 | {adversarial} | {expected} | Robustness |

### Failure Modes to Monitor
- {Failure mode 1}: {detection method}
- {Failure mode 2}: {detection method}

### Recommended Next Steps
1. Run all variants against the test suite
2. Score using the rubric
3. Select the highest-scoring variant
4. Iterate on the winner with targeted improvements

Calibration Rules

  1. One variable per variant. Each variant should change ONE thing from the baseline. Changing instruction style AND examples AND format simultaneously makes results uninterpretable.
  2. Test before declaring success. A prompt that works on 3 examples may fail on the 4th. Minimum 5 diverse test cases before concluding a variant works.
  3. Failure modes are more valuable than successes. Understanding WHY a prompt fails guides improvement more than confirming it works.
  4. Model-specific optimization. A prompt optimized for GPT-4 may not work for Claude or Llama. Always note the target model.
  5. Simplest effective prompt wins. If a zero-shot prompt scores as well as a few-shot prompt, use the zero-shot. Fewer tokens = lower cost + latency.

Error Handling

ProblemResolution
No clear objectiveAsk the user to define what "good output" looks like with 2-3 examples.
Prompt is for a task LLMs are bad at (math, counting)Flag the limitation. Suggest tool-augmented approaches or pre/post-processing.
Too many variables to testFocus on the highest-impact variable first. Iterative refinement beats combinatorial testing.
No existing prompt to analyzeStart with the simplest possible prompt. The first variant IS the baseline.
Output format requirements are strictUse structured output mode (JSON mode, function calling) instead of prompt-only constraints.

When NOT to Use

Push back if:

  • The task doesn't need an LLM (deterministic rules, regex, SQL) — use the right tool
  • The user wants prompt execution, not design — this skill designs and evaluates, it doesn't run prompts
  • The prompt is for safety-critical decisions without human review — LLM output should not be the sole input

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/prompt-lab of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml
  • references/evaluation-metrics.md
  • references/failure-modes.md
  • references/output-constraints.md
  • references/prompt-patterns.md

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Prompt Lab next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prompt Lab compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prompt Lab this skillMathews-Tom/armory329—~2.1kAutomated safety check: PassMIT
Agent Prompt Quality Barmastra-ai/mastra29k—~2kAutomated safety check: PassCustom licence
Suede AI EvalJasonColapietro/suede-creator-skills127—~3.3kAutomated safety check: PassMIT
Prompt Regressionagentscope-ai/OpenJudge871—~2.8kAutomated safety check: PassApache-2.0
Review Agent Primitivesmicrosoft/Huabu157—~2.3kAutomated safety check: PassMIT
Prompt Engineer Toolkitalirezarezvani/claude-skills28k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Agent Prompt Quality Bar

    mastra-ai/mastra

    Universal quality bar and final audit rubric for any agent system prompt.

    29k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Suede AI Eval

    JasonColapietro/suede-creator-skills

    Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates.

    127 GitHub stars~3.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Prompt Regression

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

    871 GitHub stars~2.8k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Review the quality of an agent's tool descriptions, system/agent prompts, or SKILL.md files against current agent-engineering best practices.

    157 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineer Toolkit

    alirezarezvani/claude-skills

    Turns marketing prompts into tested, versioned production assets: A/B prompt evaluation against structured test cases, immutable prompt version history with diffs, ready-to-use marketing prompt…

    28k GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Prompt Governance

    alirezarezvani/claude-skills

    A skill your agent uses when managing prompts in production at scale: versioning prompts, running A/B tests on prompts, building prompt registries, preventing prompt regressions, or creating eval…

    28k GitHub stars~2.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    329 GitHub stars~4.6k tokensUpdated 4 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    329 GitHub stars~2.6k tokensUpdated 4 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    329 GitHub stars~2.8k tokensUpdated 4 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    329 GitHub stars~3.1k tokensUpdated 4 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    329 GitHub stars~4.9k tokensUpdated 4 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    329 GitHub stars~2.7k tokensUpdated 4 days ago
    Auto-check passed

Questions about Prompt Lab

What does Prompt Lab do?

LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites. Prompt Lab is an agent skill from Mathews-Tom/armory. LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites.

When should I use Prompt Lab?

Prompt Lab fits situations like: : prompt engineering; generate prompt variants; A/B test prompts; optimize prompt.

How do I install Prompt Lab in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill prompt-lab -a claude-code`. Or copy the skill folder (skills/prompt-lab in Mathews-Tom/armory) into .claude/skills/prompt-lab in your project. Claude Code loads it when a task matches its description.

How do I install Prompt Lab in Codex?

Run `npx skills add Mathews-Tom/armory --skill prompt-lab -a codex`. Or copy the skill folder (skills/prompt-lab in Mathews-Tom/armory) into .agents/skills/prompt-lab in your project. Codex loads it when a task matches its description.

Can I use Prompt Lab in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill prompt-lab -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-lab, .gemini/skills/prompt-lab, .github/skills/prompt-lab and .opencode/skills/prompt-lab in your project.

What does Prompt Lab need to run?

SKILL.md names no scripts, command-line tools or credentials: Prompt Lab is instructions for the agent only.

Does Prompt Lab access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prompt Lab safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prompt Lab use?

Prompt Lab is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prompt Lab use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.2k tokens, read only when the agent opens those files.

What are the alternatives to Prompt Lab?

Skills that share tags, products or a category with Prompt Lab: Agent Prompt Quality Bar (mastra-ai/mastra, 29k stars), Suede AI Eval (JasonColapietro/suede-creator-skills, 127 stars), Prompt Regression (agentscope-ai/OpenJudge, 871 stars) and Review Agent Primitives (microsoft/Huabu, 157 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prompt Lab?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 329 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.