Agent skill

Prompt Engineer Toolkit

by borghei in borghei/Claude-Skills

Prompt engineering frameworks for building, testing, versioning, and evaluating prompts: chain-of-thought, few-shot, regression testing, and rubrics.

MITAuto-check passedAI & LLM Engineering

Install Prompt Engineer Toolkit

skills CLI
$ npx skills add borghei/Claude-Skills --skill prompt-engineer-toolkit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills prompt-engineer-toolkit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/prompt-engineer-toolkit .claude/skills/prompt-engineer-toolkit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prompt-engineer-toolkit
GitHub stars
886
Token cost
~1.6k tokens
SKILL.md length
657 words
Files
7 (incl. scripts, references)
Skills in repo
354
Repo updated
First seen
Licence
MIT

At a glance

Prompt engineering frameworks for building, testing, versioning, and evaluating prompts: chain-of-thought, few-shot, regression testing, and rubrics.

  • Designing production prompts
  • SKILL.md covers Core Capabilities, When to Use, Clarify First and Tools, plus 3 more sections
  • Runs Python scripts from its folder; calls python
  • Running A/B tests

What it does

Prompt Engineer Toolkit is an agent skill from borghei/Claude-Skills. Prompt engineering frameworks for building, testing, versioning, and evaluating prompts: chain-of-thought, few-shot, regression testing, and rubrics. Use when designing production prompts, running A/B tests, or building prompt libraries.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/failure-modes-and-workflows.md`, `references/prompt-patterns-catalog.md` and `references/testing-and-versioning.md`).

It sits in AI & LLM Engineering, covering Prompt engineering. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Designing production prompts
  • Running A/B tests
  • Building prompt libraries

Example prompts

  • “/prompt-engineer-toolkit”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prompt Engineer Toolkit loads about 1.6k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 657 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 657 words, ~1,613 tokens.

Download SKILL.mdSave it as .claude/skills/prompt-engineer-toolkit/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
prompt-engineer-toolkit
description
Prompt engineering frameworks for building, testing, versioning, and evaluating prompts: chain-of-thought, few-shot, regression testing, and rubrics. Use when designing production prompts, running A/B tests, or building prompt libraries.
license
MIT + Commons Clause
metadata.version
1.1.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
ai-engineering
metadata.tier
POWERFUL
metadata.updated
2026-06-17
metadata.frameworks
prompt-patterns, evaluation-rubrics, regression-testing, version-control

Prompt Engineer Toolkit — Production Prompt Engineering

The complete lifecycle for production prompts: design patterns that work, testing frameworks that catch regressions, versioning systems that track changes, and evaluation rubrics that replace subjective "looks good" with measurable quality. This treats prompts as production code with the same rigor — not clever tricks.

Tags: prompt engineering, chain-of-thought, few-shot, evaluation, testing, prompt versioning

Core Capabilities

  • Prompt patterns — 6-layer system-prompt architecture, chain-of-thought (standard, scratchpad, self-consistency), few-shot design + dynamic selection, JSON/section output structuring, decomposition pipelines, calibration (temperature + confidence levels).
  • Testing framework — test-case structure, suite composition (40/30/15/15), a 5-dimension automated scoring rubric with a weighted formula, and a regression-testing protocol.
  • Versioning — version-control layout, changelog format with rationale/baselines/rollback, and a prompt-diff risk checklist.
  • Failure-mode catalog — instruction override, format drift, sycophancy, verbosity, hallucination, anchoring, lost-in-the-middle, each with fixes.
  • Lifecycle workflows — design a prompt, debug a degraded prompt, migrate a prompt to a new model.

When to Use

  • Designing production prompts or building a prompt library.
  • Running A/B tests or regression tests on prompt variants.
  • Versioning prompts and gating changes on test scores.
  • Debugging a degraded prompt or migrating prompts across models.

Clarify First

Before designing or testing the prompt, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Task & definition of "good" — what the prompt must produce and how success is judged (drives the 5-dimension evaluation rubric)
  • Target model — calibration (temperature, few-shot count) and migration paths differ by model
  • Lifecycle stage — design new / debug a degraded prompt / migrate to a new model (selects the workflow)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Tools

ToolPurposeCommand
eval_scorer.pyScore evaluation results from JSON test cases (exact/contains/regex)python scripts/eval_scorer.py suite.json --fail-under 0.80 --json
prompt_analyzer.pyAnalyze prompt files for clarity, instruction density, few-shot coverage, tokenspython scripts/prompt_analyzer.py my_prompt.txt --json
prompt_diff.pyCompare two prompt versions for structural changes, instruction deltas, riskpython scripts/prompt_diff.py v2.txt v3.txt --show-diff --json

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

  • references/prompt-patterns-catalog.md — complete catalog of prompting techniques with examples: system-prompt architecture, chain-of-thought, few-shot, output structuring, decomposition, and calibration. Read when designing or structuring a prompt.
  • references/testing-and-versioning.md — test-case design, suite composition, the evaluation rubric and scoring formula, the regression protocol, version-control strategy, changelog format, and diff analysis. Read when building a test suite or managing versions.
  • references/failure-modes-and-workflows.md — common failure modes, the three lifecycle workflows, a quick-view integration table, the troubleshooting matrix, and success criteria. Read when debugging a prompt or running a workflow.
Show full SKILL.md (229 more words)Show less

Scope & Limitations

This skill covers:

  • Designing, structuring, and layering system prompts for production AI applications
  • Building and running test suites, evaluation rubrics, and regression tests for prompt quality
  • Versioning prompts with changelogs, baselines, and rollback plans
  • Calibration techniques including temperature tuning, confidence levels, and few-shot selection

This skill does NOT cover:

  • Fine-tuning or training models -- see engineering/model-training-pipeline for training workflows
  • Retrieval-augmented generation (RAG) pipeline design -- see engineering/context-engine for context retrieval architecture
  • Agent orchestration and multi-step tool use -- see engineering/agent-designer for agent system design
  • LLM infrastructure, hosting, or cost optimization -- see engineering/llm-gateway-design for inference infrastructure patterns

Integration Points

SkillIntegrationData Flow
agent-designerAgent system prompts are the highest-stakes prompts; use this toolkit to test and version themAgent specs → prompt layers → tested system prompts
self-improving-agentPrompt degradation signals feed into self-improvement loops for automatic correctionTest suite results → regression alerts → prompt iteration
context-engineRetrieved context quality directly impacts prompt effectiveness; coordinate retrieval tuning with prompt testingRetrieved chunks → prompt context layer → evaluation scores
ab-test-setupA/B test prompt variants in production with statistical rigor before full rolloutPrompt candidates → traffic split → scoring comparison → winner promotion
llm-gateway-designGateway handles prompt routing, versioning, and model fallback at the infrastructure layerVersioned prompts → gateway config → model routing → response logging
code-review-automationCode review prompts are high-frequency production prompts that benefit from this toolkit's testing frameworkReview criteria → prompt design → test suite → deployed reviewer prompt

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in engineering/prompt-engineer-toolkit of borghei/Claude-Skills.

  • SKILL.md
  • references/failure-modes-and-workflows.md
  • references/prompt-patterns-catalog.md
  • references/testing-and-versioning.md
  • scripts/eval_scorer.py
  • scripts/prompt_analyzer.py
  • scripts/prompt_diff.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Prompt Engineer Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prompt Engineer Toolkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prompt Engineer Toolkit this skillborghei/Claude-Skills886—~1.6kAutomated safety check: PassMIT
Agent Prompt Quality Barmastra-ai/mastra29k—~2kAutomated safety check: PassCustom licence
Prompt LabMathews-Tom/armory328—~2.1kAutomated safety check: PassMIT
Hospital Margin Forensicshh-health-AI/healthcare-equity101—~837Automated safety check: PassMIT
Module 1brevdev/workshop-build-an-agent146—~2.7kAutomated safety check: PassApache-2.0
Module 1brevdev/workshop-build-an-agent146—~2.7kAutomated safety check: PassApache-2.0

Similar skills

  • Agent Prompt Quality Bar

    mastra-ai/mastra

    Universal quality bar and final audit rubric for any agent system prompt.

    29k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Prompt Lab

    Mathews-Tom/armory

    LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites.

    328 GitHub stars~2.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Hospital Margin Forensics

    hh-health-AI/healthcare-equity

    This skill should be used when the user asks about "hospital margins", "cost reports", "HCRIS", "payer mix", "occupancy", "uncompensated care", "hospital capex capacity", "will hospitals buy", or…

    101 GitHub stars~837 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Module 1

    brevdev/workshop-build-an-agent

    This skill should be used when a learner is working through Module 1 ("Build an Agent") of the Build-an-Agent workshop and wants help understanding the concepts, notebooks, or code — e.g.

    146 GitHub stars~2.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Module 1

    brevdev/workshop-build-an-agent

    This skill should be used when a learner is working through Module 1 ("Build an Agent") of the Build-an-Agent workshop and wants help understanding the concepts, notebooks, or code — e.g.

    146 GitHub stars~2.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Dspy Bootstrap Fewshot

    OmidZamani/dspy-skills

    A skill your agent uses for BootstrapFewShot, bootstrapped demonstrations, teacher-model demos, and low-data DSPy prompt optimization.

    124 GitHub stars~1.3k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from borghei/Claude-Skills

All 354 skills in this repo
  • Agent Harness

    borghei/Claude-Skills

    Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.

    886 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    886 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    886 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    886 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    886 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    886 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Questions about Prompt Engineer Toolkit

What does Prompt Engineer Toolkit do?

Prompt engineering frameworks for building, testing, versioning, and evaluating prompts: chain-of-thought, few-shot, regression testing, and rubrics. Prompt Engineer Toolkit is an agent skill from borghei/Claude-Skills. Prompt engineering frameworks for building, testing, versioning, and evaluating prompts: chain-of-thought, few-shot, regression testing, and rubrics.

When should I use Prompt Engineer Toolkit?

Prompt Engineer Toolkit fits situations like: designing production prompts; running A/B tests; building prompt libraries.

How do I install Prompt Engineer Toolkit in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill prompt-engineer-toolkit -a claude-code`. Or copy the skill folder (engineering/prompt-engineer-toolkit in borghei/Claude-Skills) into .claude/skills/prompt-engineer-toolkit in your project. Claude Code loads it when a task matches its description.

How do I install Prompt Engineer Toolkit in Codex?

Run `npx skills add borghei/Claude-Skills --skill prompt-engineer-toolkit -a codex`. Or copy the skill folder (engineering/prompt-engineer-toolkit in borghei/Claude-Skills) into .agents/skills/prompt-engineer-toolkit in your project. Codex loads it when a task matches its description.

Can I use Prompt Engineer Toolkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill prompt-engineer-toolkit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-engineer-toolkit, .gemini/skills/prompt-engineer-toolkit, .github/skills/prompt-engineer-toolkit and .opencode/skills/prompt-engineer-toolkit in your project.

What does Prompt Engineer Toolkit need to run?

Going by SKILL.md and its folder, Prompt Engineer Toolkit needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Prompt Engineer Toolkit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prompt Engineer Toolkit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Prompt Engineer Toolkit use?

Prompt Engineer Toolkit is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prompt Engineer Toolkit use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.3k tokens, read only when the agent opens those files.

What are the alternatives to Prompt Engineer Toolkit?

Skills that share tags, products or a category with Prompt Engineer Toolkit: Agent Prompt Quality Bar (mastra-ai/mastra, 29k stars), Prompt Lab (Mathews-Tom/armory, 328 stars), Hospital Margin Forensics (hh-health-AI/healthcare-equity, 101 stars) and Module 1 (brevdev/workshop-build-an-agent, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prompt Engineer Toolkit?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 886 GitHub stars. The repository holds 354 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.