Agent skill

Skillforge

by tripleyak in tripleyak/SkillForge

A skill your agent uses when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use'…

MITAuto-check: notesAgent Workflows

Install Skillforge

skills CLI
$ npx skills add tripleyak/SkillForge --skill skillforge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tripleyak/SkillForge skillforge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skillforge
GitHub stars
905
Token cost
~2.3k tokens
SKILL.md length
883 words
Files
86 (incl. scripts, references, assets)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use'…

  • Auditing agent skills - the user says create a skill
  • SKILL.md covers Routing (Phase 0), Creation pipeline, Frontmatter and platform facts and Ecosystem maintenance, plus 4 more sections
  • Calls python3 and codex
  • Do I have a skill for X

What it does

Skillforge is an agent skill from tripleyak/SkillForge. Use when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use', asks whether a skill exists for a task, or wants to validate, test, evaluate, package, or health-check skills. Also use for skill ecosystem maintenance (duplicate detection, stale skills, trigger collisions) and advisor checkpoints.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 87 other files, including scripts, reference files and assets (for example `CONTEXT.md`, `README.md` and `SKILLFORGE_AUDIT.md`).

It sits in Agent Workflows, covering Skill authoring and LLM evaluation. The repository describes itself as: A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor… The licence is MIT.

When your agent uses it

  • Auditing agent skills - the user says create a skill
  • Do I have a skill for X
  • Improve the X skill
  • Which skill should I use

Example prompts

  • “create a skill”
  • “do I have a skill for X”
  • “improve the X skill”
  • “/skillforge”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, Bash, Write, Edit, Task

What it can do on your machine

Read from SKILL.md and the folder at commit 4fc8bb4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • Bash
    • Write
    • Edit
    • Task

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • codex

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skillforge loads about 2.3k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 105 tokens; SKILL.md has 883 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~26k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Glob, Grep, Bash, Write, Edit, Task

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tripleyak/SkillForge at commit 4fc8bb4, republished under its MIT licence (© tripleyak). 883 words, ~2,290 tokens.

Download SKILL.mdSave it as .claude/skills/skillforge/SKILL.md (or your agent's skills folder). This skill also uses 85 other files; get the full folder from GitHub.
name
skillforge
description
Use when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use', asks whether a skill exists for a task, or wants to validate, test, evaluate, package, or health-check skills. Also use for skill ecosystem maintenance (duplicate detection, stale skills, trigger collisions) and advisor checkpoints.
allowed-tools
Read, Glob, Grep, Bash, Write, Edit, Task
license
MIT
user-invocable
true
metadata.version
6.0.0
metadata.domains
meta-skill, skill-creation, skill-testing, orchestration, routing
metadata.type
orchestrator

SkillForge 6 - Skill Router, Creator & Ecosystem Maintainer

Routes any skill-related request to the right action (use, improve, create, compose), creates new skills through an evidence-driven pipeline, and maintains the health of the whole skill ecosystem. Core principle: skill quality is a property of behavior, not documents - a skill is done when a fresh agent demonstrably does better with it than without it.

Routing (Phase 0)

Always triage before creating anything:

bash
python3 scripts/discover_skills.py            # refresh index (auto-refreshes if >24h old)
python3 scripts/triage_skill_request.py "<the user's request>" --json
Triage resultAction
Strong match (existing skill)Recommend it; do not create a duplicate
Moderate matchOffer IMPROVE_EXISTING on the matched skill
Weak/no match + create intentProceed to creation pipeline
Multi-domainSuggest composing existing skills
AmbiguousAsk one clarifying question

Match bands are keyword-evidence heuristics, not calibrated probabilities - report them as "strong/moderate/weak match", never as percent confidence.

Creation pipeline

Run phases in order. Each phase's detailed procedure lives in its reference - read the reference when you reach the phase, not before.

0. Baseline gate (RED). Before designing anything, dispatch a fresh subagent (Task tool) on 1-2 representative target tasks WITHOUT the skill. Capture verbatim what it does wrong. If the baseline does not fail, stop - the skill is unnecessary. The failures become the skill's test cases and its description keywords. See references/testing-and-evals.md.

1. Analysis. Identify explicit, implicit, and discovered requirements. Apply the three load-bearing lenses - Inversion (what guarantees failure → anti-patterns), Pareto (which 20% of scope delivers 80% → cut the rest), Root Cause (is this the real problem?) - plus any others from references/multi-lens-framework.md that earn their tokens. Classify the failure type you are guarding against and match the guidance form to it (see the failure-form table in references/testing-and-evals.md). Choose instruction specificity with references/degrees-of-freedom.md. Decide scripts with references/script-integration-framework.md.

2. Specification. Write the spec using references/specification-template.md. Minimal tier (problem, requirements, decisions with WHY, success criteria, test scenarios) for most skills; full tier (temporal projection, obsolescence triggers, extension points) only for infrastructure skills. Never fill a section you cannot ground - omit it.

3. Generation in fresh context. Dispatch a subagent (Task tool) that receives ONLY the spec and the baseline failures - not the analysis transcript - to write SKILL.md and supporting files. Scaffold first: python3 scripts/init_skill.py <name> --path <skills-dir>. Description doctrine: trigger conditions only, third person, symptom keywords, never a workflow summary. Budget: SKILL.md under 1,500 words; move depth to references/; <details> tags save zero tokens for agents - do not use them.

4. Execution testing (GREEN). Re-run the baseline tasks WITH the skill via fresh subagents. Gate on behavioral delta: the with-skill runs must not exhibit the baseline failures. Then run the description-triggering check (positive and near-miss queries). Iterate description and body against observed failures, not hunches. For improvements to existing skills, use blind A/B judging. Full protocols: references/testing-and-evals.md.

5. Review = lint + one adversarial reviewer. Mechanical gates first:

bash
python3 scripts/validate_skill.py <skill-dir>     # structure, frontmatter, lint (pinned models, word budget, description shape)
python3 scripts/check_docs_safety.py <skill-dir>

Then one fresh-context subagent prompted to REFUTE the skill (find the case where it misleads, over-triggers, or fails its own scenarios), carrying the reviewer checklists in references/synthesis-protocol.md. Fix what it proves; ship what survives. Do not convene approval panels - same-model unanimity measures nothing.

6. Ship with evals. Every generated skill keeps its tests: an evals/ directory (trigger queries + behavioral scenarios + assertions) so future edits can be regression-tested with python3 scripts/run_skill_evals.py <skill-dir>. Iterate post-ship with references/iteration-guide.md.

Show full SKILL.md (355 more words)Show less

Frontmatter and platform facts

Write frontmatter against the current Claude Code field set (17 fields) documented in references/claude-code-frontmatter.md, which also covers hooks (hooks receive JSON on stdin, not env vars), context: fork/agent, $ARGUMENTS, and the agentskills.io portability limits (64-char name, 1024-char description) that validate_skill.py enforces. Never pin dated model IDs (claude-*-YYYYMMDD) - the validator rejects them.

Ecosystem maintenance

bash
python3 scripts/skillforge_doctor.py              # trigger collisions, duplicates, stale refs, token budgets, description lint
python3 scripts/compile_skill.py <dir> --target claude|codex|agentskills
python3 scripts/package_skill.py <dir> ./dist     # .skill zip, honors .skillignore
python3 scripts/mine_skill_friction.py --consent  # opt-in: mine local transcripts for skill friction

Use doctor output to drive IMPROVE_EXISTING work; use friction reports as advisor evidence.

Context Skill Advisor

Proactive suggestions are delivered through Claude Code hooks (SessionStart surfaces the queue; UserPromptSubmit scores checkpoints inline) - no daemon. Configure with python3 scripts/install_skillforge.py (interactive; hooks and Personal Context scanning are opt-in, never default). Manage the queue: python3 scripts/context_advisor.py list|use|snooze|dismiss. Suggestions are evidence-backed and never auto-invoke a skill.

Script inventory

ScriptPurpose
discover_skills.pyBuild/refresh the cross-runtime skill index
triage_skill_request.pyRoute input to use/improve/create/compose/clarify
validate_skill.pyFull structural + lint validation (quick_validate.py = fast subset)
run_skill_evals.pyRun a skill's evals/ regression suite
skillforge_doctor.pyEcosystem health report
init_skill.pyScaffold a new skill (with evals/)
compile_skill.pyCompile a skill for a target runtime
package_skill.pyPackage as .skill archive
mine_skill_friction.pyOpt-in transcript friction mining
context_advisor.py / install_skillforge.pyAdvisor queue and setup
check_docs_safety.pyUnsafe interpolation check

Script exit codes: 0 success, 1 failure, 2 usage/consent error, 10 validation failure, 11 verification/dependency failure.

Extension points: new lint checks in validate_skill.py; new doctor checks in skillforge_doctor.py; new compile targets in compile_skill.py; new lenses in references/multi-lens-framework.md.

Anti-patterns

AvoidInstead
Creating without a failing baselineRun the RED gate; no failure = no skill
Description that summarizes workflowTrigger conditions only - agents act on summaries and skip the body
Body "Triggers" sections as a mechanismOnly the frontmatter description drives invocation
Approval panels and self-scored gatesLint what is falsifiable; adversarially refute the rest
<details> blocks for "progressive disclosure"Separate reference files loaded on demand
Pinned dated model IDsFamily aliases or omit model:
Duplicating an existing skillPhase 0 triage first, always

Verification checklist

  • Baseline failure captured before writing (RED)
  • With-skill runs clear the baseline failures (GREEN)
  • Trigger check passes on positive and near-miss queries
  • validate_skill.py and check_docs_safety.py pass
  • Adversarial reviewer's proven issues fixed
  • evals/ shipped with the skill; run_skill_evals.py passes
  • SKILL.md under 1,500 words (wc -w)

© tripleyak, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 85 other files (scripts, references, assets) in the repository root of tripleyak/SkillForge.

  • SKILL.md
  • .gitignore
  • .skillignore
  • CONTEXT.md
  • LICENSE
  • README.md
  • SKILLFORGE_AUDIT.md
  • assets/images/01-title.png
  • assets/images/02-quality-gap.png
  • assets/images/03-quality-built-in.png
  • assets/images/04-four-phase-architecture.png
  • assets/images/05-phase1-thinking-lenses.png
  • assets/images/06-phases-2-3.png
  • assets/images/07-phase4-synthesis.png
  • assets/images/08-evolution-mandate.png
  • assets/images/09-core-principles.png
  • assets/images/10-agentic-capabilities.png
  • assets/images/11-directory-structure.png
  • assets/images/12-installation.png
  • … and 67 more

Open the folder on GitHubat commit 4fc8bb4

Compare with similar skills

Skillforge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skillforge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skillforge this skilltripleyak/SkillForge905—~2.3kAutomated safety check: NotesMIT
Create Skilldavekilleen/Dex493—~2.9kAutomated safety check: PassCustom licence
Author Skillericrisco/rsc-harness156—~4.3kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79489 repos~8.2kAutomated safety check: PassApache-2.0
Zach Seller Skill Creatorzach22-1999/amazon-skills2061 repos~3.9kAutomated safety check: PassApache-2.0
Skill CreatorAgentTeam-TaichuAI/ScienceClaw670—~10kAutomated safety check: PassApache-2.0

Similar skills

  • Create Skill

    davekilleen/Dex

    Author a new Dex skill that actually fires and passes the quality bar.

    493 GitHub stars~2.9k tokensUpdated 5 days ago
    Agent WorkflowsAuto-check passed
  • Author Skill

    ericrisco/rsc-harness

    A skill your agent uses when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into…

    156 GitHub stars~4.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    794 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Zach Seller Skill Creator

    zach22-1999/amazon-skills

    亚马逊卖家专用的 skill 创建器(中文)。当用户想把一个亚马逊运营/自媒体/日常工作流程变成可复用的 skill 时使用。触发场景包括但不限于:用户说"我想做一个 skill""把这个流程变成 skill""帮我写个自动化""优化我已有的 skill""给这个工作流做个自动化",即使用户没用"skill"这个词,只要在描述"以后每次都这样做"的重复性工作时也应触发。本 skill…

    206 GitHub starsUsed in 1 repo~3.9k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    AgentTeam-TaichuAI/ScienceClaw

    Create new skills, modify and improve existing skills, and measure skill performance.

    670 GitHub stars~10k tokensUpdated 5 mo ago
    Agent WorkflowsAuto-check passed
  • Skill Creator

    luongnv89/asm

    Create a skill or bring an existing one up to the same standard (validate + asm eval fix loop); run evals, tune triggering.

    953 GitHub stars~5.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

Questions about Skillforge

What does Skillforge do?

A skill your agent uses when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use'…. Skillforge is an agent skill from tripleyak/SkillForge. Use when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use', asks whether a skill exists for a task, or wants to validate, test, evaluate, package, or health-check skills.

When should I use Skillforge?

Skillforge fits situations like: auditing agent skills - the user says create a skill; do I have a skill for X; improve the X skill; which skill should I use.

How do I install Skillforge in Claude Code?

Run `npx skills add tripleyak/SkillForge --skill skillforge -a claude-code`. Or copy the skill folder (the tripleyak/SkillForge repository) into .claude/skills/skillforge in your project. Claude Code loads it when a task matches its description.

How do I install Skillforge in Codex?

Run `npx skills add tripleyak/SkillForge --skill skillforge -a codex`. Or copy the skill folder (the tripleyak/SkillForge repository) into .agents/skills/skillforge in your project. Codex loads it when a task matches its description.

Can I use Skillforge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tripleyak/SkillForge --skill skillforge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skillforge, .gemini/skills/skillforge, .github/skills/skillforge and .opencode/skills/skillforge in your project.

What does Skillforge need to run?

Going by SKILL.md and its folder, Skillforge needs the command-line tools its instructions call (python3 and codex). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, Bash, Write, Edit, Task.

Does Skillforge access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skillforge safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skillforge use?

Skillforge is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skillforge use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 23k tokens, read only when the agent opens those files.

What are the alternatives to Skillforge?

Skills that share tags, products or a category with Skillforge: Create Skill (davekilleen/Dex, 493 stars), Author Skill (ericrisco/rsc-harness, 156 stars), Skill Creator (Azure/azqr, 794 stars) and Zach Seller Skill Creator (zach22-1999/amazon-skills, 206 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skillforge?

tripleyak (a GitHub user) maintains it in tripleyak/SkillForge, which has 905 GitHub stars. The repository was last updated on July 29, 2026.

Source: tripleyak/SkillForge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.