Agent skill

Feature Factory

by glebis in glebis/claude-skills

This skill should be used when taking a single software feature from intent to shipped as a solo developer — goal-first, TDD, deterministic verification, evidence only where it earns its keep, and…

MITAuto-check passedTesting & QA

Install Feature Factory

skills CLI
$ npx skills add glebis/claude-skills --skill feature-factory -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install glebis/claude-skills feature-factory --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/feature-factory .claude/skills/feature-factory && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
feature-factory
GitHub stars
391
Token cost
~3.4k tokens
SKILL.md length
1,723 words
Files
6 (incl. references, assets)
Skills in repo
92
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when taking a single software feature from intent to shipped as a solo developer — goal-first, TDD, deterministic verification, evidence only where it earns its keep, and…

  • Works in 6 steps: Intake — create or repair a Goal Contract → Size + risk triage — decide how much… → TDD implementation loop → …
  • The user says lets build feature X
  • SKILL.md covers Core principle, When to apply vs skip, The loop (six steps) and Semantic-preservation guard…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Feature Factory is an agent skill from glebis/claude-skills. This skill should be used when taking a single software feature from intent to shipped as a solo developer — goal-first, TDD, deterministic verification, evidence only where it earns its keep, and human judgment at the two moments that matter (goal approval, merge). Trigger when the user says "let's build feature X", "ship this feature", "run this through the factory", "write a goal contract", "feature-factory", or wants a disciplined intent→merge loop that resists process bloat. NOT for whole-product planning…

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files and assets (for example `assets/goal-contract-example.md`, `assets/goal-contract.md` and `references/process-budget.md`).

It sits in Testing & QA, covering Test-driven development. The repository describes itself as: Collection of Claude Code skills for enhanced AI workflows. The licence is MIT.

When your agent uses it

  • The user says lets build feature X
  • Ship this feature
  • Run this through the factory
  • Write a goal contract

Example prompts

  • “s build feature X”
  • “ship this feature”
  • “run this through the factory”
  • “/feature-factory”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Intake — create or repair a Goal Contract
  2. Size + risk triage — decide how much process
  3. TDD implementation loop
  4. Verification & audit discipline
  5. Evidence packaging
  6. Retro deletion hook (the curator)

What it can do on your machine

Read from SKILL.md and the folder at commit 3b88261. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Feature Factory loads about 3.4k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 148 tokens; SKILL.md has 1,723 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~148
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from glebis/claude-skills at commit 3b88261, republished under its MIT licence (© glebis). 1,723 words, ~3,350 tokens.

Download SKILL.mdSave it as .claude/skills/feature-factory/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
feature-factory
description
This skill should be used when taking a single software feature from intent to shipped as a solo developer — goal-first, TDD, deterministic verification, evidence only where it earns its keep, and human judgment at the two moments that matter (goal approval, merge). Trigger when the user says "let's build feature X", "ship this feature", "run this through the factory", "write a goal contract", "feature-factory", or wants a disciplined intent→merge loop that resists process bloat. NOT for whole-product planning, multi-feature roadmaps, or autonomous multi-agent swarms.

Feature Factory

A goal-driven, local-first loop for taking one feature from intent to shipped. The whole method exists to hold a single line in tension: don't let the process outrun the feature. Keep the spine (goal → TDD → deterministic verify → human merge → evidence-when-it-matters); delete ceremony aggressively.

This is a behavior guide, not an engine. Do not build generic config, executors, telemetry, optimizers, or a universal factory verify wrapper. Run the behavior; package nothing the feature didn't earn.

Core principle

The human defines the desired system state. Agents maintain the specifications. Tests and evidence decide whether reality complied. Two human gates are always required — approve the Goal Contract (cheap-to-change moment) and review the merge (irreversible moment) — plus a conditional third (plan approval) when a size/risk trigger fires (see step 2). Everything between is a single focused agent loop.

When to apply vs skip

  • Apply to a bounded, shippable feature.
  • Skip the heavy parts for trivial changes — an S-size fix is just: short goal in your head → TDD → run the repo's checks → merge. Don't generate documents for a one-liner.
  • Refuse XL — if the feature is multi-day with shared contracts, migrations, auth/billing, or product ambiguity, split it first; do not run an XL feature through this loop whole.

The loop (six steps)

1. Intake — create or repair a Goal Contract

Make a quick size call now (step 2 formalizes it) — you need it to decide how heavy intake should be. For S-size/trivial changes, skip the file — a short goal stated in chat and confirmed by the human is enough; jump to step 3. Otherwise, copy assets/goal-contract.md into the target repo as goal.md under a feature dir (suggested: docs/factory/<date>-<slug>/goal.md); see assets/goal-contract-example.md for a filled example of the calibration expected. The template's core fields are enough for M features — the conditional half is for L or when a trigger fires. Draft it from the request first, then have the human confirm/correct each field — don't block on a blank form, and don't proceed past intake until the human has approved the wording. Enforce:

  • All <!-- required --> fields present: Smallest shippable slice and Stop condition.
  • Respect the caps (≤3, ≤5). A capped goal stays a goal, not waterfall-in-markdown.
  • Every desired outcome maps to concrete evidence. Reject vague/solution-coupled outcomes.
  • Fail rule: if a goal can't produce evidence, it's a wish with better formatting — it doesn't pass.
  • Agents may propose Goal Amendments; never silently rewrite the goal.
2. Size + risk triage — decide how much process
  • Size: S (<½ day, no public API/migration) · M (1–2 days, some UI/integration) · L (multi-day; shared contracts, migrations, auth/billing/permissions, AI behavior, data retention) · XL (split first).
  • Risk: R0 none · R1 internal dev-assist · R2 user-facing low-stakes · R3 sensitive data/recommendations/profiling · R4 prohibited/high-risk (EU AI Act Art 5, or your jurisdiction's equivalent) or needs legal review → STOP, do not implement until externally reviewed. Also screen Art 50 labelling (or local equivalent) for AI-generated/chatbot/deepfake output.
  • Default to less process. Add plan approval, a tracker, visual evidence, or an audit (plan and/or diff — see step 4) only when size/risk triggers fire (see references/process-budget.md). When in doubt, do less.
  • Plan audit (when plan approval fires): before building, have an independent fresh-context reviewer (a separate agent/model, e.g. Codex) check the goal+plan against the actual codebase for ordering, architecture, and correctness flaws. This is the cheapest place to catch blockers — fix the plan, don't discover them mid-build. Same timeout + self-review fallback as step 4.
3. TDD implementation loop

Work on a feature branch or worktree — the merge gate is only a real decision point if the work isn't already on the mainline. Then red → green → refactor, in a single focused loop. No swarm, no parallel fan-out, no speculative abstraction, no silent scope expansion. Write the failing test first. Use any available TDD skill (e.g. superpowers:test-driven-development); otherwise just follow red → green → refactor directly.

  • No test harness in the repo? Bootstrap the stack's standard runner minimally (one config, one test dir — see references/stack-discovery.md); the harness is part of the feature's cost, and if bootstrapping it is a day of work, re-triage the size.
  • Spike escape hatch: if you don't yet know enough to write the failing test (unfamiliar library, unclear API behavior), timebox a throwaway spike, discard the spike code, then start red → green with what you learned. Don't fake a test, and don't let the spike quietly become the implementation.
  • Determinism: no wall-clock/sleep-based test assertions — use synchronous barriers/callbacks.
  • Contract change ⇒ verify all call-sites: changing a shared function's contract requires enumerating every caller/parallel path and proving each honors it.
4. Verification & audit discipline

Run the target repo's real verify commands (test · lint · typecheck · build, plus secrets-scan if available) — identical locally and in CI. If the repo has a factory verify / project verify command, call it; if not, use the repo's actual commands and record them (the exact commands + output) in evidence/verify.log under the feature dir. Do not invent a universal wrapper before the repo earns it. Not sure what the repo's real commands are? Discover them — CI workflows, Makefile/justfile, contributor docs, then ecosystem manifests, in that order — per references/stack-discovery.md.

  • Flake = failure, not retry. Any intermittent fail blocks merge until root-caused or rewritten deterministically. No quarantine.
  • Audit (independent fresh-context review) is a standard step, not an afterthought. A separate agent or model with fresh context (e.g. Codex, or a different model) reviews the work at two touchpoints: (a) plan audit before building (see step 2) and (b) diff audit before the merge gate. How it scales with size/risk — when it's required, when it's skippable, timeout + self-review fallback — is defined once in references/process-budget.md; follow that table rather than re-deriving it. Persist findings in evidence/audit-*.md; fold them back into the plan/diff before proceeding (an audit you don't act on is theatre). The human gates still decide — an audit informs them, it doesn't replace goal/merge approval.
5. Evidence packaging

Persist only relevant evidence under the feature dir's evidence/: verify.log (commands + output), and — only for qualifying UI changes — screenshots in evidence/screenshots/ (what qualifies, and the one-viewport default, is defined in references/process-budget.md under "Visual evidence"). Goal-traceability table only when it adds signal. Evidence is an artifact, not a claim: "done" must be auditable. Do not let the evidence folder become the product.

  • Escape hatch — statistical/behavioral claims: if a desired outcome is a measured effect on noisy real-world data (engagement, latency distributions, refusal rates, ML metrics) rather than a pass/fail test, a green suite does not prove it. Pre-register the metric in the Goal Contract and evaluate it as a real experiment (permutation-test it, adversarially review the analysis) — using the rigorous-experiments skill if available, otherwise a held-out check. Don't assert a measured outcome you didn't actually test — that's the same gamed-proxy failure the fail rule catches, one layer down.
Show full SKILL.md (619 more words)Show less
6. Retro deletion hook (the curator)

After shipping, write exactly four lines in retro.md under the feature dir (append to goal.md only when a separate file is impractical — one greppable default beats two conventions):

  1. What slowed shipping?
  2. What caught a real bug?
  3. Which artifact was never used?
  4. What gets deleted before the next feature?

This is the entire self-improvement mechanism at small N — manual, human-readable, impossible to over-build. Do not add usage telemetry, dashboards, or counters. (Aggregate into a markdown table only after ~5 features; consider anything heavier only after ~10.)

Semantic-preservation guard (when editing this method's own artifacts)

This guard applies when editing the method itself — the Goal Contract template, the risk rubric, or the verify checks — not during ordinary feature work. When any edit, optimizer, or rewrite touches those artifacts, do not let polished prose delete load-bearing constraints. Before accepting a rewrite, confirm it preserves: required fields, the ≤N caps, the fail rule, stop condition, smallest shippable slice, risk classification, evidence mapping, no-silent-rewrite, and no-engine/config-abstraction. On conflict, preserve operational utility over readability. Details: references/semantic-preservation.md.

Issue tracking — one ledger, per feature (no abstraction)

Tracking is conditional: create an epic + issues only if the feature genuinely decomposes into >1 tracked task. A single-task S/M feature needs no tracker. The human picks one ledger per feature in the Goal Contract's ## Tracker section — a tool actually available in the environment (e.g. bd, Linear, GitHub Issues) or none — there is no adapter layer.

  • bd (beads) — good default for local/solo, git-native, dependency-aware: bd init if no .beads store; epic = parent bead, tasks = child beads, deps via bd link.
  • Hosted trackers (Linear, GitHub Issues, …) — when work must be visible to others or already lives there: use whatever access this environment provides (CLI or MCP); epic = project/parent issue, tasks = issues. If the human's explicit tracker choice isn't available here, stop and ask — switching trackers or dropping to none is a Goal Amendment, not a silent downgrade.
  • Never open two ledgers (e.g. a Linear project and a beads epic) for the same work.

What stays manual / out of scope (do not build)

GEPA/template optimization · artifact-usage telemetry · generic pipeline.config · executor abstraction · automatic tracker wiring · visual-evidence matrix · bake-off automation · risk governance beyond self-assessment prompts · auto-updating the agent-instructions file (AGENTS.md / CLAUDE.md) · anything that smells like "the engine." The method earns an engine only after 5–10 real features, not before.

Portability — running this from any agent (Codex, Copilot, Gemini, …)

This skill is plain markdown and platform-neutral: the loop is shell/CLI work, not Claude-specific tooling. Anything named here (bd, Linear, a TDD sub-skill, a second-opinion model) is optional — if it isn't available in the current environment, use the stated fallback and continue; a missing optional tool never blocks unless tracking is explicitly required.

Agents that don't auto-discover skills (e.g. the Codex CLI) won't pick this up just because the files exist. To make it discoverable in a target repo, add an entry to that repo's agent-instructions file (AGENTS.md, or CLAUDE.md):

md
## feature-factory
When asked to build/ship a single feature, or to "write a goal contract", read
<path-to>/feature-factory/SKILL.md and follow it. For non-trivial features, copy
<path-to>/feature-factory/assets/goal-contract.md to docs/factory/<date>-<slug>/goal.md.

When following this without a skill-runner, read references/process-budget.md (size/risk triggers), references/stack-discovery.md (finding the repo's verify commands; bootstrapping a missing test harness), and references/semantic-preservation.md (only when editing the method's own artifacts) directly.

When copying this skill into another repo, copy only its content files (SKILL.md, assets/, references/) — local tool state (e.g. .enzyme/, .claude/) may sit alongside them and must not ship.

Background

Distilled from the feature-factory method — public repo: https://github.com/glebis/feature-factory (README + Goal Contract template). The fuller design spec and the three external-audit research streams are kept privately; this skill is the runnable distillation. The highest-risk assumption to stay honest about: a process that worked on one bounded, logic-heavy pilot is not yet proven to stay lightweight on messy UI/integration work — pressure-test it on a deliberately different feature next.

© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references, assets) in feature-factory of glebis/claude-skills.

  • SKILL.md
  • assets/goal-contract-example.md
  • assets/goal-contract.md
  • references/process-budget.md
  • references/semantic-preservation.md
  • references/stack-discovery.md

Open the folder on GitHubat commit 3b88261

Compare with similar skills

Feature Factory next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Feature Factory compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Feature Factory this skillglebis/claude-skills391—~3.4kAutomated safety check: PassMIT
Testing Skills With Subagentsed3dai/ed3d-plugins2503 repos~3.5kAutomated safety check: PassNone
Feature Next Task Workflowmylukin/agent-foreman250—~927Automated safety check: NotesNone
Agent Teamsalinaqi/maggy707—~5kAutomated safety check: NotesMIT
Agent Task Backlog Initializermylukin/agent-foreman250—~681Automated safety check: NotesNone
Foreman VerifyVisionForge-OU/foreman443—~901Automated safety check: PassCustom licence

Similar skills

  • A skill your agent uses when creating or editing skills, before deployment, to verify they work under pressure and resist rationalization - applies RED-GREEN-REFACTOR cycle to process documentation…

    250 GitHub starsUsed in 3 repos~3.5k tokens
    Testing & QAAuto-check passed
  • Feature Next Task Workflow

    mylukin/agent-foreman

    Enforces a strict next-implement-check-done cycle through the agent-foreman CLI so an agent works one backlog task at a time, with optional TDD gating.

    250 GitHub stars~927 tokensUpdated 8 mo ago
    Testing & QAAuto-check: notes
  • Agent Teams

    alinaqi/maggy

    Claude Code Agent Teams - default team-based development with strict TDD pipeline enforcement

    707 GitHub stars~5k tokensUpdated 16 days ago
    Testing & QAAuto-check: notes
  • Agent Task Backlog Initializer

    mylukin/agent-foreman

    Builds a feature backlog, progress log, and optional strict test enforcement for agent-driven work with a single init command.

    250 GitHub stars~681 tokensUpdated 8 mo ago
    Testing & QAAuto-check: notes
  • Foreman Verify

    VisionForge-OU/foreman

    Headless self-verification gate a Foreman worker runs before it claims an issue is done.

    443 GitHub stars~901 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Superpowers Feature Workflow

    SYZ-Coder/superpowers-openspec-team-skills

    A skill your agent uses when feature work needs the Superpowers stages before or during implementation: brainstorming, design confirmation, implementation planning, worktree setup, test-driven…

    196 GitHub stars~891 tokensUpdated 4 mo ago
    Testing & QAAuto-check passed

More from glebis/claude-skills

All 92 skills in this repo
  • Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.

    391 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.

    391 GitHub stars~973 tokensUpdated 2 days ago
    Auto-check passed
  • Deep Research

    glebis/claude-skills

    This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.

    391 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check: notes
  • Elimination Research

    glebis/claude-skills

    This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…

    391 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Narrated HTML Presentations

    glebis/claude-skills

    Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.

    391 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check: notes
  • Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

    391 GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed

Questions about Feature Factory

What does Feature Factory do?

This skill should be used when taking a single software feature from intent to shipped as a solo developer — goal-first, TDD, deterministic verification, evidence only where it earns its keep, and…. Feature Factory is an agent skill from glebis/claude-skills. This skill should be used when taking a single software feature from intent to shipped as a solo developer — goal-first, TDD, deterministic verification, evidence only where it earns its keep, and human judgment at the two moments that matter (goal approval, merge).

When should I use Feature Factory?

Feature Factory fits situations like: the user says lets build feature X; ship this feature; run this through the factory; write a goal contract.

How do I install Feature Factory in Claude Code?

Run `npx skills add glebis/claude-skills --skill feature-factory -a claude-code`. Or copy the skill folder (feature-factory in glebis/claude-skills) into .claude/skills/feature-factory in your project. Claude Code loads it when a task matches its description.

How do I install Feature Factory in Codex?

Run `npx skills add glebis/claude-skills --skill feature-factory -a codex`. Or copy the skill folder (feature-factory in glebis/claude-skills) into .agents/skills/feature-factory in your project. Codex loads it when a task matches its description.

Can I use Feature Factory in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill feature-factory -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/feature-factory, .gemini/skills/feature-factory, .github/skills/feature-factory and .opencode/skills/feature-factory in your project.

What does Feature Factory need to run?

SKILL.md names no scripts, command-line tools or credentials: Feature Factory is instructions for the agent only.

Does Feature Factory access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Feature Factory safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Feature Factory use?

Feature Factory is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Feature Factory use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Feature Factory?

Skills that share tags, products or a category with Feature Factory: Testing Skills With Subagents (ed3dai/ed3d-plugins, 250 stars), Feature Next Task Workflow (mylukin/agent-foreman, 250 stars), Agent Teams (alinaqi/maggy, 707 stars) and Agent Task Backlog Initializer (mylukin/agent-foreman, 250 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Feature Factory?

glebis (a GitHub user) maintains it in glebis/claude-skills, which has 391 GitHub stars. The repository holds 92 skills in this directory. The repository was last updated on October 8, 2026.

Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.