Agent skill

Improve

by SethGammon in SethGammon/Citadel

Autonomous quality improvement loop. An agent skill from SethGammon/Citadel.

MITAuto-check passedEducation

Install Improve

skills CLI
$ npx skills add SethGammon/Citadel --skill improve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SethGammon/Citadel improve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/improve .claude/skills/improve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
improve
GitHub stars
922
Token cost
~4.3k tokens
SKILL.md length
1,971 words
Files
3
Skills in repo
48
Repo updated
First seen
Licence
MIT

At a glance

Autonomous quality improvement loop. An agent skill from SethGammon/Citadel.

  • Works in 7 steps: Rubric Bootstrap (one-time, requires… → Score → Select → …
  • Tasks that involve Quizzes and assessments
  • SKILL.md covers Orientation, Invocation, Campaign Mode and Protocol, plus 4 more sections
  • Calls node

What it does

Improve is an agent skill from SethGammon/Citadel. Autonomous quality improvement loop. Scores a target against a rubric, selects the highest-leverage axis, attacks it, verifies, documents, and loops. No pre-planning between iterations — each loop re-scores from scratch.

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `__benchmarks__/no-rubric.md` and `__benchmarks__/score-only.md`).

It sits in Education, covering Quizzes and assessments. The repository describes itself as: The operating layer for Claude Code + OpenAI Codex: persistent project memory, intent routing, safety hooks, cost telemetry, and parallel agent fleets. The licence is MIT.

When your agent uses it

  • Tasks that involve Quizzes and assessments

Example prompts

  • “/improve”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Rubric Bootstrap (one-time, requires human approval)
  2. Score
  3. Select
  4. Attack
  5. Verify
  6. Document
  7. Loop or Exit

What it can do on your machine

Read from SKILL.md and the folder at commit e41ff1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Improve loads about 4.3k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 1,971 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from SethGammon/Citadel at commit e41ff1d, republished under its MIT licence (© SethGammon). 1,971 words, ~4,260 tokens.

Download SKILL.mdSave it as .claude/skills/improve/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
improve
description
Autonomous quality improvement loop. Scores a target against a rubric, selects the highest-leverage axis, attacks it, verifies, documents, and loops. No pre-planning between iterations — each loop re-scores from scratch.
license
MIT
user-invocable
true
auto-trigger
false
trigger_keywords
improve, improvement loop, quality loop, rubric, score against, run improvement, improve citadel
last-updated
2026-03-28

/improve — Autonomous Quality Engine

Orientation

Use when: Scoring a target against a rubric and iteratively improving it. Rubric required at .planning/rubrics/{target}.md (Phase 0 creates one if missing).

Don't use when: Refactoring without a rubric (use /refactor), one-time code review (use /review), or debugging a specific bug (use /systematic-debugging).

Invocation

/improve {target}            # Loop until plateau or all axes >= 8.0
/improve {target} --n=3      # Run exactly N loops then stop
/improve {target} --axis={name}  # Force-attack a specific axis (skips scoring)
/improve {target} --score-only   # Score and report, no attack
/improve {target} --continue     # Resume from campaign state (used by daemon)
/improve citadel             # Targets Citadel itself

target is a slug that maps to .planning/rubrics/{target}.md. If no rubric exists, run Phase 0 first.

Campaign Mode

When invoked with --n or --continue, improve operates in campaign mode and maintains a campaign file that daemon can attach to.

Campaign file: .planning/campaigns/improve-{target}.md, created automatically on the first invocation with --n (full template: docs/QUALITY_LOOPS.md#campaign-file-template). Frontmatter: version, id (improve-{target}-{ISO-date-slug}), status: active, type: improve, target, total_loops ({n} or unlimited), completed_loops: 0, current_level (from rubric frontmatter), estimated_cost_per_loop: 12, started. Body: status and direction lines, a Loop History table (Loop | Axis Attacked | Outcome | Score Movement), and a Continuation State block (next_loop, last_scorecard_log, last_outcome, phase_within_loop, level_up_triggered).

Campaign lifecycle

Update phase_within_loop at each phase: scoring → selected-{axis} → attacking-{axis} → verifying → not-started.

On loop complete: increment completed_loops, update next_loop/last_scorecard_log/last_outcome, append Loop History row.

The --continue flag
  1. Read .planning/campaigns/improve-{target}.md — error if missing or status not active
  2. If completed_loops >= total_loops: mark completed, exit
  3. If phase_within_loop is not not-started: restart current loop from Phase 1 (interrupted mid-loop)
  4. Load last_scorecard_log for delta comparison, then run Phase 1 onwards

Protocol

Phase 0: Rubric Bootstrap (one-time, requires human approval)

Run only when .planning/rubrics/{target}.md does not exist.

  1. Read competitive research from .planning/research/ if available
  2. Spawn /research --parallel to survey comparable products if no research exists
  3. Draft 8-14 axes organized into 3-5 categories, each with:
    • Weight (0.0–1.0), Category, three anchors (0/5/10), verification specs (programmatic/structural/perceptual), research inputs
  4. Present draft rubric to the user with rationale for each axis
  5. STOP. Do not proceed until the user approves the rubric.
    • If the AskUserQuestion tool is available, present the gate with it. Options: "Approve as drafted", "Adjust axes", "Adjust weights". On adjust: revise, re-present.
    • If unavailable: ask as a plain text question and wait.
    • Record the answer in the campaign file as rubric_approved: {answer}.
  6. Write approved rubric to .planning/rubrics/{target}.md
Phase 1: Score

Score every axis in the rubric. No shortcuts. No cached scores from the previous loop.

1a. Programmatic checks (run first, in parallel)

Execute the programmatic verification steps from the rubric. A programmatic failure caps that axis at 5 regardless of evaluator scores. Record raw results: which checks passed, which failed, what the failure was.

1b. Structural analysis

Execute the structural checks from each axis's verification spec: file path existence, frontmatter schema consistency, benchmark coverage ratios, link rot, and cross-reference accuracy (check descriptions: docs/QUALITY_LOOPS.md#structural-check-types).

1c. Perceptual scoring panel (three independent evaluators)

Spawn three evaluator agents in parallel. Each receives the rubric with all axis definitions and anchors, read access to the target, its persona (A/B/C as defined in the rubric's Scoring Protocol), and the instruction to score every axis 0-10 with a one-sentence justification per axis (input list: docs/QUALITY_LOOPS.md#evaluator-panel).

Each evaluator scores independently. For each axis:

  • Final score = minimum of the three evaluators (plus programmatic cap if applicable)
  • If any two evaluators disagree by > 3 points: flag the axis as needs-refinement

needs-refinement axes are logged but still scored. Do not halt on evaluator disagreement.

1d. Compile scorecard

Compile a table with columns Axis | A | B | C | Prog | Final | Delta | Flag (layout: docs/QUALITY_LOOPS.md#scorecard-format). Final = min(A, B, C), then apply programmatic cap (sets Flag=cap). Delta = current − prior loop score (empty on loop 1).

Phase 2: Select

Choose the single axis to attack this loop.

Selection formula:

score(axis) = (10 - current_score) × weight × effort_multiplier × recency_penalty
  • effort_multiplier: low = 1.0, medium = 0.7, high = 0.4
  • recency_penalty: 0.5 if attacked in previous 2 loops, otherwise 1.0
  • Effort tiers: low < 1hr, medium 1-3hrs, high 3+hrs

If --axis flag was set, skip selection and attack the specified axis.

Announce the selection:

Selected: {axis_name} (score: {n}/10, weight: {w}, effort: {e}, selection score: {s})
Rationale: {one sentence on why this axis now, not another}
Phase 3: Attack

Execute the improvement. Dispatch strategy depends on the axis category (expanded per-category playbooks: docs/QUALITY_LOOPS.md#attack-dispatch-strategies).

ISOLATION MANDATE: When dispatching to /experiment, /fleet, or /research --parallel, always use the Agent tool with isolation: "worktree". Sub-agents in worktrees get their own context windows; the orchestrator only receives their HANDOFF results.

CategoryDispatchVerification
technical/experiment with before/after comparison; speculative worktrees (Agent + isolation: "worktree") for approaches that might conflictnode scripts/run-with-timeout.js 300 node scripts/test-all.js as the oracle
documentationdirect: read current docs, fix specific gaps; cross-reference every claim against sourcestructural verification before committing
experiencestructural fixes + doc updates; run the actual install flow in a clean temp dir; inject synthetic failures per the programmatic spec/qa
positioning/research to verify the competitive landscape is accurate, then update README/FAQ/demo copy/qa confirms the updated page renders
presentationtargeted changes per rubric anchors (no rewrites unless score is below 3)/live-preview or /qa confirms visual changes render
securityread the specific hooks/scripts involved, make targeted code changesrun the rubric's programmatic verification steps directly

Artifact archiving: when the attack tried multiple approaches, write a decision record to the loop log: APPROACH COMPARISON: [approach A] vs [approach B] — winner: [A] because [reason].

Phase 4: Verify

After the attack, re-score only the targeted axis (not full re-score).

Run the four verification tiers from the rubric for the targeted axis:

  1. Programmatic: execute the specific checks, confirm they now pass
  2. Structural: verify the structural requirements are met
  3. Perceptual: spawn a single evaluator agent (Evaluator B — Newcomer) and score just the targeted axis
  4. Behavioral simulation: clone the repo into a temp directory and follow INSTALL.md exactly as written — no prior knowledge, no shortcuts. Measure whether each step completes without error and record wall time to first successful /do command.
    • Required when targeted axis is: onboarding_friction, error_recovery, documentation_accuracy, command_discoverability
    • Optional for all other axes
    • Result: PASS {wall_time} or FAIL at step {n}: {what broke}
    • A behavioral FAIL overrides a passing perceptual score. Do not commit on behavioral FAIL.
    • Skip only if the targeted axis could not plausibly affect the user path. Plausible = axis governs code shown/executed in the app, or controls presence/absence of a UI element. Safe to skip: documentation axes (comments, docstrings), configuration-only axes, developer-tooling-only axes (e.g., visual_coherence, api_surface_consistency)

Regression check (run on all axes, not just targeted):

  • Re-run programmatic checks on every axis that shares files with the changes
  • If any previously passing axis now fails programmatic: abort, do not commit
  • If perceptual estimate suggests any axis dropped > 0.5 from baseline: abort, do not commit

On abort: revert the changes, log the failure, treat as "no improvement this loop".

On pass: commit the changes with a descriptive message.

Phase 5: Document

Write the loop log. Always. Even on abort.

Log path: .planning/improvement-logs/{target}/loop-{n}.md

Required sections (full template: docs/QUALITY_LOOPS.md#loop-log-template):

  • Header: date, loop number, selected axis, outcome (improved | no-change | aborted)
  • Scorecard: per-axis prior score, current score, delta
  • Attack summary: what was changed, approach (experiment / direct / research+update), files touched, plus the APPROACH COMPARISON record if multiple approaches were tried
  • Verification results: programmatic PASS/FAIL, structural PASS/FAIL, perceptual {score}/10 with one-line rationale, behavioral PASS {wall_time} | FAIL at step {n}: {reason} | SKIPPED
  • Proposed axis additions: PROPOSED AXIS: {name} | Rationale | Category | Weight | Anchors: 0=... 5=... 10=... (or: None proposed this loop.)
  • What was learned: 2-3 sentences

All proposals go to .planning/rubrics/{target}-proposals.md. Never to the live rubric.

Show full SKILL.md (793 more words)Show less
Phase 6: Loop or Exit

Exit conditions (check in order):

  1. --n flag was set and N loops have completed: exit, report scorecard
  2. All axes >= 8.0: exit with "target has reached quality ceiling"
  3. No axis improved > 0.5 in either of the last 2 loops AND no programmatic cap is active AND at least 3 loops have completed: trigger Level-Up Protocol 3a. A programmatic cap IS active AND the capped axis has not improved for 2 loops: trigger Level-Up Protocol. The cap is preventing score movement, not enforcing a ceiling — do not loop indefinitely.
  4. The user said stop: exit immediately

On Level-Up: do not exit. Escalate. See Level-Up Protocol section.

On ceiling (all >= 8.0): report the final scorecard and recommend a Level-Up run.

On normal loop: return to Phase 1. Re-score everything from scratch.

Campaign mode exit handling:

  • n-complete (all loops done): set status: completed, move to completed/
  • ceiling (all axes >= 8.0): set status: completed, move to completed/
  • level-up-triggered: set status: level-up-pending (daemon will pause, not retry)
  • aborted (security failure, unrecoverable regression): set status: parked
  • plateau (no improvement, not yet level-up): set status: parked with reason
  • user-stopped: set status: paused
Level-Up Protocol

Triggers when no axis improved > 0.5 in the last 2 consecutive loops, no programmatic cap is active, and at least 3 loops have completed.

Step 1: Freeze the snapshot

Write .planning/rubrics/{target}-level-{n}-final.md with: date, loops completed, final scorecard, axes at ceiling (≥9.0 — their 10 anchors become Level {n+1}'s 5 anchors), and axes that plateaued below 9.0 with why.

Step 2: Write proposals

For each axis: propose Level {n+1} re-anchoring (current 10 → new 5, propose new 10). For plateaued axes: re-anchor, replace with measurable proxy, or retire.

Auto-include these three process axes if not already in the rubric: decomposition_quality, scope_appropriateness, verification_depth.

Write to .planning/rubrics/{target}-proposals.md: re-anchored axes (current 10 anchor, proposed 0/5/10), proposed new axes, axes proposed for retirement.

Step 3: Halt -- human approval required

Do not self-approve. Do not continue looping.

In campaign mode: set status: level-up-pending, set level_up_triggered: true, and write awaiting: human approval of level-up proposals to Continuation State.

Report: what was achieved at this level (scorecard summary), the proposals file location, and what the expected new gains look like at the next level.

The loop resumes only when the human edits the live rubric with approved proposals and sets the campaign status back to active. Level {n+1} loops continue incrementing the loop number (they do not reset to 1).

Step 4: Historical context for future evaluators

When the loop resumes after a level-up, every evaluator in Phase 1c receives the level-{n}-final.md snapshot as a reference baseline, plus the instruction: "Scores from the previous level are the floor. A score of 5 at Level 2 means you have reached what was the ceiling at Level 1."

Fringe Cases

  • No rubric: run Phase 0, halt for human approval. Never improvise.
  • Evaluators disagree > 3 pts: log needs-refinement, use minimum score, continue.
  • Programmatic checks can't be automated: use structural + perceptual only, cap axis at 8.
  • No improvement this loop: document as "no-change", apply recency penalty.
  • Loop 1 (no prior logs): delta fields empty — expected.
  • Security axis fails programmatic: halt and report. Blocking.
  • --continue + no campaign file: error, suggest --n.
  • --continue + level-up-pending: halt, point to proposals file, require human approval then status: active.
  • --continue + completed: do not resume, report final scorecard.
  • --n + existing active campaign: treat as --continue. If completed/parked: new campaign, incremented slug.

Quality Gates

  • Phase 0 requires human approval. No exceptions.
  • Phase 4 regression check must run. No committing without it.
  • Phase 4 behavioral simulation result must appear in the loop log for applicable axes. Behavioral FAIL blocks commit regardless of perceptual score.
  • Phase 5 loop log must be written. Even on abort, even on no-change.
  • Perceptual scoring: all three evaluators required for Phase 1. Single evaluator acceptable for Phase 4 spot-check only.
  • Selection formula must be shown in output.
  • Any axis with a programmatic failure is capped at 5. Cannot be overridden.
  • The loop never writes to the live rubric. Proposals go to .planning/rubrics/{target}-proposals.md only. Human approval required.
  • Level-Up Protocol requires human approval before resuming
  • Campaign mode: campaign file must be updated after every phase transition and every loop completion.
  • Campaign mode: level-up must set status: level-up-pending, not parked or active.

Contextual Gates

Disclosure: State loop count, target, per-loop cost (~$12), total estimate. For --continue: loops remaining and spend so far. For unlimited: state exit conditions (plateau or all axes >= 8.0).

Reversibility: Green = --score-only | Amber = standard loops (each commits separately) | Red = level-up (rewrites rubric anchors permanently). Red requires explicit confirmation.

Proportionality: No rubric + no explicit request → suggest /review. All axes > 8.0 + --n=1 → suggest --axis. Cost > $50 → confirm.

Trust gating: Novice (0-4): --score-only / --n=1 only. Familiar (5-19): up to --n=5. Trusted (20+): no cap; confirm unlimited or cost > $50.

Exit Protocol

---HANDOFF---
- Target: {target} — Loop {n} of {n_total or "∞"} — Level {current_level}
- Outcome: {improved | plateau | ceiling | aborted | n-complete | level-up-triggered}
- Score movement: {axis} {before} → {after} (+{delta})
- Behavioral simulation: {PASS {wall_time} | FAIL | SKIPPED}
- Proposed rubric additions: {count} — written to .planning/rubrics/{target}-proposals.md
- Loop log: .planning/improvement-logs/{target}/loop-{n}.md
- Reversibility: amber -- each loop commits separately, revert individual loops with git revert
- Next recommended axis: {axis_name} (if not exiting)
- Level-up snapshot: .planning/rubrics/{target}-level-{n}-final.md (if level-up triggered)
---

© SethGammon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/improve of SethGammon/Citadel.

  • SKILL.md
  • __benchmarks__/no-rubric.md
  • __benchmarks__/score-only.md

Open the folder on GitHubat commit e41ff1d

Compare with similar skills

Improve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Improve compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Improve this skillSethGammon/Citadel922—~4.3kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated today
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from SethGammon/Citadel

All 48 skills in this repo
  • Create Skill

    SethGammon/Citadel

    Creates new skills from the user's repeating patterns. An agent skill from SethGammon/Citadel.

    922 GitHub stars~1.9k tokensUpdated 7 days ago
    Auto-check passed
  • Houseclean

    SethGammon/Citadel

    Cross-drive storage audit and cleanup. An agent skill from SethGammon/Citadel.

    922 GitHub stars~2.2k tokensUpdated 7 days ago
    Auto-check passed
  • Loop

    SethGammon/Citadel

    Bounded foreground repetition for the current session. An agent skill from SethGammon/Citadel.

    922 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Triage

    SethGammon/Citadel

    GitHub issue and PR investigator. An agent skill from SethGammon/Citadel.

    922 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Watch

    SethGammon/Citadel

    File sentinel that monitors the working directory for changes and marker comments, then auto-triggers appropriate skills.

    922 GitHub stars~2.9k tokensUpdated 7 days ago
    Auto-check passed
  • Archon

    SethGammon/Citadel

    Autonomous multi-session campaign agent. An agent skill from SethGammon/Citadel.

    922 GitHub stars~5.4k tokensUpdated 7 days ago
    Auto-check passed

Categories

Questions about Improve

What does Improve do?

Autonomous quality improvement loop. An agent skill from SethGammon/Citadel. Improve is an agent skill from SethGammon/Citadel. Autonomous quality improvement loop.

When should I use Improve?

Improve fits situations like: tasks that involve Quizzes and assessments.

How do I install Improve in Claude Code?

Run `npx skills add SethGammon/Citadel --skill improve -a claude-code`. Or copy the skill folder (skills/improve in SethGammon/Citadel) into .claude/skills/improve in your project. Claude Code loads it when a task matches its description.

How do I install Improve in Codex?

Run `npx skills add SethGammon/Citadel --skill improve -a codex`. Or copy the skill folder (skills/improve in SethGammon/Citadel) into .agents/skills/improve in your project. Codex loads it when a task matches its description.

Can I use Improve in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SethGammon/Citadel --skill improve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/improve, .gemini/skills/improve, .github/skills/improve and .opencode/skills/improve in your project.

What does Improve need to run?

Going by SKILL.md and its folder, Improve needs the command-line tools its instructions call (node).

Does Improve access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Improve safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Improve use?

Improve is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Improve use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Improve?

Skills that share tags, products or a category with Improve: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Improve?

SethGammon (a GitHub user) maintains it in SethGammon/Citadel, which has 922 GitHub stars. The repository holds 48 skills in this directory. The repository was last updated on October 1, 2026.

Source: SethGammon/Citadel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.