DeepTutor CLI
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
Autonomous quality improvement loop. An agent skill from SethGammon/Citadel.
$ npx skills add SethGammon/Citadel --skill improve -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install SethGammon/Citadel improve --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/improve .claude/skills/improve && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "improve" agent skill from https://github.com/SethGammon/Citadel/tree/main/skills/improve into .claude/skills/improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "improve", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/SethGammon/Citadel/tree/main/skills/improveType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add SethGammon/Citadel --skill improve -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install SethGammon/Citadel improve --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/improve .agents/skills/improve && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "improve" agent skill from https://github.com/SethGammon/Citadel/tree/main/skills/improve into .agents/skills/improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "improve", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add SethGammon/Citadel --skill improve -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install SethGammon/Citadel improve --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/improve .cursor/skills/improve && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "improve" agent skill from https://github.com/SethGammon/Citadel/tree/main/skills/improve into .cursor/skills/improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "improve", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/SethGammon/Citadel.git --path skills/improve--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add SethGammon/Citadel --skill improve -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install SethGammon/Citadel improve --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/improve .gemini/skills/improve && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "improve" agent skill from https://github.com/SethGammon/Citadel/tree/main/skills/improve into .gemini/skills/improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "improve", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install SethGammon/Citadel improveInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add SethGammon/Citadel --skill improve -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/improve .github/skills/improve && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "improve" agent skill from https://github.com/SethGammon/Citadel/tree/main/skills/improve into .github/skills/improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "improve", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add SethGammon/Citadel --skill improve -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install SethGammon/Citadel improve --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/improve .opencode/skills/improve && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "improve" agent skill from https://github.com/SethGammon/Citadel/tree/main/skills/improve into .opencode/skills/improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "improve", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
improveAutonomous quality improvement loop. An agent skill from SethGammon/Citadel.
Improve is an agent skill from SethGammon/Citadel. Autonomous quality improvement loop. Scores a target against a rubric, selects the highest-leverage axis, attacks it, verifies, documents, and loops. No pre-planning between iterations — each loop re-scores from scratch.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `__benchmarks__/no-rubric.md` and `__benchmarks__/score-only.md`).
It sits in Education, covering Quizzes and assessments. The repository describes itself as: The operating layer for Claude Code + OpenAI Codex: persistent project memory, intent routing, safety hooks, cost telemetry, and parallel agent fleets. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e41ff1d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Improve loads about 4.3k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 1,971 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from SethGammon/Citadel at commit e41ff1d, republished under its MIT licence (© SethGammon). 1,971 words, ~4,260 tokens.
.claude/skills/improve/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Use when: Scoring a target against a rubric and iteratively improving it. Rubric required at .planning/rubrics/{target}.md (Phase 0 creates one if missing).
Don't use when: Refactoring without a rubric (use /refactor), one-time code review (use /review), or debugging a specific bug (use /systematic-debugging).
/improve {target} # Loop until plateau or all axes >= 8.0
/improve {target} --n=3 # Run exactly N loops then stop
/improve {target} --axis={name} # Force-attack a specific axis (skips scoring)
/improve {target} --score-only # Score and report, no attack
/improve {target} --continue # Resume from campaign state (used by daemon)
/improve citadel # Targets Citadel itselftarget is a slug that maps to .planning/rubrics/{target}.md.
If no rubric exists, run Phase 0 first.
When invoked with --n or --continue, improve operates in campaign mode and maintains a campaign file that daemon can attach to.
Campaign file: .planning/campaigns/improve-{target}.md, created automatically on the first invocation with --n (full template: docs/QUALITY_LOOPS.md#campaign-file-template).
Frontmatter: version, id (improve-{target}-{ISO-date-slug}), status: active, type: improve, target, total_loops ({n} or unlimited), completed_loops: 0, current_level (from rubric frontmatter), estimated_cost_per_loop: 12, started.
Body: status and direction lines, a Loop History table (Loop | Axis Attacked | Outcome | Score Movement), and a Continuation State block (next_loop, last_scorecard_log, last_outcome, phase_within_loop, level_up_triggered).
Update phase_within_loop at each phase: scoring → selected-{axis} → attacking-{axis} → verifying → not-started.
On loop complete: increment completed_loops, update next_loop/last_scorecard_log/last_outcome, append Loop History row.
--continue flag.planning/campaigns/improve-{target}.md — error if missing or status not activecompleted_loops >= total_loops: mark completed, exitphase_within_loop is not not-started: restart current loop from Phase 1 (interrupted mid-loop)last_scorecard_log for delta comparison, then run Phase 1 onwardsRun only when .planning/rubrics/{target}.md does not exist.
.planning/research/ if available/research --parallel to survey comparable products if no research existsrubric_approved: {answer}..planning/rubrics/{target}.mdScore every axis in the rubric. No shortcuts. No cached scores from the previous loop.
Execute the programmatic verification steps from the rubric. A programmatic failure caps that axis at 5 regardless of evaluator scores. Record raw results: which checks passed, which failed, what the failure was.
Execute the structural checks from each axis's verification spec: file path existence, frontmatter schema consistency, benchmark coverage ratios, link rot, and cross-reference accuracy (check descriptions: docs/QUALITY_LOOPS.md#structural-check-types).
Spawn three evaluator agents in parallel. Each receives the rubric with all axis definitions and anchors, read access to the target, its persona (A/B/C as defined in the rubric's Scoring Protocol), and the instruction to score every axis 0-10 with a one-sentence justification per axis (input list: docs/QUALITY_LOOPS.md#evaluator-panel).
Each evaluator scores independently. For each axis:
needs-refinementneeds-refinement axes are logged but still scored. Do not halt on evaluator disagreement.
Compile a table with columns Axis | A | B | C | Prog | Final | Delta | Flag (layout: docs/QUALITY_LOOPS.md#scorecard-format).
Final = min(A, B, C), then apply programmatic cap (sets Flag=cap). Delta = current − prior loop score (empty on loop 1).
Choose the single axis to attack this loop.
Selection formula:
score(axis) = (10 - current_score) × weight × effort_multiplier × recency_penaltyeffort_multiplier: low = 1.0, medium = 0.7, high = 0.4recency_penalty: 0.5 if attacked in previous 2 loops, otherwise 1.0If --axis flag was set, skip selection and attack the specified axis.
Announce the selection:
Selected: {axis_name} (score: {n}/10, weight: {w}, effort: {e}, selection score: {s})
Rationale: {one sentence on why this axis now, not another}Execute the improvement. Dispatch strategy depends on the axis category (expanded per-category playbooks: docs/QUALITY_LOOPS.md#attack-dispatch-strategies).
ISOLATION MANDATE: When dispatching to /experiment, /fleet, or /research --parallel, always use the Agent tool with isolation: "worktree". Sub-agents in worktrees get their own context windows; the orchestrator only receives their HANDOFF results.
| Category | Dispatch | Verification |
|---|---|---|
| technical | /experiment with before/after comparison; speculative worktrees (Agent + isolation: "worktree") for approaches that might conflict | node scripts/run-with-timeout.js 300 node scripts/test-all.js as the oracle |
| documentation | direct: read current docs, fix specific gaps; cross-reference every claim against source | structural verification before committing |
| experience | structural fixes + doc updates; run the actual install flow in a clean temp dir; inject synthetic failures per the programmatic spec | /qa |
| positioning | /research to verify the competitive landscape is accurate, then update README/FAQ/demo copy | /qa confirms the updated page renders |
| presentation | targeted changes per rubric anchors (no rewrites unless score is below 3) | /live-preview or /qa confirms visual changes render |
| security | read the specific hooks/scripts involved, make targeted code changes | run the rubric's programmatic verification steps directly |
Artifact archiving: when the attack tried multiple approaches, write a decision record to the loop log: APPROACH COMPARISON: [approach A] vs [approach B] — winner: [A] because [reason].
After the attack, re-score only the targeted axis (not full re-score).
Run the four verification tiers from the rubric for the targeted axis:
/do command.onboarding_friction, error_recovery, documentation_accuracy, command_discoverabilityPASS {wall_time} or FAIL at step {n}: {what broke}visual_coherence, api_surface_consistency)Regression check (run on all axes, not just targeted):
On abort: revert the changes, log the failure, treat as "no improvement this loop".
On pass: commit the changes with a descriptive message.
Write the loop log. Always. Even on abort.
Log path: .planning/improvement-logs/{target}/loop-{n}.md
Required sections (full template: docs/QUALITY_LOOPS.md#loop-log-template):
APPROACH COMPARISON record if multiple approaches were triedPASS {wall_time} | FAIL at step {n}: {reason} | SKIPPEDPROPOSED AXIS: {name} | Rationale | Category | Weight | Anchors: 0=... 5=... 10=... (or: None proposed this loop.)All proposals go to .planning/rubrics/{target}-proposals.md. Never to the live rubric.
Exit conditions (check in order):
--n flag was set and N loops have completed: exit, report scorecardOn Level-Up: do not exit. Escalate. See Level-Up Protocol section.
On ceiling (all >= 8.0): report the final scorecard and recommend a Level-Up run.
On normal loop: return to Phase 1. Re-score everything from scratch.
Campaign mode exit handling:
status: completed, move to completed/status: completed, move to completed/status: level-up-pending (daemon will pause, not retry)status: parkedstatus: parked with reasonstatus: pausedTriggers when no axis improved > 0.5 in the last 2 consecutive loops, no programmatic cap is active, and at least 3 loops have completed.
Step 1: Freeze the snapshot
Write .planning/rubrics/{target}-level-{n}-final.md with: date, loops completed, final scorecard, axes at ceiling (≥9.0 — their 10 anchors become Level {n+1}'s 5 anchors), and axes that plateaued below 9.0 with why.
Step 2: Write proposals
For each axis: propose Level {n+1} re-anchoring (current 10 → new 5, propose new 10). For plateaued axes: re-anchor, replace with measurable proxy, or retire.
Auto-include these three process axes if not already in the rubric: decomposition_quality, scope_appropriateness, verification_depth.
Write to .planning/rubrics/{target}-proposals.md: re-anchored axes (current 10 anchor, proposed 0/5/10), proposed new axes, axes proposed for retirement.
Step 3: Halt -- human approval required
Do not self-approve. Do not continue looping.
In campaign mode: set status: level-up-pending, set level_up_triggered: true, and write awaiting: human approval of level-up proposals to Continuation State.
Report: what was achieved at this level (scorecard summary), the proposals file location, and what the expected new gains look like at the next level.
The loop resumes only when the human edits the live rubric with approved proposals
and sets the campaign status back to active. Level {n+1} loops continue incrementing
the loop number (they do not reset to 1).
Step 4: Historical context for future evaluators
When the loop resumes after a level-up, every evaluator in Phase 1c receives the level-{n}-final.md snapshot as a reference baseline, plus the instruction: "Scores from the previous level are the floor. A score of 5 at Level 2 means you have reached what was the ceiling at Level 1."
needs-refinement, use minimum score, continue.--continue + no campaign file: error, suggest --n.--continue + level-up-pending: halt, point to proposals file, require human approval then status: active.--continue + completed: do not resume, report final scorecard.--n + existing active campaign: treat as --continue. If completed/parked: new campaign, incremented slug..planning/rubrics/{target}-proposals.md only. Human approval required.status: level-up-pending, not parked or active.Disclosure: State loop count, target, per-loop cost (~$12), total estimate. For --continue: loops remaining and spend so far. For unlimited: state exit conditions (plateau or all axes >= 8.0).
Reversibility: Green = --score-only | Amber = standard loops (each commits separately) | Red = level-up (rewrites rubric anchors permanently). Red requires explicit confirmation.
Proportionality: No rubric + no explicit request → suggest /review. All axes > 8.0 + --n=1 → suggest --axis. Cost > $50 → confirm.
Trust gating: Novice (0-4): --score-only / --n=1 only. Familiar (5-19): up to --n=5. Trusted (20+): no cap; confirm unlimited or cost > $50.
---HANDOFF---
- Target: {target} — Loop {n} of {n_total or "∞"} — Level {current_level}
- Outcome: {improved | plateau | ceiling | aborted | n-complete | level-up-triggered}
- Score movement: {axis} {before} → {after} (+{delta})
- Behavioral simulation: {PASS {wall_time} | FAIL | SKIPPED}
- Proposed rubric additions: {count} — written to .planning/rubrics/{target}-proposals.md
- Loop log: .planning/improvement-logs/{target}/loop-{n}.md
- Reversibility: amber -- each loop commits separately, revert individual loops with git revert
- Next recommended axis: {axis_name} (if not exiting)
- Level-up snapshot: .planning/rubrics/{target}-level-{n}-final.md (if level-up triggered)
---© SethGammon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in skills/improve of SethGammon/Citadel.
Open the folder on GitHubat commit e41ff1d
Improve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Improve this skillSethGammon/Citadel | 922 | — | ~4.3k | Automated safety check: Pass | MIT | |
| DeepTutor CLIHKUDS/DeepTutor | 41k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2k | Automated safety check: Pass | MIT | |
| Codebase to Coursezarazhangrui/codebase-to-course | 5.7k | — | ~4.4k | Automated safety check: Pass | None | |
| AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Scholar EvaluationK-Dense-AI/claude-scientific-writer | 2.4k | 2 repos | ~2.9k | Automated safety check: Notes | MIT |
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
rohitg00/ai-engineering-from-scratch
Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.
zarazhangrui/codebase-to-course
Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.
rohitg00/ai-engineering-from-scratch
Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.
K-Dense-AI/claude-scientific-writer
Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.
guanyang/open-agent-hub
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…
SethGammon/Citadel
Creates new skills from the user's repeating patterns. An agent skill from SethGammon/Citadel.
SethGammon/Citadel
Cross-drive storage audit and cleanup. An agent skill from SethGammon/Citadel.
SethGammon/Citadel
Bounded foreground repetition for the current session. An agent skill from SethGammon/Citadel.
SethGammon/Citadel
GitHub issue and PR investigator. An agent skill from SethGammon/Citadel.
SethGammon/Citadel
File sentinel that monitors the working directory for changes and marker comments, then auto-triggers appropriate skills.
SethGammon/Citadel
Autonomous multi-session campaign agent. An agent skill from SethGammon/Citadel.
Categories
Autonomous quality improvement loop. An agent skill from SethGammon/Citadel. Improve is an agent skill from SethGammon/Citadel. Autonomous quality improvement loop.
Improve fits situations like: tasks that involve Quizzes and assessments.
Run `npx skills add SethGammon/Citadel --skill improve -a claude-code`. Or copy the skill folder (skills/improve in SethGammon/Citadel) into .claude/skills/improve in your project. Claude Code loads it when a task matches its description.
Run `npx skills add SethGammon/Citadel --skill improve -a codex`. Or copy the skill folder (skills/improve in SethGammon/Citadel) into .agents/skills/improve in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SethGammon/Citadel --skill improve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/improve, .gemini/skills/improve, .github/skills/improve and .opencode/skills/improve in your project.
Going by SKILL.md and its folder, Improve needs the command-line tools its instructions call (node).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Improve is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Improve: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
SethGammon (a GitHub user) maintains it in SethGammon/Citadel, which has 922 GitHub stars. The repository holds 48 skills in this directory. The repository was last updated on October 1, 2026.
Source: SethGammon/Citadel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.