Agent skill

Full Audit

by AlexZio00 in AlexZio00/sovereign-skills

Exhaustive, denominator-driven audit of an entire area (codebase, docs, memory, skills, DB, config).

MITAuto-check passedBusiness, Finance & HR

Install Full Audit

skills CLI
$ npx skills add AlexZio00/sovereign-skills --skill full-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlexZio00/sovereign-skills full-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlexZio00/sovereign-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/full-audit .claude/skills/full-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
full-audit
GitHub stars
140
Token cost
~6.4k tokens
SKILL.md length
3,298 words
Files
14 (incl. scripts)
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Exhaustive, denominator-driven audit of an entire area (codebase, docs, memory, skills, DB, config).

  • Works in 6 steps: Agree Scope + Diff Against Prior Map → Deterministic Sweep → Content Review + Rule Dry-Run… → …
  • Tasks that involve Accounting and bookkeeping
  • SKILL.md covers Dominant Variable, Trigger, Discard If and Key Assumptions, plus 14 more sections
  • Runs Python scripts from its folder; calls python and git

What it does

Full Audit is an agent skill from AlexZio00/sovereign-skills. Exhaustive, denominator-driven audit of an entire area (codebase, docs, memory, skills, DB, config). Runs a 6-phase pipeline: scope agreement + prior-map diff - deterministic sweep (counts/versions/paths/parsing plus cross-index reconciliation) - parallel read-only content review (citations forced, rule dry-run) - judgment (false-positive/UNCERTAIN triage) - fix-vs-addition split, gated by execution mode (AUDITONLY default = read-only, PROPOSE = list only, APPLYAPPROVED = apply approved items only) - coverage-map…

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts (for example `.claude-plugin/plugin.json`, `agents/openai.yaml` and `scripts/canaries/README.md`).

It sits in Business, Finance & HR, covering Accounting and bookkeeping, Code review and Citation management. The repository describes itself as: 20 production-grade skills for AI coding agents — setup, scope, discipline, code review, security, session management, governance, ops, and quality audits (eval-leakage… The licence is MIT.

When your agent uses it

  • Tasks that involve Accounting and bookkeeping
  • Tasks that involve Code review
  • Tasks that involve Citation management

Example prompts

  • “analyze”
  • “/full-audit”
  • “audit everything”
  • “/full-audit”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Agree Scope + Diff Against Prior Map
  2. Deterministic Sweep
  3. Content Review + Rule Dry-Run (Three-Layer Principle)
  4. Judgment — Dismissing False Positives
  5. Fix/Addition Split — Gated by Execution Mode
  6. Record the Coverage Map

What it can do on your machine

Read from SKILL.md and the folder at commit c062683. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 11 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Full Audit loads about 6.4k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 3,298 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~224
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from AlexZio00/sovereign-skills at commit c062683, republished under its MIT licence (© AlexZio00). 3,298 words, ~6,351 tokens.

Download SKILL.mdSave it as .claude/skills/full-audit/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
full-audit
description
Exhaustive, denominator-driven audit of an entire area (codebase, docs, memory, skills, DB, config). Runs a 6-phase pipeline: scope agreement + prior-map diff -> deterministic sweep (counts/versions/paths/parsing plus cross-index reconciliation) -> parallel read-only content review (citations forced, rule dry-run) -> judgment (false-positive/UNCERTAIN triage) -> fix-vs-addition split, gated by execution mode (AUDIT_ONLY default = read-only, PROPOSE = list only, APPLY_APPROVED = apply approved items only) -> coverage-map recording. A bare 'audit'/'analyze' request defaults to AUDIT_ONLY and never auto-advances to file writes. NOT for single-file or single-question checks (use a regular code review instead) or harness-maturity scoring against a fixed checklist (use a dedicated scoring tool instead). Triggers: '/full-audit', 'audit everything', 'full audit', 'find every gap'.
skill_type
audit-orchestrator
user-invocable
true
triggers
/full-audit, audit everything, exhaustive audit, full audit, double-check everything, find every gap
depends_on.files
scripts/canary_mix.py, scripts/canary_score.py
concurrency_profile.parallel_safe
true
concurrency_profile.parallel_phase
Phase 2 content review only (read-only reviewers, unlimited fan-out)
concurrency_profile.serialized_phases
Phase 1 sweep and Phase 4 fix-bucket edits run sequentially, not concurrently with Phase 2
not_for
Single-file or single-question checks (use a code review instead), Harness maturity scoring with a fixed checklist (that's a different, narrower tool), A…

Full Audit — Exhaustive Area Review (v1.4)

Dominant Variable

Accuracy of the coverage claim — the word "exhaustive" ships with a method label or it doesn't ship at all. The moment an unreviewed area gets reported as reviewed, this skill has failed its own purpose.

Trigger

  • /full-audit [area] · "audit everything" · "exhaustive audit" · "full audit" · "double-check everything"
  • Default mode on any of the above: AUDIT_ONLY. These are analysis requests, not execution requests — see Execution Modes below. Advancing to PROPOSE or APPLY_APPROVED in the same invocation requires the user to say so explicitly (e.g. "audit everything and apply the fixes", "전수감사하고 바로 고쳐줘").

Discard If

  • Single file / single question needs checking → use a regular code review instead
  • The goal is a harness-maturity score against a fixed checklist → use a fixed-checklist scoring tool instead
  • The goal is a single docs-vs-code drift check → too narrow a scope for this
  • An identical-scope full audit finished within the last 7 days and nothing has changed → just diff against the existing coverage map instead

Key Assumptions

  1. Target area is agreed in Phase 0 — if not, don't start without an area table.
  2. Deterministic sweeping (scripts/grep) is available for the target — if not, skip Phase 1 and never claim "exhaustive" from Phase 2 (LLM review) alone.
  3. A prior coverage map can shrink the scope via diff — if not, do a full re-scan.

Execution Modes: AUDIT_ONLY / PROPOSE / APPLY_APPROVED

This skill keeps analysis and execution in separate, explicitly-named modes. A trigger phrase like "audit everything" or "분석해줘" selects a mode — it does not, by itself, authorize any file write.

ModeWhen activeWhat runsWritable scope
AUDIT_ONLY (default)Any bare audit/analysis trigger, with no separate execution requestPhase 0-3 (scope, deterministic sweep, content review, judgment) + Phase 5 (coverage map)None — audited files are read-only. No Edit/Write to any target file, protected or not. Only audit artifacts may be written (scratch scripts, canary staging, coverage map, recall ledger).
PROPOSEUser asks for fix proposals (after an AUDIT_ONLY pass, or up front)AUDIT_ONLY output + a listed fix/addition bucket (Phase 4 framing, nothing applied)None — still read-only, output is a proposal list
APPLY_APPROVEDUser explicitly approves specific items or the fix bucket as a whole ("apply these items", "apply the fix bucket")Applies only the items the user named as approvedLimited to the approved items. A deny-listed path (rules/CLAUDE.md/settings — see Safety Layers) is reachable only here, only for that one named item, and only by drafting the edit plus an apply script for the user to run themselves — the model has no path to lift the deny (the deny rule lives in the same settings file it protects).

One-way per pass: AUDIT_ONLY never auto-advances into PROPOSE or APPLY_APPROVED within the same invocation. Advancing needs a new, explicit user statement. This closes the gap where "audit everything" silently walked all the way to Phase 4 fix-bucket execution — including a Protect-Hooks-guarded file changing — without the user ever approving execution, not just analysis.

Phase 0: Agree Scope + Diff Against Prior Map

  1. Declare the execution mode for this run in the first line of the response (e.g. "Mode: AUDIT_ONLY — read-only, no files will change"). Default AUDIT_ONLY unless the request explicitly names PROPOSE/APPLY_APPROVED or explicitly approves specific items up front.
  2. Declare the target areas as a table (e.g. codebase / docs / global skills / memory / DB / settings). For code and docs areas, build the file list as the union of git ls-files and git ls-files --others --exclude-standard - a tracked-only list drops uncommitted files, which are the ones most in need of review.
  3. If a prior coverage map exists, read it and queue its remaining gaps first.
  4. Areas the user explicitly excludes go on the map as "intentionally excluded" — never silently dropped.

Phase 1: Deterministic Sweep

Whatever a machine can count, a script counts — never eyeball it:

  • Counts (test count, DB rows, file count) / version stamps (single source of truth in N places) / path and reference existence (dead links)
  • Parsing (YAML frontmatter, JSON settings) / stale-number greps (old numbers still lingering) / expiry (TTL, aging)
  • No inline throwaway scripts — write a script to a file, run it, then delete it (guards against quoting/escaping mistakes)
  • Reuse existing checkers first (test suites, project-specific validation scripts, linters)

Cross-index contract sweep (mandatory sub-step — this is the layer most exhaustive audits skip): The layer most easily missed in "exhaustive" audits is "does the index/routing doc match the real files?" — careful reading of individual files alone will never catch this. Sweep deterministically:

  • Index vs. reality reconciliation: names listed in an inventory/index file vs. the actual directory/file listing — check both directions (ghost entries with no backing file, and real files missing from the index)
  • Routing vs. reality reconciliation: names a routing table points to vs. whether those targets actually exist (dead routes to archived/renamed targets)
  • Declared-dependency sweep: for each unit's declared dependencies (other files/skills/agents it depends on), do all of them actually exist? (including malformed declarations, e.g. a flag where a name was expected)
  • Frontmatter parsing integrity: duplicate YAML keys in frontmatter (the later one silently wins — a safety profile could flip silently)
  • Rationale: in comparable audits, most of the gap came not from "reading more carefully" but from "did we actually sweep these specific things deterministically" — a model-independent, reproducible methodology improvement.

Coverage caps intervention value (borrowed from arXiv 2608.04618): before starting Phase 2 content review, count how many items this audit could actually affect (e.g., "N files this rule change would apply to"). That count caps the maximum value of the intervention before you've even seen the results — if only 3 items are in scope, no amount of review sophistication can produce more improvement than those 3 items allow. Computing this cap up front prevents over-investing deep-review time in low-cap areas, and gives a concrete basis for shifting review effort toward higher-cap areas instead.

Canary mixing (a control group for telling "clean" apart from "the reviewer missed it"): when reviewing a code area, stage one known-clean file and one file with a planted defect from this skill's bundled example pool (scripts/canaries/{clean,seeded}/) using python scripts/canary_mix.py select --pool both --n-clean 1 --n-seeded 1 --stage-dir <scratch>/<run>/bundle --out <scratch>/<run>/canary-manifest.json, then mix the staged files into one or two real review bundles. Don't tell the reviewer dispatch that a control exists or which file it is, and don't count canaries in the Phase 1 denominator. The bundled pool ships with only 2 clean / 2 seeded example files — extend or replace it with pairs representative of the actual codebase before treating a recall number as meaningful. Docs/rules areas have no canary pool by default; skip this step there and note "canary not applied (no pool for this area)" on the map.

Phase 2: Content Review + Rule Dry-Run (Three-Layer Principle)

Structural checks (Phase 1) alone do NOT justify calling something "exhaustive" — exhaustive = structure + content + rule dry-run, three layers. Rules and guards can't be confirmed as actually working just by reading their documentation — only running them against mock input fills in the third layer. The three layers are non-substitutable: structural checks can come back clean while the content is wrong, and the content can be correct while a rule still fails to fire at runtime.

  • Fan out parallel review agents (unlimited breadth for coverage, read-only — never give reviewers edit access)
  • Neutral framing: phrase the review dispatch's goal as "judge whether this bundle satisfies policy P — with equal rigor whether it does or doesn't," not "find problems in this code." An instruction to find problems pulls something out of a clean file too.
  • Single-reviewer batch size cap (borrowed from arXiv 2609.09696 — in a controlled experiment on 150 academic papers, a review recall that held at 50–60% for 1–10 papers collapsed to 2.8% at 48 papers in a single batch): the line above is about how many parallel agents you fan out; this is a separate axis — how many files/documents one reviewer call scans at once. Even when nominal context capacity looks sufficient, the more items you put in one call, the more likely "no findings" stops being an honest negative and becomes a confident fabrication. For large scopes, lean toward splitting bundles small (the exact cap is domain-dependent, so no number is fixed here).
  • Force citations: reviewers must attach a grep/ls output as proof when they claim something is missing — "I can't find it" from memory alone is invalid
  • Anti-false-positive 4-bucket (enforce in the review dispatch's output-format instructions): classify every finding as CONFIRMED / FALSE-POSITIVE (reviewed and dismissed, with a refuting citation) / UNCERTAIN (needs inference — keep it, don't discard, to avoid false negatives) / NIT, each with a reasoning note. CONFIRMED at Critical/High needs 2+ of {condition, impact, reproduction} or it gets downgraded to Medium. FALSE-POSITIVE needs the discarded hypothesis + a refuting citation (command output or a line quote) — "no issue" in one line is not acceptable. An empty false-positive list is not a penalty (state "none dismissed" explicitly — this prevents over-suppression).
  • Kill-test (enforce in the dispatch instructions): before reporting each finding, run one command that tries to refute it, include the output, and add one line: (refutation check: output refutes the finding / output is unrelated and insufficient). If refuted, it moves to the false-positive bucket. "It's missing" claims must be backed by an exhaustive grep across the whole denominator (explicit regex and scope, grep -rn <pattern> <root> — substring matching alone doesn't count). Passing a mock test alone does not count as a refutation (remove the code and re-test instead). The conclusion needs one objective anchor: a rule/linter, an actual execution result, a direct two-point comparison within the reviewed content, or a grep-derived denominator — "it looks like" with no anchor is invalid (inference-requiring cases aren't invalid, they go to UNCERTAIN instead). Anchor-inject the countables: for anything a machine can count, hand reviewers Phase 1's deterministic values as a given anchor up front rather than asking them to re-derive it — this keeps reviewers out of the business of re-counting what a script already settled.
  • Reverse dry-run for guard/lint scripts: seeded-error injection only checks that a rule catches what it should. In the opposite direction, map every check in a guard, linter, or check script to the rule clause (file:line) that justifies it. Report any check with no mapped clause as a false-positive source candidate - checks left without a basis are what block legitimate work.
  • Rule dry-run (third layer): if the target area has rules or guards (linter configs, pre-commit hooks, validation scripts), build an actual mock input (a fixture) and run it through the rule to confirm by execution — not by reading — that it detects or blocks what its documentation claims. Static comparison (does the rule's documentation exist) and content review (does the rule's wording make sense) alone can't prove it fires at runtime — skip this layer and a dead guard (documented but inert) slips through unnoticed.

Phase 3: Judgment — Dismissing False Positives

Personally re-verify every reviewer report before classifying. Common false-positive patterns to check for:

  • Training-cutoff confusion: "this date/version can't exist yet" — re-check against the actual current date
  • "Already exists but reported missing": any reported "gap" must be re-confirmed to actually be missing via grep before being accepted
  • Historical notation mistaken for staleness: an original-version marker or changelog entry is history, not staleness — don't "fix" it
  • Number conflicts: reviewer's number vs. the Phase 1 deterministic number → deterministic wins
  • Composite-accumulation-gate (death-by-thousand-cuts guard) ([borrowed from PHP-AIO, arXiv 2607.15944v1]): even when every individual finding is separately dismissed as FALSE-POSITIVE/UNCERTAIN/NIT, if the same area (same file/module/component) accumulates 3+ UNCERTAIN findings, or 5+ combined (UNCERTAIN+NIT) findings, flag it separately as an "individually-passed, cumulatively-risky" signal — passing each individual threshold does not mean the composite threshold is also safe (structurally identical to the CRITICAL hard-cap principle in agents/code-reviewer.md). A flagged area is not promoted to CONFIRMED, but must be listed at least once in the Phase 4 addition bucket so the user sees it. [The 3/5 thresholds are initial estimates, subject to recalibration once operational data accumulates.]
  • Canary scoring (only for a run that mixed in canaries): save this run's verdicts as [{"file": <path>, "verdict": <final bucket>, "summary": <one line>}] JSON and run python scripts/canary_score.py --manifest <canary-manifest.json> --findings <that JSON>. If a clean control file comes back CONFIRMED, mark every CONFIRMED finding from this run as "needs re-verification" on the coverage map — this does not auto-downgrade them. Drop any finding on a canary file from the fix/addition buckets.
Show full SKILL.md (1,262 more words)Show less

Phase 4: Fix/Addition Split — Gated by Execution Mode

  • Runs only in PROPOSE or APPLY_APPROVED. In AUDIT_ONLY (the default for a bare audit/analysis request), stop after Phase 3 — the coverage map may still list what would land in each bucket, but nothing here executes.
  • Fix bucket (stale numbers, dead references, policy violations, broken parsing — plain factual corrections): listed under PROPOSE; applied only under APPLY_APPROVED, and only for items the user approved (a blanket "apply the fix bucket" covers non-protected paths — a deny-listed path always needs its own explicit approval, see Safety Layers)
  • Addition bucket (new features, structural changes, deletions, upgrades): listed under PROPOSE; executed under APPLY_APPROVED only per-item after explicit approval — never covered by a blanket approval
  • Re-verify after fixing: re-run any affected tests/checkers

Phase 5: Record the Coverage Map

Create or update a coverage-map file (same-day re-run = append a pass section):

  • Table: Area | Method label | Findings/actions — three method labels required: [deterministic] / [LLM judgment] / [close read]
  • State remaining gaps explicitly (what wasn't reviewed, bounded checks, intentional exclusions) — a map with zero remaining gaps deserves suspicion
  • Assumption Ledger ([borrowed from Uncertainty Ledger, arXiv 2607.16112], conditional addition): if a Phase 3 CONFIRMED verdict depends on an unverified assumption, add a separate table to the map — Assumption/Parameter | Status | Evidence needed | Materiality (would it flip the verdict?) | Owner. Five status values: externally-anchored (verified by a third party) / author-calibrated-prior (an adjusted assumption) / assertion-only (unsupported claim) / open-proposal (a TODO) / open-question. If one or more rows are assertion-only, open-proposal, or open-question AND materiality is High (flipping it changes the verdict), downgrade the final label to PARTIAL and name the owner who must resolve it (user / follow-up investigation / tooling). If the CONFIRMED conclusion does not depend on any unverified assumption, the Assumption Ledger may be omitted — if omitted, state "Assumption ledger: N/A (reason)" as one line.
  • Recall line (only if this run used canaries): append the output of python scripts/canary_score.py --summary --skill full-audit (recall: k/n (seeded), RECALL_UNMEASURED while n<5) to the map verbatim.
  • End with 1-3 lines on what methodology was established or fixed during this audit

Scope Boundary

DoesDoes NOT
[BASH] Deterministic sweep (counts/versions/paths/parsing)Compute a harness maturity score (a different tool's job)
[AGENT] Dispatch parallel content review (read-only)Grant reviewers edit access
[EDIT] Apply fix-bucket edits (stale/dead-refs/violations) — only in APPLY_APPROVED modeApply any edit while in AUDIT_ONLY or PROPOSE mode; execute addition-bucket changes without per-item approval (propose-only)
[WRITE] Record the coverage mapDeclare "100% done" while hiding remaining gaps
[READ] Read the prior coverage map and diffSilently include areas the user excluded

Safety Layers

Risky ActionReversibilityApplied Layers
Fix-bucket edit (existing file)high (git)L1+L3 (the mode-gate itself is the L3 confirmation — switching to APPLY_APPROVED is the approval)
Delete/move a file (addition bucket)mediumL1+L3 (explicit per-item user approval required) + mode-gate (APPLY_APPROVED only)
Editing a protected config/rules path (deny-listed)mediumL2 (deny — the model cannot edit it directly) + L3 (the user runs the apply script themselves) + mode-gate: reachable only in APPLY_APPROVED, and only for the specific item the user named, by drafting an edit plus an apply script. There is no path for the model to lift the deny.

Guard-degradation observability: if a Phase 1 checker can't run, don't silently skip it — record ⚠️ check unavailable: [reason] on the coverage map.

Invariants (never violate)

  1. Three-layer exhaustiveness: never report "exhaustive" from structural checks alone. Any area where content review or rule dry-run was skipped gets labeled "[deterministic] only — content/dry-run not executed" on the map. Violation → overstated coverage; the user trusts an area that was never actually reviewed or exercised.
  2. Deterministic wins: when an LLM's count/existence claim conflicts with a script's result, the script wins. Violation → hallucinated numbers get written into the source of truth.
  3. Fixes and additions stay separate: only apply plain factual corrections immediately; everything else is propose-only. Violation → scope creep, and the user's decision rights get bypassed.
  4. Coverage map is mandatory: never declare completion without recording it. A map with zero remaining gaps needs re-review. Violation → nobody in a future session can tell how far the last audit actually went.
  5. Analysis and execution stay in separate modes: a bare audit/analysis request defaults to AUDIT_ONLY and never auto-advances into Phase 4 execution, and a deny-listed path is never touched outside APPLY_APPROVED with that specific item named. Violation → a request to "look at X" silently becomes a request that changed X, including a Protect-Hooks-guarded file.

Error Recovery

FailureDetectionRecovery
Checker script execution failureScript errors outRewrite and retry once → on repeat failure, record "⚠️ check unavailable" on the map, never treat it as passed
Ambiguous scope ("everything" — but everything up to where?)Scope unclear at Phase 0Make the area table explicit and confirm with the user
Reviewers disagree with each other, or disagree with the deterministic sweepContradictory reportsDeterministic wins → adversarial re-check → escalate to the user if still unresolved

Truthful Reporting

  1. No mock deception: "clean" is reported only from an actually-run check's output. Never mark a check that wasn't run as passed.
  2. Bounded checks are labeled as bounded: a sample-only or grep-only check gets its bound stated explicitly on the map.
  3. No silent brokenness: final status is one of WORKING / PARTIAL / BROKEN / BLOCKED (stalled on an external unresolved dependency or pending approval) + a bullet list of remaining gaps or blockers.

Rationalization Table

RationalizationCounter
"The grep came back clean, so this area is done"A clean grep means "pattern not found," not "content is sound." Label it [deterministic] only until content review happens (Invariant 1)
"Three reviewers found the same thing, so it must be true"Agreement from reviewers sharing the same blind spot is a false-consensus trap, not confirmation. Verify with citations before accepting
"I checked this area last week, skip it this time"Diff-based scope reduction is fine; silently skipping with no label is coverage inflation
"It's a small addition, let's just fix it along the way"Violates fix/addition separation (Invariant 3). Batch additions and propose them together
"I'll do the map later if there's time"An audit with no map resets to zero for the next session (Invariant 4)
"A few UNCERTAINs here and there do not matter"Individual passes do not hide accumulated risk — 3+ in the same area triggers the composite-accumulation-gate flag (Phase 3)
"The user said 'audit', so finding an issue and just fixing it right there is helpful"An audit/analysis trigger defaults to AUDIT_ONLY — fixing without an explicit mode-advance conflates analysis with execution (Invariant 5). Report it in the coverage map instead and wait for PROPOSE/APPLY_APPROVED

Output

  • Updated coverage map file
  • Chat report: execution mode used, stated first (AUDIT_ONLY/PROPOSE/APPLY_APPROVED) / list of applied fixes with line anchors (APPLY_APPROVED only) / list of proposed fixes and additions (PROPOSE/APPLY_APPROVED) / verification results (✅⚠️❌) / final status label / remaining gaps / Assumption ledger (if applicable)

Changelog: v1.0 (initial release) → v1.1 (added the coverage-caps-intervention-value step to Phase 1) → v1.2 (added the AUDIT_ONLY/PROPOSE/APPLY_APPROVED execution-mode gate — a bare audit/analysis trigger now defaults to read-only and never auto-advances to Phase 4 fix-bucket execution or a Protect-Hooks deny-lift; those now require an explicit mode-advance from the user) → v1.3 (added the single-reviewer batch-size cap to Phase 2 — batch size is a separate axis from parallel-agent count) → v1.4 (added an opt-in canary-mixing step to Phase 1/2/3/5 — a known-clean and a known-defective file from a bundled example pool are mixed unlabeled into a review bundle, then scored after the fact, as a control for whether "zero findings" means clean or means the reviewer missed it; corrected the deny-lift claim — the model drafts an edit plus an apply script for the user to run, it never lifts a deny itself; added BLOCKED to the Truthful Reporting status enum)

© AlexZio00, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts) in full-audit of AlexZio00/sovereign-skills.

  • SKILL.md
  • .claude-plugin/plugin.json
  • agents/openai.yaml
  • scripts/canaries/README.md
  • scripts/canaries/clean/account_repo.py
  • scripts/canaries/clean/settings_loader.py
  • scripts/canaries/seeded/billing_client.json
  • scripts/canaries/seeded/billing_client.py
  • scripts/canaries/seeded/user_lookup.json
  • scripts/canaries/seeded/user_lookup.py
  • scripts/canary_mix.py
  • scripts/canary_score.py
  • scripts/test_canary_mix.py
  • scripts/test_canary_score.py

Open the folder on GitHubat commit c062683

Compare with similar skills

Full Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Full Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Full Audit this skillAlexZio00/sovereign-skills140—~6.4kAutomated safety check: PassMIT
Interlinked Spec AuditQuentinCody/interlinked-cli178—~3.3kAutomated safety check: PassMIT
Sandbox Lifecycle Debugvercel-labs/vercel-openclaw-archived117—~809Automated safety check: PassMIT
Agentic System Designooiyeefei/ccc495—~7.3kAutomated safety check: PassMIT
Merge ReviewSethGammon/Citadel924—~1.7kAutomated safety check: PassMIT
Canva Core Workflow Bjeremylongshore/tons-of-skills-marketplace2.8k—~950Automated safety check: PassMIT

Similar skills

  • Interlinked Spec Audit

    QuentinCody/interlinked-cli

    Keep prose specs and design docs honest against the code using Interlinked's spec-audit system.

    178 GitHub stars~3.3k tokensUpdated yesterday
    Business, Finance & HRAuto-check passed
  • Sandbox Lifecycle Debug

    vercel-labs/vercel-openclaw-archived

    Official

    Sandbox lifecycle debugging for vercel-openclaw: create, resume, stop, snapshotting, reset, stale-running reconciliation, persistent Sandbox v2 behavior, hot spares, and lifecycle locks.

    117 GitHub stars~809 tokensUpdated 4 mo ago
    Business, Finance & HRAuto-check passed
  • Prescriptive Q&A workflow for designing agentic pipelines, multi-model councils, sub-agent hierarchies, and tool-loop hardening for any domain.

    495 GitHub stars~7.3k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Merge Review

    SethGammon/Citadel

    Reviews pending fleet worktree merges before they're accepted.

    924 GitHub stars~1.7k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Canva Core Workflow B

    jeremylongshore/tons-of-skills-marketplace

    Operate Canva assets, brand-template autofill, and folders with explicit rights and asynchronous reconciliation.

    2.8k GitHub stars~950 tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Mistral Reference Architecture

    jeremylongshore/tons-of-skills-marketplace

    Design a Mistral integration from policy gateway through adapter, queues, state reconciliation, and audited output.

    2.8k GitHub stars~932 tokensUpdated today
    Business, Finance & HRAuto-check passed

More from AlexZio00/sovereign-skills

All 17 skills in this repo
  • Project Overview

    AlexZio00/sovereign-skills

    A skill your agent uses when the user wants a deterministic cross-project status map generated from registered projects' session handoffs.

    140 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Scope

    AlexZio00/sovereign-skills

    Scope definition before implementation — two modes. An agent skill from AlexZio00/sovereign-skills.

    140 GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Project Init

    AlexZio00/sovereign-skills

    Interview-based project setup — generates CLAUDE.md, ROADMAP, .gitignore, .env.example from scratch.

    140 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check: notes
  • Collab Audit

    AlexZio00/sovereign-skills

    This skill should be used when the user types /collab-audit or requests AI collaboration diagnosis.

    140 GitHub stars~8k tokensUpdated yesterday
    Auto-check passed
  • Doc Drift

    AlexZio00/sovereign-skills

    A skill your agent uses when the user wants to audit the memory and documents Claude Code loads into context — CLAUDE.md (user global + project + nested), MEMORY.md, @imports, .claude/skills…

    140 GitHub stars~6.2k tokensUpdated yesterday
    Auto-check passed
  • Session Checkpoint

    AlexZio00/sovereign-skills

    A skill your agent uses when saving session state before context compaction, switching tasks, or ending a session.

    140 GitHub stars~14k tokensUpdated yesterday
    Auto-check passed

Questions about Full Audit

What does Full Audit do?

Exhaustive, denominator-driven audit of an entire area (codebase, docs, memory, skills, DB, config). Full Audit is an agent skill from AlexZio00/sovereign-skills. Exhaustive, denominator-driven audit of an entire area (codebase, docs, memory, skills, DB, config).

When should I use Full Audit?

Full Audit fits situations like: tasks that involve Accounting and bookkeeping; tasks that involve Code review; tasks that involve Citation management.

How do I install Full Audit in Claude Code?

Run `npx skills add AlexZio00/sovereign-skills --skill full-audit -a claude-code`. Or copy the skill folder (full-audit in AlexZio00/sovereign-skills) into .claude/skills/full-audit in your project. Claude Code loads it when a task matches its description.

How do I install Full Audit in Codex?

Run `npx skills add AlexZio00/sovereign-skills --skill full-audit -a codex`. Or copy the skill folder (full-audit in AlexZio00/sovereign-skills) into .agents/skills/full-audit in your project. Codex loads it when a task matches its description.

Can I use Full Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexZio00/sovereign-skills --skill full-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/full-audit, .gemini/skills/full-audit, .github/skills/full-audit and .opencode/skills/full-audit in your project.

What does Full Audit need to run?

Going by SKILL.md and its folder, Full Audit needs Python for the scripts in its folder and the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Full Audit access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Full Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Full Audit use?

Full Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Full Audit use?

About 6.4k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Full Audit?

Skills that share tags, products or a category with Full Audit: Interlinked Spec Audit (QuentinCody/interlinked-cli, 178 stars), Sandbox Lifecycle Debug (vercel-labs/vercel-openclaw-archived, 117 stars), Agentic System Design (ooiyeefei/ccc, 495 stars) and Merge Review (SethGammon/Citadel, 924 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Full Audit?

AlexZio00 (a GitHub user) maintains it in AlexZio00/sovereign-skills, which has 140 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: AlexZio00/sovereign-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.