Agent skill

Audit

by leanEthereum in leanEthereum/leanSpec

Read-only, multi-agent audit of the leanSpec codebase. An agent skill from leanEthereum/leanSpec.

MITAuto-check passedAgent Workflows

Install Audit

skills CLI
$ npx skills add leanEthereum/leanSpec --skill audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install leanEthereum/leanSpec audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/leanEthereum/leanSpec.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/audit .claude/skills/audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit
GitHub stars
143
Token cost
~2.5k tokens
SKILL.md length
1,023 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Read-only, multi-agent audit of the leanSpec codebase. An agent skill from leanEthereum/leanSpec.

  • Works in 8 steps: Dead code — unused functions, classes,… → Over-abstraction & helper sprawl — a… → Stdlib & idiom simplification —… → …
  • Tasks that involve Multi-agent orchestration
  • SKILL.md covers Hard rules, What counts as a finding, The lenses (sub-agents) and Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Audit is an agent skill from leanEthereum/leanSpec. Read-only, multi-agent audit of the leanSpec codebase. Fans out the py-architect, consensus-researcher, code-tester, and doc-writer agents across the source tree to find dead code, over-abstraction, stdlib simplifications, test gaps, safety/security defects, and documentation rot, then synthesizes a precise, prioritized AUDITREPORT.md. Never modifies code.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Multi-agent orchestration. It works with Ethereum. The repository describes itself as: Lean Ethereum protocol specifications and test vector generation framework. The licence is MIT.

When your agent uses it

  • Tasks that involve Multi-agent orchestration

Example prompts

  • “/audit”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Dead code — unused functions, classes, methods, variables, parameters, imports,
  2. Over-abstraction & helper sprawl — a function/class with exactly one caller, a
  3. Stdlib & idiom simplification — hand-rolled code that a single standard-library
  4. Modernization — Python 3.12 / Pydantic v2 features that read better: match,
  5. Refactoring & reorganization — oversized modules, mixed concerns, misplaced types,
  6. Test gaps — uncovered branches, missing boundary/adversarial/error-path cases,
  7. Safety & security — the gravest class. Consensus safety/liveness hazards,
  8. Naming & documentation — abbreviations, vague placeholder names, stale or missing

What it can do on your machine

Read from SKILL.md and the folder at commit 0b7d33e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit loads about 2.5k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 1,023 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from leanEthereum/leanSpec at commit 0b7d33e, republished under its MIT licence (© leanEthereum). 1,023 words, ~2,519 tokens.

Download SKILL.mdSave it as .claude/skills/audit/SKILL.md (or your agent's skills folder).
name
audit
description
Read-only, multi-agent audit of the leanSpec codebase. Fans out the py-architect, consensus-researcher, code-tester, and doc-writer agents across the source tree to find dead code, over-abstraction, stdlib simplifications, test gaps, safety/security defects, and documentation rot, then synthesizes a precise, prioritized AUDIT_REPORT.md. Never modifies code.

/audit — multi-agent codebase audit

Produce a rigorous, evidence-backed audit of leanSpec and write it to AUDIT_REPORT.md at the repository root. This skill never edits source, tests, or docs — its only artifact is the report. Findings become follow-up pull requests.

The bar is the one used for life-critical software: cryptography, spacecraft flight software (NASA/JPL "Power of Ten", DO-178C), and formally-verified kernels (seL4). A reference Ethereum specification is read by implementers across the ecosystem; an ambiguous line or a dead abstraction here propagates into every client. Audit as if a client bug would trace back to this exact line.

Hard rules

  • Read-only. Do not call Edit, Write, or NotebookEdit on anything under src/, tests/, packages/, or docs/. The only file you create is AUDIT_REPORT.md.
  • Evidence or it does not exist. Every finding cites file:line and quotes the offending code. No vague "consider improving error handling" entries.
  • Recommend, do not rewrite. Show the fix as a short illustrative snippet, not a finished patch. The report drives PRs; it is not itself a PR.
  • Respect the project laws. The audit's own standard is CLAUDE.md plus .claude/rules/*. Flag violations of those rules; never recommend anything that breaks them (no backward-compat shims, no abbreviations, no TYPE_CHECKING, test tree mirrors source, forks tested by vectors not pytest, full-equality assertions).

What counts as a finding

Hunt across these dimensions. Each maps to a lens (and an agent) below.

  1. Dead code — unused functions, classes, methods, variables, parameters, imports, constants, type aliases, or whole modules. The recommendation is always delete (no deprecation shims — see CLAUDE.md). Confirm with a repo-wide usage search before reporting; a symbol used only in tests, re-exported in __init__.py, or referenced by string is not dead.
  2. Over-abstraction & helper sprawl — a function/class with exactly one caller, a wrapper that only forwards, an indirection the reader must chase. Recommend inlining. This is a top priority: helpers tax every future reader.
  3. Stdlib & idiom simplification — hand-rolled code that a single standard-library call replaces (itertools, functools, collections, bisect, math.ceil/floor, enumerate, zip, comprehensions, dataclasses, Enum). Quantify the lines saved. (Respect repo taste: math.ceil/math.floor over (a+b-1)//b — see project memory.)
  4. Modernization — Python 3.12 / Pydantic v2 features that read better: match, Self, @override, PEP 695 type params, StrEnum, frozen models, computed fields, field validators. Only where it genuinely clarifies; never churn for fashion.
  5. Refactoring & reorganization — oversized modules, mixed concerns, misplaced types, a class living in the wrong file, inconsistent module shape across siblings.
  6. Test gaps — uncovered branches, missing boundary/adversarial/error-path cases, weak assertions (substring instead of full-equality), over-mocking, and especially test-tree mirroring violations and any pytest that should be a consensus vector.
  7. Safety & security — the gravest class. Consensus safety/liveness hazards, non-determinism (set/dict iteration order, Math.random-style, float in consensus), unvalidated external input, integer over/underflow, unbounded allocation from wire data, panics on attacker-controlled paths, crypto misuse (non-constant-time compares on secrets, nonce reuse, weak domain separation, RFC/test-vector deviations).
  8. Naming & documentation — abbreviations, vague placeholder names, stale or missing docstrings, comments that restate code, missing "why" on a non-obvious constant.

The lenses (sub-agents)

Launch these four project agents. Each owns a lens but reads the whole tree; the orchestrator sharded the tree only so the work parallelizes.

  • py-architect → dimensions 1–5 (dead code, helper sprawl, stdlib, modernization, refactoring). It is the simplicity guardian; trust its instinct to inline and delete.
  • consensus-researcher → dimension 7 (safety & security) plus spec-fidelity of the consensus and crypto logic. It reasons about adversaries, finality, and soundness.
  • code-tester → dimension 6 (test gaps, mirroring, assertion quality, vector vs. pytest placement).
  • doc-writer → dimension 8 (naming clarity, docstring accuracy, comment noise, reorganization for comprehension).
Show full SKILL.md (434 more words)Show less

Workflow

  1. Scope. Read the /audit argument:

    • No argument → audit the full tree (src/lean_spec/ and packages/).
    • A path → audit only that subtree.
    • A subsystem name (e.g. crypto, ssz, forks, networking, chain, packages) → map it to its directory and audit that.
  2. Shard. Split the in-scope tree into coherent subsystems so agents run in parallel. The natural shards:

    • src/lean_spec/spec/crypto/ (XMSS, hashing, signatures, aggregation)
    • src/lean_spec/spec/forks/ (state transition, fork choice, containers, validator duties, aggregation)
    • src/lean_spec/node/networking/ (gossipsub, reqresp, quic, discovery)
    • src/lean_spec/node/chain/ (clock, store, sync) and the rest of node/
    • src/lean_spec/cli/, base.py, config.py, log.py, __main__.py, spec/observability/
    • packages/testing/ (the consensus-testing framework)
  3. Fan out. For each shard, spawn the relevant lenses in parallel (one message, multiple Agent calls). Give every agent the same contract:

    • It is read-only. It produces findings, not edits.
    • It returns findings in the exact schema below — nothing else.
    • It must run a usage search before claiming any symbol is dead.
    • It must read CLAUDE.md and .claude/rules/* first and judge against them.

    For a large audit, prefer the Workflow tool: a pipeline of (shard × lens) → adversarial verification → synthesis scales better than manual fan-out and keeps each finding independently checked. Manual Agent fan-out is fine for a single subsystem.

  4. Verify (anti-false-positive pass). Before a finding lands in the report, confirm it. A "dead" symbol must have zero non-test references (grep/usage search across src, tests, packages, and string-keyed registries). A "simplification" must be behavior-preserving — call out any semantic difference (e.g. itertools laziness, exception types). Discard anything you cannot stand behind. False positives destroy the report's credibility faster than missing findings.

  5. Synthesize. Merge agent outputs, dedupe overlaps, drop unverified items, and write AUDIT_REPORT.md in the format below. Sort the backlog by severity then effort so the top of the list is the obvious first PR.

Finding schema

Every agent returns findings as a list of records with these fields:

ID:           <SHARD>-<NN>            e.g. CRYPTO-03
Location:     path/to/file.py:LINE    (range if multi-line)
Category:     dead-code | over-abstraction | stdlib | modernization |
              refactor | test-gap | safety | security | naming | docs
Severity:     Critical | Major | Minor
Title:        one line, imperative ("Inline single-use `_pad_chunk` helper")
Evidence:     a short quoted snippet of the offending code
Why:          the concrete cost — reader confusion, attack surface, drift risk,
              lines wasted. Tie safety items to a property (safety/liveness/soundness).
Recommendation: the fix, with a minimal before/after snippet. State lines saved.
Effort:       Trivial | Small | Medium | Large
References:   optional — paper, RFC, CPython docs, or a sibling file that does it right

Severity rubric:

  • Critical — can cause consensus divergence, fund loss, a crypto break, a panic on attacker-controlled input, or non-determinism in state transition / fork choice.
  • Major — meaningful correctness, test-coverage, or maintainability risk: a real dead-code path, an untested error branch, an abstraction that obscures protocol logic.
  • Minor — clarity, naming, idiom, single-line simplification.

Report format (AUDIT_REPORT.md)

markdown
# leanSpec Audit Report

_Scope: <what was audited> · Commit: <git short sha> · Date: <date>_

## Executive summary

3–6 sentences: overall health, the single most important finding, and the themes
(e.g. "helper sprawl in networking", "untested error paths in SSZ decode").

## Findings by severity

A table: ID · Severity · Category · Location · Title. Sorted Critical → Minor.

## Detailed findings

One subsection per finding, rendered from the schema above. Group by subsystem.
Lead with Critical and Major.

## Recommended PR backlog

An ordered checklist mapping findings to PRs, grouped so each line is one shippable
change. Note dependencies ("do X before Y"). This is the section the maintainer acts on.

## Coverage & limitations

What was audited, what was skipped and why, and any finding the auditors could not fully
verify (flagged so a human checks it). No silent gaps.

Quality gates for the report itself

  • Lean. No filler, no restating the code, no "it is recommended that". One idea per line.
  • Every finding is independently actionable and independently verifiable.
  • Concrete over abstract: show the snippet, name the stdlib call, count the lines.
  • Honest about uncertainty: a low-confidence finding says so rather than overclaiming.
  • The reader should be able to open the report, pick the top item, and start a PR with no further investigation.

© leanEthereum, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/audit of leanEthereum/leanSpec.

Open the folder on GitHubat commit 0b7d33e

Compare with similar skills

Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit this skillleanEthereum/leanSpec143—~2.5kAutomated safety check: PassMIT
NookplotBankrBot/skills1.2k—~3.2kAutomated safety check: PassMIT
Orca CLIstablyai/orca87k2 repos~593Automated safety check: PassMIT
Paseo Advisor Second Opiniongetpaseo/paseo20k1 repos~756Automated safety check: PassCustom licence
O2 Review Loopopenobserve/openobserve22k—~3.7kAutomated safety check: PassAGPL-3.0
Paseo Committeegetpaseo/paseo20k1 repos~496Automated safety check: PassCustom licence

Similar skills

  • Nookplot

    BankrBot/skills

    Decentralized coordination network for AI agents on Base (Ethereum L2).

    1.2k GitHub stars~3.2k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Orca CLI

    stablyai/orca

    Operate Orca-managed worktrees, folder contexts, terminals, repos, automations, artifacts, skill sharing, worktree comments, and Orca's embedded browser…

    87k GitHub starsUsed in 2 repos~593 tokens
    Agent WorkflowsAuto-check passed
  • Launches one separate agent through Paseo to give a second opinion on the current task, with a self-contained briefing and no permission to edit files.

    20k GitHub starsUsed in 1 repo~756 tokens
    Agent WorkflowsAuto-check passed
  • O2 Review Loop

    openobserve/openobserve

    Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.

    22k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Paseo Committee

    getpaseo/paseo

    Forms a two-agent committee with contrasting profiles to analyze a stuck problem in parallel, reconcile their views and return a consensus plan without editing files.

    20k GitHub starsUsed in 1 repo~496 tokens
    Agent WorkflowsAuto-check passed
  • Mission Control Agent API

    builderz-labs/mission-control

    Teaches an agent to use the Mission Control dashboard API: register, send heartbeats, fetch assigned tasks, report progress and disconnect, with API key auth.

    6.3k GitHub stars~2.1k tokensUpdated 9 days ago
    Agent WorkflowsAuto-check passed

More from leanEthereum/leanSpec

All 9 skills in this repo
  • Client Test

    leanEthereum/leanSpec

    Run leanSpec fixtures against a client implementation. An agent skill from leanEthereum/leanSpec.

    143 GitHub stars~594 tokensUpdated 1 mo ago
    Auto-check passed
  • Review

    leanEthereum/leanSpec

    Quick-reference checklist for code review conventions in leanSpec

    143 GitHub stars~468 tokensUpdated 1 mo ago
    Auto-check passed
  • Spec Diff

    leanEthereum/leanSpec

    Show what changed in leanSpec between devnet versions or HEAD

    143 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Workflows

    leanEthereum/leanSpec

    Common developer workflows, commands, and troubleshooting for leanSpec

    143 GitHub stars~496 tokensUpdated 1 mo ago
    Auto-check passed
  • Doc

    leanEthereum/leanSpec

    Document code in this repository.

    143 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Fill

    leanEthereum/leanSpec

    Generate consensus layer test fixtures

    143 GitHub stars~202 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Categories

Questions about Audit

What does Audit do?

Read-only, multi-agent audit of the leanSpec codebase. An agent skill from leanEthereum/leanSpec. Audit is an agent skill from leanEthereum/leanSpec. Read-only, multi-agent audit of the leanSpec codebase.

When should I use Audit?

Audit fits situations like: tasks that involve Multi-agent orchestration.

How do I install Audit in Claude Code?

Run `npx skills add leanEthereum/leanSpec --skill audit -a claude-code`. Or copy the skill folder (.claude/skills/audit in leanEthereum/leanSpec) into .claude/skills/audit in your project. Claude Code loads it when a task matches its description.

How do I install Audit in Codex?

Run `npx skills add leanEthereum/leanSpec --skill audit -a codex`. Or copy the skill folder (.claude/skills/audit in leanEthereum/leanSpec) into .agents/skills/audit in your project. Codex loads it when a task matches its description.

Can I use Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add leanEthereum/leanSpec --skill audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit, .gemini/skills/audit, .github/skills/audit and .opencode/skills/audit in your project.

What does Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Audit is instructions for the agent only. Our summary lists: Python 3.

Does Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit use?

Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit?

Skills that share tags, products or a category with Audit: Nookplot (BankrBot/skills, 1.2k stars), Orca CLI (stablyai/orca, 87k stars), Paseo Advisor Second Opinion (getpaseo/paseo, 20k stars) and O2 Review Loop (openobserve/openobserve, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit?

leanEthereum (a GitHub organization) maintains it in leanEthereum/leanSpec, which has 143 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 2, 2026.

Source: leanEthereum/leanSpec on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.