Agent skill

Ax Audit

by mblode in mblode/agent-skills

Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 24 rules and a ship verdict.

MITAuto-check passed

Install Ax Audit

skills CLI
$ npx skills add mblode/agent-skills --skill ax-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mblode/agent-skills ax-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ax-audit .claude/skills/ax-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ax-audit
GitHub stars
143
Token cost
~2.8k tokens
SKILL.md length
1,331 words
Files
37 (incl. references)
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 24 rules and a ship verdict.

  • Asked for an AX audit
  • SKILL.md covers Contents, Audit workflow, Two rule layers and Tiers and verdict, plus 5 more sections
  • Calls rg
  • Review an agent approval flow

What it does

Ax Audit is an agent skill from mblode/agent-skills. Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 24 rules and a ship verdict. Use when asked for an "AX audit", to review an agent approval flow, or whether an agent can operate the product.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 39 other files, including reference files (for example `evals/evals.json`, `evals/evaluation-scenarios.md` and `references/ax-evolution-curve.md`).

The repository describes itself as: Nobody ships AI slop on purpose. These skills make sure you don’t. The licence is MIT.

When your agent uses it

  • Asked for an AX audit
  • Review an agent approval flow
  • Whether an agent can operate the product

Example prompts

  • “AX audit”
  • “Use the ax-audit skill to audit agentic products for tool parity, authority, approval payloads, recovery, and trust using 24 rules and a ship verdict”
  • “/ax-audit”

What it can do on your machine

Read from SKILL.md and the folder at commit cef4cfa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ax Audit loads about 2.8k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 1,331 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mblode/agent-skills at commit cef4cfa, republished under its MIT licence (© mblode). 1,331 words, ~2,770 tokens.

Download SKILL.mdSave it as .claude/skills/ax-audit/SKILL.md (or your agent's skills folder). This skill also uses 36 other files; get the full folder from GitHub.
name
ax-audit
description
Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 24 rules and a ship verdict. Use when asked for an "AX audit", to review an agent approval flow, or whether an agent can operate the product.

AX Audit

Feature-level reviewer for apps where an agent acts for the user. One question: does it earn trust, and where does it break?

  • IS: rules-based audit of agentic surfaces (chat, tool execution, config, dashboards) across architecture (rules-arch/) and trust (rules-ax/), ending in a ship-readiness verdict plus an AX Relationship Summary.
  • IS NOT: traditional frontend UX (use ui-design Audit mode); developer-facing API, CLI, or type ergonomics (use dx-audit); public site or docs agent scores (use agent-ready); agent instruction files (use agents-md); what the product should do before it exists (use product-design).

No agentic features in scope? Stop. AX rules against forms and lists are noise.

Contents

Audit workflow

text
AX Audit progress:
- [ ] Step 1: Scope, via the diff against the PR base merge-base (PR mode) or explicit path (full sweep)
- [ ] Step 2: Detect agentic features per references/feature-playbooks.md
- [ ] Step 3: Run each detected feature's playbook in order, plus the diff-wide checks (PR mode only)
- [ ] Step 4: For each check, load the rule file and follow its detection recipe
- [ ] Step 5: Tier each finding per references/ship-readiness.md (rule override table wins)
- [ ] Step 6: Render verdict + findings + AX Relationship Summary per references/output-format.md
- [ ] Step 7: Run the audit self-check and report its evidence counts

PR-mode scope is the diff plus the tool definitions and orchestrator it touches. Findings in untouched files belong in a full sweep, not this verdict. Playbook annotations are a scan copy; the rule file is authoritative. parity-orphan-ui-action runs on every PR-mode audit and never in a full sweep, where there is no diff for it to read.

Rule greps name the most common identifiers, not every framework's spelling. When a grep misses in code that plainly does the thing (a gate, a stream, a tool result), check references/framework-signals.md for the stack's name for it before recording unknown.

Two rule layers

LayerFolderRulesLoad when a playbook names
1: Architecturerules-arch/11rules-arch/<category>-<slug>.md
2: Experiencerules-ax/13rules-ax/<category>-<slug>.md

Categories: arch = parity, granularity, context, comm; ax = trust, control, context, comm. Shared prefixes are different rules: rules-arch/comm-no-approval-gate.md (no gate on the execution path) is not rules-ax/control-no-approval-gate.md (gate exists, stakes are wrong).

Run Layer 1 comm/parity and Layer 2 control/trust first. They hold the blockers. Category map: rules-arch/_sections.md, rules-ax/_sections.md.

PriorityLayerCategoryPrefixRules
1archCommunicationcomm-3
2archParityparity-4
3axControlcontrol-4
4axTrusttrust-3
5archContextcontext-2
6axCommunicationcomm-4
7axContextcontext-2
8archGranularitygranularity-2

Tiers and verdict

Three tiers: release-blocker, fix-this-sprint, backlog. Definitions, the generic surface bump, and verdict logic live in references/ship-readiness.md.

Precedence: the rule's own surface-override table > the generic bump > defaultTier. Apply at most one adjustment.

Verdict: ✅ READY (0 blockers, ≤3 sprint) · ⚠️ READY WITH FOLLOW-UP (0 blockers, ≥4 sprint) · ❌ NOT READY (≥1 blocker) · 🚫 INCOMPLETE (self-check failed).

Blockers outrank an incomplete audit. With ≥1 release-blocker and a failed self-check, report ❌ NOT READY and note the self-check failure beneath it: the blockers are established findings and stay actionable, while 🚫 reads as "nothing was learned" and sends the reader away. Reserve 🚫 for an audit with no blockers whose coverage you cannot vouch for.

AX Relationship Summary

Render after findings when any agentic feature was detected. Findings serve engineers; this serves designers and PMs. Four fields: evolution stage (behavior, not a label), trust signal (high/moderate/low plus one-line reason), key gap (one actionable sentence), trust question (one question only research can answer).

Reference files

FileRead when
references/feature-playbooks.mdSteps 2-3: detection heuristics, per-feature ordered checks, diff-wide checks
references/framework-signals.mdStep 4, when the code uses AI SDK, MCP, the Claude Agent SDK, or AG-UI: where the gate, the stream, the completion signal, and the structured result live in each, with the spec defaults the rules lean on
references/ship-readiness.mdStep 5: tier definitions, precedence, generic surface bump, verdict logic
references/output-format.mdStep 6: findings JSON schema, summary schema, terminal rendering
references/ax-evolution-curve.mdWriting the AX Relationship Summary: stage, action depth, costume vs intelligence, and the arguments with no rule that land in keyGap
rules-arch/_sections.mdLayer 1 categories, default tiers, co-firing pairs
rules-ax/_sections.mdLayer 2 categories, default tiers, co-firing pairs
Show full SKILL.md (726 more words)Show less

Gotchas

  • Scope before rules. Running all 24 rules repo-wide on a 3-file PR buries a new release-blocker under pre-existing backlog noise; the verdict stops meaning "can this PR merge."
  • The rule's override table is authoritative. comm-no-intent-handshake defaults to fix-this-sprint but its table says release-blocker on tool execution. Stacking the generic "+1 tier on tool execution" bump on an explicit override double-upgrades backlog findings into blockers.
  • A stop button not wired to AbortController.abort() is a false affordance. control-no-escape-hatch still fails: verify the abort() call, not the button label, or the audit passes a UI that lies to users.
  • A client stop() that only closes the stream leaves the executor running. useChat().stop() aborts the fetch. Unless the route passes req.signal into streamText({ abortSignal }) and tool execute reads it, the server finishes every remaining tool call after the user pressed Stop. Trace the signal to the loop, not to the button.
  • Tool annotations are hints, not stakes. MCP tells clients to treat annotations from untrusted servers as untrusted; a gate that auto-approves on a third-party server's readOnlyHint: true has handed the gate to that server. comm-no-approval-gate fails it. The spec defaults (destructiveHint: true, readOnlyHint: false) are the fail-closed baseline.
  • A framework approval flag is the gate's input, not the gate. AI SDK toolApproval: "user-approval" emits a tool-approval-request part and waits. A UI that never renders state === "approval-requested", or answers it with addToolApprovalResponse({ approved: true }) on arrival, has a gate in the type system and none for the user. Check the renderer and the response call, not the option.
  • Absence checks need a recorded file list. "Find components lacking X" greps return nothing both when everything passes and when nothing was scanned. List candidate files first (rg -l <feature-pattern>), check each for the counter-pattern, and cite the file list as evidence.
  • detection: observational rules cannot fail on grep evidence alone. granularity-static-api-mapping, control-over-conversational, comm-no-generative-momentum, and the uncertainty-gradient half of trust-no-confidence-cues need interaction-flow judgment; on static evidence alone, return unknown with a reason, not fail.
  • Gates fail in three separate places. Absent from the path (comm-no-approval-gate), present but mismatched to the stakes (control-no-approval-gate), or correct and unreadable (control-thin-approval-payload). Report the first that holds and fix in that order.
  • Interactive gates do not cover unattended runs. Cron, webhook, and queue entry points reach the same executor with nobody to prompt. comm-unrequested-action-no-consent audits that path; evidence names the entry point, not the executor.
  • ax-audit-ignore:<slug> comments count as suppressed, not pass. Report the count in the verdict block; a suppression with no reason is itself a warn.
  • Don't inflate tiers. comm-no-generative-momentum and granularity-static-api-mapping default to backlog. One finding promoted to release-blocker flips the whole PR to ❌ NOT READY, so promoting cosmetic ones trains the team to ignore the verdict entirely.
  • Don't duplicate ui-design Audit mode findings. "Missing loading state" and "form clears on error" are its territory; duplicating them trains engineers to dismiss the whole AX report.
  • A Personally Intelligent agent that only ever suggests has plateaued. Memory stage is not trust. Name the highest action rung in evolutionStage.behavior or the summary flatters a polite chatbot.

Audit self-check

Flag the audit INCOMPLETE if any of these hold, and include the counts as evidence (planned vs. run rules per playbook, unknown rate, suppressed count):

  • Fewer rules ran than the playbooks planned
  • More than 30% of rules returned unknown. Count only unknown here, never out-of-scope: a rule whose layer is absent from the scope you were given was answered correctly, and a narrow diff is the scope Step 1 asks for. Marking a correctly scoped audit INCOMPLETE buries its real blockers under a verdict that reads as "we learned nothing".
  • Any fail/warn finding lacks file:line evidence or a fix snippet
  • Every finding landed in the same tier (suspect blanket assignment)
  • AX Relationship Summary is missing despite detected agentic features
  • ui-design Audit mode: traditional frontend UX around agentic surfaces; run both on agentic feature PRs
  • dx-audit: same files, different reader. This skill asks whether an agent can operate and recover; dx-audit asks whether a human adopting the API, CLI, or types finds it ergonomic
  • agent-ready: whether public docs and HTTP APIs are discoverable to coding agents; this skill audits in-product agent UX
  • product-design: what the agentic feature should do, before this audit
  • agents-md: CLAUDE.md / AGENTS.md instruction files

Maintenance only: evals/evals.json and evals/evaluation-scenarios.md hold regression scenarios for changes to this skill; neither loads during a user task.

© mblode, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 36 other files (references) in skills/ax-audit of mblode/agent-skills.

  • SKILL.md
  • evals/evals.json
  • evals/evaluation-scenarios.md
  • references/ax-evolution-curve.md
  • references/feature-playbooks.md
  • references/framework-signals.md
  • references/output-format.md
  • references/ship-readiness.md
  • rules-arch/_sections.md
  • rules-arch/_template.md
  • rules-arch/comm-no-approval-gate.md
  • rules-arch/comm-no-completion-signal.md
  • rules-arch/comm-no-progress-visibility.md
  • rules-arch/context-no-checkpoint-resume.md
  • rules-arch/context-starvation.md
  • rules-arch/granularity-static-api-mapping.md
  • rules-arch/granularity-workflow-shaped-tool.md
  • rules-arch/parity-crud-incomplete.md
  • … and 19 more

Open the folder on GitHubat commit cef4cfa

Compare with similar skills

Ax Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ax Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ax Audit this skillmblode/agent-skills143—~2.8kAutomated safety check: PassMIT
Production Auditaffaan-m/ECC275k1 repos~1.9kAutomated safety check: PassMIT
Production Auditsickn33/agentic-awesome-skills47k1 repos~2.7kAutomated safety check: PassMIT
Payloadpayloadcms/payload45k5 repos~6.2kAutomated safety check: PassMIT
Auditing LLM Gateway ParityPostHog/posthog40k—~1.2kAutomated safety check: PassCustom licence
Production Code Auditdavila7/claude-code-templates32k7 repos~3.9kAutomated safety check: WarnMIT

Similar skills

  • Production Audit

    affaan-m/ECC

    Local-evidence production readiness audit for shipped apps, pre-launch reviews, post-merge checks, and "what breaks in prod?" questions without sending repo data to an external audit service.

    275k GitHub starsUsed in 1 repo~1.9k tokens
    Product & Project ManagementAuto-check passed
  • Production Audit

    sickn33/agentic-awesome-skills

    Audit a shipped repo for production-readiness gaps across RLS, webhooks, secrets, grants, Stripe idempotency, mobile UX, and deployment health.

    47k GitHub starsUsed in 1 repo~2.7k tokens
    Backend & APIsAuto-check passed
  • Payload

    payloadcms/payload

    A skill your agent uses when working with Payload projects (payload.config.ts, collections, fields, hooks, access control, Payload API).

    45k GitHub starsUsed in 5 repos~6.2k tokens
    Backend & APIsAuto-check passed
  • Official

    Audits services/llm-gateway against PostHog/ai-gateway and updates services/llm-gateway/PARITY.md from current implementation evidence.

    40k GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Production Code Audit

    davila7/claude-code-templates

    Autonomously deep-scan entire codebase line-by-line, understand architecture and patterns, then systematically transform it to production-grade, corporate-level professional quality with optimizations

    32k GitHub starsUsed in 7 repos~3.9k tokens
    SecurityAuto-check: warnings
  • Product

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing e-commerce product pages or implementing structured data for a shop.

    74k GitHub stars~592 tokensUpdated 2 days ago
    Marketing & SEOAuto-check passed

More from mblode/agent-skills

All 28 skills in this repo
  • Agent Ready

    mblode/agent-skills

    Implements agent-readiness on public sites and docs from Mintlify Agent Score, AFDocs, Is Agentic, Is It Agent Ready, or url-discovery-bench reports, or from server logs of agents 404ing on guessed…

    143 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • Agent Skills Creator

    mblode/agent-skills

    Creates and improves portable Agent Skills with a validator, routing scenarios, and evidence-based keep, cut, merge, or retire decisions.

    143 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed
  • Chat History

    mblode/agent-skills

    Recovers decisions, previous fixes, research, and what followed a prompt from past AI conversations, with source evidence.

    143 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • CI Speedup

    mblode/agent-skills

    Cuts the wait from push to green by measuring a pipeline's critical path from run timestamps, then splitting, sharding, trimming setup and sharing test module state, with a before/after ledger.

    143 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • PR Babysitter

    mblode/agent-skills

    Monitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes.

    143 GitHub stars~3.4k tokensUpdated 2 days ago
    Auto-check passed
  • App Verification

    mblode/agent-skills

    Builds and maintains a repo's own verification harness (verify CLI, doctor, worktree isolation, feature map, seed data) and a reproduce-first bug handoff.

    143 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed

Questions about Ax Audit

What does Ax Audit do?

Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 24 rules and a ship verdict. Ax Audit is an agent skill from mblode/agent-skills. Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 24 rules and a ship verdict.

When should I use Ax Audit?

Ax Audit fits situations like: asked for an AX audit; review an agent approval flow; whether an agent can operate the product.

How do I install Ax Audit in Claude Code?

Run `npx skills add mblode/agent-skills --skill ax-audit -a claude-code`. Or copy the skill folder (skills/ax-audit in mblode/agent-skills) into .claude/skills/ax-audit in your project. Claude Code loads it when a task matches its description.

How do I install Ax Audit in Codex?

Run `npx skills add mblode/agent-skills --skill ax-audit -a codex`. Or copy the skill folder (skills/ax-audit in mblode/agent-skills) into .agents/skills/ax-audit in your project. Codex loads it when a task matches its description.

Can I use Ax Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mblode/agent-skills --skill ax-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ax-audit, .gemini/skills/ax-audit, .github/skills/ax-audit and .opencode/skills/ax-audit in your project.

What does Ax Audit need to run?

Going by SKILL.md and its folder, Ax Audit needs the command-line tools its instructions call (rg).

Does Ax Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ax Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ax Audit use?

Ax Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ax Audit use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.2k tokens, read only when the agent opens those files.

What are the alternatives to Ax Audit?

Skills that share tags, products or a category with Ax Audit: Production Audit (affaan-m/ECC, 275k stars), Production Audit (sickn33/agentic-awesome-skills, 47k stars), Payload (payloadcms/payload, 45k stars) and Auditing LLM Gateway Parity (PostHog/posthog, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ax Audit?

mblode (a GitHub user) maintains it in mblode/agent-skills, which has 143 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 6, 2026.

Source: mblode/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.