Agent skill

Agentic System Audit

by boadij in boadij/pi-herdsman

Audit an agentic software system end to end for contradictions, gaps, inconsistent contracts, capability mismatches, instruction conflicts, ambiguous results, stale documentation, unsafe lifecycle…

Apache-2.0Auto-check passedAgent Workflows

Install Agentic System Audit

skills CLI
$ npx skills add boadij/pi-herdsman --skill agentic-system-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install boadij/pi-herdsman agentic-system-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/boadij/pi-herdsman.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/agentic-system-audit .claude/skills/agentic-system-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentic-system-audit
GitHub stars
131
Token cost
~5.8k tokens
SKILL.md length
2,274 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit an agentic software system end to end for contradictions, gaps, inconsistent contracts, capability mismatches, instruction conflicts, ambiguous results, stale documentation, unsafe lifecycle…

  • Works in 8 steps: Verdict → Executive summary → Findings → …
  • Reviewing the overall health and coherence of an agent/tool system
  • SKILL.md covers Preferred topology, Leaf fallback, Provenance and Reachability, plus 17 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agentic System Audit is an agent skill from boadij/pi-herdsman. Audit an agentic software system end to end for contradictions, gaps, inconsistent contracts, capability mismatches, instruction conflicts, ambiguous results, stale documentation, unsafe lifecycle behavior, redundant guidance, missing enforcement, test blind spots, and architectural drift. Use when reviewing the overall health and coherence of an agent/tool system, especially systems with prompts, tools, subagents, lifecycle state, dynamic results, configuration, documentation, and multiple instruction layers.

Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Subagents. The repository describes itself as: Asynchronous Pi subagents and agent fleet orchestration for parallel coding agents with nested delegation, background work, and supervision in herdr. The licence is Apache-2.0.

When your agent uses it

  • Reviewing the overall health and coherence of an agent/tool system
  • Especially systems with prompts
  • Lifecycle state
  • Dynamic results

Example prompts

  • “/agentic-system-audit”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Verdict
  2. Executive summary
  3. Findings
  4. Instruction ownership map
  5. Contract coverage
  6. Remediation order
  7. Unverified risks
  8. Non-blocking follow-up

What it can do on your machine

Read from SKILL.md and the folder at commit 3603d9f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentic System Audit loads about 5.8k tokens when it runs. Until then it costs about 134 tokens; SKILL.md has 2,274 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~5.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from boadij/pi-herdsman at commit 3603d9f, republished under its Apache-2.0 licence (© boadij). 2,274 words, ~5,804 tokens.

Download SKILL.mdSave it as .claude/skills/agentic-system-audit/SKILL.md (or your agent's skills folder).
name
agentic-system-audit
description
Audit an agentic software system end to end for contradictions, gaps, inconsistent contracts, capability mismatches, instruction conflicts, ambiguous results, stale documentation, unsafe lifecycle behavior, redundant guidance, missing enforcement, test blind spots, and architectural drift. Use when reviewing the overall health and coherence of an agent/tool system, especially systems with prompts, tools, subagents, lifecycle state, dynamic results, configuration, documentation, and multiple instruction layers.

Agentic System Audit

Audit the whole agentic contract, not isolated files.

Find where code, prompts, tools, state, results, docs, tests, configuration, packaging, or runtime capabilities disagree or leave the next safe action ambiguous.

Prefer deletion, consolidation, one canonical owner, and the smallest durable correction.

This is a read-only audit. Do not modify the repository.

Operating model

Delegate the audit. The orchestrator coordinates and returns the result; it should not perform a competing repository review.

Use:

  • reviewer: lead auditor, evidence reconciliation, synthesis;
  • scout: focused local evidence;
  • researcher: external/API verification when material conclusions depend on it;
  • implementer: never during the audit;
  • generalist: fallback only when the required role is unavailable and its effective policy is read-only.

Start with:

json
{ "action": "list" }

Use effective capabilities from the current context. Do not assume definition-level delegation, tools, skills, or extensions survive overrides or depth.

Preferred topology

text
orchestrator
└── reviewer
    ├── scout A: runtime/contracts/reachability
    ├── scout B: instructions/capabilities
    ├── scout C: presentation/docs/tests/package glue
    └── researcher: only when the research trigger is met

Use one lead reviewer. Do not spawn a second generic reviewer by default.

Leaf fallback

If the lead reviewer cannot delegate:

  1. run the three scout lenses directly in parallel;
  2. run researcher only when triggered;
  3. pass exact scout/researcher resultPath values to one reviewer through files;
  4. have that reviewer verify, reconcile, prioritize, and synthesize.

The orchestrator must not redo the scouts' analysis.

Lead reviewer contract

The reviewer may not receive this skill directly, so its assignment must carry the rules required for a valid audit.

Give the reviewer this objective:

Audit the repository as one agentic system. Remain strictly read-only.

Delegate three focused scouts:

  • A: runtime producers, validation, state, identity, lifecycle, and production reachability;
  • B: instruction ownership and effective capabilities;
  • C: model/UI projections, docs, tests, and repository/package glue.

Use researcher when a candidate material finding or its impact depends on external API/runtime truth.

Before reporting any P0-P2 finding:

  • identify the semantic source and relevant producers;
  • prove the disputed state/value is reachable on a production path when behavior is at issue;
  • trace it through every consumer boundary relevant to the claim;
  • resolve material upstream/runtime assumptions for the repository's supported versions using primary evidence when available;
  • verify that a claimed regression check would fail for the old behavior.

Do not report representable-but-unreachable states as defects. Do not treat differently named fields as equivalent without tracing semantics. Do not leave an answerable material uncertainty as follow-up.

If required evidence genuinely cannot be established, report a bounded Unverified risk, not a definitive finding.

Compare root/parent/child, definition/session/fork, model/UI, success/failure/recovery, configured/effective capability, and static/dynamic guidance.

Independently verify material conclusions and return one de-duplicated report using the required output contract with the smallest durable fix for each finding.

Pass supplied canonical diffs, snapshots, specs, or source artifacts to scouts instead of making them rediscover the same material.

Material finding closure

Scouts return candidate evidence. The reviewer closes candidates before severity is assigned.

Trace every boundary used by the claim. For result/state findings, prefer:

text
semantic source
→ producer
→ validation/normalization
→ runtime/durable state
→ structured/public projection
→ model consumer
→ human/API consumer
→ docs/tests

Not every step applies to every finding.

Provenance

Find where the disputed fact originates and how it changes.

Do not infer semantic equivalence from similar field names. Trace producers and scope first.

Reachability

Do not confuse representability with reachability.

A type, schema, fixture, or doc proves a value can be described. A behavioral finding requires a production path that can actually produce the material state/value.

If the harmful state cannot occur on the claimed path, dismiss or narrow the candidate.

Consumer contract

Establish what the actual consumer receives.

For model-visible claims, determine whether the model gets content, details, both, or another transformation.

For human claims, distinguish compact TUI, expanded TUI, plain/RPC/JSON, and logs/files.

Never assume renderer-visible metadata is model-visible.

External contract

If a material conclusion depends on upstream behavior, resolve it for the repository's supported version range.

Prefer:

text
supported-version primary source/API/type contract
→ supported-version official docs
→ installed dependency source
→ targeted runtime smoke only if still unresolved

Do not substitute latest behavior for the supported range. Do not require a smoke test when source plus local integration already establishes the behavior.

Test discrimination

When claiming a test gap or proposing a regression:

Would this test fail for the old bug and pass for the corrected behavior?

For lifecycle/recovery bugs, pair the harmful case with its nearest legitimate opposite when that distinction matters, such as a stale working parent versus a parent correctly waiting on unresolved child work.

Do not count assertion volume as coverage.

Closure outcomes

Every material candidate becomes exactly one of:

  • Verified finding: evidence establishes defect and consequence.
  • Dismissed/narrowed: intentional, unreachable, semantically different, or immaterial.
  • Unverified risk: evidence genuinely cannot be established with available read-only capabilities.

An Unverified risk must state known evidence, missing evidence, version/runtime scope, why it matters, and one exact resolver.

Do not use it for an answerable question that was not investigated.

Scout A: runtime contracts and reachability

Own:

Can this state/value actually happen, where is it produced, and what runtime contract governs it?

Prove this liveness invariant for unresolved directly-owned assignments:

Every unresolved directly-owned assignment must either remain capable of making progress without its owner, or have a reliable reconciliation path that brings the exact owner back when action is required.

Trace:

text
schema/input
→ validation
→ effective configuration
→ execution
→ runtime/durable state
→ public result
→ next legal action

Check:

  • schema vs runtime acceptance/rejection;
  • defaults, normalization, mutually exclusive fields;
  • validation before mutation;
  • sibling-path validation consistency;
  • logical label, definition, Pi session, owner, request, ask, run, workspace, tab, pane, Herdr identity;
  • stale/replacement identity binding;
  • public state vs actual action eligibility;
  • working, blocked, settling, unknown, steerability, staleness, pending ask/result, child gating, cleanup convergence;
  • unresolved-work liveness across owner idle/busy state, failed attention delivery, reminder recurrence, episode termination, and restart;
  • parent/child distinctions where a working parent may itself be stale while a legitimately waiting parent projects blocked;
  • fresh delegation, continuation, fork, replacement, completion, parent-child completion, ask/reply, close, rollback, restart/recovery;
  • source identity, request correlation, result paths, truncation, cleanup warnings, primary/cleanup errors, rollback/retry state.

Return exact production paths, reachability evidence, and candidate mismatches. Do not decide architecture from local evidence alone.

Scout B: instruction and capability stack

Own:

Who owns this rule, and is it truthful for the effective execution context?

Map:

text
system/base prompt
tool description
controller scope
shared agent prompt
role body
body @file
skills
dynamic success/error guidance
human docs

For each normative rule, find its intended single owner.

Flag:

  • zero owner: instruction gap;
  • multiple authoritative owners: DUPLICATE_AUTHORITY;
  • conflicting owners: CONTRADICTION;
  • wrong owner: AUTHORITY_LEAK.

Compare:

text
definition
→ override composition
→ effective tools/skills/extensions/context
→ launch args
→ actual role/depth
→ model instructions
→ presentation

Find:

  • GHOST_CAPABILITY;
  • HIDDEN_CAPABILITY;
  • root/parent/leaf capability mismatches;
  • metadata advertising removed capabilities;
  • infrastructure accidentally removable by normal policy;
  • user-configurable policy treated as mandatory infrastructure;
  • impossible or underspecified instructions;
  • global policy embedded in role prompts;
  • API mechanics redundantly taught by prompts/skills.

Return exact instruction text, owner, effective runtime evidence, and candidate conflicts.

Scout C: presentation, docs, tests, and repository glue

Own:

Where does runtime truth go, what does each consumer actually see, and do docs/tests describe that contract?

Model-visible presentation

Inspect list, assignment/steer/reply/close acknowledgements, completion delivery, errors, and truncation.

Ask:

Using only what the model actually receives, can it identify the source, understand the state, and choose the next legal action?

Flag missing correlation, hidden material details, ambiguous terminology, duplicated IDs, static policy repeated on every result, or missing dynamic next-action evidence.

For opacity candidates, trace projection semantics, not just field names.

Human presentation

Compare status widget, definitions view, completion renderer, warnings, compact/expanded views, and any plain/API surface.

Find hidden failures, duplicated data, display identity reused as machine identity, or UI claims stronger than runtime evidence.

Documentation

Find shipped-but-undocumented behavior, documented-but-unshipped behavior, duplicate canonical owners, stale terminology, rejected examples, and implementation details presented as public contracts.

Tests

Find public contracts without discriminating regressions, tests that pass for correct and broken behavior, mocks that cannot detect claimed integration failures, missing sibling invariants, and stale fixtures.

Repository/package glue

Check package files, discovery, README links, AGENTS/maintainer guidance, validation scripts, dead compatibility/configuration, and removed concepts that still ship or remain referenced.

Return projection chains, consumer visibility, docs/tests evidence, and candidate mismatches.

Research trigger

Research is conditional. Once triggered, verification is required when capability exists.

Trigger researcher when:

text
a candidate likely to affect a P0-P2 finding, verdict, or remediation
depends on external API/runtime behavior

Research input should contain only:

text
exact disputed assumption
supported version range
preferred primary source/repository
required direct answer

Require primary/official evidence, exact version relevance, and only material findings.

If web capability is unavailable:

  1. inspect installed/local upstream source;
  2. inspect pinned/supported dependency source where reachable;
  3. use other authoritative read-only evidence;
  4. otherwise return Unverified risk.

Do not add dependencies or require optional web tooling to run the audit.

Cross-surface contract matrix

Cover each applicable chain.

Input

text
schema ↔ description ↔ validation ↔ errors ↔ docs ↔ tests

Capability

text
definition ↔ override ↔ effective config ↔ launch ↔ role/depth ↔ tools/prompts ↔ presentation

Instruction

text
controller ↔ shared prompt ↔ role body ↔ skills ↔ dynamic results

Identity

text
label ↔ definition ↔ session ↔ request ↔ owner ↔ Herdr identity ↔ result/error/UI

State

text
runtime evidence ↔ public state ↔ steerability ↔ eligibility ↔ displayed state ↔ next action

Completion

text
completion ↔ durable result ↔ result file ↔ owner message ↔ model content ↔ human rendering ↔ cleanup

Configuration

text
syntax ↔ parsing ↔ merge ↔ validation ↔ effective metadata ↔ runtime ↔ persistence/UI ↔ docs

Recovery

text
failure ↔ retained evidence ↔ autonomous progress/attention owner ↔ wake/delivery ↔ recurrence/termination ↔ structured error ↔ model/human guidance ↔ retry boundary

Distribution

text
repository ↔ package manifest ↔ installed files ↔ discovery ↔ maintainer guidance

Mandatory boundary comparisons

Check where applicable:

Controller depth

text
root | direct agent/parent | nested child/leaf | unmanaged session

Agent generation

text
fresh | exact-session continuation | fork

Visibility

text
model content | structured details | compact TUI | expanded TUI | plain/RPC/JSON | logs/files

Outcome

text
success | blocked | failure | rollback failure | close | cleanup pending | overflow | restart recovery

Configuration

text
bundled | partial override | standalone global | empty fields | false | [] | invalid
Show full SKILL.md (932 more words)Show less

Defect taxonomy

Use consistently:

  • CONTRADICTION: authoritative surfaces prescribe incompatible behavior.
  • GAP: required knowledge/behavior has no appropriate owner or result.
  • GHOST_CAPABILITY: instructions/metadata claim unavailable capability.
  • HIDDEN_CAPABILITY: required capability exists without enough safe guidance.
  • DUPLICATE_AUTHORITY: one normative rule has multiple authoritative owners.
  • AUTHORITY_LEAK: policy is owned by a layer that should not control it.
  • CONTEXT_MISMATCH: rule is correct in one context and wrong in another.
  • IDENTITY_AMBIGUITY: identity loses exact or single semantic meaning.
  • STATE_AMBIGUITY: displayed/instructed state does not map cleanly to legal actions.
  • RESULT_OPACITY: receiver lacks material source, state, correlation, outcome, or next-action evidence.
  • FAIL_OPEN: missing/ambiguous evidence causes unsafe continuation.
  • DRIFT: code, tests, docs, metadata, examples, or prompts describe different generations.
  • TEST_BLIND_SPOT: material public contract lacks a discriminating regression.
  • TOKEN_WASTE: repeated model context adds no decision value.
  • DEAD_COMPLEXITY: compatibility/abstraction/state/configuration has no justified current use.
  • UX_AMBIGUITY: human presentation obscures correct system behavior.

Create another label only when none fits.

Severity

  • P0: unsafe/destructive behavior, wrong-target control, trust-boundary violation, data loss, or durable corruption.
  • P1: materially wrong action, stuck lifecycle, broken delegation/control, or misunderstanding of an authoritative result.
  • P2: meaningful confusion, drift risk, redundant authority, weak diagnostics, avoidable context cost, or recurring maintainer mistakes.
  • P3: small cleanup, naming, presentation, or maintainability issue with little behavioral risk.

Do not inflate severity because a finding is interesting.

Evidence standard

A material cross-surface finding needs:

text
A: semantic behavior/claim + exact source
B: conflicting/missing behavior/claim + exact source
Connection: why A and B are the same contract crossing a boundary
Reachability: production path when behavior is at issue
Consequence: what can actually go wrong
Minimal correction: narrowest root owner

A scout's conclusion is evidence, not authority.

The lead reviewer independently verifies every P0/P1, every P2 affecting architecture/verdict/external assumptions, and any disputed conclusion.

Do not report architectural preference as a defect.

Audit heuristics

Use these to find candidates. They do not replace closure.

  • Who owns this sentence? One normative rule should normally have one authoritative owner.
  • Can the recipient act on this? Check only what that consumer actually receives.
  • Can the agent do what it is told? Compare prompts to effective capability at every depth.
  • Does runtime enforce the important part? Safety, ownership, identity, validation, and data-loss boundaries should not rely only on prompts.
  • Does code know something the model does not? Inspect details, hidden metadata, cleanup evidence, capability inference, state subconditions, correlation IDs.
  • Does the model know something code does not enforce? Separate strategy from invariants such as exact ownership, read-only, no overlap, or one-shot rules.
  • What changes at continuation? Generation-time facts described as dynamically refreshed are suspect.
  • What changes at root/child depth? Global policy projected differently by depth is suspect.
  • What was deleted but still has a shadow? Search docs, tests, package files, comments, compatibility branches, fields, examples.
  • Would this test fail for the old bug? Prefer one discriminating regression over many weak assertions.
  • Can unresolved work disappear from attention forever? Prove either autonomous progress remains possible or reconciliation eventually wakes the exact responsible owner. Check hierarchy boundaries, owner idle/busy state, failed delivery, recurring versus one-shot attention, restart behavior, and episode termination.
  • Is the same fact sent twice? Check model/human identity, paths, model, state, warnings, next-action text.
  • Is this field mismatch semantic drift? Find all producers, scope, reachability, and consumers first.

Efficiency rules

  1. Start from entry points and contract surfaces.
  2. Search ownership before opening large files.
  3. Trace callers and producers before declaring a bug.
  4. Keep scouts inside their lenses.
  5. Pass canonical artifacts instead of rediscovering them.
  6. Research only external assumptions that can change a material conclusion.
  7. Reviewer verifies candidate findings, not every inspected line.
  8. Do not spawn agents merely to confirm No findings.
  9. Use a second reviewer only for an uncertain P0/P1 or disputed architecture.
  10. Never assign an implementer during diagnosis.
  11. Stop when the contract matrix is covered and new searches produce no materially new issue class.

Optimize for useful evidence per agent, not agent count.

Required output

1. Verdict

Exactly one:

text
PASS
PASS WITH FINDINGS
FAIL

PASS means no material P0-P2 findings after required surfaces were covered and material candidates were closed.

2. Executive summary

At most five bullets.

3. Findings

Order by severity, then confidence.

text
[P1][CONTEXT_MISMATCH] Short title

Evidence:
- path:line ...
- path:line ...

Why it matters:
...

Minimal durable fix:
...

Confidence: high

Do not bury findings in prose.

4. Instruction ownership map

Summarize meaningful ownership gaps/collisions across:

text
tool schema
tool description
controller scope
shared agent prompt
role body
skill
dynamic success guidance
dynamic error guidance
human docs

5. Contract coverage

For:

text
input
capability
instruction
identity
state
completion
configuration
recovery
distribution

report:

text
status: verified | finding | not applicable | not verified
scope: what was actually traced

Example:

text
recovery | verified | startup rollback, retained cleanup evidence, retry markers
completion | finding | compact cleanup-warning projection

verified applies only to the stated scope. Do not claim whole-system PASS with unexplained not verified.

6. Remediation order

Give the shortest dependency-aware sequence.

Prefer:

text
fix root contract once
→ delete duplicate authority
→ add/update one discriminating regression
→ update canonical docs

7. Unverified risks

Include only genuinely unresolved material claims after the closure workflow. State known evidence, missing evidence, version/runtime scope, consequence, and exact resolver.

Omit when empty.

8. Non-blocking follow-up

Include only useful work not required to validate a material finding, such as a separate implementation task or non-material integration smoke.

Never put answerable material uncertainty here.

Omit when empty.

Orchestrator finalization

Normally return the lead reviewer's report instead of writing a second review.

Before returning, verify:

  • every required contract chain has scoped coverage;
  • P0/P1 findings have exact evidence;
  • material P2 findings have closed evidence;
  • result/state findings trace producer, reachability, and consumer semantics;
  • material external assumptions were resolved for supported versions when possible;
  • duplicate scout candidates were consolidated;
  • fixes target root owners, not symptoms;
  • only genuinely unresolved claims are Unverified risk;
  • no answerable material question was deferred to follow-up;
  • implementation remains separate from diagnosis.

If a check fails, return the report to the reviewer for reconciliation.

Otherwise stop.

After the audit

Do not automatically fix findings.

If implementation is requested later:

  1. use the audit report as canonical scope;
  2. assign the smallest capable implementer;
  3. keep unrelated cleanup out unless it simplifies the same proven root cause;
  4. independently review the final diff against the audit findings;
  5. run focused and repository validation appropriate to the change.

Success criteria

A successful audit can answer without material guessing:

text
What can each agent actually do?
What is each agent told to do?
Who owns each normative rule?
Which identity is authoritative for each operation?
What does each public state permit next?
What changes by depth and generation?
What does the model see versus human/API surfaces?
Can each result be attributed to source/request?
Can each failure be acted on safely?
Can unresolved directly-owned work remain unnoticed forever?
Does effective configuration match prompts/presentation?
Are reported states and error fields reachable?
Do tests discriminate the important contracts?
Do docs/package contents describe what ships?
Is important authority duplicated or missing?
Is unjustified complexity present?

If a material answer still requires guessing and the evidence was available, the audit is incomplete.

© boadij, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/agentic-system-audit of boadij/pi-herdsman.

Open the folder on GitHubat commit 3603d9f

Compare with similar skills

Agentic System Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentic System Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentic System Audit this skillboadij/pi-herdsman131—~5.8kAutomated safety check: PassApache-2.0
Claude Code Agent Developmentanthropics/claude-plugins-official38k8 repos~2.8kAutomated safety check: PassApache-2.0
Subagent Driven DevelopmentAsvarox/allkaraoke26138 repos~1.2kAutomated safety check: PassNone
Dispatching Parallel Agentsultralisp/ultralisp25841 repos~1.5kAutomated safety check: PassNone
Paseo Advisor Second Opiniongetpaseo/paseo20k1 repos~756Automated safety check: PassCustom licence
Task Observerrebelytics/one-skill-to-rule-them-all3.2k1 repos~12kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 8 repos~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Subagent Driven Development

    Asvarox/allkaraoke

    A skill your agent uses when executing implementation plans with independent tasks in the current session

    261 GitHub starsUsed in 38 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Dispatching Parallel Agents

    ultralisp/ultralisp

    A skill your agent uses when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies

    258 GitHub starsUsed in 41 repos~1.5k tokens
    Agent WorkflowsAuto-check passed
  • Launches one separate agent through Paseo to give a second opinion on the current task, with a self-contained briefing and no permission to edit files.

    20k GitHub starsUsed in 1 repo~756 tokens
    Agent WorkflowsAuto-check passed
  • Task Observer

    rebelytics/one-skill-to-rule-them-all

    Monitors task execution for skill improvement opportunities.

    3.2k GitHub starsUsed in 1 repo~12k tokens
    Agent WorkflowsAuto-check passed
  • O2 Review Loop

    openobserve/openobserve

    Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.

    22k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from boadij/pi-herdsman

  • Testing

    boadij/pi-herdsman

    Write and review tests that protect meaningful Herdsman behavior without creating redundant, brittle, or low-value coverage.

    131 GitHub stars~863 tokensUpdated today
    Auto-check passed
  • Agent Definitions

    boadij/pi-herdsman

    Manage Pi Herdsman Agent definitions and managed Lead configuration.

    131 GitHub stars~800 tokensUpdated today
    Auto-check passed
  • Agents

    boadij/pi-herdsman

    Optional reinforcement and strategy for orchestrating managed agents.

    131 GitHub stars~6.9k tokensUpdated today
    Auto-check passed

Categories

Questions about Agentic System Audit

What does Agentic System Audit do?

Audit an agentic software system end to end for contradictions, gaps, inconsistent contracts, capability mismatches, instruction conflicts, ambiguous results, stale documentation, unsafe lifecycle…. Agentic System Audit is an agent skill from boadij/pi-herdsman. Audit an agentic software system end to end for contradictions, gaps, inconsistent contracts, capability mismatches, instruction conflicts, ambiguous results, stale documentation, unsafe lifecycle behavior, redundant guidance, missing enforcement, test blind spots, and architectural drift.

When should I use Agentic System Audit?

Agentic System Audit fits situations like: reviewing the overall health and coherence of an agent/tool system; especially systems with prompts; lifecycle state; dynamic results.

How do I install Agentic System Audit in Claude Code?

Run `npx skills add boadij/pi-herdsman --skill agentic-system-audit -a claude-code`. Or copy the skill folder (.agents/skills/agentic-system-audit in boadij/pi-herdsman) into .claude/skills/agentic-system-audit in your project. Claude Code loads it when a task matches its description.

How do I install Agentic System Audit in Codex?

Run `npx skills add boadij/pi-herdsman --skill agentic-system-audit -a codex`. Or copy the skill folder (.agents/skills/agentic-system-audit in boadij/pi-herdsman) into .agents/skills/agentic-system-audit in your project. Codex loads it when a task matches its description.

Can I use Agentic System Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add boadij/pi-herdsman --skill agentic-system-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentic-system-audit, .gemini/skills/agentic-system-audit, .github/skills/agentic-system-audit and .opencode/skills/agentic-system-audit in your project.

What does Agentic System Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentic System Audit is instructions for the agent only.

Does Agentic System Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentic System Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentic System Audit use?

Agentic System Audit is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentic System Audit use?

About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agentic System Audit?

Skills that share tags, products or a category with Agentic System Audit: Claude Code Agent Development (anthropics/claude-plugins-official, 38k stars), Subagent Driven Development (Asvarox/allkaraoke, 261 stars), Dispatching Parallel Agents (ultralisp/ultralisp, 258 stars) and Paseo Advisor Second Opinion (getpaseo/paseo, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentic System Audit?

boadij (a GitHub user) maintains it in boadij/pi-herdsman, which has 131 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.

Source: boadij/pi-herdsman on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.