Agent skill

Agent Architecture Audit

by affaan-m in affaan-m/ECC

Full-stack diagnostic for agent and LLM applications. An agent skill from affaan-m/ECC.

MITAuto-check passedAgent Workflows

Install Agent Architecture Audit

skills CLI
$ npx skills add affaan-m/ECC --skill agent-architecture-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC agent-architecture-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-architecture-audit .claude/skills/agent-architecture-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-architecture-audit
GitHub stars
277k
Used in
1 other repo
Token cost
~2.5k tokens
SKILL.md length
1,071 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

Full-stack diagnostic for agent and LLM applications. An agent skill from affaan-m/ECC.

  • Works in 9 steps: Wrapper Regression → Memory Contamination → Tool Discipline Failure → …
  • LLM feature misbehaves and the failing layer is unknown
  • SKILL.md covers When to Activate, The 12-Layer Stack, Common Failure Patterns and Audit Workflow, plus 6 more sections
  • Calls rg

What it does

Agent Architecture Audit is an agent skill from affaan-m/ECC. Full-stack diagnostic for agent and LLM applications. Audits the 12-layer agent stack for wrapper regression, memory pollution, tool discipline failures, hidden repair loops, and rendering corruption. Produces severity-ranked findings with code-first fixes. Essential for developers building agent applications, autonomous loops, or any LLM-powered feature. Use when an agent or LLM feature misbehaves and the failing layer is unknown, or before shipping an agent stack.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Building AI agents and Autonomous loops. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • LLM feature misbehaves and the failing layer is unknown
  • Before shipping an agent stack

Example prompts

  • “/agent-architecture-audit”

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Wrapper Regression
  2. Memory Contamination
  3. Tool Discipline Failure
  4. Rendering/Transport Corruption
  5. Hidden Agent Layers
  6. Scope
  7. Evidence Collection
  8. Failure Mapping
  9. Fix Strategy

What it can do on your machine

Read from SKILL.md and the folder at commit 2d515e4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Architecture Audit loads about 2.5k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 1,071 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 2d515e4, republished under its MIT licence (© affaan-m). 1,071 words, ~2,550 tokens.

Download SKILL.mdSave it as .claude/skills/agent-architecture-audit/SKILL.md (or your agent's skills folder).
name
agent-architecture-audit
description
Full-stack diagnostic for agent and LLM applications. Audits the 12-layer agent stack for wrapper regression, memory pollution, tool discipline failures, hidden repair loops, and rendering corruption. Produces severity-ranked findings with code-first fixes. Essential for developers building agent applications, autonomous loops, or any LLM-powered feature. Use when an agent or LLM feature misbehaves and the failing layer is unknown, or before shipping an agent stack.
metadata.origin
oh-my-agent-check
tools
Read, Write, Edit, Bash, Grep, Glob

Agent Architecture Audit

A diagnostic workflow for agent systems that hide failures behind wrapper layers, stale memory, retry loops, or transport/rendering mutations.

When to Activate

MANDATORY for:

  • Releasing any agent or LLM-powered application to production
  • Shipping features with tool calling, memory, or multi-step workflows
  • Agent behavior degrades after adding wrapper layers
  • User reports "the agent is getting worse" or "tools are flaky"
  • Same model works in playground but breaks inside your wrapper
  • Debugging agent behavior for more than 15 minutes without finding root cause

Especially critical when:

  • You've added new prompt layers, tool definitions, or memory systems
  • Different agents in your system behave inconsistently
  • The model was fine yesterday but is hallucinating today
  • You suspect hidden repair/retry loops silently mutating responses

Do not use for:

  • General code debugging — use agent-introspection-debugging
  • Code review — use language-specific reviewer agents
  • Security scanning — use security-review or security-review/scan
  • Agent performance benchmarking — use agent-eval
  • Writing new features — use the appropriate workflow skill

The 12-Layer Stack

Every agent system has these layers. Any of them can corrupt the answer:

#LayerWhat Goes Wrong
1System promptConflicting instructions, instruction bloat
2Session historyStale context injection from previous turns
3Long-term memoryPollution across sessions, old topics in new conversations
4DistillationCompressed artifacts re-entering as pseudo-facts
5Active recallRedundant re-summary layers wasting context
6Tool selectionWrong tool routing, model skips required tools
7Tool executionHallucinated execution — claims to call but doesn't
8Tool interpretationMisread or ignored tool output
9Answer shapingFormat corruption in final response
10Platform renderingTransport-layer mutation (UI, API, CLI mutates valid answers)
11Hidden repair loopsSilent fallback/retry agents running second LLM pass
12PersistenceExpired state or cached artifacts reused as live evidence

Common Failure Patterns

1. Wrapper Regression

The base model produces correct answers, but the wrapper layers make it worse.

Symptoms:

  • Model works fine in playground or direct API call, breaks in your agent
  • Added a new prompt layer, existing behavior degraded
  • Agent sounds confident but is confidently wrong
  • "It was working before the last update"
2. Memory Contamination

Old topics leak into new conversations through history, memory retrieval, or distillation.

Symptoms:

  • Agent brings up unrelated past topics
  • User corrections don't stick (old memory overwrites new)
  • Same-session artifacts re-enter as pseudo-facts
  • Memory grows without bound, degrading response quality over time
3. Tool Discipline Failure

Tools are declared in the prompt but not enforced in code. The model skips them or hallucinates execution.

Symptoms:

  • "Must use tool X" in prompt, but model answers without calling it
  • Tool results look correct but were never actually executed
  • Different tools fight over the same responsibility
  • Model uses tool when it shouldn't, or skips it when it must
4. Rendering/Transport Corruption

The agent's internal answer is correct, but the platform layer mutates it during delivery.

Symptoms:

  • Logs show correct answer, user sees broken output
  • Markdown rendering, JSON parsing, or streaming fragments corrupt valid responses
  • Hidden fallback agent quietly replaces the answer before delivery
  • Output differs between terminal and UI
5. Hidden Agent Layers

Silent repair, retry, summarization, or recall agents run without explicit contracts.

Symptoms:

  • Output changes between internal generation and user delivery
  • "Auto-fix" loops run a second LLM pass the user doesn't know about
  • Multiple agents modify the same output without coordination
  • Answers get "smoothed" or "corrected" by invisible layers

Audit Workflow

Phase 1: Scope

Define what you're auditing:

  • Target system — what agent application?
  • Entrypoints — how do users interact with it?
  • Model stack — which LLM(s) and providers?
  • Symptoms — what does the user report?
  • Time window — when did it start?
  • Layers to audit — which of the 12 layers apply?
Phase 2: Evidence Collection

Gather evidence from the codebase:

  • Source code — agent loop, tool router, memory admission, prompt assembly
  • Logs — historical session traces, tool call records
  • Config — prompt templates, tool schemas, provider settings
  • Memory files — SOPs, knowledge bases, session archives

Use rg to search for anti-patterns:

bash
# Tool requirements expressed only in prompt text (not code)
rg "must.*tool|必须.*工具|required.*call" --type md

# Tool execution without validation
rg "tool_call|toolCall|tool_use" --type py --type ts

# Hidden LLM calls outside main agent loop
rg "completion|chat\.create|messages\.create|llm\.invoke"

# Memory admission without user-correction priority
rg "memory.*admit|long.*term.*update|persist.*memory" --type py --type ts

# Fallback loops that run additional LLM calls
rg "fallback|retry.*llm|repair.*prompt|re-?prompt" --type py --type ts

# Silent output mutation
rg "mutate|rewrite.*response|transform.*output|shap" --type py --type ts
Show full SKILL.md (426 more words)Show less
Phase 3: Failure Mapping

For each finding, document:

  • Symptom — what the user sees
  • Mechanism — how the wrapper causes it
  • Source layer — which of the 12 layers
  • Root cause — the deepest cause
  • Evidence — file:line or log:row reference
  • Confidence — 0.0 to 1.0
Phase 4: Fix Strategy

Default fix order (code-first, not prompt-first):

  1. Code-gate tool requirements — enforce in code, not just prompt text
  2. Remove or narrow hidden repair agents — make fallback explicit with contracts
  3. Reduce context duplication — same info through prompt + history + memory + distillation
  4. Tighten memory admission — user corrections > agent assertions
  5. Tighten distillation triggers — don't compress what shouldn't be compressed
  6. Reduce rendering mutation — pass-through, don't transform
  7. Convert to typed JSON envelopes — structured internal flow, not freeform prose

Severity Model

LevelMeaningAction
criticalAgent can confidently produce wrong operational behaviorFix before next release
highAgent frequently degrades correctness or stabilityFix this sprint
mediumCorrectness usually survives but output is fragile or wastefulPlan for next cycle
lowMostly cosmetic or maintainability issuesBacklog

Output Format

Present findings to the user in this order:

  1. Severity-ranked findings (most critical first)
  2. Architecture diagnosis (which layer corrupted what, and why)
  3. Ordered fix plan (code-first, not prompt-first)

Do not lead with compliments or summaries. If the system is broken, say so directly.

Quick Diagnostic Questions

When auditing an agent system, answer these:

#QuestionIf Yes →
1Can the model skip a required tool and still answer?Tool not code-gated
2Does old conversation content appear in new turns?Memory contamination
3Is the same info in system prompt AND memory AND history?Context duplication
4Does the platform run a second LLM pass before delivery?Hidden repair loop
5Does the output differ between internal generation and user delivery?Rendering corruption
6Are "must use tool X" rules only in prompt text?Tool discipline failure
7Can the agent's own monologue become persistent memory?Memory poisoning

Anti-Patterns to Avoid

  • Avoid blaming the model before falsifying wrapper-layer regressions.
  • Avoid blaming memory without showing the contamination path.
  • Do not let a clean current state erase a dirty historical incident.
  • Do not treat markdown prose as a trustworthy internal protocol.
  • Do not accept "must use tool" in prompt text when code never enforces it.
  • Keep findings direct, evidence-backed, and severity-ranked.

Report Schema

Audits should produce structured reports following this shape:

json
{
  "schema_version": "ecc.agent-architecture-audit.report.v1",
  "executive_verdict": {
    "overall_health": "high_risk",
    "primary_failure_mode": "string",
    "most_urgent_fix": "string"
  },
  "scope": {
    "target_name": "string",
    "model_stack": ["string"],
    "layers_to_audit": ["string"]
  },
  "findings": [
    {
      "severity": "critical|high|medium|low",
      "title": "string",
      "mechanism": "string",
      "source_layer": "string",
      "root_cause": "string",
      "evidence_refs": ["file:line"],
      "confidence": 0.0,
      "recommended_fix": "string"
    }
  ],
  "ordered_fix_plan": [
    { "order": 1, "goal": "string", "why_now": "string", "expected_effect": "string" }
  ]
}
  • agent-introspection-debugging — Debug agent runtime failures (loops, timeouts, state errors)
  • agent-eval — Benchmark agent performance head-to-head
  • security-review — Security audit for code and configuration
  • autonomous-agent-harness — Set up autonomous agent operations
  • agent-harness-construction — Build agent harnesses from scratch

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agent-architecture-audit of affaan-m/ECC.

Open the folder on GitHubat commit 2d515e4

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Agent Architecture Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Architecture Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Architecture Audit this skillaffaan-m/ECC277k1 repos~2.5kAutomated safety check: PassMIT
Topologychmod777john/swarm-ide1.5k—~447Automated safety check: PassNone
AI Agent Developmentaiskillstore/marketplace4333 repos~1kAutomated safety check: PassNone
Author Skillericrisco/rsc-harness180—~4.3kAutomated safety check: PassMIT
Swarms Multi-Agent Frameworkkyegomez/swarms7.2k—~5.5kAutomated safety check: PassApache-2.0
Uipath FunctionsUiPath/skills167—~3.6kAutomated safety check: NotesMIT

Similar skills

  • Topology

    chmod777john/swarm-ide

    Explain the IM+Agent framework: create+send as minimal primitives, IM system vs agent loop separation, message vs llmHistory, and the recursive property.

    1.5k GitHub stars~447 tokensUpdated 7 mo ago
    Agent WorkflowsAuto-check passed
  • AI Agent Development

    aiskillstore/marketplace

    AI agent development workflow for building autonomous agents, multi-agent systems, and agent orchestration with CrewAI, LangGraph, and custom agents.

    433 GitHub starsUsed in 3 repos~1k tokens
    Agent WorkflowsAuto-check passed
  • Author Skill

    ericrisco/rsc-harness

    A skill your agent uses when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into…

    180 GitHub stars~4.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Teaches the Swarms Python framework: the Agent class, tools, loops, memory and multi-agent structures such as sequential, concurrent and graph workflows.

    7.2k GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Uipath Functions

    UiPath/skills

    UiPath Coded Functions — deterministic Python or TypeScript/JavaScript units built with the uip function CLI (new -l py|ts|js, init, serve, run, pack, publish); the functions map in uipath.json…

    167 GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Strands

    strands-agents/harness-sdk

    Build, extend, evaluate, or migrate applications with Strands Agents in Python or TypeScript.

    8.8k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Agent Architecture Audit

What does Agent Architecture Audit do?

Full-stack diagnostic for agent and LLM applications. An agent skill from affaan-m/ECC. Agent Architecture Audit is an agent skill from affaan-m/ECC. Full-stack diagnostic for agent and LLM applications.

When should I use Agent Architecture Audit?

Agent Architecture Audit fits situations like: LLM feature misbehaves and the failing layer is unknown; before shipping an agent stack.

How do I install Agent Architecture Audit in Claude Code?

Run `npx skills add affaan-m/ECC --skill agent-architecture-audit -a claude-code`. Or copy the skill folder (skills/agent-architecture-audit in affaan-m/ECC) into .claude/skills/agent-architecture-audit in your project. Claude Code loads it when a task matches its description.

How do I install Agent Architecture Audit in Codex?

Run `npx skills add affaan-m/ECC --skill agent-architecture-audit -a codex`. Or copy the skill folder (skills/agent-architecture-audit in affaan-m/ECC) into .agents/skills/agent-architecture-audit in your project. Codex loads it when a task matches its description.

Can I use Agent Architecture Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill agent-architecture-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-architecture-audit, .gemini/skills/agent-architecture-audit, .github/skills/agent-architecture-audit and .opencode/skills/agent-architecture-audit in your project.

What does Agent Architecture Audit need to run?

Going by SKILL.md and its folder, Agent Architecture Audit needs the command-line tools its instructions call (rg).

Does Agent Architecture Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Architecture Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Architecture Audit use?

Agent Architecture Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Architecture Audit use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Architecture Audit?

Skills that share tags, products or a category with Agent Architecture Audit: Topology (chmod777john/swarm-ide, 1.5k stars), AI Agent Development (aiskillstore/marketplace, 433 stars), Author Skill (ericrisco/rsc-harness, 180 stars) and Swarms Multi-Agent Framework (kyegomez/swarms, 7.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Architecture Audit?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,673 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 11, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.