Agent skill

Agent Prompt Audit

by avibe-bot in avibe-bot/avibe

Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch…

MITAuto-check passedAI & LLM Engineering

Install Agent Prompt Audit

skills CLI
$ npx skills add avibe-bot/avibe --skill agent-prompt-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install avibe-bot/avibe agent-prompt-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/avibe-bot/avibe.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-prompt-audit .claude/skills/agent-prompt-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-prompt-audit
GitHub stars
622
Token cost
~2.9k tokens
SKILL.md length
1,558 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch…

  • An Agent misbehaves (stalls
  • SKILL.md covers Core ideas, Where the surface lives, Finding evidence and From symptom to likely cause, plus 1 more section
  • Calls git
  • Over-applies a rule)

What it does

Agent Prompt Audit is an agent skill from avibe-bot/avibe. Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch messages — using real run evidence. Use when an Agent misbehaves (stalls, over-asks, over-reaches, ignores or over-applies a rule), after a model or backend change, or when the user asks to review, clean up, or tighten prompts.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Prompt engineering. The repository describes itself as: The local-first Agent OS — your AI partner lives on your own machine. Drive the official Claude Code, Codex & OpenCode from your browser or any chat app. The licence is MIT.

When your agent uses it

  • An Agent misbehaves (stalls
  • Over-applies a rule)
  • The user asks to review
  • Tighten prompts

Example prompts

  • “/agent-prompt-audit”

What it can do on your machine

Read from SKILL.md and the folder at commit b3fc733. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Prompt Audit loads about 2.9k tokens when it runs. Until then it costs about 110 tokens; SKILL.md has 1,558 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from avibe-bot/avibe at commit b3fc733, republished under its MIT licence (© avibe-bot). 1,558 words, ~2,875 tokens.

Download SKILL.mdSave it as .claude/skills/agent-prompt-audit/SKILL.md (or your agent's skills folder).
name
agent-prompt-audit
description
Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch messages — using real run evidence. Use when an Agent misbehaves (stalls, over-asks, over-reaches, ignores or over-applies a rule), after a model or backend change, or when the user asks to review, clean up, or tighten prompts.
slug
agent-prompt-audit
version
0.1.0

Agent Prompt Audit

An Agent's behavior comes from every piece of text that reaches it over its lifecycle, not just its system prompt. The audit's job is to find which text causes the behavior the user sees, on the backend and model that actually ran it, and to propose the smallest change that fixes it. Judge each instruction by what it does to behavior, not by its length: sometimes the fix is adding a missing reason or exit, and a clean surface is a valid result.

Deliver a report of findings — each with its evidence, confidence, and a concrete proposed change — and apply changes only when asked.

Core ideas

Evidence over reading. What Agents actually did beats what the text seems to say. Start from real runs and user corrections ("you stopped", "why didn't you report", "don't ask me that"), trace each symptom to the line that caused it or the missing line that would have prevented it, and use git blame to learn what incident a rule was written for and whether it still happens. A finding without evidence or documented model behavior is a flag, not a fix.

Context is kept; constraints must earn their place. Facts only the author knows — environment, contracts, ownership, quality bar, the reason behind a rule — are what prompts are for. Behavioral constraints are what go stale. Keep exact scripts where one sequence is safe (destructive commands, auth, merge gates, key custody), prohibitions against failures that still reproduce, and the scope bounds that make autonomy safe.

Say intent and reason, not pressure or method. Caps, MUST/NEVER, and emphasis without a reason make current models rigid; step scripts for judgment work and strategy coaching usually do worse than the model's own plan; fixed formats, word caps, and "don't narrate" rules produce silence or starved answers. Fossils — named-model workarounds, incident numbers as authority, "now/no longer" phrasing, one session's stumble made permanent — should become the current rule they stand for.

Every stop needs an exit. Agents run across turns, wake on callbacks, and hand work to each other, so the costliest defects are lifecycle gaps: a "stop/wait" with no statement of what the turn produces instead, asking permission for reversible in-scope steps, continuing without bounds after repeated failure, waiting with no durable waiter or expiry meaning, briefs missing the goal or report target, callbacks that say "done" without the result, and Task/Watch messages that restate rules on every fire.

One home per rule, at the layer whose timing fits. Always-loaded and recurring text has the most leverage and deserves the most scrutiny. Duplicates that disagree force the Agent to guess; keep the mechanism in one place and a principle or pointer elsewhere. Agreeing fallbacks are fine. Long procedures belong in on-demand Skills, not always-loaded rules.

Shared text runs on every backend. GPT/Codex tend to follow a bare prohibition or stop literally, so they need scope and exit conditions; strong Claude models tend to over-reach, so they need scope bounds and a definition of done; tool names and native mechanics dangle on other backends. Take model-specific behavior from the vendor's current docs, and lower confidence when you cannot reach them.

A removal is a hypothesis. For contested changes, compare behavior before and after with a scratch run on the target that produced the failure, and read the transcript rather than asking the model whether it needs the rule.

Where the surface lives

Verify against the current machine; these are starting points. Each backend also reads its own native configuration — config directories moved by environment variables (CLAUDE_CONFIG_DIR, CODEX_HOME, OpenCode's config path), and native subagent definitions such as .claude/agents/, .codex/agents/, or OpenCode agents — so resolve what the target backend actually loads rather than assuming default paths.

LayerWhereHow it changes
Avibe runtime promptvibe debug prompt export --format json lists every source; vibe debug prompt export --format json --context-file <file> renders a composition from the inputs you supply (backend, Agent instructions, Skill directory, context), so it approximates the target only as well as those inputs match (history in the Avibe repo core/prompts/, if checked out)Proposal to the Avibe repository
Global rules and native backend config~/.claude/CLAUDE.md, ~/.codex/AGENTS.md, Codex developer_instructions in $CODEX_HOME/config.toml (default ~/.codex), OpenCode instructions in global or project opencode.json[c], …Edit the source if the file is generated or imports others
Project rulesnearest AGENTS.md / CLAUDE.md chainThe repository's own delivery process
Agent system prompt, model, effortvibe agent show <name> --jsonvibe agent update <name> --system-prompt-file <file>
Skillsuser skill dirs (follow symlinks), Avibe skills/, project .agents/skills/The directory's owner
Task and Watch messages (re-sent every fire)vibe task list / vibe watch list for ids, then vibe task show <id> / vibe watch show <id> for the full textvibe task update, vibe watch update
User preferences (read on demand)~/.avibe/state/user_preferences.md; inspect only the reported user's part, and only when the transcript shows it was readThe user
Delegation briefs and callbacksagent_runs.message / result_textThe prompt or Skill that writes them

Only Skill descriptions on the first catalog page are loaded every turn; later pages, disable-model-invocation Skills, Skill bodies, and references load on demand. Project Skills in that catalog resolve from the target Session's working directory, not yours, so read them from there.

Show full SKILL.md (692 more words)Show less

Finding evidence

Resolve the actual target (backend, model, effort) from the run record, not the Agent's current definition. A Session's model and effort can change during its life and ordinary IM turns have no run record, so treat the Session row as the current setting and mark the target unconfirmed if it may have changed. Attribute a symptom only to prompt text that existed when it ran. Each run's prompt and message in vibe runs show snapshot what a Task or Watch actually sent; file-owned text needs Git or release history matching the run; Agent system prompts have no history. A long-lived Codex thread also keeps earlier injected prompt snapshots in its native history, so text since removed may still have been in view. Where the text may have changed since the run and no history covers it, say the attribution is unconfirmed.

vibe runs show <id> gives one run's prompt, result, and callback state; vibe data query is read-only SQLite over agent_sessions, agent_runs, and messages. Keep evidence to the Session the user reported, plus any others they point to — a channel's scope_id can hold other people's threads; the user's own corrections in messages are usually the sharpest evidence. For a recurring Task or Watch, vibe runs list --definition-id <id> gathers its fires across per-run Sessions. A stuck delegated turn stays running, so include long-running rows when the complaint is a stall; ordinary IM turns have no agent_runs row, so read that Session's messages instead. A delegated run can succeed while its report never arrives; check callback_status and callback_error when the complaint is a missing result. Two starting points, each run with vibe data query --sql-file <file> (or --sql-file - for stdin):

sql
-- The reported session and its current backend, model, and effort
select id, scope_id, agent_name, agent_backend, model, reasoning_effort, status
from agent_sessions where id = '<session>';

-- Recent failed, cancelled, or silent Agent runs in that session
select id, run_type, status, model, created_at
from agent_runs
where session_id = '<session>'
  and run_type in ('agent_run','scheduled','watch','webhook','task_escalation')
  and exit_code is null  -- command-backed Tasks record an exit code instead
  and (status in ('failed','canceled','cancelled')
       or (status in ('succeeded','completed') and coalesce(trim(result_text),'') = ''))
order by created_at desc;

Quote the minimum excerpt and redact secrets and unrelated private content.

A before/after probe spends the user's account and writes session state, so propose it in the report unless the user asked for verification. A useful probe reproduces the original conditions — same backend, model, and effort, and only the context before the failing turn — rather than forking a session that already holds the failure and its correction.

From symptom to likely cause

User complaints map to recurring prompt defects. Treat these as leads to check against the transcript, not verdicts.

What the user seesWhere to look first
Agent stopped or went quiet mid-taskA "stop / wait / do not proceed" with no stated exit; "don't narrate" or "report only at the end"; a wait with no durable Watch or expiry meaning
Keeps asking for permission"Ask before…" with no threshold separating reversible in-scope steps from irreversible or outward-facing ones
Did far more than askedAutonomy with no scope bound or definition of done, most often on strong Claude models
Followed a rule where it made no senseA bare prohibition with no reason or scope, most often on GPT/Codex; pressure language (caps, MUST/NEVER)
Behaves differently across Agents or backendsThe same rule at different strengths in different layers; backend-specific tool names in shared text
Delegated work came back unusableA brief missing goal, acceptance evidence, or report target; a callback that says "done" without the result
Recurring Task or Watch runs drift or repeat themselvesThe fire message restates loaded rules, names finished work, or asks for output the recipient cannot act on
Stale commands, paths, or answersFacts that no longer match the CLI or code; fossils like named-model workarounds or "now / no longer" phrasing

Report

Open with counts and up to three findings that matter most; zero findings is a valid report. For each finding: location, the evidence excerpt, which idea above it violates and why on which target, confidence (high: reproduced in transcripts or documented; medium: consistent known behavior; low: heuristic, flag only), and the proposed change — a file hunk, or a before/after payload plus the update command for text stored in Avibe state. Rewrite rather than delete when the concern is still live, and complete each removal across duplicates, tests, and mirrors.

State each finding's confidence once and the audit's overall limits once; repeating caveats in every paragraph buries the findings. Read-only checks, such as --help or reading a file, settle a doubt faster than flagging it.

© avibe-bot, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agent-prompt-audit of avibe-bot/avibe.

Open the folder on GitHubat commit b3fc733

Compare with similar skills

Agent Prompt Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Prompt Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Prompt Audit this skillavibe-bot/avibe622—~2.9kAutomated safety check: PassMIT
Prompt Improverseverity1/claude-code-prompt-improver1.9k1 repos~1.7kAutomated safety check: PassMIT
Prompt Engineering Patternsynulihao/AgentSkillOS61814 repos~1.7kAutomated safety check: PassNone
Patch CreationPiebald-AI/tweakcc2.5k—~1.6kAutomated safety check: PassMIT
Senior Prompt Engineermaslennikov-ig/claude-code-orchestrator-kit2603 repos~1.4kAutomated safety check: PassCustom licence
Codex Fable5baskduf/FableCodex437—~1.6kAutomated safety check: PassAGPL-3.0

Similar skills

  • Prompt Improver

    severity1/claude-code-prompt-improver

    This skill enriches vague prompts with targeted research and clarification before execution.

    1.9k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineering Patterns

    ynulihao/AgentSkillOS

    Master advanced prompt engineering techniques to maximize LLM performance, reliability, and controllability in production.

    618 GitHub starsUsed in 14 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Patch Creation

    Piebald-AI/tweakcc

    Create and register new patches for tweakcc. An agent skill from Piebald-AI/tweakcc.

    2.5k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Senior Prompt Engineer

    maslennikov-ig/claude-code-orchestrator-kit

    Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.

    260 GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Codex Fable5

    baskduf/FableCodex

    Apply a Claude Fable 5 inspired operating style inside Codex.

    437 GitHub stars~1.6k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed

More from avibe-bot/avibe

  • Use Avibe

    avibe-bot/avibe

    Safely inspect and modify local Avibe configuration, routing, runtime settings, watches, scheduled tasks, Avibe Cloud remote access, and operational state.

    622 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • PR Delivery Loop

    avibe-bot/avibe

    Deliver implementation PRs across Avibe, avibe-backend, avibe-docs, avault, and vault-sandbox.

    622 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Background Watch Hook

    avibe-bot/avibe

    Use vibe watch to run a managed Harness waiter that returns to the same conversation later.

    622 GitHub stars~7.7k tokensUpdated today
    Auto-check passed
  • Use Show Pages

    avibe-bot/avibe

    Build, inspect, update, restore, or share Avibe Show Pages for visual explanations, diagrams, reports, or interactive prototypes.

    622 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Use Avibe Harness

    avibe-bot/avibe

    Use Avibe Harness for durable Agent delegation, Sessions, scheduled Tasks, Watches, Runs, queues, and work that must continue beyond the current turn.

    622 GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Use Avibe Vault

    avibe-bot/avibe

    Use Avibe Vault for API keys, tokens, passwords, protected credentials, authenticated HTTP requests, or digest signing without exposing secret values to the agent.

    622 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Agent Prompt Audit

What does Agent Prompt Audit do?

Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch…. Agent Prompt Audit is an agent skill from avibe-bot/avibe. Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch messages — using real run evidence.

When should I use Agent Prompt Audit?

Agent Prompt Audit fits situations like: an Agent misbehaves (stalls; over-applies a rule); the user asks to review; tighten prompts.

How do I install Agent Prompt Audit in Claude Code?

Run `npx skills add avibe-bot/avibe --skill agent-prompt-audit -a claude-code`. Or copy the skill folder (skills/agent-prompt-audit in avibe-bot/avibe) into .claude/skills/agent-prompt-audit in your project. Claude Code loads it when a task matches its description.

How do I install Agent Prompt Audit in Codex?

Run `npx skills add avibe-bot/avibe --skill agent-prompt-audit -a codex`. Or copy the skill folder (skills/agent-prompt-audit in avibe-bot/avibe) into .agents/skills/agent-prompt-audit in your project. Codex loads it when a task matches its description.

Can I use Agent Prompt Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add avibe-bot/avibe --skill agent-prompt-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-prompt-audit, .gemini/skills/agent-prompt-audit, .github/skills/agent-prompt-audit and .opencode/skills/agent-prompt-audit in your project.

What does Agent Prompt Audit need to run?

Going by SKILL.md and its folder, Agent Prompt Audit needs the command-line tools its instructions call (git).

Does Agent Prompt Audit access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Agent Prompt Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Prompt Audit use?

Agent Prompt Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Prompt Audit use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Prompt Audit?

Skills that share tags, products or a category with Agent Prompt Audit: Prompt Improver (severity1/claude-code-prompt-improver, 1.9k stars), Prompt Engineering Patterns (ynulihao/AgentSkillOS, 618 stars), Patch Creation (Piebald-AI/tweakcc, 2.5k stars) and Senior Prompt Engineer (maslennikov-ig/claude-code-orchestrator-kit, 260 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Prompt Audit?

avibe-bot (a GitHub organization) maintains it in avibe-bot/avibe, which has 622 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 10, 2026.

Source: avibe-bot/avibe on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.