Agent skill

Audit Agent Run Evidence

by sickn33 in sickn33/agentic-awesome-skills

A skill your agent uses when an agent, harness, gateway, MCP workflow, or multi-step automation claims completion and the available traces, checkpoints, approvals, tool calls, or deployment records…

MITAuto-check passedDevOps & Cloud

Install Audit Agent Run Evidence

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill audit-agent-run-evidence -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills audit-agent-run-evidence --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/audit-agent-run-evidence .claude/skills/audit-agent-run-evidence && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-agent-run-evidence
GitHub stars
47k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
996 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when an agent, harness, gateway, MCP workflow, or multi-step automation claims completion and the available traces, checkpoints, approvals, tool calls, or deployment records…

  • Works in 7 steps: Order events by causal links and… → Build the state-transition path and mark… → Link each retry chain by logical… → …
  • Multi-step automation claims completion and the available traces
  • SKILL.md covers Overview, When to Use, Establish the Contract and Build a Claim Ledger, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Audit Agent Run Evidence is an agent skill from sickn33/agentic-awesome-skills. Use when an agent, harness, gateway, MCP workflow, or multi-step automation claims completion and the available traces, checkpoints, approvals, tool calls, or deployment records must be judged without trusting self-reported success.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Deployment. It works with Model Context Protocol. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Multi-step automation claims completion and the available traces
  • Deployment records must be judged without trusting self-reported success

Example prompts

  • “/audit-agent-run-evidence”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Order events by causal links and per-source sequence; use timestamps only as supporting evidence.
  2. Build the state-transition path and mark every gap or illegal transition.
  3. Link each retry chain by logical operation, request ID, and idempotency key.
  4. Link checkpoints to the state they contain and the resume event that consumes them.
  5. Preserve every parallel branch outcome; apply the declared all_required, quorum, first_success, or other join rule.
  6. Track remaining budgets at each transition. A late success after budget exhaustion is a budget violation.
  7. Bind approvals and deployment records to exact artifact digests and targets.

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Agent Run Evidence loads about 2.1k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 996 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 996 words, ~2,111 tokens.

Download SKILL.mdSave it as .claude/skills/audit-agent-run-evidence/SKILL.md (or your agent's skills folder).
name
audit-agent-run-evidence
description
Use when an agent, harness, gateway, MCP workflow, or multi-step automation claims completion and the available traces, checkpoints, approvals, tool calls, or deployment records must be judged without trusting self-reported success.
risk
safe
source
self
date_added
2026-08-19

Audit Agent Run Evidence

Overview

Turn an end-to-end success statement into independently decidable claims. Reconstruct what happened from available records, grade each claim against the strongest witness, and keep missing evidence distinct from failure.

This is a read-only audit. Do not rerun tools, approve actions, resume workers, deploy artifacts, or modify evidence unless the user separately authorizes those actions.

When to Use

  • Auditing a completed or interrupted agent run from traces and artifacts.
  • Checking whether an agent's end-to-end success claim is actually supported.
  • Reviewing MCP, gateway, sandbox, checkpoint, retry, memory, approval, or deployment evidence.
  • Separating autonomous success from human-assisted or merely requested outcomes.

Do not use this skill to design instrumentation for a future run or to perform the missing actions. It evaluates evidence that already exists.

Establish the Contract

Record these inputs before judging the run:

  • declared goal and terminal success criteria;
  • run, workflow, task, and parent identifiers;
  • immutable code, configuration, model, prompt, tool-schema, and artifact revisions when available;
  • actors and trust boundaries: orchestrator, worker, sandbox, MCP server, gateway, human approver, CI, and deployment platform;
  • retry, deadline, token, cost, concurrency, and human-escalation budgets;
  • supplied evidence inventory and known collection gaps.

Do not silently strengthen the original success criteria. Do not weaken them to match the evidence that happens to exist.

Build a Claim Ledger

Split the overall claim into atomic predicates. Give every row a stable claim ID.

FieldRequired content
claim_idStable identifier
predicateOne falsifiable statement
required_witnessSource that can independently prove it
evidence_refsExact event, log, artifact, or record IDs
counterevidence_refsConflicting records
coverageRequired instances versus observed instances
verdictproven, partially_proven, contradicted, or not_proven
gapMissing field, actor, interval, or verification

Typical predicates include:

  • every required step reached its terminal postcondition;
  • sandbox isolation held for every executing worker;
  • each required MCP/tool call has a correlated response;
  • retries respected idempotency and did not duplicate committed effects;
  • a checkpoint was durably written, verified, and actually used for resume;
  • parallel branches satisfied the declared join policy;
  • memory reads cite a versioned source rather than an untracked summary;
  • retry, deadline, token, cost, and escalation budgets were respected;
  • approval was granted by an authorized human for the exact artifact and target;
  • the platform deployed that same artifact and passed the declared health checks.

Normalize Evidence

Preserve original records and create a normalized event view with:

json
{
  "run_id": "run-123",
  "event_id": "evt-42",
  "sequence": 42,
  "observed_at": "RFC3339 timestamp",
  "actor": {"type": "worker", "id": "worker-2"},
  "operation": "mcp.search",
  "state_before": "researching",
  "state_after": "researching",
  "attempt": 2,
  "request_id": "req-9",
  "idempotency_key": "task-7:search:2",
  "input_digest": "sha256:...",
  "output_digest": "sha256:...",
  "checkpoint_seq": 3,
  "parent_event_id": "evt-41",
  "status": "succeeded",
  "evidence_ref": "tool-log:991"
}

Use null or unknown for absent values. Never synthesize IDs, timestamps, digests, costs, approvals, or outcomes.

Verify bundle hashes or signatures when supplied. Check duplicate IDs, broken parent links, non-monotonic per-source sequences, impossible state transitions, unaccounted clock skew, and unexplained trace gaps. Treat an integrity failure as counterevidence for claims that depend on the affected records.

Rank Witnesses

Prefer the witness closest to the effect:

ClaimStrong witnessInsufficient alone
Code changedCommit/tree and diffAgent narration
Test passedComplete test result bound to revisionCommand invocation
MCP effect occurredServer or provider audit recordClient request
Checkpoint resumedDurable checkpoint plus verified load eventCheckpoint file exists
Human approvedAuthorization-system decision bound to artifact and targetApproval requested
Deployment succeededPlatform record plus required health checksDeployment started
Memory grounded a decisionVersioned memory read and citationFinal answer resembles memory

An orchestrator and its child worker are not independent witnesses when they repeat the same unverified result. A cryptographic digest proves byte identity, not semantic correctness.

Show full SKILL.md (443 more words)Show less

Reconstruct the Run

  1. Order events by causal links and per-source sequence; use timestamps only as supporting evidence.
  2. Build the state-transition path and mark every gap or illegal transition.
  3. Link each retry chain by logical operation, request ID, and idempotency key.
  4. Link checkpoints to the state they contain and the resume event that consumes them.
  5. Preserve every parallel branch outcome; apply the declared all_required, quorum, first_success, or other join rule.
  6. Track remaining budgets at each transition. A late success after budget exhaustion is a budget violation.
  7. Bind approvals and deployment records to exact artifact digests and targets.

Do not infer successful completion from a final state label when required intermediate predicates are missing.

Assign Verdicts

  • proven: authentic evidence covers every instance of the predicate and no reliable counterevidence remains.
  • partially_proven: some required instances or fields are proven and the uncovered portion is named.
  • contradicted: reliable evidence conflicts with the predicate.
  • not_proven: evidence is absent, circular, unverifiable, or only self-reported.

Use not_proven, not contradicted, for missing logs. Use contradicted when the trace shows a failed health check, duplicate effect, unauthorized approver, corrupt checkpoint, skipped required branch, or exhausted budget.

The end-to-end verdict cannot be stronger than its weakest required predicate. Optional diagnostics may remain unproven without failing the run if they were never part of the declared contract.

Report

Return sections in this order:

  1. Scope and evidence inventory — run identity, declared criteria, records inspected, integrity checks.
  2. Claim ledger — one row per predicate with verdict and exact references.
  3. Reconstructed timeline — only state-changing, fault, retry, checkpoint, join, approval, and deployment events.
  4. Gaps and counterevidence — identify the affected claims and whether collection can still recover the evidence.
  5. Overall verdict — one sentence plus the blocking claim IDs.

Example conclusion:

partially_proven: repository steps C1-C18 and checkpoint recovery C22 are proven, but deployment success is not proven because C31 has only a client-side start event and no platform health result.

Common Mistakes

  • Treating a successful process exit as proof of the business postcondition.
  • Counting retries as separate successful logical operations.
  • Accepting a child agent's summary as independent corroboration.
  • Calling a checkpoint recoverable without observing a verified reload.
  • Calling an approval request an approval grant.
  • Reporting percentages without listing the denominator and missing instances.
  • Recommending instrumentation as though it were evidence from the completed run.

Limitations

  • An audit cannot recover facts that no trusted source recorded.
  • Provider logs may establish external effects without proving the agent's internal reasoning.
  • Redaction may be necessary for secrets and personal data; record the redaction scope and preserve stable references.
  • If evidence collection would mutate external state or expose sensitive data, stop and request authorization.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/audit-agent-run-evidence of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Audit Agent Run Evidence next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Agent Run Evidence compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Agent Run Evidence this skillsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT
Prepare Cloudflare Production DeploymentLubomirGeorgiev/cloudflare-workers-nextjs-saas-template786—~5.9kAutomated safety check: NotesMIT
Deploy Observabilityaliyun/alibabacloud-observability-mcp-server166—~2.6kAutomated safety check: NotesNone
Release Allpaperboytm/spool592—~1.1kAutomated safety check: PassCustom licence
Deploynoskillish/bankmcp277—~744Automated safety check: PassMIT

Similar skills

  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Prepare Cloudflare Production Deployment

    LubomirGeorgiev/cloudflare-workers-nextjs-saas-template

    Source-of-truth runbook for preparing this Vinext Cloudflare Workers SaaS template for production deployment.

    786 GitHub stars~5.9k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Deploy Observability

    aliyun/alibabacloud-observability-mcp-server

    Deploy, start, and update the Alibaba Cloud Observability MCP Server (阿里云可观测 MCP Server).

    166 GitHub stars~2.6k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Release All

    paperboytm/spool

    Publish the complete Spool CLI release train: synchronized versions, npm packages, the GitHub release, and the matching production web deployment.

    592 GitHub stars~1.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Deploy

    noskillish/bankmcp

    Deploy BankMCP™ to a small server so it works in claude.ai and on the phone: Railway or Fly.io, volume, domain, setup page, connector.

    277 GitHub stars~744 tokensUpdated 12 days ago
    DevOps & CloudAuto-check passed
  • Hcls Deploy Agent

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    A skill your agent uses when a developer wants to deploy an HCLS agent to Amazon Bedrock AgentCore, configure Gateway tools as MCP endpoints, set up authentication with Cognito, configure memory, or…

    274 GitHub stars~813 tokensUpdated 9 days ago
    DevOps & CloudAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Categories

Questions about Audit Agent Run Evidence

What does Audit Agent Run Evidence do?

A skill your agent uses when an agent, harness, gateway, MCP workflow, or multi-step automation claims completion and the available traces, checkpoints, approvals, tool calls, or deployment records…. Audit Agent Run Evidence is an agent skill from sickn33/agentic-awesome-skills. Use when an agent, harness, gateway, MCP workflow, or multi-step automation claims completion and the available traces, checkpoints, approvals, tool calls, or deployment records must be judged without trusting self-reported success.

When should I use Audit Agent Run Evidence?

Audit Agent Run Evidence fits situations like: multi-step automation claims completion and the available traces; deployment records must be judged without trusting self-reported success.

How do I install Audit Agent Run Evidence in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill audit-agent-run-evidence -a claude-code`. Or copy the skill folder (skills/audit-agent-run-evidence in sickn33/agentic-awesome-skills) into .claude/skills/audit-agent-run-evidence in your project. Claude Code loads it when a task matches its description.

How do I install Audit Agent Run Evidence in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill audit-agent-run-evidence -a codex`. Or copy the skill folder (skills/audit-agent-run-evidence in sickn33/agentic-awesome-skills) into .agents/skills/audit-agent-run-evidence in your project. Codex loads it when a task matches its description.

Can I use Audit Agent Run Evidence in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill audit-agent-run-evidence -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-agent-run-evidence, .gemini/skills/audit-agent-run-evidence, .github/skills/audit-agent-run-evidence and .opencode/skills/audit-agent-run-evidence in your project.

What does Audit Agent Run Evidence need to run?

SKILL.md names no scripts, command-line tools or credentials: Audit Agent Run Evidence is instructions for the agent only.

Does Audit Agent Run Evidence access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audit Agent Run Evidence safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit Agent Run Evidence use?

Audit Agent Run Evidence is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Agent Run Evidence use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Agent Run Evidence?

Skills that share tags, products or a category with Audit Agent Run Evidence: AWS Cdk Development (zxkane/aws-skills, 367 stars), Prepare Cloudflare Production Deployment (LubomirGeorgiev/cloudflare-workers-nextjs-saas-template, 786 stars), Deploy Observability (aliyun/alibabacloud-observability-mcp-server, 166 stars) and Release All (paperboytm/spool, 592 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Agent Run Evidence?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.