Agent skill

Hk Autonomy Audit

by deepklarity in deepklarity/harness-kit

Audit whether an AI agent can autonomously close the loop on problems in a given area — from discovering a symptom to verifying a fix — without human intervention.

MITAuto-check: notesDevelopment

Install Hk Autonomy Audit

skills CLI
$ npx skills add deepklarity/harness-kit --skill hk-autonomy-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install deepklarity/harness-kit hk-autonomy-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/deepklarity/harness-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/hk-autonomy-audit .claude/skills/hk-autonomy-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hk-autonomy-audit
GitHub stars
100
Token cost
~2.5k tokens
SKILL.md length
1,096 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Audit whether an AI agent can autonomously close the loop on problems in a given area — from discovering a symptom to verifying a fix — without human intervention.

  • Works in 4 steps: Scope the audit → Walk the agent journey → Rate each stage → …
  • Someone wants to assess debugging readiness
  • SKILL.md covers Target, The Mental Model, Process and Output, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Hk Autonomy Audit is an agent skill from deepklarity/harness-kit. Audit whether an AI agent can autonomously close the loop on problems in a given area — from discovering a symptom to verifying a fix — without human intervention. Evaluates documentation, diagnostic tools, commands, logs, and flows for completeness and actionability. Generates a gap-focused report with ratings. Use this skill whenever someone wants to assess debugging readiness, check if docs are agent-sufficient, audit a workflow for autonomous solvability, evaluate operational tooling coverage, or wants to…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Debugging. The repository describes itself as: A kit for building with AI agents and also the engineering patterns around it. The licence is MIT.

When your agent uses it

  • Someone wants to assess debugging readiness
  • Check if docs are agent-sufficient
  • Audit a workflow for autonomous solvability
  • Evaluate operational tooling coverage

Example prompts

  • “could an agent fix this on its own?”
  • “loop audit”
  • “audit this flow”
  • “/hk-autonomy-audit”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read, Edit, Write, Task, Grep, Glob

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Scope the audit
  2. Walk the agent journey
  3. Rate each stage
  4. Identify the critical gaps

What it can do on your machine

Read from SKILL.md and the folder at commit 87305cd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Edit
    • Write
    • Task
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hk Autonomy Audit loads about 2.5k tokens when it runs. Until then it costs about 186 tokens; SKILL.md has 1,096 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~186
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Edit, Write, Task, Grep, Glob

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from deepklarity/harness-kit at commit 87305cd, republished under its MIT licence (© deepklarity). 1,096 words, ~2,462 tokens.

Download SKILL.mdSave it as .claude/skills/hk-autonomy-audit/SKILL.md (or your agent's skills folder).
name
hk-autonomy-audit
description
Audit whether an AI agent can autonomously close the loop on problems in a given area — from discovering a symptom to verifying a fix — without human intervention. Evaluates documentation, diagnostic tools, commands, logs, and flows for completeness and actionability. Generates a gap-focused report with ratings. Use this skill whenever someone wants to assess debugging readiness, check if docs are agent-sufficient, audit a workflow for autonomous solvability, evaluate operational tooling coverage, or wants to know 'could an agent fix this on its own?' Triggers on: 'loop audit', 'audit this flow', 'is this debuggable', 'agent readiness', 'can an agent solve this', 'autonomous debugging check', or /hk-autonomy-audit.
allowed-tools
Bash, Read, Edit, Write, Task, Grep, Glob
argument-hint
[area/flow/doc to audit, e.g. 'odin exec flow' or 'task failure debugging']

/hk-autonomy-audit — Autonomous Loop-Closing Readiness Audit

You're auditing whether the tooling, docs, commands, and flows in a given area are sufficient for an AI agent to autonomously solve problems — from first symptom to verified fix — without stopping to ask a human.

This is not a documentation quality check. It's an operational readiness assessment. The question isn't "do docs exist?" but "if an agent hit a wall here at 3am, could it get itself unstuck?"

Target

<audit_target> $ARGUMENTS </audit_target>

If the target is empty or vague, ask the user:

  1. What area or flow should be audited? (e.g., "odin task execution", "taskit API debugging", "reflection quality issues")
  2. Is there a specific scenario that prompted this? (a recent failure where an agent got stuck is the best input)

If the user provides a doc path, start there but don't stop there — trace outward to the commands, tools, and flows the doc references.

The Mental Model

An AI agent closing the loop on a problem goes through six stages. A gap at any stage breaks the chain:

DISCOVER → DIAGNOSE → HYPOTHESIZE → FIX → VERIFY → DOCUMENT
   ↓          ↓           ↓          ↓       ↓          ↓
 "Something  "The root   "Changing  "Apply  "Confirm   "Record what
  is wrong"   cause is    X should   the     it works   happened and
              Y because   fix it     change  end-to-    why"
              Z"          because W" itself  end"

Each stage needs specific resources. The audit checks whether those resources exist, are discoverable, and are actually usable by an agent (not just by a human who knows where to look).

Process

Step 1: Scope the audit

Read the target area's CLAUDE.md, AGENTS.md, and any referenced docs. Build a mental map of:

  • What problems can occur here? (error types, failure modes, misconfigurations)
  • What tools exist for this area? (diagnostic scripts, CLI commands, log files)
  • What docs cover this area? (guides, patterns, solutions)

Don't read everything — scan headings and structure first. Depth comes in Step 2 when you know where to look.

Step 2: Walk the agent journey

For each stage, evaluate from the perspective of an AI agent that has access to the repo's CLAUDE.md files and tools but no prior tribal knowledge. Use parallel subagents to check multiple stages simultaneously.

DISCOVER — Can the agent detect that something is wrong?

  • Are error messages actionable? (Do they say what failed, or just "error"?)
  • Are logs accessible and parseable? (Where are they? What format? Can an agent tail them?)
  • Are there health checks or status commands? (Quick "is this working?" checks)
  • Is there monitoring that surfaces problems before a human reports them?

DIAGNOSE — Can the agent find the root cause?

  • Are diagnostic scripts/commands available? (Not just "look at the code")
  • Do diagnostic tools explain what they find? (Auto-detected problems, not just raw data)
  • Is the data flow traceable? (Can the agent follow data from input to symptom?)
  • Are common failure patterns documented with their signatures?

HYPOTHESIZE — Can the agent form a theory?

  • Do docs explain the WHY behind design decisions? (Not just what the code does)
  • Are edge cases and gotchas documented? (The non-obvious things)
  • Are there solution docs from past incidents? (Searchable by symptom)
  • Is there enough architectural context to reason about side effects?

FIX — Can the agent make the change?

  • Are modification commands documented? (Not just read-only inspection)
  • Are there guard rails? (Tests that catch regressions, linters, type checks)
  • Is the change surface well-bounded? (Can the agent know which files to touch?)
  • Are there examples of similar past fixes? (Patterns to follow)

VERIFY — Can the agent confirm the fix works?

  • Are test commands documented and runnable? (Not just "run the tests")
  • Is there a live verification path? (Beyond unit tests — can the agent check end-to-end?)
  • Are success criteria defined? (How does "working" look, specifically?)
  • Can the agent verify without human eyes? (No "check the UI visually" without tooling)

DOCUMENT — Can the agent record what happened?

  • Is there a documentation workflow? (Where to put learnings, what format)
  • Are there templates for incident docs? (Solution docs, RCA reports)
  • Is the compounding mechanism discoverable? (Would an agent know to use /hk-compound?)
Show full SKILL.md (474 more words)Show less
Step 3: Rate each stage

For each of the six stages, assign a readiness level:

  • GREEN — Agent can handle this autonomously. Tools exist, are documented, and are discoverable.
  • YELLOW — Agent can probably handle this but might waste time or miss things. Partial tooling, unclear docs, or undiscoverable resources.
  • RED — Agent will get stuck here. Missing tools, no docs, or requires human knowledge that isn't written down.

The rating is about the weakest realistic scenario, not the happy path. If the diagnostic script works great for task failures but there's no way to debug harness timeouts, the DIAGNOSE stage is YELLOW (not GREEN just because one path works).

Step 4: Identify the critical gaps

For each YELLOW and RED stage, identify the specific gaps. A gap is:

  • Something an agent would need but can't find
  • Something that exists but isn't discoverable (buried in code, not in docs)
  • Something that requires human judgment that could be codified
  • Something that works for one scenario but not others in the same area

Prioritize gaps by impact: which ones would block the agent most often?

Output

Create the report

Write to docs/loop_audits/<area-slug>-<date>.md:

markdown
# Loop Audit: [Area Name]

**Audited**: [date]
**Target**: [what was audited]
**Trigger**: [what prompted this audit, if known]

## Readiness Summary

| Stage | Rating | Key Gap |
|-------|--------|---------|
| Discover | GREEN/YELLOW/RED | [one-line gap or "—"] |
| Diagnose | GREEN/YELLOW/RED | [one-line gap or "—"] |
| Hypothesize | GREEN/YELLOW/RED | [one-line gap or "—"] |
| Fix | GREEN/YELLOW/RED | [one-line gap or "—"] |
| Verify | GREEN/YELLOW/RED | [one-line gap or "—"] |
| Document | GREEN/YELLOW/RED | [one-line gap or "—"] |

**Overall**: [RED/YELLOW/GREEN — the weakest stage determines the overall rating]

## Gaps (ranked by agent-blocking impact)

### GAP-1: [Short title]
**Stage**: [which stage this blocks]
**Impact**: [what happens when an agent hits this — be specific]
**What exists**: [what's already there, briefly]
**What's missing**: [the specific gap]
**Suggested fix**: [concrete action — a doc to write, a script to add, a command to document]
**Effort**: [small/medium/large]

### GAP-2: ...
[repeat for each gap, ranked by impact]

## What Works Well
[2-3 sentences max. Not a list of everything that's fine — just notable strengths that other areas should learn from. Skip this section entirely if nothing stands out.]

## Recommendations
[Ordered list of the top 3-5 actions that would most improve autonomous solvability. Each should be concrete enough to act on without further clarification.]
Report principles

The report is the product. It should be:

  • Gap-focused: Don't catalog what works. An agent that reads this report should immediately know what's broken and what to do about it. Strengths get at most 2-3 sentences — only if they're worth replicating elsewhere.
  • Specific: "Docs are insufficient" is not a finding. "There's no way to diagnose harness timeout failures — task_inspect.py shows task metadata but doesn't surface the harness subprocess stderr, which is where timeout errors appear" is a finding.
  • Actionable: Every gap must have a suggested fix that someone could execute without asking follow-up questions.
  • Honest about severity: If an area is genuinely well-covered, say so with a GREEN and move on. Don't manufacture gaps to justify the audit. An all-GREEN report with "no significant gaps found" is a valid and useful outcome.
Print summary

After writing the report, print to conversation:

  • The readiness summary table
  • The top 3 gaps with their suggested fixes
  • The report file path

Don't paste the entire report — the user can read the file.

Edge Cases

What if the target is too broad? (e.g., "audit everything") — Pick the area with the most recent failures or the most complex flow. Audit that deeply rather than auditing everything shallowly. Suggest follow-up audits for other areas.

What if the target is already well-covered? — That's a valid finding. Write a short report confirming GREEN across stages, note any minor improvements, and move on. Don't inflate minor issues.

What if the audit reveals a gap you can fix right now? — Don't fix it. The audit's job is to produce the report. Fixing gaps is a separate task that the user should prioritize. Mention "this could be fixed now" in the effort field if it's truly trivial.

© deepklarity, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/hk-autonomy-audit of deepklarity/harness-kit.

Open the folder on GitHubat commit 87305cd

Compare with similar skills

Hk Autonomy Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hk Autonomy Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hk Autonomy Audit this skilldeepklarity/harness-kit100—~2.5kAutomated safety check: NotesMIT
Trellis Session Insightmindfold-ai/Trellis15k4 repos~1.7kAutomated safety check: PassAGPL-3.0
Trellis Channelmindfold-ai/Trellis15k4 repos~1.2kAutomated safety check: PassAGPL-3.0
Vibe Coding PartnershareAI-lab/Kode-CLI5.2k—~5.6kAutomated safety check: PassApache-2.0
Leon Coding Agentleon-ai/leon18k—~1.1kAutomated safety check: PassMIT
Octocode Code Researchbgauryy/octocode949—~1.5kAutomated safety check: PassMIT

Similar skills

  • Trellis Session Insight

    mindfold-ai/Trellis

    Reach into past AI conversation history through the trellis mem CLI.

    15k GitHub starsUsed in 4 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Trellis Channel

    mindfold-ai/Trellis

    Use Trellis channel for live multi-agent collaboration, spawned workers, cross-agent review, progress inspection, forum channels, and channel log debugging.

    15k GitHub starsUsed in 4 repos~1.2k tokens
    DevelopmentAuto-check passed
  • Vibe Coding Partner

    shareAI-lab/Kode-CLI

    Gives an agent a set of working rules for any development task: understand first, surface decisions, verify results, and load deeper reference files per scenario.

    5.2k GitHub stars~5.6k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Leon Coding Agent

    leon-ai/leon

    Has Leon's agent investigate, change and verify code in a repository with its file, search and shell tools, staying inside the scope the owner authorized.

    18k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Octocode Code Research

    bgauryy/octocode

    Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label.

    949 GitHub stars~1.5k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Misakanet Failure Memory

    Ikalus1988/MisakaNet

    Search and record failure-recovery lessons from real engineering sessions; submit and verify debugging lessons across the MisakaNet network.

    526 GitHub stars~1.9k tokensUpdated today
    DevelopmentAuto-check passed

More from deepklarity/harness-kit

All 18 skills in this repo
  • Hk Skill Creator

    deepklarity/harness-kit

    Create new skills, modify and improve existing skills, and measure skill performance.

    100 GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Arch Audit

    deepklarity/harness-kit

    Run comprehensive agent-native architecture review with scored principles.

    100 GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Mock First

    deepklarity/harness-kit

    Mock-first, layer-by-layer feature development. An agent skill from deepklarity/harness-kit.

    100 GitHub stars~3.9k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Breadcrumb Creator

    deepklarity/harness-kit

    Traces a workflow end-to-end through the harness-kit monorepo and creates a breadcrumb analysis doc in docs/breadcrumbanalysis/.

    100 GitHub stars~3k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Changelog

    deepklarity/harness-kit

    Generate changelog entries from git diffs, prepend to CHANGELOG.md, and optionally commit + PR.

    100 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Compound

    deepklarity/harness-kit

    Compound a learning into a reusable pattern. An agent skill from deepklarity/harness-kit.

    100 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check: notes

Questions about Hk Autonomy Audit

What does Hk Autonomy Audit do?

Audit whether an AI agent can autonomously close the loop on problems in a given area — from discovering a symptom to verifying a fix — without human intervention. Hk Autonomy Audit is an agent skill from deepklarity/harness-kit. Audit whether an AI agent can autonomously close the loop on problems in a given area — from discovering a symptom to verifying a fix — without human intervention.

When should I use Hk Autonomy Audit?

Hk Autonomy Audit fits situations like: someone wants to assess debugging readiness; check if docs are agent-sufficient; audit a workflow for autonomous solvability; evaluate operational tooling coverage.

How do I install Hk Autonomy Audit in Claude Code?

Run `npx skills add deepklarity/harness-kit --skill hk-autonomy-audit -a claude-code`. Or copy the skill folder (.claude/skills/hk-autonomy-audit in deepklarity/harness-kit) into .claude/skills/hk-autonomy-audit in your project. Claude Code loads it when a task matches its description.

How do I install Hk Autonomy Audit in Codex?

Run `npx skills add deepklarity/harness-kit --skill hk-autonomy-audit -a codex`. Or copy the skill folder (.claude/skills/hk-autonomy-audit in deepklarity/harness-kit) into .agents/skills/hk-autonomy-audit in your project. Codex loads it when a task matches its description.

Can I use Hk Autonomy Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add deepklarity/harness-kit --skill hk-autonomy-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hk-autonomy-audit, .gemini/skills/hk-autonomy-audit, .github/skills/hk-autonomy-audit and .opencode/skills/hk-autonomy-audit in your project.

What does Hk Autonomy Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Hk Autonomy Audit is instructions for the agent only. Its frontmatter pre-approves these tools: Bash, Read, Edit, Write, Task, Grep, Glob.

Does Hk Autonomy Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hk Autonomy Audit safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Hk Autonomy Audit use?

Hk Autonomy Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hk Autonomy Audit use?

About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hk Autonomy Audit?

Skills that share tags, products or a category with Hk Autonomy Audit: Trellis Session Insight (mindfold-ai/Trellis, 15k stars), Trellis Channel (mindfold-ai/Trellis, 15k stars), Vibe Coding Partner (shareAI-lab/Kode-CLI, 5.2k stars) and Leon Coding Agent (leon-ai/leon, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hk Autonomy Audit?

deepklarity (a GitHub organization) maintains it in deepklarity/harness-kit, which has 100 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on July 15, 2026.

Source: deepklarity/harness-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.