Agent skill

Hk Refine

by deepklarity in deepklarity/harness-kit

Iteratively improve any output by running a structured observe-hypothesize-change-rerun loop.

MITAuto-check: notesDevelopment

Install Hk Refine

skills CLI
$ npx skills add deepklarity/harness-kit --skill hk-refine -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install deepklarity/harness-kit hk-refine --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/deepklarity/harness-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/hk-refine .claude/skills/hk-refine && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hk-refine
GitHub stars
100
Token cost
~2.4k tokens
SKILL.md length
716 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Iteratively improve any output by running a structured observe-hypothesize-change-rerun loop.

  • Works in 7 steps: SET UP (first invocation only) → OBSERVE → HYPOTHESIZE → …
  • An output (reflection
  • SKILL.md covers Context, The Scratch Directory, The Loop and Context Management Rules
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Hk Refine is an agent skill from deepklarity/harness-kit. Iteratively improve any output by running a structured observe-hypothesize-change-rerun loop. Uses an organized scratch directory to prevent context blowup — the conversation stays thin while iterations accumulate on disk. Use when an output (reflection, plan, prompt, pipeline result) isn't good enough and needs systematic refinement. Triggers on: 'close the loop', 'this output isn't good enough', 'iterate on this', 'refine this output', 'improve this reflection', or /hk-refine.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development. The repository describes itself as: A kit for building with AI agents and also the engineering patterns around it. The licence is MIT.

When your agent uses it

  • An output (reflection
  • Pipeline result) isnt good enough and needs systematic refinement
  • : close the loop
  • This output isnt good enough

Example prompts

  • “t good enough and needs systematic refinement. Triggers on:”
  • “this output isn”
  • “iterate on this”
  • “/hk-refine”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read, Edit, Write, Task, Grep, Glob

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. SET UP (first invocation only)
  2. OBSERVE
  3. HYPOTHESIZE
  4. CHANGE
  5. RUN
  6. COMPARE
  7. VERDICT

What it can do on your machine

Read from SKILL.md and the folder at commit 87305cd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Edit
    • Write
    • Task
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hk Refine loads about 2.4k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 716 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Edit, Write, Task, Grep, Glob

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from deepklarity/harness-kit at commit 87305cd, republished under its MIT licence (© deepklarity). 716 words, ~2,371 tokens.

Download SKILL.mdSave it as .claude/skills/hk-refine/SKILL.md (or your agent's skills folder).
name
hk-refine
description
Iteratively improve any output by running a structured observe-hypothesize-change-rerun loop. Uses an organized scratch directory to prevent context blowup — the conversation stays thin while iterations accumulate on disk. Use when an output (reflection, plan, prompt, pipeline result) isn't good enough and needs systematic refinement. Triggers on: 'close the loop', 'this output isn't good enough', 'iterate on this', 'refine this output', 'improve this reflection', or /hk-refine.
allowed-tools
Bash, Read, Edit, Write, Task, Grep, Glob
argument-hint
[scratch_dir] [description of what to improve]

/hk-refine — Iterative Output Refinement

You have an open loop: something produced output, the output isn't good enough, and you need to systematically improve it. This skill closes that loop.

The core discipline: everything lives on disk, not in conversation context. The scratch directory is the organized record of what was tried, what worked, and why. Subagents read from disk, write to disk, and report back in 2-3 line summaries. The main conversation only holds the current hypothesis and verdict — never full outputs.

Context

<loop_context> $ARGUMENTS </loop_context>

If the context above is empty or unclear, ask the user:

  1. What produced the output? (command, prompt, pipeline step)
  2. What's wrong with it? (vague is ok — "it's not good enough" is a valid start)
  3. Where should the scratch directory live? (suggest temp_reflections/ or temp_loop/)

The Scratch Directory

This is the product. Not temp files — the organized log of the refinement process.

<scratch_dir>/
├── loop.md                # Live loop state (see template below)
├── baseline/
│   ├── output.md          # Original output that needs improvement
│   ├── critique.md        # Structured critique of what's wrong
│   └── run_command.txt    # Exact command/process that produced the output
├── iter-1/
│   ├── hypothesis.md      # What to change and why
│   ├── changes.md         # What was actually changed (with file paths and diffs)
│   ├── output.md          # New output after changes
│   ├── comparison.md      # Before vs after, structured
│   └── verdict.md         # Better / worse / mixed — with evidence
├── iter-2/
│   └── ...
└── summary.md             # Written when loop closes
loop.md template

This file is the single source of truth for where the loop is. Read it at the start of every iteration. Update it after every verdict.

markdown
# Close-the-Loop: [short description]

## Target
What we're improving: [one line]
Run command: [the exact command to re-run]
Quality signal: [how we know it's better — specific, measurable if possible]

## Current State
Iteration: [N]
Best so far: [baseline | iter-N]
Status: [observing | hypothesizing | changing | running | comparing | closed]

## Hypothesis Log
- iter-1: [hypothesis] → [verdict: better/worse/mixed]
- iter-2: [hypothesis] → [verdict]
- ...

## What We've Learned
- [Accumulated insights that carry forward — things that definitely help or definitely don't]

## Next
[What to try next, or "CLOSED: [reason]"]

The Loop

Phase 0: SET UP (first invocation only)
  1. Create the scratch directory structure
  2. Capture the baseline output — either from the user's clipboard/description, or by running the command
  3. Write run_command.txt with the exact command that produces the output
  4. Write loop.md with initial state
  5. Proceed to Phase 1
Phase 1: OBSERVE

Delegate critique to a subagent (sonnet). The subagent reads the current output and produces a structured critique. The main conversation does NOT read the full output — only the critique summary.

Task(model: sonnet, subagent_type: general-purpose)

Read <scratch_dir>/[baseline or iter-N]/output.md

Produce a structured critique:
1. STRENGTHS: What's working well (keep these)
2. WEAKNESSES: What's not working, ranked by impact
3. MISSING: What should be there but isn't
4. EXCESS: What's there but shouldn't be (noise, fluff, wrong focus)
5. ROOT ISSUE: The single biggest thing to fix (not a list — pick one)

Write your critique to <scratch_dir>/[baseline or iter-N]/critique.md
Return a 3-line summary: the root issue, the top weakness, and one strength to preserve.
Phase 2: HYPOTHESIZE

Based on the critique summary (not the full output), form a hypothesis. This happens in the main conversation — it's a judgment call, not mechanical work.

Write to <scratch_dir>/iter-N/hypothesis.md:

markdown
# Hypothesis for Iteration N

## What to change
[Specific change — which file, which prompt section, which config value]

## Why this should help
[Connect the change to the root issue from the critique]

## What to watch for
[Side effects — things that might get worse when this gets better]

## Estimated impact
[High / Medium / Low — on the specific quality signal defined in loop.md]

The hypothesis must be specific enough that someone else could apply the change without seeing the conversation. "Make the prompt better" is not a hypothesis. "Add a structured output format requirement to the reflection prompt because the current output is unstructured prose that's hard to evaluate" is a hypothesis.

Phase 3: CHANGE

Apply the changes described in the hypothesis. This could be:

  • Editing a prompt file
  • Changing a config value
  • Modifying code that processes/generates the output
  • Adjusting parameters (model, temperature, max tokens)

Log what changed in <scratch_dir>/iter-N/changes.md:

markdown
# Changes for Iteration N

## Files modified
- `path/to/file.py` — [what changed, 1 line]
- `path/to/prompt.md` — [what changed, 1 line]

## Diffs
[Actual diffs or before/after snippets for each change]
Phase 4: RUN

Execute the command from run_command.txt to produce new output. Capture the output to <scratch_dir>/iter-N/output.md.

If the run command involves odin exec, odin reflect, or similar commands that can't run inside Claude Code, provide the user with copy-paste commands and wait for them to paste the output back. Note this in loop.md's status.

If the run command is something that CAN run (a Python script, a test, an API call), run it directly.

Show full SKILL.md (260 more words)Show less
Phase 5: COMPARE

Delegate comparison to a subagent (sonnet). The subagent reads ONLY the two outputs — it does not see the hypothesis or changes. This keeps the comparison unbiased.

Task(model: sonnet, subagent_type: general-purpose)

Compare these two outputs for quality. You do not know which is "old" or "new."

Output A: <scratch_dir>/[previous best]/output.md
Output B: <scratch_dir>/iter-N/output.md

Quality signal: [from loop.md]

Produce:
1. WINNER: A or B or TIE (on the specific quality signal)
2. EVIDENCE: 3-5 specific examples showing why
3. TRADE-OFFS: Did anything get worse in the winner?
4. CONFIDENCE: How clear is the difference? (obvious / marginal / unclear)

Write to <scratch_dir>/iter-N/comparison.md
Return: winner + confidence + one-line evidence summary
Phase 6: VERDICT

Based on the comparison summary, update loop state:

Write <scratch_dir>/iter-N/verdict.md:

markdown
# Verdict: Iteration N

Result: [BETTER / WORSE / MIXED]
Confidence: [obvious / marginal / unclear]
Evidence: [1-2 lines from comparison]
Keep: [what to preserve from this iteration]
Revert: [what to undo if anything]

Update loop.md:

  • Increment iteration
  • Update "best so far"
  • Add to hypothesis log
  • Add to "what we've learned"
  • Set "next" — either another hypothesis or CLOSED
When to close the loop

Close when any of these are true:

  • The quality signal is met (output is good enough)
  • 3 consecutive iterations show no improvement (diminishing returns)
  • The user says to stop
  • The cost of another iteration exceeds the expected improvement

Write <scratch_dir>/summary.md:

markdown
# Loop Summary: [description]

## Result
Started: [date]
Iterations: [N]
Best: [iter-N]
Status: [closed — quality met / closed — diminishing returns / closed — user stopped]

## What worked
- [Changes that improved output, with evidence]

## What didn't work
- [Changes that didn't help or made things worse]

## Final state
Run command: [the command with all improvements applied]
Output quality: [assessment against the original quality signal]

## If reopening later
Read iter-[best]/output.md for the current best.
The key changes that got us here: [1-2 sentences].
The remaining weakness: [if any].

Context Management Rules

These are non-negotiable — they're the entire point of using a scratch directory:

  1. Never paste full outputs into the conversation. They live on disk. Subagents read them from disk. The main conversation sees only summaries.

  2. Never hold more than one iteration's hypothesis + verdict in conversation. If you need to reference earlier iterations, re-read loop.md — it has the condensed history.

  3. Subagents are stateless. Each subagent gets pointed at specific files on disk. They don't inherit conversation context. This is a feature — it prevents context buildup.

  4. loop.md is the resumption point. If the conversation compacts or a new session starts, loop.md + the iteration folders contain everything needed to continue.

  5. The user sees summaries, not data. After each phase, report to the user in 2-3 lines: what happened, what the verdict was, what's next. They can dig into the scratch directory if they want details.

© deepklarity, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/hk-refine of deepklarity/harness-kit.

Open the folder on GitHubat commit 87305cd

Compare with similar skills

Hk Refine next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hk Refine compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hk Refine this skilldeepklarity/harness-kit100—~2.4kAutomated safety check: NotesMIT
Vercel Composition Patternssupabase/supabase111k58 repos~726Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k25 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 58 repos~726 tokens
    DevelopmentAuto-check passed
  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 25 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed

More from deepklarity/harness-kit

All 18 skills in this repo
  • Hk Skill Creator

    deepklarity/harness-kit

    Create new skills, modify and improve existing skills, and measure skill performance.

    100 GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Arch Audit

    deepklarity/harness-kit

    Run comprehensive agent-native architecture review with scored principles.

    100 GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Mock First

    deepklarity/harness-kit

    Mock-first, layer-by-layer feature development. An agent skill from deepklarity/harness-kit.

    100 GitHub stars~3.9k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Autonomy Audit

    deepklarity/harness-kit

    Audit whether an AI agent can autonomously close the loop on problems in a given area — from discovering a symptom to verifying a fix — without human intervention.

    100 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Breadcrumb Creator

    deepklarity/harness-kit

    Traces a workflow end-to-end through the harness-kit monorepo and creates a breadcrumb analysis doc in docs/breadcrumbanalysis/.

    100 GitHub stars~3k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Changelog

    deepklarity/harness-kit

    Generate changelog entries from git diffs, prepend to CHANGELOG.md, and optionally commit + PR.

    100 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check: notes

Categories

Questions about Hk Refine

What does Hk Refine do?

Iteratively improve any output by running a structured observe-hypothesize-change-rerun loop. Hk Refine is an agent skill from deepklarity/harness-kit. Iteratively improve any output by running a structured observe-hypothesize-change-rerun loop.

When should I use Hk Refine?

Hk Refine fits situations like: an output (reflection; pipeline result) isnt good enough and needs systematic refinement; : close the loop; this output isnt good enough.

How do I install Hk Refine in Claude Code?

Run `npx skills add deepklarity/harness-kit --skill hk-refine -a claude-code`. Or copy the skill folder (.claude/skills/hk-refine in deepklarity/harness-kit) into .claude/skills/hk-refine in your project. Claude Code loads it when a task matches its description.

How do I install Hk Refine in Codex?

Run `npx skills add deepklarity/harness-kit --skill hk-refine -a codex`. Or copy the skill folder (.claude/skills/hk-refine in deepklarity/harness-kit) into .agents/skills/hk-refine in your project. Codex loads it when a task matches its description.

Can I use Hk Refine in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add deepklarity/harness-kit --skill hk-refine -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hk-refine, .gemini/skills/hk-refine, .github/skills/hk-refine and .opencode/skills/hk-refine in your project.

What does Hk Refine need to run?

SKILL.md names no scripts, command-line tools or credentials: Hk Refine is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Edit, Write, Task, Grep, Glob.

Does Hk Refine access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hk Refine safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Hk Refine use?

Hk Refine is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hk Refine use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hk Refine?

Skills that share tags, products or a category with Hk Refine: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 297k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hk Refine?

deepklarity (a GitHub organization) maintains it in deepklarity/harness-kit, which has 100 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on July 15, 2026.

Source: deepklarity/harness-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.