Agent skill

Harness Feedback

by AnastasiyaW in AnastasiyaW/codex-claude-code-config

A skill your agent uses when an agent says a test, VM, proof, evaluator, or release gate is overloaded, too strict, blocking staging, or causing false positives; split checks by profile, measure the…

MITAuto-check passedSecurity

Install Harness Feedback

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill harness-feedback -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config harness-feedback --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/development/harness-feedback .claude/skills/harness-feedback && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-feedback
GitHub stars
154
Token cost
~1k tokens
SKILL.md length
472 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when an agent says a test, VM, proof, evaluator, or release gate is overloaded, too strict, blocking staging, or causing false positives; split checks by profile, measure the…

  • Works in 6 steps: requested profile and change boundary; → gate that blocked or dominated the run; → command, elapsed time, failure count,… → …
  • An agent says a test
  • SKILL.md covers Profiles, Feedback Loop, Required Report and Gotchas, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Harness Feedback is an agent skill from AnastasiyaW/codex-claude-code-config. Use when an agent says a test, VM, proof, evaluator, or release gate is overloaded, too strict, blocking staging, or causing false positives; split checks by profile, measure the burden, preserve high-risk evidence, and verify the smallest corrected workflow. Do not use for ordinary test selection, a single test failure, or a full security audit without a harness-scope question.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Security, covering Failing and flaky tests and Security review. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • An agent says a test
  • Release gate is overloaded
  • Blocking staging
  • Causing false positives

Example prompts

  • “/harness-feedback”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. requested profile and change boundary;
  2. gate that blocked or dominated the run;
  3. command, elapsed time, failure count, and evidence actually produced;
  4. whether the gate was relevant, duplicated, flaky, or misplaced;
  5. the smallest profile split or deletion of duplicate coverage;
  6. a before/after run of the affected profile and a fresh review of the rule.

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Feedback loads about 1k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 472 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 472 words, ~1,041 tokens.

Download SKILL.mdSave it as .claude/skills/harness-feedback/SKILL.md (or your agent's skills folder).
name
harness-feedback
description
Use when an agent says a test, VM, proof, evaluator, or release gate is overloaded, too strict, blocking staging, or causing false positives; split checks by profile, measure the burden, preserve high-risk evidence, and verify the smallest corrected workflow. Do not use for ordinary test selection, a single test failure, or a full security audit without a harness-scope question.

Harness Feedback

Treat "the harness is too strict" as an engineering finding, not as permission to disable a safety check. Find the boundary that owns the mismatch and move the check to the narrowest profile that actually needs its evidence.

Profiles

Use these profiles unless the project has a more specific, documented contract:

ProfilePurposeTypical blocking checks
staging-smokeFast proof that the changed build starts and the critical path worksbuild, focused regression, one stable smoke/contract check
security-proofProve an adversarial or trust-boundary claimhostile tests, source/collector proof, fresh-context evaluator
release-attestationProve the exact releasable artifact and its identitysigning, Authenticode/tool identity, installer/package checks
nightly-stressFind intermittent and capacity failuresrace, stress, AV/OS matrix, long-running evals

staging-smoke must not require signing, production credentials, a release certificate, or a long VM stress run. security-proof may run on an unsigned staging build when its claim is source or runtime behavior. A release check may remain blocking for release promotion without becoming a per-edit gate.

Feedback Loop

For every overload signal, record:

  1. requested profile and change boundary;
  2. gate that blocked or dominated the run;
  3. command, elapsed time, failure count, and evidence actually produced;
  4. whether the gate was relevant, duplicated, flaky, or misplaced;
  5. the smallest profile split or deletion of duplicate coverage;
  6. a before/after run of the affected profile and a fresh review of the rule.

Use the deterministic harness-load-advisor.py signal as an intake event. It stores metadata outside the repository and forces the final report to name the mismatch. Durable policy changes belong in Git; raw session traces do not.

Required Report

Do not write "overkill" and move on. Report:

text
Harness feedback: OVERLOAD | CLEAR
Requested profile: staging-smoke | security-proof | release-attestation | nightly-stress
Mis-scoped gate: <name>
Evidence: <command, result, elapsed time, or explicit missing proof>
Correction: <profile split or rule change>
Verification: <before/after commands and result>
Residual risk: <what remains intentionally gated and where>
Show full SKILL.md (198 more words)Show less

Gotchas

  • A fresh evaluator is an independence control, not a release-signing check.
  • A VM can be a reusable execution environment without forcing release identity checks into every VM smoke.
  • A green fast gate does not prove release readiness; a red release-only gate does not invalidate a staging smoke unless the staging claim depends on it.
  • Do not replace a misplaced gate with retries, sleeps, or a bypass marker.
  • Do not infer overload from one slow run; distinguish environment failure from a profile contract error.

Troubleshooting

SymptomLikely causeAction
Staging smoke asks for signingRelease gate leaked into staging profilesplit release-attestation and run the smoke on the unsigned staging artifact
Security proof blocks on a production VMRuntime environment and release identity are coupledkeep the VM, remove release-only assertions from the security profile
Same gate fails repeatedlyWrong scope, flaky boundary, or missing fixtureclassify the failure and add a focused reproducer; never silently retry
Agent says "tests passed" with no profileEvidence contract is incompleterequire the report fields above and the exact command/result
Fix removes a safety checkCausal ownership was not tracedrestore the check, document the narrower boundary, and re-verify it there

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/development/harness-feedback of AnastasiyaW/codex-claude-code-config.

Open the folder on GitHubat commit 67709af

Compare with similar skills

Harness Feedback next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Feedback compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Feedback this skillAnastasiyaW/codex-claude-code-config154—~1kAutomated safety check: PassMIT
Deepsec Documentation Guidevercel-labs/deepsec8.1k—~956Automated safety check: PassApache-2.0
Kubernetes Network Security Auditkubeshark/kubeshark12k—~7.3kAutomated safety check: NotesApache-2.0
Native Dependency Updatemono/SkiaSharp5.6k—~4.1kAutomated safety check: PassMIT
Semgrep Security Scantrailofbits/skills7.5k—~3.7kAutomated safety check: NotesCC-BY-SA-4.0
Skillward AuditFangcun-AI/SkillWard143—~2.9kAutomated safety check: PassCustom licence

Similar skills

  • Deepsec Documentation Guide

    vercel-labs/deepsec

    Official

    Points the agent at deepsec's own docs to answer questions about initializing, configuring, resuming, scanning with and extending the vulnerability scanner.

    8.1k GitHub stars~956 tokensUpdated 12 days ago
    SecurityAuto-check passed
  • Hunts for compromised workloads and malicious traffic in a Kubernetes cluster by sweeping network data through Kubeshark MCP, mapped to MITRE ATT&CK.

    12k GitHub stars~7.3k tokensUpdated yesterday
    SecurityAuto-check: notes
  • Update native dependencies (libpng, libexpat, zlib, libwebp, harfbuzz, freetype, libjpeg-turbo, etc.) in SkiaSharp's Skia fork.

    5.6k GitHub stars~4.1k tokensUpdated yesterday
    SecurityAuto-check passed
  • Semgrep Security Scan

    trailofbits/skills

    Official

    Detects languages, proposes rulesets for approval, then runs the approved Semgrep scan across a codebase and merges the output into one SARIF file.

    7.5k GitHub stars~3.7k tokensUpdated yesterday
    SecurityAuto-check: notes
  • Skillward Audit

    Fangcun-AI/SkillWard

    Security-audit a third-party skill bundle (folder with SKILL.md, or .zip / .tar.gz archive) before installing it, using the SkillWard cloud scanner.

    143 GitHub stars~2.9k tokensUpdated 2 mo ago
    SecurityAuto-check passed
  • Security Audit

    TheDecipherist/claude-code-mastery

    Checks a codebase for hardcoded secrets, vulnerable dependencies, weak input handling, weak authentication and unsafe transport settings before deployment or merge.

    551 GitHub stars~1.3k tokensUpdated 5 mo ago
    SecurityAuto-check: notes

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Categories

Questions about Harness Feedback

What does Harness Feedback do?

A skill your agent uses when an agent says a test, VM, proof, evaluator, or release gate is overloaded, too strict, blocking staging, or causing false positives; split checks by profile, measure the…. Harness Feedback is an agent skill from AnastasiyaW/codex-claude-code-config. Use when an agent says a test, VM, proof, evaluator, or release gate is overloaded, too strict, blocking staging, or causing false positives; split checks by profile, measure the burden, preserve high-risk evidence, and verify the smallest corrected workflow.

When should I use Harness Feedback?

Harness Feedback fits situations like: an agent says a test; release gate is overloaded; blocking staging; causing false positives.

How do I install Harness Feedback in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-feedback -a claude-code`. Or copy the skill folder (skills/development/harness-feedback in AnastasiyaW/codex-claude-code-config) into .claude/skills/harness-feedback in your project. Claude Code loads it when a task matches its description.

How do I install Harness Feedback in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-feedback -a codex`. Or copy the skill folder (skills/development/harness-feedback in AnastasiyaW/codex-claude-code-config) into .agents/skills/harness-feedback in your project. Codex loads it when a task matches its description.

Can I use Harness Feedback in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-feedback -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-feedback, .gemini/skills/harness-feedback, .github/skills/harness-feedback and .opencode/skills/harness-feedback in your project.

What does Harness Feedback need to run?

SKILL.md names no scripts, command-line tools or credentials: Harness Feedback is instructions for the agent only.

Does Harness Feedback access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harness Feedback safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Harness Feedback use?

Harness Feedback is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Feedback use?

About 1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Harness Feedback?

Skills that share tags, products or a category with Harness Feedback: Deepsec Documentation Guide (vercel-labs/deepsec, 8.1k stars), Kubernetes Network Security Audit (kubeshark/kubeshark, 12k stars), Native Dependency Update (mono/SkiaSharp, 5.6k stars) and Semgrep Security Scan (trailofbits/skills, 7.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Feedback?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.