Agent skill

Agentic Harness Design and Review

by NateBJones-Projects in NateBJones-Projects/OB1

Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.

Custom licenceAuto-check passedAgent Workflows

Install Agentic Harness Design and Review

skills CLI
$ npx skills add NateBJones-Projects/OB1 --skill n-agentic-harnesses -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NateBJones-Projects/OB1 n-agentic-harnesses --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NateBJones-Projects/OB1.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/n-agentic-harnesses .claude/skills/n-agentic-harnesses && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
n-agentic-harnesses
GitHub stars
4.7k
Token cost
~1.8k tokens
SKILL.md length
790 words
Files
16 (incl. references)
Skills in repo
24
Repo updated
First seen
Licence
Custom licence

At a glance

Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.

  • Works in 4 steps: Gather Context → Classify The Request → Classify The Product Shape → …
  • Designing the tool-use and permission layer for a new agent
  • SKILL.md covers Problem, Trigger Conditions, Default Posture and Step 0: Gather Context, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill treats most AI product failures as harness problems rather than model weakness: unclear tool boundaries, missing approval policy, brittle state, sloppy context assembly, no evaluation loop and little operator visibility. It triggers on design or rebuild requests and on symptoms such as tools firing without permission, sessions dying on a crash, stale context and unexplained costs.

The agent first gathers context, or reads the real codebase, configs and hooks when evaluating, then picks a design or evaluation mode and loads matching files from a numbered reference set covering principles, architectures, tool permissions, state and durability, memory and evaluation, extensibility, UX and observability, plus build and improvement playbooks. Defaults favor lean, solo-maintainable, single-agent designs, an evaluation plan even for new builds, and advice turned into phases, success criteria and failure tests.

When your agent uses it

  • Designing the tool-use and permission layer for a new agent
  • Reviewing an existing agent for missing approval gates or durability gaps
  • Planning memory, state and resumability for long-running sessions
  • Building an evaluation and observability plan for an AI product

Example prompts

  • “Design the harness for a support copilot that can issue refunds, with approval gates and an eval plan.”
  • “Evaluate our agent's codebase for session persistence and retry problems.”
  • “Our agent costs keep climbing and nobody can see why, so find the harness gaps.”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Gather Context
  2. Classify The Request
  3. Classify The Product Shape
  4. Read The Smallest Useful Reference Set

What it can do on your machine

Read from SKILL.md and the folder at commit 238df6c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentic Harness Design and Review loads about 1.8k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 144 tokens; SKILL.md has 790 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 790 words (~1,833 tokens).

“Most AI products do not break because the model is too weak. They break at the harness layer: unclear tool boundaries, missing approval policy, brittle state, sloppy context assembly, no evaluation loop, and weak operator visibility. This skill turns those…”

— opening of SKILL.md by NateBJones-Projects, Custom licence
name
n-agentic-harnesses
author
Jonathan Edwards
version
1.0.0

Read the full SKILL.md on GitHub

Files

SKILL.md and 15 other files (references) in skills/n-agentic-harnesses of NateBJones-Projects/OB1.

  • SKILL.md
  • README.md
  • agents/openai.yaml
  • metadata.json
  • references/01-principles-and-solo-dev-defaults.md
  • references/02-harness-shapes-and-architecture.md
  • references/03-tools-execution-and-permissions.md
  • references/04-state-sessions-and-durability.md
  • references/05-context-memory-and-evaluation.md
  • references/06-agents-and-extensibility.md
  • references/07-ux-observability-and-operations.md
  • references/08-design-and-build-playbook.md
  • references/09-evaluation-and-improvement-playbook.md
  • references/10-example-requests-and-output-patterns.md
  • references/11-codex-translation-notes.md
  • variants

Open the folder on GitHubat commit 238df6c

Compare with similar skills

Agentic Harness Design and Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentic Harness Design and Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentic Harness Design and Review this skillNateBJones-Projects/OB14.7k—~1.8kAutomated safety check: PassCustom licence
Deep Agents Corelangchain-ai/langchain-skills1.3k—~3.1kAutomated safety check: PassMIT
Dive Into LangGraphluochang212/dive-into-langgraph457—~837Automated safety check: NotesCustom licence
Create Agentvectorize-io/hindsight48k—~1.1kAutomated safety check: PassMIT
Neo4j Agent Memory Skillneo4j-contrib/neo4j-skills114—~5.8kAutomated safety check: PassMIT
Agent Developmentsundial-org/awesome-openclaw-skills663—~2.4kAutomated safety check: PassMIT

Similar skills

  • Deep Agents Core

    langchain-ai/langchain-skills

    Official

    Explains how to build agents with the Deep Agents framework: create_deep_agent, the built-in middleware, the harness, SKILL.md format and configuration options.

    1.3k GitHub stars~3.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Dive Into LangGraph

    luochang212/dive-into-langgraph

    A Chinese-language guide and reference for building agents with LangGraph 1.0, from a first ReAct agent through middleware, memory, MCP, RAG and web search.

    457 GitHub stars~837 tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • Create Agent

    vectorize-io/hindsight

    Create a new Hindsight-powered subagent with long-term memory.

    48k GitHub stars~1.1k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Neo4j Agent Memory Skill

    neo4j-contrib/neo4j-skills

    Authoritative reference for the neo4j-agent-memory Python package — a graph-native memory system for AI agents built on Neo4j — and for the hosted service (NAMS) at memory.neo4jlabs.com.

    114 GitHub stars~5.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Agent Development

    sundial-org/awesome-openclaw-skills

    Design and build custom Claude Code agents with effective descriptions, tool access patterns, and self-documenting prompts.

    663 GitHub stars~2.4k tokensUpdated 7 mo ago
    Agent WorkflowsAuto-check passed
  • Adds persistent memory to AI apps with the Mem0 Python and TypeScript SDKs: store, search, update and delete user memories, with framework integrations.

    67k GitHub starsUsed in 1 repo~2.2k tokens
    AI & LLM EngineeringAuto-check passed

More from NateBJones-Projects/OB1

All 24 skills in this repo
  • Heavy File Ingestion

    NateBJones-Projects/OB1

    Converts large PDF, DOCX, PPTX, XLSX and CSV files into markdown or CSV plus an index before the agent reads them, so tokens go to the compressed copy.

    4.7k GitHub stars~995 tokensUpdated yesterday
    Auto-check passed
  • Aiception Skill Extraction

    NateBJones-Projects/OB1

    Pulls reusable knowledge out of work sessions and turns it into new skills, checking existing notes and skills first to avoid duplicates.

    4.7k GitHub stars~2k tokensUpdated yesterday
    Auto-check: notes
  • Claude Code Auto-Capture Hook

    NateBJones-Projects/OB1

    Fires a Stop hook that captures a Claude Code session transcript automatically when the session ends without a verbal wrap-up.

    4.7k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check: notes
  • Open Brain Local HTTP

    NateBJones-Projects/OB1

    Captures and searches personal notes in a self-hosted Open Brain using plain curl calls over HTTP, for setups where MCP is disabled or blocked.

    4.7k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check: notes
  • Panning for Gold

    NateBJones-Projects/OB1

    Processes voice transcripts and brain dumps in three phases, extracting every idea thread, evaluating the strongest and saving a permanent inventory and synthesis.

    4.7k GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Weekly Signal Diff

    NateBJones-Projects/OB1

    Turns a noisy week of AI or software market news into a short list of structural changes, weighted by what the user already tracks in Open Brain memory.

    4.7k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed

Questions about Agentic Harness Design and Review

What does Agentic Harness Design and Review do?

Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans. The skill treats most AI product failures as harness problems rather than model weakness: unclear tool boundaries, missing approval policy, brittle state, sloppy context assembly, no evaluation loop and little operator visibility. It triggers on design or rebuild requests and on symptoms such as tools firing without permission, sessions dying on a crash, stale context and unexplained costs.

When should I use Agentic Harness Design and Review?

Agentic Harness Design and Review fits situations like: designing the tool-use and permission layer for a new agent; reviewing an existing agent for missing approval gates or durability gaps; planning memory, state and resumability for long-running sessions; building an evaluation and observability plan for an AI product.

How do I install Agentic Harness Design and Review in Claude Code?

Run `npx skills add NateBJones-Projects/OB1 --skill n-agentic-harnesses -a claude-code`. Or copy the skill folder (skills/n-agentic-harnesses in NateBJones-Projects/OB1) into .claude/skills/n-agentic-harnesses in your project. Claude Code loads it when a task matches its description.

How do I install Agentic Harness Design and Review in Codex?

Run `npx skills add NateBJones-Projects/OB1 --skill n-agentic-harnesses -a codex`. Or copy the skill folder (skills/n-agentic-harnesses in NateBJones-Projects/OB1) into .agents/skills/n-agentic-harnesses in your project. Codex loads it when a task matches its description.

Can I use Agentic Harness Design and Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NateBJones-Projects/OB1 --skill n-agentic-harnesses -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/n-agentic-harnesses, .gemini/skills/n-agentic-harnesses, .github/skills/n-agentic-harnesses and .opencode/skills/n-agentic-harnesses in your project.

What does Agentic Harness Design and Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentic Harness Design and Review is instructions for the agent only.

Does Agentic Harness Design and Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentic Harness Design and Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentic Harness Design and Review use?

Agentic Harness Design and Review has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Agentic Harness Design and Review use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9k tokens, read only when the agent opens those files.

What are the alternatives to Agentic Harness Design and Review?

Skills that share tags, products or a category with Agentic Harness Design and Review: Deep Agents Core (langchain-ai/langchain-skills, 1.3k stars), Dive Into LangGraph (luochang212/dive-into-langgraph, 457 stars), Create Agent (vectorize-io/hindsight, 48k stars) and Neo4j Agent Memory Skill (neo4j-contrib/neo4j-skills, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentic Harness Design and Review?

NateBJones-Projects (a GitHub organization) maintains it in NateBJones-Projects/OB1, which has 4,710 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 9, 2026.

Source: NateBJones-Projects/OB1 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.