Agent skill

Harness Engineering

by Mark393295827 in Mark393295827/third-brain-v7-skills

A skill your agent uses when an agent workflow needs production-like runtime controls for context, tools, permissions, observability, scheduling, evaluation, recovery, or maintenance.

MITAuto-check passedDevOps & Cloud

Install Harness Engineering

skills CLI
$ npx skills add Mark393295827/third-brain-v7-skills --skill harness-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mark393295827/third-brain-v7-skills harness-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mark393295827/third-brain-v7-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/harness-engineering .claude/skills/harness-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-engineering
GitHub stars
141
Token cost
~2.2k tokens
SKILL.md length
975 words
Files
5 (incl. scripts, references)
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when an agent workflow needs production-like runtime controls for context, tools, permissions, observability, scheduling, evaluation, recovery, or maintenance.

  • Works in 12 steps: Pass Four-C: Context truth/retrieval,… → Compile the reviewed intent plan into a… → Map runtime: stored program, control… → …
  • An agent workflow needs production-like runtime controls for context
  • SKILL.md covers Usage Template, Workflow, Failure Protocol and Output Contract, plus 3 more sections
  • Runs Python scripts from its folder

What it does

Harness Engineering is an agent skill from Mark393295827/third-brain-v7-skills. Use when an agent workflow needs production-like runtime controls for context, tools, permissions, observability, scheduling, evaluation, recovery, or maintenance.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/runtime-control-patterns.md`, `references/runtime-envelope-example.json` and `references/runtime-envelope-plan.md`).

It sits in DevOps & Cloud, covering Observability. The repository describes itself as: agent wiki +engineering skills. The licence is MIT.

When your agent uses it

  • An agent workflow needs production-like runtime controls for context
  • Tasks that involve Observability

Example prompts

  • “/harness-engineering”

Requirements

  • Python 3

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Pass Four-C: Context truth/retrieval, Connections scoped accounts/APIs, Capabilities versioned skills/scripts/evals, Cadence…
  2. Compile the reviewed intent plan into a versioned runtime envelope. Validate
  3. Map runtime: stored program, control unit, hot context, durable disk, event bus, I/O tools, verifier, and garbage collector.
  4. Choose the lowest-context primitive: deterministic script/hook, skill,
  5. Define each tool as a narrow host-owned system call with purpose, explicit inputs, bounds, timeout, idempotency, failure path, evidence…
  6. Enforce zero trust and least privilege in the environment, not only prose
  7. Normalize each model termination_reason into complete, tool request,
  8. For delegated action require mandate, scope, limit, preview, receipt, and rollback. Human approval governs…
  9. Define allowed output types, maximum external outputs, and a verifiable
  10. Add deterministic feedback (tests, lint, LSP, policy checks) outside context when possible; add independent evaluator/red team for…
  11. Persist an append-only session event log and checkpoint; define alerts, fallback, incident response, cleanup, permission review, and…
  12. For scheduled work define Trigger, Context, Steering, Receipt, budget, stop, recovery, and executor health. A schedule firing is not task…

What it can do on your machine

Read from SKILL.md and the folder at commit 5a64514. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Engineering loads about 2.2k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 975 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Mark393295827/third-brain-v7-skills at commit 5a64514, republished under its MIT licence (© Mark393295827). 975 words, ~2,174 tokens.

Download SKILL.mdSave it as .claude/skills/harness-engineering/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
harness-engineering
description
Use when an agent workflow needs production-like runtime controls for context, tools, permissions, observability, scheduling, evaluation, recovery, or maintenance.
metadata.version
8.1.0
metadata.updated
2026-08-18
metadata.profile
high-risk
metadata.assumes
The workflow uses tools or delegated actions whose environment, permissions, and event trail can be controlled.
metadata.conflicts_with
Prompt-only safety, broad credentials, hidden tool effects, or autonomous routines without finite budgets and rollback.

Harness Engineering

<skill_contract> <input>An agent workflow, runtime environment, tools, data sensitivity, effects, cadence, risk, and operator constraints.</input> <output>An auditable runtime kernel with scoped permissions, scheduling, observability, recovery, and eval controls.</output> <done>An end-to-end trace and failure-path tests prove bounded, replayable, recoverable delegated action.</done> <non_goals>Business-task decomposition, prompt-only safety, broad credentials, or unbounded scheduled autonomy.</non_goals>

Treat the harness as the kernel around an LLM OS: context is RAM, durable state is disk, tools are system calls, skills are programs, the scheduler is control, and evals are verifiers. Load references/runtime-control-patterns.md for matrices and schemas. Start guarded automation from references/runtime-envelope-example.json and validate it with scripts/validate_runtime_envelope.py --strict.

Usage Template

Provide: workflow, users, agent roles, environment, tools/connections, data sensitivity, delegated actions, cadence, throughput/SLA, failure history, and risk tolerance.

Workflow

<intake>

Run the trace gate: the harness must be able to show what the agent saw, proposed, called, changed, and verified. Separate Intent Plan (human-reviewable source), Compiled Contract (validated runtime envelope), Agent (instructions/capabilities), Environment (network/files/credential broker), and Session (mounted context/events/state). Define one auditable control path for high-risk intent and final joins. Model output is never execution authority.

</intake>

<unknowns_gate>

If state ownership, credential scope, external side effects, retention, or approval authority is unclear, return NEEDS_INPUT. Probe tools with read-only discovery where possible; unknown side effects default to denied.

</unknowns_gate>

<execute>
  1. Pass Four-C: Context truth/retrieval, Connections scoped accounts/APIs, Capabilities versioned skills/scripts/evals, Cadence trigger/receipt/anomaly/stop.
  2. Compile the reviewed intent plan into a versioned runtime envelope. Validate plan hash, tool_execution_owner: host, filesystem/network/secret boundaries, output cardinality, legal no-op, budgets, approvals, audit paths, and rollback before execution.
  3. Map runtime: stored program, control unit, hot context, durable disk, event bus, I/O tools, verifier, and garbage collector.
  4. Choose the lowest-context primitive: deterministic script/hook, skill, static Graph, connector, dynamic workflow, or agent team. Load capabilities lazily. Graph Engineering owns dependency semantics; the harness owns the ready queue, leases, duplicate delivery, concurrency, and executor health.
  5. Define each tool as a narrow host-owned system call with purpose, explicit inputs, bounds, timeout, idempotency, failure path, evidence, and audit location. Validate model-proposed arguments before dispatch.
  6. Enforce zero trust and least privilege in the environment, not only prose: bind access to task, resource, operation, and time; use exact network allowlists and opaque secret handles; never expose raw credentials to model context. Stage and vet writes before external commit.
  7. Normalize each model termination_reason into complete, tool request, checkpoint/truncation, refusal/error, or unknown. The host decides whether to execute, continue, checkpoint, or escalate; success prose cannot override the control signal or verifier.
  8. For delegated action require mandate, scope, limit, preview, receipt, and rollback. Human approval governs irreversible/shared/financial/published/credentialed actions.
  9. Define allowed output types, maximum external outputs, and a verifiable NO_OP condition. Quiet execution is success only when eligibility was checked and no side effect occurred.
  10. Add deterministic feedback (tests, lint, LSP, policy checks) outside context when possible; add independent evaluator/red team for high-risk semantic output.
  11. Persist an append-only session event log and checkpoint; define alerts, fallback, incident response, cleanup, permission review, and stale-context/rule review.
  12. For scheduled work define Trigger, Context, Steering, Receipt, budget, stop, recovery, and executor health. A schedule firing is not task success.
  13. For Graph execution, persist node/edge/join transitions before releasing successors, make delivery idempotent, recover from the last verified checkpoint, and test permission denial, worker loss, duplicate events, and compensation without relying on in-memory scheduler state.
</execute>
<evaluate>

Threat-model normal, denied, timeout, partial-write, stale-context, duplicate-trigger, compromised-input, evaluator-disagreement, and rollback paths. Verify permissions with actual environment boundaries and run an independent end-to-end trace. Adoption fails if operators cannot inspect or recover the system.

</evaluate>

<retry_policy>

max_attempts: 3 per tool/failure class. Retry only idempotent or compensated actions after changing diagnosis/strategy. Use exponential delay for transient dependencies. Stop on repeated signature, permission denial, ambiguous side effect, or NO_PROGRESS.

</retry_policy>

<state_contract>

Persist {run_id, status, attempt, budget, evidence, unknowns, last_error, next_action} plus intent-plan hash, compiled-contract hash, agent/environment/session versions, context manifest, normalized termination reason, tool/permission/credential-lease registry, output quota, no-op receipt, event offsets, approvals, evals, alerts, recovery point, and rollback receipts. State transitions are auditable and replayable.

</state_contract>

Show full SKILL.md (304 more words)Show less

Failure Protocol

  • NEEDS_INPUT: ownership, effect, retention, or approval authority is ambiguous.
  • BLOCKED_PERMISSION: deny the call and continue with a safer read-only path when useful.
  • BLOCKED_DEPENDENCY: checkpoint, back off, and expose executor health.
  • VERIFY_FAILED: trace, eval, guardrail, or rollback test fails; block autonomy escalation.
  • NO_PROGRESS: changed attempts repeat the signature. max_attempts: 3.
  • BUDGET_STOP: stop scheduler/workers, checkpoint, and emit a recovery receipt.

Output Contract

Return status, result (runtime architecture and controls), evidence (trace/eval/failure tests), unknowns, and next_action including approval or rollback.

Edge Cases

  • A connector exposes broad account access for a narrow task: create a scoped proxy/allowlist or keep the workflow manual; instructions alone are insufficient.
  • A model emits plausible success text with a tool request: validate and execute the tool in host code, then verify its result; do not treat the text as task completion.
  • A maintenance run finds no eligible change: emit a verified NO_OP receipt and zero external outputs rather than creating notification or PR noise.
  • The scheduler ran on time but the state store was stale: block mutation, mark the run failed, and recover from the last verified checkpoint.
  • A Graph worker completes twice after lease expiry: deduplicate by node/run identity and release successors only from one verified transition.

Success Metrics

  • Every delegated action is bounded, observable, independently verifiable, and recoverable.
  • Runtime state can be replayed across sessions without relying on chat history.
  • Autonomy level follows measured reliability and operator review capacity.

Quality Gates

  • Agent, Environment, Session, and serial control ownership are explicit.
  • Intent plan and compiled runtime envelope are hash-bound and pass strict validation.
  • Tool effects are enforced by real permissions and contracts.
  • Termination routing, secret isolation, output limits, and legal no-op are host-enforced.
  • Independent eval, approval, manual fallback, and rollback match risk.
  • Schedules have executor receipts, hard budgets, stop, and recovery.
  • Maintenance includes cleanup, permission audit, and rule half-life.

</skill_contract>

© Mark393295827, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/harness-engineering of Mark393295827/third-brain-v7-skills.

  • SKILL.md
  • references/runtime-control-patterns.md
  • references/runtime-envelope-example.json
  • references/runtime-envelope-plan.md
  • scripts/validate_runtime_envelope.py

Open the folder on GitHubat commit 5a64514

Compare with similar skills

Harness Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Engineering this skillMark393295827/third-brain-v7-skills141—~2.2kAutomated safety check: PassMIT
Vercel Optimize Auditvercel-labs/agent-skills32k9 repos~4.3kAutomated safety check: PassNone
Kubeshark Installerkubeshark/kubeshark12k—~3.6kAutomated safety check: NotesApache-2.0
Kubeshark KFL2 Filter Referencekubeshark/kubeshark12k—~3.6kAutomated safety check: PassApache-2.0
KubeSphere ServiceMesh Managerkubesphere/kubesphere17k—~2.4kAutomated safety check: PassCustom licence
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0

Similar skills

  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 9 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Syntax reference for KFL2, the CEL-based display filter language used to search Kubernetes network traffic captured by Kubeshark, loaded before any filter is written.

    12k GitHub stars~3.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • KubeSphere ServiceMesh Manager

    kubesphere/kubesphere

    Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.

    17k GitHub stars~2.4k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Caveman Gateway Setup

    JuliusBrussee/caveman

    Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

    110k GitHub starsUsed in 1 repo~2.6k tokens
    DevOps & CloudAuto-check: warnings

More from Mark393295827/third-brain-v7-skills

All 21 skills in this repo
  • Loop Engineering

    Mark393295827/third-brain-v7-skills

    A skill your agent uses when a repeatable task must become a bounded Trigger - Execute - Verify - State loop, scheduled automation, goal agent, or metric-driven research cycle.

    141 GitHub stars~1.9k tokensUpdated 19 days ago
    Auto-check passed
  • Graph Engineering

    Mark393295827/third-brain-v7-skills

    A skill your agent uses when a workflow has explicit data dependencies, independently executable branches, typed joins, or node-local recovery needs that justify a bounded static dependency graph.

    141 GitHub stars~1.9k tokensUpdated 19 days ago
    Auto-check passed
  • Agent Teams Command

    Mark393295827/third-brain-v7-skills

    A skill your agent uses when work has genuinely independent streams or distinct builder, evaluator, domain, and integration roles that require bounded multi-agent command scaled from 5 to 100+ agents.

    141 GitHub stars~1.7k tokensUpdated 19 days ago
    Auto-check passed
  • AI Six Sigma Property Os

    Mark393295827/third-brain-v7-skills

    A skill your agent uses when property-service operations need an AI plus ontology plus DMAIC design for work orders, dispatch, quotes, evidence, CTQ metrics, and control dashboards.

    141 GitHub stars~1.4k tokensUpdated 19 days ago
    Auto-check passed
  • Anthropic Os

    Mark393295827/third-brain-v7-skills

    A skill your agent uses when a personal or team operating system needs a bounded redesign using Four-C, closed-loop controls, 70/30 allocation, 3B creativity, experiments, and prediction-error…

    141 GitHub stars~1.6k tokensUpdated 19 days ago
    Auto-check passed
  • Deep Research

    Mark393295827/third-brain-v7-skills

    A skill your agent uses when a decision-relevant question needs multi-source search, claim-level citations, contradiction handling, uncertainty, or a durable wiki handoff.

    141 GitHub stars~1.5k tokensUpdated 19 days ago
    Auto-check passed

Categories

Questions about Harness Engineering

What does Harness Engineering do?

A skill your agent uses when an agent workflow needs production-like runtime controls for context, tools, permissions, observability, scheduling, evaluation, recovery, or maintenance. Harness Engineering is an agent skill from Mark393295827/third-brain-v7-skills. Use when an agent workflow needs production-like runtime controls for context, tools, permissions, observability, scheduling, evaluation, recovery, or maintenance.

When should I use Harness Engineering?

Harness Engineering fits situations like: an agent workflow needs production-like runtime controls for context; tasks that involve Observability.

How do I install Harness Engineering in Claude Code?

Run `npx skills add Mark393295827/third-brain-v7-skills --skill harness-engineering -a claude-code`. Or copy the skill folder (skills/harness-engineering in Mark393295827/third-brain-v7-skills) into .claude/skills/harness-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Harness Engineering in Codex?

Run `npx skills add Mark393295827/third-brain-v7-skills --skill harness-engineering -a codex`. Or copy the skill folder (skills/harness-engineering in Mark393295827/third-brain-v7-skills) into .agents/skills/harness-engineering in your project. Codex loads it when a task matches its description.

Can I use Harness Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mark393295827/third-brain-v7-skills --skill harness-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-engineering, .gemini/skills/harness-engineering, .github/skills/harness-engineering and .opencode/skills/harness-engineering in your project.

What does Harness Engineering need to run?

Going by SKILL.md and its folder, Harness Engineering needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Harness Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harness Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Harness Engineering use?

Harness Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Engineering use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Harness Engineering?

Skills that share tags, products or a category with Harness Engineering: Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars), Kubeshark Installer (kubeshark/kubeshark, 12k stars), Kubeshark KFL2 Filter Reference (kubeshark/kubeshark, 12k stars) and KubeSphere ServiceMesh Manager (kubesphere/kubesphere, 17k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Engineering?

Mark393295827 (a GitHub user) maintains it in Mark393295827/third-brain-v7-skills, which has 141 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on September 19, 2026.

Source: Mark393295827/third-brain-v7-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.