Agent skill

Agent Harness Fault Injection

by sickn33 in sickn33/agentic-awesome-skills

A skill your agent uses when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures.

MITAuto-check passedAgent Workflows

Install Agent Harness Fault Injection

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill agent-harness-fault-injection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills agent-harness-fault-injection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-harness-fault-injection .claude/skills/agent-harness-fault-injection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-harness-fault-injection
GitHub stars
47k
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
1,291 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures.

  • Works in 4 steps: Freeze the workflow revision,… → Run in a disposable sandbox with… → Make every injected failure an in-memory… → …
  • An agent workflow needs deterministic recovery evidence for sandbox
  • SKILL.md covers Overview, When to Use This Skill, Safety and Boundary… and Recovery Contract, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Harness Fault Injection is an agent skill from sickn33/agentic-awesome-skills. Use when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Chaos engineering and MCP servers. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • An agent workflow needs deterministic recovery evidence for sandbox
  • Orchestration failures

Example prompts

  • “/agent-harness-fault-injection”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Freeze the workflow revision, model/prompt configuration, tool schemas, seed,
  2. Run in a disposable sandbox with synthetic inputs and stubbed tools. Keep
  3. Make every injected failure an in-memory or fixture-controlled event. Never
  4. Record the test scope and a run identifier before starting. A missing scope,

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Harness Fault Injection loads about 2.9k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 1,291 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 1,291 words, ~2,880 tokens.

Download SKILL.mdSave it as .claude/skills/agent-harness-fault-injection/SKILL.md (or your agent's skills folder).
name
agent-harness-fault-injection
description
Use when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures.
category
development
risk
safe
source
self
source_type
self
date_added
2026-08-19
author
Whxuan0701
tags
agent-harness, fault-injection, recovery, state-machine, mcp, multi-agent
tools
claude, cursor, gemini, codex-cli

Agent Harness Fault Injection

Overview

Use a deterministic, non-production fault schedule to test whether an agent workflow preserves state, budgets, safety boundaries, and evidence when a dependency fails. The output is a small fault matrix, an event timeline, and a verdict that distinguishes recovered, contained, unrecoverable, and inconclusive runs.

When to Use This Skill

  • Use when a multi-step agent, state machine, loop, or multi-agent workflow has a new recovery path.
  • Use when sandbox execution, an MCP/tool call, a worker, a checkpoint store, or memory can time out or disappear.
  • Use before claiming retry, resume, deadline, isolation, or partial-failure behavior is production-ready.
  • Use when a regression needs reproducible failure evidence instead of a random chaos run.

Do not use this skill against a production target, real user data, live credentials, or an unbounded external service. Convert those cases to a local simulator or an authorized staging harness first.

Safety and Boundary Preconditions

  1. Freeze the workflow revision, model/prompt configuration, tool schemas, seed, input fixture, timeout, retry budget, deadline, and expected terminal states.
  2. Run in a disposable sandbox with synthetic inputs and stubbed tools. Keep network disabled unless the test explicitly needs a local test server.
  3. Make every injected failure an in-memory or fixture-controlled event. Never delete real data, revoke real credentials, kill an unrelated process, or mutate a live service to create a failure.
  4. Record the test scope and a run identifier before starting. A missing scope, fixture, or recovery contract makes the verdict inconclusive.

Recovery Contract

Write the invariant before injecting a fault. A useful contract names the state that must survive and the side effects that must not repeat:

text
After recovery, resume from the latest durable checkpoint, preserve the task
identity and safety policy, spend no more than the remaining retry/deadline
budget, and commit each externally visible effect at most once.

Model the workflow with explicit states. For example:

text
created -> running -> checkpointed -> waiting_for_tool
                       |                |
                       v                v
                    failed <--------- recovering -> resumed -> completed

For each transition, define the owner, durable fields, allowed retry count, and terminal behavior. In-memory values are not checkpoints unless the harness proves they survive the simulated restart.

Fault Matrix

Select the smallest set of faults that covers the new recovery logic. Do not randomize the schedule until a deterministic schedule has passed.

FaultInjection boundaryRequired observationExpected containment
sandbox denialbefore a tool startsno unsafe side effect; reason is retainedretry only when policy allows
MCP/tool timeoutafter request id is assignedtimeout is attributed to that requestbounded retry with same idempotency key
worker restartafter checkpoint writeworker reloads the same task versionresume from latest checkpoint
missing/stale checkpointbefore resumestale data is rejected or markedstop safely; never invent progress
parallel branch failureone branch after fan-outsibling status is preservedjoin policy decides retry, degrade, or stop
memory lossclear ephemeral contextdurable facts are reconstructedask or stop when required facts are absent
retry/deadline exhaustionon the final attemptno extra call is scheduledterminal failed or timed_out

Deterministic Injection Schedule

Use event numbers rather than wall-clock randomness. A schedule should be portable across harnesses:

json
{
  "seed": "harness-fixture-07",
  "faults": [
    {"event": "tool.call", "ordinal": 2, "kind": "timeout", "tool": "search"},
    {"event": "worker.start", "ordinal": 2, "kind": "restart"},
    {"event": "branch.join", "ordinal": 1, "kind": "partial_failure", "branch": "summarize"}
  ]
}

The harness should emit the schedule, not merely the seed. Keep fault identity separate from the observed error so a wrapper cannot accidentally turn a timeout into a generic failure. Run the same schedule twice and compare the normalized timeline before trying a different schedule.

Recovery Rules by Boundary

Sandbox and MCP/tool failures
  • Assign a request id and idempotency key before the call.
  • Distinguish timeout, explicit tool error, invalid output, and policy denial.
  • Retry only the declared retryable classes; preserve the original error and attempt count in the evidence.
  • Do not retry a side effect unless the tool contract says the key is safe to replay. A read timeout is not proof that a write did not happen.
  • When the deadline or retry budget is exhausted, emit one terminal event and stop scheduling work.
Worker restart and checkpoints
  • Persist task id, workflow version, state name, completed effects, remaining budgets, and the checkpoint sequence before a restart test.
  • Reload the newest valid checkpoint and reject a future-version or corrupted checkpoint instead of guessing.
  • Verify that resumption does not replay a committed effect. If exactly-once cannot be proven, downgrade the verdict and require reconciliation.
Parallel branches

Represent each branch as its own child attempt. The join record must retain success, failure, timeout, and not-started states. Choose one predeclared join policy:

  • all_required: any required branch failure stops the join;
  • best_effort: continue with an explicit degraded marker;
  • compensate: run a bounded compensating action and then stop or resume.

Never let a successful sibling erase a failed branch from the final ledger.

Memory loss

Clear only the ephemeral context named in the schedule. Rebuild from the checkpoint and durable evidence, then check that the agent does not fabricate missing user intent, tool output, or approval. If a required fact is absent, the safe result is inconclusive or a human clarification state.

Show full SKILL.md (510 more words)Show less

Budgets and Terminal Verdicts

Track remaining attempts and remaining time after every event. Do not reset a budget on a worker restart or branch retry. Use these verdicts:

VerdictMeaning
recoveredThe declared invariant held and the workflow completed within budget.
contained_failureThe fault was isolated and the workflow stopped safely as designed.
unrecoverableRecovery violated an invariant, repeated a side effect, crossed a boundary, or exceeded budget.
inconclusiveThe fixture, checkpoint, contract, or evidence was insufficient to judge.

contained_failure is not autonomous success. Report it separately from completed work and include the terminal reason.

Evidence Output

Produce one machine-readable record and one concise human summary. Every event should include run_id, monotonic seq, logical time, state_before, state_after, actor, event, fault_id (when injected), attempt, checkpoint_seq, retry_remaining, deadline_remaining_ms, and a redacted evidence_ref.

json
{
  "run_id": "fi-2026-08-19-07",
  "verdict": "recovered",
  "invariants": {"resume_from_checkpoint": "pass", "effect_at_most_once": "pass", "budget": "pass"},
  "faults": [{"id": "f1", "kind": "tool_timeout", "at": "tool.call#2", "handled": true}],
  "timeline": [
    {"seq": 4, "event": "checkpoint.write", "checkpoint_seq": 3},
    {"seq": 5, "event": "tool.timeout", "fault_id": "f1", "retry_remaining": 1},
    {"seq": 8, "event": "workflow.completed", "checkpoint_seq": 4}
  ],
  "limitations": ["Tool output was synthetic; no deployed MCP was exercised."]
}

The human summary should state the frozen contract, injected schedule, verdict, failed invariants, budget consumption, and the narrowest next verification. Redact prompts, tokens, private records, and tool payloads; stable references are enough for replay.

Example: Local Harness Run

text
Fixture: checkout planner / seed harness-fixture-07
Schedule: search timeout on call 2; worker restart after checkpoint 3
Policy: one retry, 2s deadline, all_required branch join

Result: recovered
Proof: checkpoint 3 reloaded, search request key replayed once, no duplicate
commit, deadline remaining 640ms, final ledger contains both branch outcomes.

Best Practices

  • Freeze inputs and schedules so a failure can be replayed from the evidence.
  • Test one boundary at a time, then add a combined schedule for interaction risk.
  • Assert invariants after every recovery transition, not only at final output.
  • Keep attempt-level faults and task-level outcomes in separate ledgers.
  • Treat missing evidence as inconclusive, never as a passing recovery.

Limitations

  • A local stub cannot prove behavior of a deployed model, MCP server, scheduler, filesystem, or network.
  • Deterministic schedules cover named paths; they do not estimate random-fault frequency or discover unknown failure modes.
  • At-most-once effects require an idempotent, observable contract; a timeline alone cannot prove an external write was not duplicated.
  • This skill does not select production SLOs, repair broken workflows, or grant permission to test systems outside the declared sandbox.

Security & Safety Notes

  • Keep tests local-only and read-only by default; use synthetic fixtures and fake credentials that cannot access a real account.
  • Require explicit authorization and a disposable staging boundary before any test that could contact a non-local service.
  • Do not include destructive commands, exploit payloads, credential material, or automatic cleanup of user data in a harness or report.
  • Redact secrets and personal data before storing timelines or attaching them to a pull request.

Common Pitfalls

  • Problem: A retry clears the original timeout and hides the fault. Solution: Keep fault id, original class, attempt, and retry lineage in the ledger.
  • Problem: A restart passes because the test reused in-memory state. Solution: Serialize, clear, and reload only the declared checkpoint fields.
  • Problem: A partial fan-out is reported as success. Solution: Preserve every branch state and apply the predeclared join policy.
  • Problem: A missing checkpoint is replaced with guessed progress. Solution: Stop safely and return inconclusive or unrecoverable with evidence.
  • Problem: A green final answer hides a deadline or duplicate-effect violation. Solution: Gate the verdict on invariants and remaining budget, not output text alone.
  • @agent-evaluation-reporting - Report autonomous, assisted, failed, timed-out, and invalid outcomes.
  • @cross-platform-contract-propagation-audit - Trace recovery fields and status contracts across consumers.
  • @multi-agent-patterns - Choose a multi-agent topology before testing its failure behavior.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agent-harness-fault-injection of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Agent Harness Fault Injection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Harness Fault Injection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Harness Fault Injection this skillsickn33/agentic-awesome-skills47k1 repos~2.9kAutomated safety check: PassMIT
Chaos Experimentharness/harness-skills115—~1.6kAutomated safety check: PassApache-2.0
App AuthoringOtoDock/oto-dock190—~23kAutomated safety check: PassCustom licence
Claude Docs Consultantcentminmod/my-claude-code-setup2.7k—~959Automated safety check: PassMIT
K8s Agent Sandbox MCPkubernetes-sigs/agent-sandbox4.2k—~1.3kAutomated safety check: PassApache-2.0
Unraiddinglebear-ai/unraid135—~5.4kAutomated safety check: NotesMIT

Similar skills

  • Chaos Experiment

    harness/harness-skills

    A skill your agent uses when the user asks to create, edit, update, design, or configure a Harness Chaos Experiment — including faults, probes, actions, experiment YAML, fault injection, pod-delete…

    115 GitHub stars~1.6k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • App Authoring

    OtoDock/oto-dock

    Authoring contracts for interactive artifacts, app dashboards, and Dock panels — displayui backchannel and theme events, pinapp action buttons (firetask / sendprompt / mcptool / platform, per-action…

    190 GitHub stars~23k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Claude Docs Consultant

    centminmod/my-claude-code-setup

    Consult official Claude Code documentation from code.claude.com using selective fetching.

    2.7k GitHub stars~959 tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • K8s Agent Sandbox MCP

    kubernetes-sigs/agent-sandbox

    Official

    An MCP server skill for managing Kubernetes sandboxes. An agent skill from kubernetes-sigs/agent-sandbox.

    4.2k GitHub stars~1.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Unraid

    dinglebear-ai/unraid

    This skill should be used when the user mentions Unraid, asks to check server health, monitor array or disk status, list or restart Docker containers, start or stop VMs, read system logs, check…

    135 GitHub stars~5.4k tokensUpdated 7 days ago
    DevOps & CloudAuto-check: notes
  • Devsy

    devsy-org/devsy

    Operate Devsy workspaces and providers for end users. An agent skill from devsy-org/devsy.

    113 GitHub stars~1.7k tokensUpdated today
    DevOps & CloudAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Agent Harness Fault Injection

What does Agent Harness Fault Injection do?

A skill your agent uses when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures. Agent Harness Fault Injection is an agent skill from sickn33/agentic-awesome-skills. Use when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures.

When should I use Agent Harness Fault Injection?

Agent Harness Fault Injection fits situations like: an agent workflow needs deterministic recovery evidence for sandbox; orchestration failures.

How do I install Agent Harness Fault Injection in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill agent-harness-fault-injection -a claude-code`. Or copy the skill folder (skills/agent-harness-fault-injection in sickn33/agentic-awesome-skills) into .claude/skills/agent-harness-fault-injection in your project. Claude Code loads it when a task matches its description.

How do I install Agent Harness Fault Injection in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill agent-harness-fault-injection -a codex`. Or copy the skill folder (skills/agent-harness-fault-injection in sickn33/agentic-awesome-skills) into .agents/skills/agent-harness-fault-injection in your project. Codex loads it when a task matches its description.

Can I use Agent Harness Fault Injection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill agent-harness-fault-injection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-harness-fault-injection, .gemini/skills/agent-harness-fault-injection, .github/skills/agent-harness-fault-injection and .opencode/skills/agent-harness-fault-injection in your project.

What does Agent Harness Fault Injection need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Harness Fault Injection is instructions for the agent only.

Does Agent Harness Fault Injection access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Harness Fault Injection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Harness Fault Injection use?

Agent Harness Fault Injection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Harness Fault Injection use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Harness Fault Injection?

Skills that share tags, products or a category with Agent Harness Fault Injection: Chaos Experiment (harness/harness-skills, 115 stars), App Authoring (OtoDock/oto-dock, 190 stars), Claude Docs Consultant (centminmod/my-claude-code-setup, 2.7k stars) and K8s Agent Sandbox MCP (kubernetes-sigs/agent-sandbox, 4.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Harness Fault Injection?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.