Agent skill

Error Handling

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when designing the reaction to a class of failures — typed error taxonomies, retry/backoff/timeout policy, circuit breakers, React/Next error boundaries, and the user-message…

MITAuto-check passedBackend & APIs

Install Error Handling

skills CLI
$ npx skills add ericrisco/rsc-harness --skill error-handling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness error-handling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/error-handling .claude/skills/error-handling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
error-handling
GitHub stars
156
Token cost
~3.2k tokens
SKILL.md length
1,277 words
Files
6 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing the reaction to a class of failures — typed error taxonomies, retry/backoff/timeout policy, circuit breakers, React/Next error boundaries, and the user-message…

  • Works in 5 steps: Model failure as a taxonomy → Decide retryability → Contain blast radius → …
  • Designing the reaction to a class of failures — typed error taxonomies
  • SKILL.md covers Step 1 — Model failure as a…, Step 2 — Decide retryability, Step 3 — Contain blast radius and Step 4 — Boundaries, plus 3 more sections
  • Runs Shell scripts from its folder

What it does

Error Handling is an agent skill from ericrisco/rsc-harness. Use when designing the reaction to a class of failures — typed error taxonomies, retry/backoff/timeout policy, circuit breakers, React/Next error boundaries, and the user-message vs operator-log split. NOT diagnosing one specific crash (that is debug), NOT logs/metrics/traces (that is observability), NOT the wire error envelope (that is api-design).

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/boundaries-and-messaging.md`).

It sits in Backend & APIs, covering Error handling, API design and Observability. It works with React. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Designing the reaction to a class of failures — typed error taxonomies
  • Retry/backoff/timeout policy
  • Circuit breakers
  • React/Next error boundaries

Example prompts

  • “/error-handling”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Model failure as a taxonomy
  2. Decide retryability
  3. Contain blast radius
  4. Boundaries
  5. Surface it

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Error Handling loads about 3.2k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 92 tokens; SKILL.md has 1,277 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,277 words, ~3,169 tokens.

Download SKILL.mdSave it as .claude/skills/error-handling/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
error-handling
description
Use when designing the reaction to a class of failures — typed error taxonomies, retry/backoff/timeout policy, circuit breakers, React/Next error boundaries, and the user-message vs operator-log split. NOT diagnosing one specific crash (that is debug), NOT logs/metrics/traces (that is observability), NOT the wire error envelope (that is api-design).
tags
error-handling, resilience, retry, circuit-breaker, error-boundary, typed-errors, fault-tolerance
recommends
debug, observability, monitoring, api-design, secure-coding, testing-web, nextjs
origin
risco

Error handling — classify, contain, surface

You are designing what happens whenever anything in a class breaks, not chasing one crash (that is debug). Every failure gets classified, contained, and surfaced — never swallowed. The deliverable, in that order: a typed error taxonomy, a retry policy with caps and jitter, boundary placement, and a two-audience message contract — never a pile of try { … } catch {}.

Step 1 — Model failure as a taxonomy

Bucket every failure into one of three kinds. The bucket dictates the reaction; get the bucket wrong and every downstream decision is wrong too.

BucketExamplesRetry?Tell the userTell the operator
Domain / expectedinsufficient funds, slot taken, validation failedNoYes, actionableinfo — it is normal
Infrastructure / transienttimeout, 503, connection reset, 429Yes (capped)"temporary, retrying"warn — watch the rate
Programmer error / bugnull deref, bad assertion, type errorNogeneric "something broke" + iderror — page if frequent
Result vs throw

Decide per call site, not per codebase:

  • Result<T, E> for expected domain failures the caller must handle. The type checker forces a branch — the failure cannot be ignored by accident.
  • throw for exceptional / programmer errors. These should crash up to the nearest boundary, not be threaded through every signature.

In TypeScript, neverthrow (current) is the instrument for the Result path:

ts
import { ok, err, Result } from "neverthrow";

type ChargeError = "insufficient_funds" | "card_declined";

// Expected domain failure → Result. The caller MUST handle both arms.
function charge(cents: number, balance: number): Result<number, ChargeError> {
  if (cents > balance) return err("insufficient_funds");
  return ok(balance - cents);
}

const r = charge(500, 200);
if (r.isErr()) {
  // r.error is the typed union — exhaustive, no `any`.
}
Stable codes and cause chaining

Every error carries a stable code (a string the UI and logs key off, never the human message) and never drops the original cause.

ts
// BAD — string error, loses the original, nothing to branch on.
throw new Error("payment failed");

// GOOD — typed class, stable code, cause preserved.
class PaymentError extends Error {
  constructor(public code: "provider_down" | "declined", cause?: unknown) {
    super(code);
    this.name = "PaymentError";
    this.cause = cause; // the original error/stack survives for the log
  }
}
try {
  await provider.charge();
} catch (e) {
  throw new PaymentError("provider_down", e); // wrap, do not erase
}

Cross-language error-class skeletons (Python, Java, Go, .NET) live in references/retry-and-resilience.md.

Step 2 — Decide retryability

Retry only transient failures, and only on idempotent operations. Retrying the wrong thing turns one slow dependency into a self-inflicted outage.

Retry these (transient)Never retry these (permanent)
Network error, connection reset400 bad request, 422 unprocessable
Timeout401 / 403 (auth/permission)
429 too many requests (honor Retry-After)404 not found
503 / 502 / 504Any business-rule rejection (insufficient funds)
500 on a GET (idempotent)500 on a non-idempotent POST without a key

Idempotency is a precondition, not a nicety. A retried POST that creates a charge can double-charge. Retry only operations that are idempotent by nature (GET, PUT, DELETE) or that carry an idempotency key so the server dedupes. Key design itself belongs to ../api-design/SKILL.md; here you just require one before you retry a mutation.

Caps (industry-converged — AWS Builders' Library, REL05-BP03):

  • Max 3–5 total attempts.
  • Base delay 100–200ms, doubling per attempt.
  • Per-delay cap 10–30s; total retry budget 10–60s then give up.
  • Full jitter to spread load — beats fixed and equal jitter:
text
delay = random_between(0, min(cap, base * 2 ** attempt))

Set a per-attempt timeout first, then retry — a retry on a call that never times out just stacks hung requests.

ts
// GOOD — classify before retrying; cap; full jitter; per-attempt timeout.
async function withRetry<T>(fn: () => Promise<T>, max = 4): Promise<T> {
  for (let attempt = 0; ; attempt++) {
    try {
      return await fn(); // fn must enforce its own per-attempt timeout
    } catch (e) {
      if (!isTransient(e) || attempt >= max - 1) throw e; // permanent or budget spent
      const cap = 10_000, base = 150;
      const delay = Math.random() * Math.min(cap, base * 2 ** attempt); // full jitter
      await new Promise((r) => setTimeout(r, delay));
    }
  }
}

Per-language withRetry (Python tenacity, Java Resilience4j, .NET Polly) is in references/retry-and-resilience.md.

Step 3 — Contain blast radius

Retries alone make a struggling dependency worse. Contain it.

  • Timeout every outbound call. No timeout is a bug, not a default. An un-timed call holds a connection until the OS gives up — minutes you do not have.
  • Circuit breaker — stop hammering a dead dependency. Three states:
    • Closed: requests flow; count failures.
    • Open: trip at ~50% failure over a ~20-request window; reject fast for 30–60s without calling downstream.
    • Half-Open: after the cooldown, let a probe through; success → Closed, failure → Open again.
    • Critical services trip tighter (~30%); tolerant ones up to ~70%.
  • Instruments (current): Opossum (Node — defaults timeout 3000ms / errorThresholdPercentage 50 / resetTimeout 30000ms), Polly 8.6.6 (.NET fluent pipelines), Resilience4j 2.3.0 (Java 17+ 2.x line; a 3.x line targets Java 21). The config matrix is in references/retry-and-resilience.md.
  • Fallback / graceful degradation — when the breaker is Open, serve a stale cache, a safe default, or an honest "this feature is temporarily unavailable". Degrade; do not 500 the whole page.
  • Bulkhead — isolate resource pools (separate connection pool / worker queue per dependency) so one saturated downstream cannot starve the rest. Sizing in the reference.

Step 4 — Boundaries

A boundary is where an unhandled failure is caught and converted into a contained reaction. Place one at each level that can fail independently.

React / Next.js App Router
  • error.tsx is a route-segment boundary. It MUST be a Client Component ('use client') and receives { error, reset }. An error in a segment bubbles to the nearest parent error.tsx.
  • error.tsx does NOT catch an error thrown in its own segment's layout.tsx or template.tsx. Those run outside the boundary — move the boundary to the parent segment to cover them.
  • global-error.tsx wraps the whole app and must render its own <html> and <body> (it replaces the root layout when the root itself fails).
  • Boundaries only catch errors during render. Errors in event handlers, async callbacks, setTimeout, or server-side data fetching are invisible to them — handle those with explicit try/catch + state.
tsx
"use client"; // app/dashboard/error.tsx — REQUIRED

export default function Error({
  error,
  reset,
}: {
  error: Error & { digest?: string };
  reset: () => void;
}) {
  // Log to your telemetry sink (see observability); show the user the digest id.
  return (
    <div role="alert">
      <p>Something went wrong. Quote id {error.digest} to support.</p>
      <button onClick={reset}>Try again</button>
    </div>
  );
}

The full segment-tree placement map and the global-error.tsx skeleton are in references/boundaries-and-messaging.md. Framework specifics: ../nextjs/SKILL.md.

Show full SKILL.md (496 more words)Show less
Server and process
  • Request boundary — one error handler / middleware that maps the taxonomy → HTTP status, attaches a correlation id, and emits the operator log. Every route funnels through it instead of formatting errors ad hoc.
  • Process boundary — a top-level unhandledRejection / uncaughtException handler (Node) or equivalent. Log the cause chain, then exit and let the supervisor restart. A process that keeps running after an unhandled error is running corrupted.

Step 5 — Surface it

Two audiences. Never conflate them — that is how stack traces reach end users and how logs become useless.

User messageOperator log
Goaltell them what to do nextlet you reconstruct what happened
Contentplain language, one action, a correlation idcode, cause chain, request context, structured fields
Neverstack trace, SQL, internal hostnames, PIIa swallowed/lost cause

Taxonomy → HTTP status (the in-process map; the wire envelope shape — RFC 9457 problem+json — belongs to ../api-design/SKILL.md; shipping the log to a sink belongs to ../observability/SKILL.md):

Bucket / codeStatus
validation / bad input400 / 422
unauthenticated / forbidden401 / 403
not found404
domain conflict (slot taken)409
transient downstream / breaker open503 (+ Retry-After)
programmer error / unknown500
text
BAD  → alert("TypeError: cannot read 'id' of undefined")
GOOD → "We couldn't load your orders. Try again in a moment — id a1b2c3."

BAD  → 500 { "error": "ECONNREFUSED 10.0.3.12:5432" }   // leaks topology
GOOD → 503 { "code": "upstream_unavailable", "correlationId": "a1b2c3" }

More before/after rewrites and the copy contract are in references/boundaries-and-messaging.md.

Anti-patterns

Anti-patternWhy it bitesFix
Empty catch {} / except: passfailure vanishes; you debug blind laterhandle, or rethrow with context
Bare except: (Python)swallows KeyboardInterrupt/SystemExit toocatch the specific type
Retry everything, including 4xxretrying a 400 just burns budget; never succeedsretry only the transient table
Retry a non-idempotent POST without a keydouble-charges, duplicate rowsrequire an idempotency key first
Infinite retry, no cap or jitterthundering herd; turns a blip into an outagecap attempts + full jitter + budget
Leak stack trace / SQL to the userhands attackers your internalsgeneric message + id; detail to the log
Expect error.tsx to catch its own segment's layout errorit runs outside the boundary; nothing catches itmove the boundary to the parent
alert(e.message) as the handlerblocks the UI, leaks internals, no recoveryrender an error.tsx with reset
Catch-and-rethrow that drops causethe root error is gone; logs are a dead endwrap, set cause, preserve the chain
Log the error and rethrowdouble-logged at every layer; noise buries signallog at the boundary, or rethrow — not both
Swallow, then return nullcallers deref null later, far from the causereturn a typed Result error
Outbound call with no timeoutone hung dependency exhausts the pooltimeout every call, then retry
One giant try around 200 linesyou cannot tell which call failedscope try to the fallible call
Treat every error as a retryable transientmasks real bugs as "flaky"classify first (Step 1), then react

The gate

Before claiming the failure path is handled, read your diff against the table above — those rows are the checklist. scripts/verify.sh <path> scans the working tree for the highest-signal ones. It is advisory (exit 0 with warnings); pass --strict to make any hit fail. It is heuristic — it flags, you judge.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/error-handling of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/boundaries-and-messaging.md
  • references/retry-and-resilience.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Error Handling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Error Handling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Error Handling this skillericrisco/rsc-harness156—~3.2kAutomated safety check: PassMIT
Rust SkillsJMBeresford/retrom2.1k1 repos~9.5kAutomated safety check: PassMIT
Code PatternsAedelon/claude-code-blueprint120—~1.2kAutomated safety check: PassCustom licence
Inngest Durable FunctionsAsymmetric-al/core382—~4kAutomated safety check: PassAGPL-3.0
Sentry v8 Error Trackingdiet103/claude-code-infrastructure-showcase10k2 repos~2.3kAutomated safety check: NotesMIT
API Error Design Patternsrevfactory/harness-1001.3k—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Rust Skills

    JMBeresford/retrom

    Comprehensive Rust coding guidelines with 265 rules across 26 categories.

    2.1k GitHub starsUsed in 1 repo~9.5k tokens
    DevelopmentAuto-check passed
  • Code Patterns

    Aedelon/claude-code-blueprint

    Reference patterns for REST APIs, pytest/vitest testing, Docker multi-stage builds, GitHub Actions CI/CD, PostgreSQL, TypeScript generics, Python async, and React Server Components.

    120 GitHub stars~1.2k tokensUpdated 7 mo ago
    DevOps & CloudAuto-check passed
  • Inngest Durable Functions

    Asymmetric-al/core

    A skill your agent uses when building functions that must survive process crashes, retry automatically on failure, run on a schedule, react to events, or maintain state across infrastructure…

    382 GitHub stars~4k tokensUpdated today
    Backend & APIsAuto-check passed
  • Sentry v8 Error Tracking

    diet103/claude-code-infrastructure-showcase

    Enforces that every error in a service is captured to Sentry v8, with patterns for controllers, routes, cron jobs and database performance spans instead of console logging alone.

    10k GitHub starsUsed in 2 repos~2.3k tokens
    Backend & APIsAuto-check: notes
  • API Error Design Patterns

    revfactory/harness-100

    Reference for designing how an API reports failures: structured error codes, response shapes, client-friendly messages, an error catalog and retry or fallback advice.

    1.3k GitHub stars~1.6k tokensUpdated 6 mo ago
    Backend & APIsAuto-check passed
  • Official

    Comprehensive technology-agnostic prompt generator for documenting end-to-end application workflows.

    40k GitHub starsUsed in 1 repo~2.7k tokens
    Backend & APIsAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Works with

Questions about Error Handling

What does Error Handling do?

A skill your agent uses when designing the reaction to a class of failures — typed error taxonomies, retry/backoff/timeout policy, circuit breakers, React/Next error boundaries, and the user-message…. Error Handling is an agent skill from ericrisco/rsc-harness. Use when designing the reaction to a class of failures — typed error taxonomies, retry/backoff/timeout policy, circuit breakers, React/Next error boundaries, and the user-message vs operator-log split.

When should I use Error Handling?

Error Handling fits situations like: designing the reaction to a class of failures — typed error taxonomies; retry/backoff/timeout policy; circuit breakers; React/Next error boundaries.

How do I install Error Handling in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill error-handling -a claude-code`. Or copy the skill folder (skills/error-handling in ericrisco/rsc-harness) into .claude/skills/error-handling in your project. Claude Code loads it when a task matches its description.

How do I install Error Handling in Codex?

Run `npx skills add ericrisco/rsc-harness --skill error-handling -a codex`. Or copy the skill folder (skills/error-handling in ericrisco/rsc-harness) into .agents/skills/error-handling in your project. Codex loads it when a task matches its description.

Can I use Error Handling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill error-handling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/error-handling, .gemini/skills/error-handling, .github/skills/error-handling and .opencode/skills/error-handling in your project.

What does Error Handling need to run?

Going by SKILL.md and its folder, Error Handling needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does Error Handling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Error Handling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Error Handling use?

Error Handling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Error Handling use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Error Handling?

Skills that share tags, products or a category with Error Handling: Rust Skills (JMBeresford/retrom, 2.1k stars), Code Patterns (Aedelon/claude-code-blueprint, 120 stars), Inngest Durable Functions (Asymmetric-al/core, 382 stars) and Sentry v8 Error Tracking (diet103/claude-code-infrastructure-showcase, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Error Handling?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.