Agent skill

Error Handler

by EliasOulkadi in EliasOulkadi/shokunin

Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

MITAuto-check: notesDevOps & Cloud

Install Error Handler

skills CLI
$ npx skills add EliasOulkadi/shokunin --skill error-handler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EliasOulkadi/shokunin error-handler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.pack/skills/error-handler .claude/skills/error-handler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
error-handler
GitHub stars
114
Token cost
~3.6k tokens
SKILL.md length
986 words
Files
5 (incl. scripts, references, assets)
Skills in repo
49
Repo updated
First seen
Licence
MIT

At a glance

Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

  • Works in 5 steps: Classify errors → Set up OpenTelemetry (exact configuration) → Implement structured error middleware → …
  • User asks to implement error handling
  • SKILL.md covers Sub-Commands, The 3 Signals (OpenTelemetry), Workflow and Structured Logging Schema, plus 5 more sections
  • Runs TypeScript and Shell scripts from its folder; calls curl

What it does

Error Handler is an agent skill from EliasOulkadi/shokunin. Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead, timeout), error budgets/SLOs with burn rate alerts, and production incident triage. Use when user asks to implement error handling, logging, monitoring, observability, OpenTelemetry, error boundaries, circuit breakers, retry logic, or SLO tracking. Do NOT use for incident runbooks (use runbook-gen), vendor-specific APM setup…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/error-middleware.template.ts`, `references/error-budgets.md` and `references/opentelemetry-deep.md`). Compatibility notes: opencode

It sits in DevOps & Cloud, covering Site reliability engineering, Observability and Error handling. It works with OpenTelemetry, Datadog, Sentry and Kubernetes. The repository describes itself as: 職人 Shokunin 62 AI agent skills for OpenCode, Claude Code, Cursor, Windsurf. ChromaDB memory, MCP servers, declarative self-updates. Multi-model, open source, zero cost. The licence is MIT.

When your agent uses it

  • User asks to implement error handling
  • Error boundaries
  • Circuit breakers
  • Incident runbooks (use runbook-gen)

Example prompts

  • “/error-handler”

Requirements

  • Node.js
  • A Bash shell
  • Compatibility (from SKILL.md): opencode
  • Pre-approved tools (allowed-tools): Read, Bash, Write, Grep

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Classify errors
  2. Set up OpenTelemetry (exact configuration)
  3. Implement structured error middleware
  4. Recovery patterns (exact implementations)
  5. Error budgets + SLO

What it can do on your machine

Read from SKILL.md and the folder at commit 4c68e5b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash
    • Write
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (TypeScript and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    opencode

    From compatibility in the SKILL.md frontmatter.

Context cost

Error Handler loads about 3.6k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 146 tokens; SKILL.md has 986 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash, Write, Grep

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from EliasOulkadi/shokunin at commit 4c68e5b, republished under its MIT licence (© EliasOulkadi). 986 words, ~3,586 tokens.

Download SKILL.mdSave it as .claude/skills/error-handler/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
error-handler
description
Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead, timeout), error budgets/SLOs with burn rate alerts, and production incident triage. Use when user asks to implement error handling, logging, monitoring, observability, OpenTelemetry, error boundaries, circuit breakers, retry logic, or SLO tracking. Do NOT use for incident runbooks (use runbook-gen), vendor-specific APM setup (Datadog, Sentry agent config), or K8s debugging.
allowed-tools
Read, Bash, Write, Grep
compatibility
opencode
triggers
error handling, logging, observability, OpenTelemetry, error boundary, circuit breaker, retry logic, SLO, error budget, structured logging, monitoring, tracing
negatives
incident runbook, on-call, Datadog setup, Sentry config, K8s debugging, APM vendor setup
license
MIT
metadata.workflow
backend
metadata.audience
developers
metadata.version
4.0.0
metadata.author
shokunin

Error Handler

Observability that makes debugging fast and production predictable. Based on Google SRE (Site Reliability Engineering), OpenTelemetry semantic conventions, and production patterns from Sentry, Honeycomb, and Datadog.

Sub-Commands

CommandDescription
setupSet up OpenTelemetry SDK + instrumentation for a service
classifyClassify errors by HTTP status, severity, and alert priority
recoverImplement recovery patterns: retry with jitter, circuit breaker, timeout
sloDefine error budget, SLO, and burn rate alerts
auditAudit existing error handling against best practices

The 3 Signals (OpenTelemetry)

SignalPurposeExample
TracesFollow a request across servicesUser request ? API Gateway ? Auth ? DB
MetricsAggregate measurements over timeRequest count, error rate, latency p50/p95/p99
LogsDiscrete events with context"User login failed: invalid credentials for user_abc123"

All three signals share trace_id + span_id for correlation.

Workflow

Step 1: Classify errors
CategoryHTTPSeverityLog levelSLO impactAlert?
Validation (user error)400LowinfoNoNo
Authentication401MediumwarnNoOnly if spike detected
Authorization403MediumwarnNoOnly if spike detected
Not Found404LowdebugNoNo
Conflict409MediuminfoNoNo
Rate Limit429LowwarnNoNo
Internal500HigherrorYesYes
Downstream failure502/503HigherrorYesYes
Timeout504HigherrorYesYes
Step 2: Set up OpenTelemetry (exact configuration)
typescript
import { NodeSDK } from '@opentelemetry/sdk-node'
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http'
import { Resource } from '@opentelemetry/resources'
import { ATTR_SERVICE_NAME } from '@opentelemetry/semantic-conventions'

const sdk = new NodeSDK({
  resource: new Resource({
    [ATTR_SERVICE_NAME]: 'api-service',
    'deployment.environment': process.env.NODE_ENV,
  }),
  traceExporter: new OTLPTraceExporter({
    url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT,
  }),
  instrumentations: [
    new HttpInstrumentation(),
    new ExpressInstrumentation(),
    new PgInstrumentation(),
    new RedisInstrumentation(),
  ],
})

sdk.start()

process.on('SIGTERM', async () => {
  await sdk.shutdown()
})

Sampling strategy:

  • Dev/staging: AlwaysOn (100%)
  • Production: ParentBased + TraceIdRatioBased(0.1) (10% head-based)
  • Never sample errors out. Use AlwaysOn for traces with status=ERROR.

Import note: ParentBasedSampler is imported from @opentelemetry/sdk-trace-base, NOT from @opentelemetry/sdk-node. If using @opentelemetry/sdk-node, configure it via the sampler option in NodeSDK constructor.

Step 3: Implement structured error middleware
typescript
import { trace, SpanStatusCode } from '@opentelemetry/api'
import { randomUUID } from 'crypto'

app.use((err: Error, req: Request, res: Response, next: NextFunction) => {
  const span = trace.getActiveSpan()
  const requestId = req.headers['x-request-id'] || randomUUID()

  span?.recordException(err)
  span?.setStatus({ code: SpanStatusCode.ERROR, message: err.message })

  logger.error({
    message: err.message,
    errorName: err.name,
    stack: err.stack,
    requestId,
    trace_id: span?.spanContext().traceId,
    span_id: span?.spanContext().spanId,
    path: req.path,
    method: req.method,
    userId: req.user?.id,
  })

  if (err instanceof ValidationError) {
    return res.status(400).json({
      error: {
        code: 'VALIDATION_ERROR',
        message: err.message,
        details: err.details,
        requestId,
      },
    })
  }

  if (err instanceof NotFoundError) {
    return res.status(404).json({
      error: { code: 'NOT_FOUND', message: err.message, requestId },
    })
  }

  // Generic fallback - never leak internal state
  res.status(500).json({
    error: {
      code: 'INTERNAL_ERROR',
      message: 'An unexpected error occurred',
      requestId,
    },
  })
})
Step 4: Recovery patterns (exact implementations)
Retry with exponential backoff + jitter
typescript
async function withRetry<T>(
  fn: () => Promise<T>,
  options: { maxRetries?: number; baseMs?: number; maxMs?: number } = {}
): Promise<T> {
  const { maxRetries = 3, baseMs = 200, maxMs = 10000 } = options

  for (let attempt = 0; attempt <= maxRetries; attempt++) {
    try {
      return await fn()
    } catch (error) {
      if (attempt === maxRetries) throw error
      if (!isRetryable(error)) throw error

      const exponentialDelay = baseMs * Math.pow(2, attempt)
      const jitter = Math.random() * 0.3 * exponentialDelay
      const delay = Math.min(exponentialDelay + jitter, maxMs)

      await new Promise(r => setTimeout(r, delay))
    }
  }
  throw new Error('Unreachable')
}

function isRetryable(error: any): boolean {
  const retryableStatuses = [408, 429, 502, 503, 504]
  if (error?.status && retryableStatuses.includes(error.status)) return true
  if (error?.code === 'ETIMEDOUT' || error?.code === 'ECONNRESET') return true
  return false
}

// Never retry: 400, 401, 403, 404, 409, 422
Circuit breaker
typescript
class CircuitBreaker {
  private failures = 0
  private state: 'closed' | 'open' | 'half-open' = 'closed'
  private lastFailureTime = 0

  constructor(
    private failureThreshold: number = 5,
    private resetTimeout: number = 30000
  ) {}

  async call<T>(fn: () => Promise<T>): Promise<T> {
    if (this.state === 'open') {
      if (Date.now() - this.lastFailureTime > this.resetTimeout) {
        this.state = 'half-open'
      } else {
        throw new Error('Circuit breaker is open')
      }
    }

    try {
      const result = await fn()
      if (this.state === 'half-open') {
        this.state = 'closed'
        this.failures = 0
      }
      return result
    } catch (error) {
      this.failures++
      this.lastFailureTime = Date.now()

      if (this.failures >= this.failureThreshold) {
        this.state = 'open'
      }
      throw error
    }
  }
}
Timeout
typescript
function withTimeout<T>(promise: Promise<T>, timeoutMs: number): Promise<T> {
  const timeout = new Promise<T>((_, reject) =>
    setTimeout(() => reject(new Error(`Operation timed out after ${timeoutMs}ms`)), timeoutMs)
  )
  return Promise.race([promise, timeout])
}

// Usage
const result = await withTimeout(fetch('https://api.example.com/data'), 5000)
Step 5: Error budgets + SLO
SLOError budget (monthly)Burn rate criticalBurn rate warning
99.9%43m 50s6x (exhausted in 4.6d)3x (exhausted in 9.3d)
99.5%3h 39m10x (exhausted in 1.1d)5x (exhausted in 2.2d)
99%7h 18m15x (exhausted in 0.5d)7.5x (exhausted in 1d)

Burn rate alert (PromQL for 99.9% SLO):

promql
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
> 0.001 * 6

Multi-window, multi-burn-rate approach (Google SRE):

  • Short window (5m) + high burn rate (6-10x) ? Page on-call immediately
  • Long window (1h) + low burn rate (2-3x) ? Ticket for next business day

Structured Logging Schema

Every log line must be JSON with these fields:

json
{
  "timestamp": "2026-05-16T10:21:27.000Z",
  "level": "error",
  "message": "Failed to process payment",
  "trace_id": "0af7651916cd43dd8448eb211c80319c",
  "span_id": "b7ad6b7169203331",
  "request_id": "req_a1b2c3d4",
  "service": "payment-service",
  "environment": "production",
  "userId": "user_abc123",
  "path": "/api/payments",
  "method": "POST",
  "duration_ms": 234,
  "error": {
    "name": "PaymentFailedError",
    "message": "Card declined",
    "code": "CARD_DECLINED"
  }
}

Rules:

  • Never log PII, tokens, passwords, or full credit card numbers
  • Always include trace_id, span_id, request_id for correlation
  • duration_ms on every request. p95 > 200ms ? investigate.

Production Checklist

  • OpenTelemetry SDK initialized at app startup (before any imports)
  • Auto-instrumentations for HTTP, DB, messaging, Redis
  • Every route handler wrapped in try/catch or error middleware
  • Structured JSON logging (never plain text)
  • trace_id + span_id + request_id in every log line
  • Recovery: retry with jitter on external calls. Timeout on every external call.
  • Circuit breaker for critical downstream services
  • SLO defined for every service (99.9% critical, 99.5% standard)
  • Burn rate alerts: multi-window (5m short, 1h long)
  • No PII, tokens, or secrets in logs or error messages
  • Error classification by type (not generic catch-all)
  • Sampling: 10% production, 100% for errors

Anti-Patterns

Anti-patternFix
Log in catch AND rethrowLog OR throw. Not both.
Generic error messagesInclude code, request_id, details.
No trace context in async flowsPropagate context via context.active().
Silent catch with no logEvery catch logs, handles, or rethrows.
Retry non-retryable errors (400, 401)Check status before retry: isRetryable().
No timeout on external callsAlways set connection (2s) + read (10s) timeouts.
No error classificationClassify by type and handle accordingly.
Plain text logsStructured JSON with consistent fields.
Same retry delay every timeExponential backoff + jitter.
Circuit breaker never testedTest in staging by killing dependencies.
Error budget without alertBurn rate alerts are mandatory.
Show full SKILL.md (389 more words)Show less

Error Handling

ErrorCauseFix
OpenTelemetry SDK not exporting tracesExporter endpoint unreachable or OTEL_EXPORTER_OTLP_ENDPOINT not setVerify endpoint with curl $OTEL_EXPORTER_OTLP_ENDPOINT/v1/traces. Set env var before SDK initialization.
ParentBasedSampler import fails from @opentelemetry/sdk-nodeWrong import pathImport from @opentelemetry/sdk-trace-base, not @opentelemetry/sdk-node. Configure via sampler option in NodeSDK constructor.
Traces missing in production but appear in stagingSampling rate too low (TraceIdRatioBased(0.1) drops 90%)Use AlwaysOn for traces with status=ERROR. Never sample errors out.
Structured logs have null trace_id/span_idContext not propagated into async flow or logger not context-awareUse context.active() before async operations. Inject trace_id/span_id into every log line via context propagation.
Circuit breaker stays open indefinitelyresetTimeout is set but half-open state logic never triggersVerify Date.now() - this.lastFailureTime > this.resetTimeout comparison. Add logging when state transitions occur.
Retry exhausts budget on non-retryable errorsisRetryable() returns true for 400/401/403 status codesHard-gate retryable statuses: only [408, 429, 502, 503, 504] and network errors (ETIMEDOUT, ECONNRESET).
Retry with jitter produces thundering herdAll instances retry at the same time due to identical base delayRandomize baseMs per instance or use Math.random() * baseMs as the starting backoff.
Burn rate alert fires for low-traffic servicesDivision by small request count amplifies noiseSet minimum request thresholds: only evaluate SLO when rate(http_requests_total[5m]) > 1.
Error middleware exposes stack traces to usersGeneric INTERNAL_ERROR handler includes err.stack in response bodyNever return stack in API responses. Log it server-side; respond with { error: { code, message, requestId } }.
PII leaks into structured logsFull email, IP, SSN, or credit card numbers in log fieldsMask sensitive fields: email.replace(/(.{3}).*(@.*)/, '$1***$2'). Configure PII redaction in log pipeline.
Timeout wrapped around fetch never firesPromise.race with setTimeout races against a promise that catches and suppresses errorsAlways reject in the timeout handler, not resolve. The timed-out promise must reject, not resolve.

Sources

  • Google SRE Book - Monitoring Distributed Systems (Chapter 6)
  • Google SRE Workbook - Implementing SLOs (Chapter 5)
  • OpenTelemetry documentation (opentelemetry.io)
  • OpenTelemetry semantic conventions
  • Sentry error handling best practices
  • AWS Well-Architected Framework - Reliability Pillar
  • Microsoft Polly - Circuit breaker patterns
  • Hystrix - Netflix circuit breaker patterns

Checklist

  • Skill loads without errors in the AI agent
  • YAML frontmatter is valid (description, compatibility, audience)
  • Workflow section provides clear step-by-step instructions
  • Error handling section covers common failure modes
  • All referenced files (references/, scripts/, assets/) exist
  • Skill triggers correctly for intended use cases
  • No broken links or missing resources

© EliasOulkadi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in .pack/skills/error-handler of EliasOulkadi/shokunin.

  • SKILL.md
  • assets/error-middleware.template.ts
  • references/error-budgets.md
  • references/opentelemetry-deep.md
  • scripts/setup-opentelemetry.sh

Open the folder on GitHubat commit 4c68e5b

Compare with similar skills

Error Handler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Error Handler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Error Handler this skillEliasOulkadi/shokunin114—~3.6kAutomated safety check: NotesMIT
Monitoring Observabilityahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNone
Tsh Implementing ObservabilityTheSoftwareHouse/copilot-collections284—~2kAutomated safety check: PassMIT
Observability Sre Triageelastic/agent-skills592—~7.4kAutomated safety check: PassApache-2.0
Logging Observabilitygetsentry/toolkit917—~2.6kAutomated safety check: PassCustom licence
Temps Best Practicesgotempsh/temps822—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Tsh Implementing Observability

    TheSoftwareHouse/copilot-collections

    Observability patterns for logging, monitoring, alerting, and distributed tracing.

    284 GitHub stars~2k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Observability Sre Triage

    elastic/agent-skills

    Official

    Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health…

    592 GitHub stars~7.4k tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • Logging Observability

    getsentry/toolkit

    Official

    Review code for correct logging and error handling patterns.

    917 GitHub stars~2.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Temps Best Practices

    gotempsh/temps

    Best-practices reference for preparing and instrumenting applications on Temps.

    822 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Logfire Infrastructure

    pydantic/skills

    Official

    Monitor hosts, Docker containers, Kubernetes clusters, database/queue/cache servers, and cloud-provider metrics with Pydantic Logfire — no application code required.

    140 GitHub stars~1.8k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed

More from EliasOulkadi/shokunin

All 49 skills in this repo
  • CI CD

    EliasOulkadi/shokunin

    Design CI/CD pipelines for GitHub Actions, GitLab CI, and CircleCI with matrix builds, test sharding, caching, Docker layer caching, OIDC auth, deployment strategies (rolling, blue-green, canary)…

    114 GitHub stars~3.4k tokensUpdated 3 days ago
    Auto-check: notes
  • Component Forge

    EliasOulkadi/shokunin

    Build production-grade components for React, Vue 3, and Svelte 5 with all states (loading, empty, error, success, idle), TypeScript strict, WCAG 2.2 accessibility, server components (RSC), and…

    114 GitHub stars~3.6k tokensUpdated 3 days ago
    Auto-check: notes
  • DB Admin

    EliasOulkadi/shokunin

    PostgreSQL database administration — backup/restore (pgdump, PITR, WAL archiving), health monitoring (connections, bloat, cache hit ratio, dead tuples), connection pooling (PgBouncer), replication…

    114 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check: notes
  • DB Sculptor

    EliasOulkadi/shokunin

    Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…

    114 GitHub stars~3.1k tokensUpdated 3 days ago
    Auto-check: notes
  • Docker

    EliasOulkadi/shokunin

    Optimize Docker images with multi-stage builds, distroless bases, BuildKit cache mounts, multi-arch builds, compose watch, security hardening (non-root, seccomp, capabilities drop), and…

    114 GitHub stars~3.8k tokensUpdated 3 days ago
    Auto-check: notes
  • Kubernetes

    EliasOulkadi/shokunin

    Deploy, manage, and debug Kubernetes in production — Deployments, Services, Gateway API, Service Mesh (Istio/Linkerd/Cilium), eBPF observability (Cilium Hubble), security hardening (Pod Security…

    114 GitHub stars~3.3k tokensUpdated 3 days ago
    Auto-check: notes

Categories

Questions about Error Handler

What does Error Handler do?

Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…. Error Handler is an agent skill from EliasOulkadi/shokunin. Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead, timeout), error budgets/SLOs with burn rate alerts, and production incident triage.

When should I use Error Handler?

Error Handler fits situations like: user asks to implement error handling; error boundaries; circuit breakers; incident runbooks (use runbook-gen).

How do I install Error Handler in Claude Code?

Run `npx skills add EliasOulkadi/shokunin --skill error-handler -a claude-code`. Or copy the skill folder (.pack/skills/error-handler in EliasOulkadi/shokunin) into .claude/skills/error-handler in your project. Claude Code loads it when a task matches its description.

How do I install Error Handler in Codex?

Run `npx skills add EliasOulkadi/shokunin --skill error-handler -a codex`. Or copy the skill folder (.pack/skills/error-handler in EliasOulkadi/shokunin) into .agents/skills/error-handler in your project. Codex loads it when a task matches its description.

Can I use Error Handler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EliasOulkadi/shokunin --skill error-handler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/error-handler, .gemini/skills/error-handler, .github/skills/error-handler and .opencode/skills/error-handler in your project.

What does Error Handler need to run?

Going by SKILL.md and its folder, Error Handler needs TypeScript and a shell for the scripts in its folder and the command-line tools its instructions call (curl). Our summary lists: Node.js; A Bash shell. Its frontmatter pre-approves these tools: Read, Bash, Write, Grep. Compatibility (from SKILL.md): opencode.

Does Error Handler access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Error Handler safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Error Handler use?

Error Handler is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Error Handler use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.4k tokens, read only when the agent opens those files.

What are the alternatives to Error Handler?

Skills that share tags, products or a category with Error Handler: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Tsh Implementing Observability (TheSoftwareHouse/copilot-collections, 284 stars), Observability Sre Triage (elastic/agent-skills, 592 stars) and Logging Observability (getsentry/toolkit, 917 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Error Handler?

EliasOulkadi (a GitHub user) maintains it in EliasOulkadi/shokunin, which has 114 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 5, 2026.

Source: EliasOulkadi/shokunin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.