Monitoring Observability
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…
$ npx skills add EliasOulkadi/shokunin --skill error-handler -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install EliasOulkadi/shokunin error-handler --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.pack/skills/error-handler .claude/skills/error-handler && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "error-handler" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/error-handler into .claude/skills/error-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-handler", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/error-handlerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add EliasOulkadi/shokunin --skill error-handler -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install EliasOulkadi/shokunin error-handler --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.pack/skills/error-handler .agents/skills/error-handler && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "error-handler" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/error-handler into .agents/skills/error-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-handler", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add EliasOulkadi/shokunin --skill error-handler -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install EliasOulkadi/shokunin error-handler --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.pack/skills/error-handler .cursor/skills/error-handler && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "error-handler" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/error-handler into .cursor/skills/error-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-handler", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/EliasOulkadi/shokunin.git --path .pack/skills/error-handler--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add EliasOulkadi/shokunin --skill error-handler -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install EliasOulkadi/shokunin error-handler --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.pack/skills/error-handler .gemini/skills/error-handler && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "error-handler" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/error-handler into .gemini/skills/error-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-handler", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install EliasOulkadi/shokunin error-handlerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add EliasOulkadi/shokunin --skill error-handler -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .github/skills && cp -r skills-src/.pack/skills/error-handler .github/skills/error-handler && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "error-handler" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/error-handler into .github/skills/error-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-handler", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add EliasOulkadi/shokunin --skill error-handler -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install EliasOulkadi/shokunin error-handler --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/EliasOulkadi/shokunin.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.pack/skills/error-handler .opencode/skills/error-handler && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "error-handler" agent skill from https://github.com/EliasOulkadi/shokunin/tree/master/.pack/skills/error-handler into .opencode/skills/error-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "error-handler", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
error-handlerDesign error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…
Error Handler is an agent skill from EliasOulkadi/shokunin. Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead, timeout), error budgets/SLOs with burn rate alerts, and production incident triage. Use when user asks to implement error handling, logging, monitoring, observability, OpenTelemetry, error boundaries, circuit breakers, retry logic, or SLO tracking. Do NOT use for incident runbooks (use runbook-gen), vendor-specific APM setup…
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/error-middleware.template.ts`, `references/error-budgets.md` and `references/opentelemetry-deep.md`). Compatibility notes: opencode
It sits in DevOps & Cloud, covering Site reliability engineering, Observability and Error handling. It works with OpenTelemetry, Datadog, Sentry and Kubernetes. The repository describes itself as: 職人 Shokunin 62 AI agent skills for OpenCode, Claude Code, Cursor, Windsurf. ChromaDB memory, MCP servers, declarative self-updates. Multi-model, open source, zero cost. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4c68e5b. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashWriteGrepFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (TypeScript and Shell), which the agent can run.
Shell commands in SKILL.md call:
curlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
opencode
From compatibility in the SKILL.md frontmatter.
Error Handler loads about 3.6k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 146 tokens; SKILL.md has 986 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Bash, Write, GrepAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from EliasOulkadi/shokunin at commit 4c68e5b, republished under its MIT licence (© EliasOulkadi). 986 words, ~3,586 tokens.
.claude/skills/error-handler/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Observability that makes debugging fast and production predictable. Based on Google SRE (Site Reliability Engineering), OpenTelemetry semantic conventions, and production patterns from Sentry, Honeycomb, and Datadog.
| Command | Description |
|---|---|
setup | Set up OpenTelemetry SDK + instrumentation for a service |
classify | Classify errors by HTTP status, severity, and alert priority |
recover | Implement recovery patterns: retry with jitter, circuit breaker, timeout |
slo | Define error budget, SLO, and burn rate alerts |
audit | Audit existing error handling against best practices |
| Signal | Purpose | Example |
|---|---|---|
| Traces | Follow a request across services | User request ? API Gateway ? Auth ? DB |
| Metrics | Aggregate measurements over time | Request count, error rate, latency p50/p95/p99 |
| Logs | Discrete events with context | "User login failed: invalid credentials for user_abc123" |
All three signals share trace_id + span_id for correlation.
| Category | HTTP | Severity | Log level | SLO impact | Alert? |
|---|---|---|---|---|---|
| Validation (user error) | 400 | Low | info | No | No |
| Authentication | 401 | Medium | warn | No | Only if spike detected |
| Authorization | 403 | Medium | warn | No | Only if spike detected |
| Not Found | 404 | Low | debug | No | No |
| Conflict | 409 | Medium | info | No | No |
| Rate Limit | 429 | Low | warn | No | No |
| Internal | 500 | High | error | Yes | Yes |
| Downstream failure | 502/503 | High | error | Yes | Yes |
| Timeout | 504 | High | error | Yes | Yes |
import { NodeSDK } from '@opentelemetry/sdk-node'
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http'
import { Resource } from '@opentelemetry/resources'
import { ATTR_SERVICE_NAME } from '@opentelemetry/semantic-conventions'
const sdk = new NodeSDK({
resource: new Resource({
[ATTR_SERVICE_NAME]: 'api-service',
'deployment.environment': process.env.NODE_ENV,
}),
traceExporter: new OTLPTraceExporter({
url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT,
}),
instrumentations: [
new HttpInstrumentation(),
new ExpressInstrumentation(),
new PgInstrumentation(),
new RedisInstrumentation(),
],
})
sdk.start()
process.on('SIGTERM', async () => {
await sdk.shutdown()
})Sampling strategy:
ParentBased + TraceIdRatioBased(0.1) (10% head-based)AlwaysOn for traces with status=ERROR.Import note:
ParentBasedSampleris imported from@opentelemetry/sdk-trace-base, NOT from@opentelemetry/sdk-node. If using@opentelemetry/sdk-node, configure it via thesampleroption inNodeSDKconstructor.
import { trace, SpanStatusCode } from '@opentelemetry/api'
import { randomUUID } from 'crypto'
app.use((err: Error, req: Request, res: Response, next: NextFunction) => {
const span = trace.getActiveSpan()
const requestId = req.headers['x-request-id'] || randomUUID()
span?.recordException(err)
span?.setStatus({ code: SpanStatusCode.ERROR, message: err.message })
logger.error({
message: err.message,
errorName: err.name,
stack: err.stack,
requestId,
trace_id: span?.spanContext().traceId,
span_id: span?.spanContext().spanId,
path: req.path,
method: req.method,
userId: req.user?.id,
})
if (err instanceof ValidationError) {
return res.status(400).json({
error: {
code: 'VALIDATION_ERROR',
message: err.message,
details: err.details,
requestId,
},
})
}
if (err instanceof NotFoundError) {
return res.status(404).json({
error: { code: 'NOT_FOUND', message: err.message, requestId },
})
}
// Generic fallback - never leak internal state
res.status(500).json({
error: {
code: 'INTERNAL_ERROR',
message: 'An unexpected error occurred',
requestId,
},
})
})async function withRetry<T>(
fn: () => Promise<T>,
options: { maxRetries?: number; baseMs?: number; maxMs?: number } = {}
): Promise<T> {
const { maxRetries = 3, baseMs = 200, maxMs = 10000 } = options
for (let attempt = 0; attempt <= maxRetries; attempt++) {
try {
return await fn()
} catch (error) {
if (attempt === maxRetries) throw error
if (!isRetryable(error)) throw error
const exponentialDelay = baseMs * Math.pow(2, attempt)
const jitter = Math.random() * 0.3 * exponentialDelay
const delay = Math.min(exponentialDelay + jitter, maxMs)
await new Promise(r => setTimeout(r, delay))
}
}
throw new Error('Unreachable')
}
function isRetryable(error: any): boolean {
const retryableStatuses = [408, 429, 502, 503, 504]
if (error?.status && retryableStatuses.includes(error.status)) return true
if (error?.code === 'ETIMEDOUT' || error?.code === 'ECONNRESET') return true
return false
}
// Never retry: 400, 401, 403, 404, 409, 422class CircuitBreaker {
private failures = 0
private state: 'closed' | 'open' | 'half-open' = 'closed'
private lastFailureTime = 0
constructor(
private failureThreshold: number = 5,
private resetTimeout: number = 30000
) {}
async call<T>(fn: () => Promise<T>): Promise<T> {
if (this.state === 'open') {
if (Date.now() - this.lastFailureTime > this.resetTimeout) {
this.state = 'half-open'
} else {
throw new Error('Circuit breaker is open')
}
}
try {
const result = await fn()
if (this.state === 'half-open') {
this.state = 'closed'
this.failures = 0
}
return result
} catch (error) {
this.failures++
this.lastFailureTime = Date.now()
if (this.failures >= this.failureThreshold) {
this.state = 'open'
}
throw error
}
}
}function withTimeout<T>(promise: Promise<T>, timeoutMs: number): Promise<T> {
const timeout = new Promise<T>((_, reject) =>
setTimeout(() => reject(new Error(`Operation timed out after ${timeoutMs}ms`)), timeoutMs)
)
return Promise.race([promise, timeout])
}
// Usage
const result = await withTimeout(fetch('https://api.example.com/data'), 5000)| SLO | Error budget (monthly) | Burn rate critical | Burn rate warning |
|---|---|---|---|
| 99.9% | 43m 50s | 6x (exhausted in 4.6d) | 3x (exhausted in 9.3d) |
| 99.5% | 3h 39m | 10x (exhausted in 1.1d) | 5x (exhausted in 2.2d) |
| 99% | 7h 18m | 15x (exhausted in 0.5d) | 7.5x (exhausted in 1d) |
Burn rate alert (PromQL for 99.9% SLO):
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
> 0.001 * 6Multi-window, multi-burn-rate approach (Google SRE):
Every log line must be JSON with these fields:
{
"timestamp": "2026-05-16T10:21:27.000Z",
"level": "error",
"message": "Failed to process payment",
"trace_id": "0af7651916cd43dd8448eb211c80319c",
"span_id": "b7ad6b7169203331",
"request_id": "req_a1b2c3d4",
"service": "payment-service",
"environment": "production",
"userId": "user_abc123",
"path": "/api/payments",
"method": "POST",
"duration_ms": 234,
"error": {
"name": "PaymentFailedError",
"message": "Card declined",
"code": "CARD_DECLINED"
}
}Rules:
trace_id, span_id, request_id for correlationduration_ms on every request. p95 > 200ms ? investigate.trace_id + span_id + request_id in every log line| Anti-pattern | Fix |
|---|---|
| Log in catch AND rethrow | Log OR throw. Not both. |
| Generic error messages | Include code, request_id, details. |
| No trace context in async flows | Propagate context via context.active(). |
| Silent catch with no log | Every catch logs, handles, or rethrows. |
| Retry non-retryable errors (400, 401) | Check status before retry: isRetryable(). |
| No timeout on external calls | Always set connection (2s) + read (10s) timeouts. |
| No error classification | Classify by type and handle accordingly. |
| Plain text logs | Structured JSON with consistent fields. |
| Same retry delay every time | Exponential backoff + jitter. |
| Circuit breaker never tested | Test in staging by killing dependencies. |
| Error budget without alert | Burn rate alerts are mandatory. |
| Error | Cause | Fix |
|---|---|---|
| OpenTelemetry SDK not exporting traces | Exporter endpoint unreachable or OTEL_EXPORTER_OTLP_ENDPOINT not set | Verify endpoint with curl $OTEL_EXPORTER_OTLP_ENDPOINT/v1/traces. Set env var before SDK initialization. |
ParentBasedSampler import fails from @opentelemetry/sdk-node | Wrong import path | Import from @opentelemetry/sdk-trace-base, not @opentelemetry/sdk-node. Configure via sampler option in NodeSDK constructor. |
| Traces missing in production but appear in staging | Sampling rate too low (TraceIdRatioBased(0.1) drops 90%) | Use AlwaysOn for traces with status=ERROR. Never sample errors out. |
Structured logs have null trace_id/span_id | Context not propagated into async flow or logger not context-aware | Use context.active() before async operations. Inject trace_id/span_id into every log line via context propagation. |
| Circuit breaker stays open indefinitely | resetTimeout is set but half-open state logic never triggers | Verify Date.now() - this.lastFailureTime > this.resetTimeout comparison. Add logging when state transitions occur. |
| Retry exhausts budget on non-retryable errors | isRetryable() returns true for 400/401/403 status codes | Hard-gate retryable statuses: only [408, 429, 502, 503, 504] and network errors (ETIMEDOUT, ECONNRESET). |
| Retry with jitter produces thundering herd | All instances retry at the same time due to identical base delay | Randomize baseMs per instance or use Math.random() * baseMs as the starting backoff. |
| Burn rate alert fires for low-traffic services | Division by small request count amplifies noise | Set minimum request thresholds: only evaluate SLO when rate(http_requests_total[5m]) > 1. |
| Error middleware exposes stack traces to users | Generic INTERNAL_ERROR handler includes err.stack in response body | Never return stack in API responses. Log it server-side; respond with { error: { code, message, requestId } }. |
| PII leaks into structured logs | Full email, IP, SSN, or credit card numbers in log fields | Mask sensitive fields: email.replace(/(.{3}).*(@.*)/, '$1***$2'). Configure PII redaction in log pipeline. |
Timeout wrapped around fetch never fires | Promise.race with setTimeout races against a promise that catches and suppresses errors | Always reject in the timeout handler, not resolve. The timed-out promise must reject, not resolve. |
© EliasOulkadi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts, references, assets) in .pack/skills/error-handler of EliasOulkadi/shokunin.
Open the folder on GitHubat commit 4c68e5b
Error Handler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Error Handler this skillEliasOulkadi/shokunin | 114 | — | ~3.6k | Automated safety check: Notes | MIT | |
| Monitoring Observabilityahmedasmar/devops-claude-skills | 203 | — | ~3.9k | Automated safety check: Pass | None | |
| Tsh Implementing ObservabilityTheSoftwareHouse/copilot-collections | 284 | — | ~2k | Automated safety check: Pass | MIT | |
| Observability Sre Triageelastic/agent-skills | 592 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | |
| Logging Observabilitygetsentry/toolkit | 917 | — | ~2.6k | Automated safety check: Pass | Custom licence | |
| Temps Best Practicesgotempsh/temps | 822 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 |
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
TheSoftwareHouse/copilot-collections
Observability patterns for logging, monitoring, alerting, and distributed tracing.
elastic/agent-skills
Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health…
getsentry/toolkit
Review code for correct logging and error handling patterns.
gotempsh/temps
Best-practices reference for preparing and instrumenting applications on Temps.
pydantic/skills
Monitor hosts, Docker containers, Kubernetes clusters, database/queue/cache servers, and cloud-provider metrics with Pydantic Logfire — no application code required.
EliasOulkadi/shokunin
Design CI/CD pipelines for GitHub Actions, GitLab CI, and CircleCI with matrix builds, test sharding, caching, Docker layer caching, OIDC auth, deployment strategies (rolling, blue-green, canary)…
EliasOulkadi/shokunin
Build production-grade components for React, Vue 3, and Svelte 5 with all states (loading, empty, error, success, idle), TypeScript strict, WCAG 2.2 accessibility, server components (RSC), and…
EliasOulkadi/shokunin
PostgreSQL database administration — backup/restore (pgdump, PITR, WAL archiving), health monitoring (connections, bloat, cache hit ratio, dead tuples), connection pooling (PgBouncer), replication…
EliasOulkadi/shokunin
Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…
EliasOulkadi/shokunin
Optimize Docker images with multi-stage builds, distroless bases, BuildKit cache mounts, multi-arch builds, compose watch, security hardening (non-root, seccomp, capabilities drop), and…
EliasOulkadi/shokunin
Deploy, manage, and debug Kubernetes in production — Deployments, Services, Gateway API, Service Mesh (Istio/Linkerd/Cilium), eBPF observability (Cilium Hubble), security hardening (Pod Security…
Works with
Categories
Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…. Error Handler is an agent skill from EliasOulkadi/shokunin. Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead, timeout), error budgets/SLOs with burn rate alerts, and production incident triage.
Error Handler fits situations like: user asks to implement error handling; error boundaries; circuit breakers; incident runbooks (use runbook-gen).
Run `npx skills add EliasOulkadi/shokunin --skill error-handler -a claude-code`. Or copy the skill folder (.pack/skills/error-handler in EliasOulkadi/shokunin) into .claude/skills/error-handler in your project. Claude Code loads it when a task matches its description.
Run `npx skills add EliasOulkadi/shokunin --skill error-handler -a codex`. Or copy the skill folder (.pack/skills/error-handler in EliasOulkadi/shokunin) into .agents/skills/error-handler in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EliasOulkadi/shokunin --skill error-handler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/error-handler, .gemini/skills/error-handler, .github/skills/error-handler and .opencode/skills/error-handler in your project.
Going by SKILL.md and its folder, Error Handler needs TypeScript and a shell for the scripts in its folder and the command-line tools its instructions call (curl). Our summary lists: Node.js; A Bash shell. Its frontmatter pre-approves these tools: Read, Bash, Write, Grep. Compatibility (from SKILL.md): opencode.
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Error Handler is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Error Handler: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Tsh Implementing Observability (TheSoftwareHouse/copilot-collections, 284 stars), Observability Sre Triage (elastic/agent-skills, 592 stars) and Logging Observability (getsentry/toolkit, 917 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
EliasOulkadi (a GitHub user) maintains it in EliasOulkadi/shokunin, which has 114 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 5, 2026.
Source: EliasOulkadi/shokunin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.