UModel Root Cause Analysis
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
A skill your agent uses when instrumenting a service from the inside so an incident can be explained from telemetry alone — wiring OpenTelemetry logs, metrics and traces, standing up a Collector…
$ npx skills add ericrisco/rsc-harness --skill observability -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness observability --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/observability .claude/skills/observability && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "observability" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/observability into .claude/skills/observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/observabilityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill observability -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness observability --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/observability .agents/skills/observability && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "observability" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/observability into .agents/skills/observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill observability -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness observability --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/observability .cursor/skills/observability && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "observability" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/observability into .cursor/skills/observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/observability--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill observability -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness observability --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/observability .gemini/skills/observability && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "observability" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/observability into .gemini/skills/observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness observabilityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill observability -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/observability .github/skills/observability && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "observability" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/observability into .github/skills/observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill observability -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness observability --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/observability .opencode/skills/observability && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "observability" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/observability into .opencode/skills/observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
observabilityA skill your agent uses when instrumenting a service from the inside so an incident can be explained from telemetry alone — wiring OpenTelemetry logs, metrics and traces, standing up a Collector…
Observability is an agent skill from ericrisco/rsc-harness. Use when instrumenting a service from the inside so an incident can be explained from telemetry alone — wiring OpenTelemetry logs, metrics and traces, standing up a Collector, exporting via OTLP, and defining telemetry-driven alerts. NOT outside-in uptime probes, on-call rotation, or who-gets-paged (that is monitoring).
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/collector-config.md`).
It sits in DevOps & Cloud, covering Observability and Incident response. It works with OpenTelemetry. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
npmnodepipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm and pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Observability loads about 3.8k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 1,327 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,327 words, ~3,753 tokens.
.claude/skills/observability/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.You are wiring the inside view of a service: when something breaks at 3am, an engineer must be able to answer "what happened, where, and why" from telemetry alone — without adding a console.log and redeploying into the fire. This skill emits a concrete artifact: SDK init code, instrumentation (spans/metrics/structured logs), a Collector config, and alert rules that the instrumentation makes possible. The outside-in half — is it up, who gets paged — is ../monitoring/SKILL.md.
Every signal carries the same correlation identity: trace_id, service.name, deployment.environment. A log line you cannot pivot to its trace, or a spiking metric you cannot pivot to an exemplar span, doubles your mean-time-to-resolution — you are back to grepping. Three signals that don't share keys are three disconnected tools; three that do are one queryable system. Set the resource once at SDK init, inject trace_id/span_id into every log, and never emit a metric you can't tie back to a service and environment.
Before choosing signals, write two to four questions on-call must answer during the likely incident. For example: “Are payment retries recovering?”, “Which dependency and failure class drives exhaustion?”, “Can one payment be charged twice?”, “Which customer-visible operations need intervention now?” Then assign the cheapest signal: metrics say that/how much, traces show where/causal path, logs explain why for this event. If a proposed event or label answers none of the questions, do not emit it.
Don't emit all three of everything. Each signal answers a different question at a different cost.
| Signal | Answers | Cost | Alert on it? | Main gotcha |
|---|---|---|---|---|
| Logs | "what exactly happened in this one event" | high per-event, cheap to skip | rarely (noisy) | high-cardinality fields belong in the body, not in stream labels |
| Metrics | "what's the rate/aggregate over time" | cheap, pre-aggregated | yes — this is your alert source | cardinality explosion if a label is unbounded |
| Traces | "what was the causal path across hops, and where did time go" | medium; sample it | indirectly (via derived RED metrics) | one giant span = no causality; sample or you pay for noise |
The "fourth pillar," continuous profiling (CPU/heap flame graphs over time), is now a first-class OTel signal — reach for it only when traces say "the time is inside this function" and you need to know which line.
Instrument once, route anywhere. The app talks OTLP to a Collector; the Collector fans the firehose out to backends.
┌─────────────┐ OTLP/gRPC :4317 ┌───────────────────────────┐
│ app + SDK │ OTLP/HTTP :4318 ───▶ │ OTel Collector │
│ (resource: │ /v1/traces │ receivers → processors │
│ service.name│ /v1/metrics │ → exporters (per signal) │
│ +env+ver) │ /v1/logs │ wired in service.pipelines│
└─────────────┘ └───────────────────────────┘
│ │ │
logs │ traces │ metrics│
▼ ▼ ▼
Loki Tempo Mimir/Prom (+ Grafana to view)
└── or a single vendor: Datadog / Honeycomb ──┘OTLP is the wire format: gRPC on 4317 (TLS + gzip by default), HTTP on 4318 with per-signal paths /v1/traces, /v1/metrics, /v1/logs. Always export to a Collector, never straight to the vendor. Why: the Collector gives you one place to batch (fewer round-trips), retry (survive a backend blip), redact PII, and swap or add a backend without redeploying the app. App SDKs should be dumb pipes; policy lives in the Collector.
Never hand-roll a span for something auto-instrumentation already covers (HTTP servers, DB clients, queues). You will miss edges and waste effort. Turn on zero-code instrumentation, confirm traces flow, then add manual spans only where your business logic lives.
# Node — zero-code, no app changes. SDK is 2.0+; the register hook is compatible.
npm i @opentelemetry/api @opentelemetry/auto-instrumentations-node
OTEL_SERVICE_NAME=checkout-api \
OTEL_RESOURCE_ATTRIBUTES=deployment.environment=prod,service.version=1.4.2 \
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \
node --require '@opentelemetry/auto-instrumentations-node/register' app.js# Python — zero-code via the launcher; it patches known libraries on import.
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install
OTEL_SERVICE_NAME=checkout-api \
OTEL_RESOURCE_ATTRIBUTES=deployment.environment=prod,service.version=1.4.2 \
opentelemetry-instrument python app.pyThen add a manual span only around a meaningful business operation — and give it attributes and a status, or it tells you nothing:
// Bad — a span with no attributes and no status. You learn that "something ran."
const span = tracer.startSpan('work');
await chargeCard(order);
span.end();// Good — named for the business op, carries the inputs you'd filter by, records outcome.
const { trace, SpanStatusCode } = require('@opentelemetry/api');
const tracer = trace.getTracer('checkout');
await tracer.startActiveSpan('charge_card', async (span) => {
span.setAttribute('order.id', order.id); // searchable dimension
span.setAttribute('payment.provider', 'stripe');
try {
await chargeCard(order);
span.setStatus({ code: SpanStatusCode.OK });
} catch (err) {
span.recordException(err); // attaches the stack as an event
span.setStatus({ code: SpanStatusCode.ERROR, message: err.message });
throw err;
} finally {
span.end(); // a span you never end leaks forever
}
});Per-stack init (Node SDK 2.0 manual setup, Go SDK, context propagation across HTTP/queue hops, the GenAI span template) lives in references/instrumentation-recipes.md.
Use the standard attribute names; never invent bespoke keys. The whole correlation story and every prebuilt backend dashboard assume http.*, db.*, gen_ai.*, service.*. A homegrown mycompany.endpoint attribute is invisible to every tool that expects http.route.
service.name (required — unset means telemetry lands as unknown_service), service.version, deployment.environment.k8s.*) reached release candidate (2026-03); DB conventions are on their 2nd RC. Prefer the standard names even while an area is still stabilizing.gen_ai.* so the same span feeds your spend view:// LLM call as a span: model + token attrs. These attributes feed ../cost-tracking/SKILL.md.
await tracer.startActiveSpan('chat gpt-4o', async (span) => {
span.setAttribute('gen_ai.system', 'openai');
span.setAttribute('gen_ai.request.model', 'gpt-4o');
const res = await openai.chat.completions.create({ /* ... */ });
span.setAttribute('gen_ai.usage.input_tokens', res.usage.prompt_tokens);
span.setAttribute('gen_ai.usage.output_tokens', res.usage.completion_tokens);
span.end();
});This skill emits the signal (token attributes on a span). Turning those tokens into a dollar figure and a budget is ../cost-tracking/SKILL.md.
Logs are JSON, carry the trace context, and never carry PII. Free-text print lines cost you twice: you can't query them, and you can't jump from the log to its trace.
# Bad — unstructured, unsearchable, unlinkable to a trace.
print("charged user " + email + " amount " + str(amount))# Good — structured, leveled, correlated, no PII (id not email).
import logging, json
from opentelemetry import trace
def log_charge(order_id, amount):
ctx = trace.get_current_span().get_span_context()
logging.info(json.dumps({
"event": "charge.succeeded",
"level": "info",
"order_id": order_id, # an opaque id, not the customer's email
"amount_cents": amount,
"trace_id": format(ctx.trace_id, "032x"), # ← the pivot back to the trace
"span_id": format(ctx.span_id, "016x"),
}))Level discipline: error = a human should look, warn = degraded but handled, info = business milestones, debug = off in prod. If everything is error, nothing is.
RED for request-driven services, USE for finite resources. Alerts come from metrics — logs and traces are for investigating the alert, not firing it.
The cardinality rule — this is the #1 way an observability stack falls over. A metric's total series count is the product of its label cardinalities. Put an unbounded value on a label and you create a near-infinite series count; the time-series database (Prometheus/Mimir) is most often restarted because of exactly this. Loki indexes labels only, not log contents — so the same rule binds its stream labels.
# Bad — user_id is unbounded; 5M users = 5M series per metric. OOMs the TSDB.
http_requests_total{route="/checkout", user_id="u_8f3a...", status="200"}# Good — only bounded, low-cardinality dimensions on the metric.
http_requests_total{route="/checkout", method="POST", status="200"}
# Need to slice by user? That's a trace attribute or a log field, never a metric label.A minimal valid pipeline. An exporter is inert until it appears in a pipeline — defining one under exporters: does nothing on its own.
# otel-collector.yaml — Collector v0.153.0 shape
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
processors:
memory_limiter: # first line of defense: shed load before OOM
check_interval: 1s
limit_percentage: 80
batch: {} # batch before export — fewer, bigger round-trips
redaction: # strip PII before it leaves your network
allow_all_keys: true
blocked_values: ["[0-9]{13,16}", "\\b[\\w.]+@[\\w.]+\\b"] # PANs, emails
exporters:
otlphttp/traces: { endpoint: http://tempo:4318 }
otlphttp/logs: { endpoint: http://loki:3100/otlp }
otlphttp/metrics: { endpoint: http://mimir:9009/otlp }
service:
pipelines:
traces: { receivers: [otlp], processors: [memory_limiter, redaction, batch], exporters: [otlphttp/traces] }
logs: { receivers: [otlp], processors: [memory_limiter, redaction, batch], exporters: [otlphttp/logs] }
metrics: { receivers: [otlp], processors: [memory_limiter, batch], exporters: [otlphttp/metrics] }Tail sampling, gateway-vs-agent topology, multi-backend fan-out (LGTM and a vendor in parallel), and resourcedetection live in references/collector-config.md. Shipping the Collector as a container or in CI is ../docker/SKILL.md.
The instrumentation above makes these possible — defining them is your job; routing the page to a human is ../monitoring/SKILL.md.
| Anti-pattern | Why it bites | Do instead |
|---|---|---|
Unbounded label (user_id, request_id, email) on a metric or Loki stream | Cardinality explosion → TSDB OOM/restart, cost blowup | Keep labels low-cardinality; put the high-cardinality field on a span attribute or log body |
| Logging PII / secrets (email, card, token) | Compliance breach + the leak is now in every log backend | Log opaque ids; redact in the Collector before export |
| 100% trace sampling in prod, no policy | Pay to store noise; backend throttles and drops the traces you needed | Head/tail sampling — keep all errors + slow traces, sample the rest |
| App exports straight to the vendor, no Collector | Can't batch, retry, redact, or swap backends without a redeploy | Always route through a Collector |
| Instrument everything before deciding the question | Noise with no signal; nobody opens the dashboard | Start from "what would I ask during an incident," instrument that path |
| Alert on a raw error count | Fires on traffic spikes, silent during a low-traffic outage | Alert on error rate / SLO burn |
| One giant span per request (or a thousand contentless ones) | No causality, or context with no detail — both useless | Span per meaningful operation, each with attributes + status |
Run scripts/verify.sh against the directory holding your Collector config + SDK init. It checks the config is valid, that every defined exporter is actually wired into a pipeline (the classic "defined but unused" footgun), that service.name is set, and warns on high-cardinality metric labels. It is read-only and exits 0 when there's nothing to check.
Then prove the wire with one safe induced failure in a test/staging path. Confirm the expected error-rate metric changes, the trace records the failing operation and error status, and the structured log carries the same trace_id without PII. Query the backend/Collector output; “the instrumentation code ran” is not evidence that usable telemetry arrived. Live paging and escalation proof remains ../monitoring/SKILL.md.
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/observability of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Observability this skillericrisco/rsc-harness | 156 | — | ~3.8k | Automated safety check: Pass | MIT | |
| UModel Root Cause Analysisalibaba/UnifiedModel | 412 | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Monitoring EngineerFerroxLabs/wayland | 608 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| Kubernetes Network Root Cause Analysiskubeshark/kubeshark | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | |
| Motel Debugkitlangton/motel | 298 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Tempsgotempsh/temps | 822 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 |
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
FerroxLabs/wayland
Observability and monitoring. An agent skill from FerroxLabs/wayland.
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
kitlangton/motel
Debug applications with motel, a local OpenTelemetry ingest and query server.
gotempsh/temps
Manage, deploy, operate, and instrument applications with Temps.
openclaw/clawhub
Explores and queries OpenTelemetry metrics in Axiom MetricsDB, listing datasets, metrics and tags first and picking the right aggregation for each metric's type.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when instrumenting a service from the inside so an incident can be explained from telemetry alone — wiring OpenTelemetry logs, metrics and traces, standing up a Collector…. Observability is an agent skill from ericrisco/rsc-harness. Use when instrumenting a service from the inside so an incident can be explained from telemetry alone — wiring OpenTelemetry logs, metrics and traces, standing up a Collector, exporting via OTLP, and defining telemetry-driven alerts.
Observability fits situations like: instrumenting a service from the inside so an incident can be explained from telemetry alone — wiring OpenTelemetry logs; metrics and traces; standing up a Collector; exporting via OTLP.
Run `npx skills add ericrisco/rsc-harness --skill observability -a claude-code`. Or copy the skill folder (skills/observability in ericrisco/rsc-harness) into .claude/skills/observability in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill observability -a codex`. Or copy the skill folder (skills/observability in ericrisco/rsc-harness) into .agents/skills/observability in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability, .gemini/skills/observability, .github/skills/observability and .opencode/skills/observability in your project.
Going by SKILL.md and its folder, Observability needs a shell for the scripts in its folder and the command-line tools its instructions call (npm, node and pip). Our summary lists: Python 3; Node.js; A Bash shell; Docker.
SKILL.md contains no URLs. Its commands use npm and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Observability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Observability: UModel Root Cause Analysis (alibaba/UnifiedModel, 412 stars), Monitoring Engineer (FerroxLabs/wayland, 608 stars), Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars) and Motel Debug (kitlangton/motel, 298 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.