Agent skill

Observability Triage

by every-app in every-app/open-seo

Triage OpenSEO production errors in Cloudflare Workers Observability — verified query recipes, counting gotchas, and a known-noise filter list applied automatically.

MITAuto-check passedDevOps & Cloud

Install Observability Triage

skills CLI
$ npx skills add every-app/open-seo --skill observability-triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install every-app/open-seo observability-triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/every-app/open-seo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/observability-triage .claude/skills/observability-triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
observability-triage
GitHub stars
23k
Token cost
~1.7k tokens
SKILL.md length
760 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
MIT

At a glance

Triage OpenSEO production errors in Cloudflare Workers Observability — verified query recipes, counting gotchas, and a known-noise filter list applied automatically.

  • Works in 2 steps: SAM chat Durable Object lifecycle (close… → Network connection lost.
  • Asked to review Cloudflare logs/observability
  • SKILL.md covers Access, Query recipes (verified shapes), Counting gotchas and Known noise — filter these…, plus 1 more section
  • Calls npx

What it does

Observability Triage is an agent skill from every-app/open-seo. Triage OpenSEO production errors in Cloudflare Workers Observability — verified query recipes, counting gotchas, and a known-noise filter list applied automatically. Use when asked to review Cloudflare logs/observability, count OOMs or worker errors, compare error rates between periods, or investigate a prod error spike.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Observability. It works with Cloudflare Workers, Cloudflare and Model Context Protocol. The repository describes itself as: Open source alternative to Semrush and Ahrefs. The licence is MIT.

When your agent uses it

  • Asked to review Cloudflare logs/observability
  • Compare error rates between periods
  • Investigate a prod error spike

Example prompts

  • “/observability-triage”

Requirements

  • Node.js

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. SAM chat Durable Object lifecycle (close code 1006)
  2. Network connection lost.

What it can do on your machine

Read from SKILL.md and the folder at commit deb4491. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Observability Triage loads about 1.7k tokens when it runs. Until then it costs about 86 tokens; SKILL.md has 760 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from every-app/open-seo at commit deb4491, republished under its MIT licence (© every-app). 760 words, ~1,735 tokens.

Download SKILL.mdSave it as .claude/skills/observability-triage/SKILL.md (or your agent's skills folder).
name
observability-triage
description
Triage OpenSEO production errors in Cloudflare Workers Observability — verified query recipes, counting gotchas, and a known-noise filter list applied automatically. Use when asked to review Cloudflare logs/observability, count OOMs or worker errors, compare error rates between periods, or investigate a prod error spike.
metadata.internal
true

Observability triage

Query Cloudflare Workers Observability for prod errors, count them correctly, and skip the noise that has already been investigated to a dead end. Apply the known-noise list below without re-investigating those entries.

Access

  • Resolve the account at runtime — never hardcode it: the cloudflare-api MCP server pre-binds accountId in mcp__cloudflare-api__execute, and npx wrangler whoami prints it. The workers to triage are the ones this repo deploys (see deploy/alchemy/alchemy.run.ts): the main app worker plus the aux workers (audit engine, landing, self-host).
  • Query via the cloudflare-api MCP server (mcp__cloudflare-api__execute). If its tools are absent, run its authenticate flow and give the user the URL — wrangler's OAuth token gets a 403 on the observability API (missing scope), so don't burn time on curl-with-wrangler-token.
  • PostHog is the second error source but cannot see exceededMemory / canceled / responseStreamDisconnected outcomes — worker-outcome questions are answerable only here.

Query recipes (verified shapes)

POST /accounts/{account_id}/workers/observability/telemetry/query. All of these are load-bearing; the API's 400s are opaque:

  • queryId: "adhoc" is required. Timeframe is epoch milliseconds: timeframe: { from, to }.

  • Count invocations by outcome:

    json
    {
      "queryId": "adhoc",
      "timeframe": { "from": 0, "to": 0 },
      "parameters": {
        "datasets": ["cloudflare-workers"],
        "filters": [
          { "key": "$metadata.type", "operation": "eq", "value": "cf-worker-event", "type": "string" }
        ],
        "calculations": [{ "operator": "count", "alias": "count" }],
        "groupBys": [
          { "value": "$workers.scriptName", "type": "string" },
          { "value": "$workers.outcome", "type": "string" }
        ],
        "limit": 100
      },
      "view": "calculations",
      "limit": 100
    }
  • The cf-worker-event filter is what makes counts mean invocations; without it you count every log line.

  • Raw samples: view: "events", empty calculations/groupBys, small limit. Events carry $metadata (service, trigger, level, message, fingerprint, requestId) and $workers (outcome, scriptName, scriptVersion, wallTimeMs, cpuTimeMs, eventType).

  • Error-level app logs grouped by message: filter $metadata.level eq error, groupBy $metadata.message.

  • Verified filter operations: eq, neq. Percentiles: the API wants "median", not "p50". Filter for substrings client-side on fetched events.

  • Useful drill-downs: groupBy $metadata.trigger (route), $workers.scriptVersion (did a deploy change the rate mid-window), $workers.event.request.headers.user-agent.

Counting gotchas

  • Grouped results are unsorted and effectively capped (~10 rows returned regardless of limit). A missing group ≠ zero. To see error outcomes, add $workers.outcome neq ok instead of hoping the error rows make the cut; sort client-side.
  • One isolate death fans out. An OOM kills every request pinned to the isolate at the same instant — cluster raw OOM events by timestamp before reading the count as user impact.
  • Message prefixes fragment groups. Logs with leading timestamps (better-auth's format) split one error into N single-count groups; grouped-by-message counts badly understate them. Sample events and merge client-side.
  • Prod deploys are manual — main being fixed doesn't mean prod runs the fix. Check $workers.scriptVersion and the scripts' modified_on before concluding a fix didn't work.

Known noise — filter these out, do not re-investigate

Entries land here only after an investigation proved there is no first-party emit site to fix or demote. Each keeps the one condition that would make it real signal again.

Show full SKILL.md (344 more words)Show less
1. SAM chat Durable Object lifecycle (close code 1006)
  • Messages (one phenomenon, counted three ways): Connection closed: this Durable Object instance is no longer active. Reconnect or retry the request. (hibernation/eviction), Durable Object reset because its code was updated. (deploy), plus the paired invocation summary whose $metadata.error is close.
  • Identify by: eventType: "hibernatableWebSocket", entrypoint SamChatAgent/OnboardingChatAgent, webSocketType: "close", code: 1006, wasClean: false, outcome: "exception", single-digit wallTimeMs, cpuTimeMs: 0, no stack. Fingerprints 3aa4cac26653d09a0a41100a33d413ae (exception), 0ae15457af49b4d9a117eeecf66b040a (summary).
  • Why unfixable: the DO is destroyed under the JS — its IoContext is already aborted when workerd delivers webSocketClose, so the first await never settles. partyserver already try/catches the whole close path; an onClose override would catch nothing.
  • Nothing breaks: transcripts persist per message in DO SQLite, PartySocket reconnects unconditionally, and credit metering (onChatResponse) never fires on an aborted turn.
  • Real signal: a sustained rise that does not correlate with a deploy (would mean mid-conversation evictions beyond hibernation).
2. Network connection lost.
  • Identify by: fingerprint be89d4ff7a64cb4d4dceae0f51cfe708, or the message verbatim. Runtime-generated shape: source.level absent, source.exception present, $metadata.origin: "fetch" — the opposite of every app log (source.level present, no exception).
  • Why unfixable: workerd's own record of a client disconnecting mid-stream; the exception never enters app code — the sibling request event for the same requestId has outcome: "ok". No emit site exists; wrangler.jsonc observability config has no per-message filter. Do not add try/catch around the stream handlers (dead ceremony).
  • Real signal: if the /mcp share of this group grows, treat it as a tool-call-latency symptom (clients timing out and cancelling) and route it to the /mcp performance track — not to this log group.
Runtime-vs-app litmus test

Before investigating any unfamiliar error event: source.level absent + source.exception present ⇒ the Workers runtime wrote it, not the app. There is no call site to grep for; judge it by the sibling request's outcome.

Adding an entry

Add to this list only after establishing there is no first-party emit site (grep the message; check the runtime-vs-app litmus above) and nothing is left in a wrong state. Every entry must include identify-by markers (fingerprint if stable) and its real-signal condition.

© every-app, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/observability-triage of every-app/open-seo.

Open the folder on GitHubat commit deb4491

Compare with similar skills

Observability Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Observability Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Observability Triage this skillevery-app/open-seo23k—~1.7kAutomated safety check: PassMIT
Prepare Cloudflare Production DeploymentLubomirGeorgiev/cloudflare-workers-nextjs-saas-template786—~5.9kAutomated safety check: NotesMIT
Debug Cloudflare Workers ObservabilityConsensys/c0105—~1.2kAutomated safety check: NotesLGPL-3.0
Workers Best Practiceshodgef/apiker1276 repos~1.8kAutomated safety check: PassMIT
Agents SDKhodgef/apiker1273 repos~3kAutomated safety check: PassMIT
Codflow Setupbighadj22/codflow346—~7.8kAutomated safety check: NotesApache-2.0

Similar skills

  • Prepare Cloudflare Production Deployment

    LubomirGeorgiev/cloudflare-workers-nextjs-saas-template

    Source-of-truth runbook for preparing this Vinext Cloudflare Workers SaaS template for production deployment.

    786 GitHub stars~5.9k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Debug deployed Cloudflare Workers using the cfobservability MCP, Wrangler, D1/R2 state, repo evidence, and safe live reproduction.

    105 GitHub stars~1.2k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check: notes
  • Reviews and authors Cloudflare Workers code against production best practices.

    127 GitHub starsUsed in 6 repos~1.8k tokens
    DevOps & CloudAuto-check passed
  • Agents SDK

    hodgef/apiker

    Build AI agents on Cloudflare Workers using the Agents SDK. An agent skill from hodgef/apiker.

    127 GitHub starsUsed in 3 repos~3k tokens
    Backend & APIsAuto-check passed
  • Codflow Setup

    bighadj22/codflow

    Setup runbook for CodFlow — an AI agent following it authenticates with Cloudflare, creates the required resources (D1, R2, KV) in the developer's account, binds their real IDs into both…

    346 GitHub stars~7.8k tokensUpdated 2 days ago
    DevOps & CloudAuto-check: notes
  • Build, refactor, or review Effect v4 HttpApi services on Cloudflare Workers.

    105 GitHub stars~1.4k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from every-app/open-seo

All 19 skills in this repo
  • Papercuts

    every-app/open-seo

    Log genuine, recurring repository friction to .agents/PAPERCUTS.md — confusing setup, a flaky repo command or script, a misleading in-repo error, stale generated files, or a non-obvious gotcha that…

    23k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Evaluate Skill

    every-app/open-seo

    Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save.

    23k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check: notes
  • Simple Issue Description

    every-app/open-seo

    Turn a rough bug report, feature request, support note, or pull request into a short, plain-language issue focused on the problem and desired behavior.

    23k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Create Repo Skill

    every-app/open-seo

    Create or update a skill in this repository the right way — canonical home in .agents/skills, internal-vs-public marking, symlink mirroring into .claude/skills, and public docs registration for…

    23k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Deslop

    every-app/open-seo

    Remove AI writing patterns from prose so it reads like a person wrote it.

    23k GitHub stars~602 tokensUpdated yesterday
    Auto-check passed
  • SEO Report

    every-app/open-seo

    Write and save an OpenSEO report as one self-contained HTML page.

    23k GitHub stars~5.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Observability Triage

What does Observability Triage do?

Triage OpenSEO production errors in Cloudflare Workers Observability — verified query recipes, counting gotchas, and a known-noise filter list applied automatically. Observability Triage is an agent skill from every-app/open-seo. Triage OpenSEO production errors in Cloudflare Workers Observability — verified query recipes, counting gotchas, and a known-noise filter list applied automatically.

When should I use Observability Triage?

Observability Triage fits situations like: asked to review Cloudflare logs/observability; compare error rates between periods; investigate a prod error spike.

How do I install Observability Triage in Claude Code?

Run `npx skills add every-app/open-seo --skill observability-triage -a claude-code`. Or copy the skill folder (.agents/skills/observability-triage in every-app/open-seo) into .claude/skills/observability-triage in your project. Claude Code loads it when a task matches its description.

How do I install Observability Triage in Codex?

Run `npx skills add every-app/open-seo --skill observability-triage -a codex`. Or copy the skill folder (.agents/skills/observability-triage in every-app/open-seo) into .agents/skills/observability-triage in your project. Codex loads it when a task matches its description.

Can I use Observability Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add every-app/open-seo --skill observability-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-triage, .gemini/skills/observability-triage, .github/skills/observability-triage and .opencode/skills/observability-triage in your project.

What does Observability Triage need to run?

Going by SKILL.md and its folder, Observability Triage needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Observability Triage access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Observability Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Observability Triage use?

Observability Triage is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Observability Triage use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Observability Triage?

Skills that share tags, products or a category with Observability Triage: Prepare Cloudflare Production Deployment (LubomirGeorgiev/cloudflare-workers-nextjs-saas-template, 786 stars), Debug Cloudflare Workers Observability (Consensys/c0, 105 stars), Workers Best Practices (hodgef/apiker, 127 stars) and Agents SDK (hodgef/apiker, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Observability Triage?

every-app (a GitHub organization) maintains it in every-app/open-seo, which has 22,680 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 6, 2026.

Source: every-app/open-seo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.