Agent skill

Triage Agent Events

by netdata in netdata/netdata

Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries.

GPL-3.0Auto-check: notesDevOps & Cloud

Install Triage Agent Events

skills CLI
$ npx skills add netdata/netdata --skill triage-agent-events -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install netdata/netdata triage-agent-events --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/netdata/netdata.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/triage-agent-events .claude/skills/triage-agent-events && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triage-agent-events
GitHub stars
81k
Token cost
~2.4k tokens
SKILL.md length
990 words
Files
18 (incl. scripts)
Skills in repo
27
Repo updated
First seen
Licence
GPL-3.0

At a glance

Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries.

  • Works in 7 steps: The dataset: fleet-scale status events,… → Index-friendly queries (HARD RULE): use… → Three transports (priority order) → …
  • Restart/dedup timing
  • SKILL.md covers Choose the task, Why this skill exists, Workflow and Key concepts (read first), plus 6 more sections
  • Runs Shell and Python scripts from its folder; needs NETDATA_CLOUD_TOKEN

What it does

Triage Agent Events is an agent skill from netdata/netdata. Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries. Use for AE fields, restart/dedup timing, structured filters, version comparisons and reviews of these investigation helpers. Ordinary logs use the Agent/Cloud query skills.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including scripts (for example `AE_FIELDS.md`, `finding-crashes.md` and `finding-fatals.md`).

It sits in DevOps & Cloud. The repository describes itself as: The fastest path to AI-powered full stack observability, even for lean teams. The licence is GPL-3.0.

When your agent uses it

  • Restart/dedup timing
  • Structured filters
  • Version comparisons and reviews of these investigation helpers

Example prompts

  • “/triage-agent-events”

Requirements

  • Python 3
  • A Bash shell
  • A credential in NETDATA_CLOUD_TOKEN

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. The dataset: fleet-scale status events, not a complete
  2. Index-friendly queries (HARD RULE): use multi-value
  3. Three transports (priority order)
  4. After-the-fact event model: agents POST events ONLY on
  5. 23h client-side dedup (src/daemon/status-file-dedup.c:11)
  6. Default time + version filters: 24h time window;
  7. AE_* field naming: every JSON path in the producer's

What it can do on your machine

Read from SKILL.md and the folder at commit 9fe30d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Shell and Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NETDATA_CLOUD_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triage Agent Events loads about 2.4k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 990 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:173
    and the synthetic self-test do not load `.env`.
  • NoteMentions a .env fileSKILL.md:183
    All values live in `<repo>/.env` (gitignored). See

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from netdata/netdata at commit 9fe30d9, republished under its GPL-3.0 licence (© netdata). 990 words, ~2,393 tokens.

Download SKILL.mdSave it as .claude/skills/triage-agent-events/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
triage-agent-events
description
Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries. Use for AE_* fields, restart/dedup timing, structured filters, version comparisons and reviews of these investigation helpers. Ordinary logs use the Agent/Cloud query skills.

triage-agent-events

Private developer skill for triaging crashes, panics, and fatals across the Netdata fleet. Reads the agent-events systemd-journal namespace via the Netdata systemd-journal Function (Cloud-proxied or direct-agent transport) and ships scripts that bake in index-friendly query patterns.

Choose the task

TaskRead or use
Analyze a supplied captureAE_FIELDS.md, relevant crash/fatal/recipe guidance, analyze-events.sh --input PATH; no credential setup
Explain fields, timing or query constructionAE_FIELDS.md, update-cadence.md, query-discipline.md; transport reference only as needed
Fetch live evidenceSelected transport and query discipline, then configured environment and get-events.sh within existing authorization
Review helper or investigation changesAffected contracts/source and existing tests; examples do not authorize live queries or bug fixes

The no-leak self-test uses synthetic transport in a subshell, without environment loading or network access. It checks Cloud/direct dispatch success and visible request masking; arbitrary event response contents remain private evidence.

Why this skill exists

Historical observations found 40k-200k daily status events and a fleet around 1.5M agents. These are not current size guarantees; the dataset can be large and noisy (many unupdated agents report crashes that have been fixed). Naive "grep all" queries are slow and wasteful. This skill teaches the maintainer (and any AI assistant helping them) how to slice the dataset efficiently and how to interpret what comes back.

Workflow

+-------------------------+        +---------------------+
|  get-events.sh          |  -->   |  <run>.json         |
|  (cloud or agent API)   |        |  in .local/audits   |
+-------------------------+        +---------------------+
                                          |
                                          v
                              +------------------------+
                              |  analyze-events.sh     |
                              |  --by signal|version|  |
                              |       function|...     |
                              +------------------------+
                                          |
                                          v
                              +------------------------+
                              |  cluster + read source |
                              |  + report the finding  |
                              +------------------------+

The skill is a bug-investigation tool, not a generic logs query tool. The two existing query-netdata-cloud and query-netdata-agents skills already cover transport mechanics; this skill EXTENDS them with the agent-events specifics (what fields are present, what predicates are index-friendly, what each enum value means for triage).

Key concepts (read first)

  1. The dataset: fleet-scale status events, not a complete census of crashes at occurrence time. Naive full-namespace queries with bare FTS are slow.

  2. Index-friendly queries (HARD RULE): use multi-value field filters FIRST. The Netdata systemd-journal plugin supports the syntax:

    (FIELD1 in A, B, C) AND (FIELD2 in D, E, F) AND ...

    Between fields = AND. Between values = OR. This is a facet-engine feature, NOT raw journalctl. Use FTS via query= only as a residual narrower over the structured slice. See query-discipline.md.

  3. Three transports (priority order):

    • Cloud API -- proxied through Netdata Cloud at the agent-events space. Primary for the team.
    • Direct agent API -- against the agent-events node's /api/v3/function?function=systemd-journal. Primary for scripts.
    • ssh to the host -- operator-only path; mentioned in transports.md but no scripted ssh transport.
  4. After-the-fact event model: agents POST events ONLY on start (the previous session's exit reason). They commit status to disk on start, stop, and at most every 10 minutes. So the meaningful query unit is "events posted in the last 24 hours"; "the last hour" misses real crashes that haven't restarted yet.

  5. 23h client-side dedup (src/daemon/status-file-dedup.c:11): same agent + same event-content hash within 23h -> suppressed at the producer. So 1 record per agent per event-signature per day is the natural unit. Different agents posting the same crash signature -> both arrive (server does not dedup).

  6. Default time + version filters: 24h time window; highest numeric stable + up to three nightlies observed in the discovery response, not the published release catalog. This focuses triage on bugs that still matter. Wide windows (--since '7d ago' or longer) are reserved for rare crashes (1-per-few-days class) and for "when did this start / get fixed" investigations.

  7. AE_ field naming*: every JSON path in the producer's status document becomes an AE_-prefixed journal field (per log2journal --prefix 'AE_' on the ingestion server). See AE_FIELDS.md for the verified map and enum meanings.

get-events.sh fetches one page (default 500 rows), not a paginated census. Before count or absence claims, inspect status, partial/sampling flags and matched/returned limits as in ./how-tos/trace-stack-symbol-regression-to-mutator.md#3-prove-that-the-response-is-complete. Narrow or paginate through the transport API when needed. Client-side version regexes change rows only; facets/totals still describe the server response. Prefer an explicit --input when analyzing a particular run; the default latest-file choice is a convenience.

Show full SKILL.md (370 more words)Show less

Table of contents

DocPurpose
AE_FIELDS.mdVerified field map (~80 rows) + enum meanings for triage. Indispensable.
transports.mdCloud API + direct agent API call patterns; ssh footnote.
update-cadence.mdAfter-the-fact model, dedup, push timing, disk commits, query implications.
query-discipline.mdThe multi-value filter syntax, structured-filters-first rule, anti-patterns.
finding-crashes.mdRecipe: signal crashes (SIGSEGV / SIGBUS / SIGFPE / SIGABRT) on stable.
finding-fatals.mdRecipe: deliberate fatals (OOM, disk full, asserts).
recipes/INDEX.mdLive catalog of recipes (find-by-function, find-by-version, find-related-to-work).
how-tos/INDEX.mdCatalog of reusable investigation how-tos.

Knowledge Capture

Capture timing and authorization follow AGENTS.md#knowledge-capture. Investigation recipes live in how-tos/ and are listed in ./how-tos/INDEX.md; check the per-domain guides and recipes/ before adding another recipe.

Scripts (in scripts/)

ScriptPurpose
_lib.shHelpers (agentevents_* prefix). Sources query-netdata-agents/scripts/_lib.sh. Token-safe; ships a no-leak self-test.
get-events.shFetch events of interest. Index-friendly defaults. Fresh private JSON output to .local/audits/query-agent-events/; explicit output paths must be new.
analyze-events.shGroup-by stats over a downloaded dump (signal, version, fatal_function, architecture, etc.).
redact-events.shOpt-in redaction (machine_guid / claim_id / host_id / ephemeral_id -> placeholders). For sharing only.

Path discipline

This skill follows <repo>/.agents/sensitive-data-discipline.md:

  • Repo files: repo-relative (<repo>/src/...).
  • Sibling Netdata-org repos: ${NETDATA_REPOS_DIR}/<repo>/....
  • agent-events host / namespace / machine GUID / node ID: ALWAYS via env keys. Never literal values in any committed file.
  • Producer ingest URL: NEVER quoted literally. Reference only as src/daemon/status-file.c:988.
  • Fetched event payloads land under <repo>/.local/audits/query-agent-events/<run>.json (gitignored). Do NOT paste raw event JSON into committed artifacts.

Required env keys

get-events.sh requires Bash 4 or later for associative arrays; select a compatible Bash on systems with an older default. Only live get-events.sh calls load these settings. Its current loader requires all listed values for either transport; offline analysis/redaction, explanation, source review and the synthetic self-test do not load .env.

KeyRole
NETDATA_CLOUD_TOKENCloud REST token (long-lived).
NETDATA_CLOUD_HOSTNAMECloud REST API host.
AGENT_EVENTS_HOSTNAMEDual-duty: ssh host AND direct-HTTP host of the ingestion node. Can be IP or DNS name. NOT the journalctl namespace (hardcoded agent-events); NOT the Cloud room name (also hardcoded agent-events).
AGENT_EVENTS_MACHINE_GUIDAgent machine GUID for direct-agent transport.
AGENT_EVENTS_NODE_IDCloud node UUID for cloud-proxy transport.

All values live in <repo>/.env (gitignored). See <repo>/.agents/ENV.md for setup (where each value comes from, sample formats, common mistakes).

  • query-netdata-cloud -- transport: Cloud REST API.
  • query-netdata-agents -- transport: direct agent REST + bearer auto-mint.
  • This skill consumes both via their _lib.sh helpers.

© netdata, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts) in .agents/skills/triage-agent-events of netdata/netdata.

  • SKILL.md
  • AE_FIELDS.md
  • finding-crashes.md
  • finding-fatals.md
  • how-tos/INDEX.md
  • how-tos/trace-stack-symbol-regression-to-mutator.md
  • query-discipline.md
  • recipes/INDEX.md
  • recipes/find-by-function.md
  • recipes/find-by-version.md
  • recipes/find-related-to-work.md
  • scripts/_lib.sh
  • scripts/analyze-events.sh
  • scripts/get-events.sh
  • scripts/redact-events.sh
  • tests/test_helpers.py
  • transports.md
  • … and 1 more

Open the folder on GitHubat commit 9fe30d9

Compare with similar skills

Triage Agent Events next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triage Agent Events compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triage Agent Events this skillnetdata/netdata81k—~2.4kAutomated safety check: NotesGPL-3.0
Monitor CInrwl/nx29k5 repos~4.7kAutomated safety check: PassMIT
Terraform and OpenTofu Guideagentscope-ai/QwenPaw35k6 repos~4.2kAutomated safety check: PassApache-2.0
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Analyze GitHub Action Logswithastro/astro63k1 repos~1.3kAutomated safety check: PassCustom licence
Iron Proxy Gateway for NanoClawnanocoai/nanoclaw31k—~4.6kAutomated safety check: NotesMIT

Similar skills

  • Monitor CI

    nrwl/nx

    Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.

    29k GitHub starsUsed in 5 repos~4.7k tokens
    DevOps & CloudAuto-check passed
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    35k GitHub starsUsed in 6 repos~4.2k tokens
    DevOps & CloudAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Official

    Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.

    63k GitHub starsUsed in 1 repo~1.3k tokens
    DevOps & CloudAuto-check passed
  • Installs or refreshes Iron Proxy and its Iron Control web console for NanoClaw, with a local Docker setup, database, credentials and a human approval bridge.

    31k GitHub stars~4.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Terraform Skill

    antonbabenko/terraform-skill

    A skill your agent uses when writing, reviewing, or debugging Terraform/OpenTofu modules, tests, CI, scans, or state ops - diagnoses failure mode (identity churn, secrets, blast radius, CI drift…

    2.4k GitHub starsUsed in 1 repo~5.1k tokens
    DevOps & CloudAuto-check passed

More from netdata/netdata

All 27 skills in this repo
  • Docs Learn PR Preview

    netdata/netdata

    Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.

    81k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Repo Mirror Sources

    netdata/netdata

    Inspect Netdata-org source checkouts under NETDATAREPOSDIR, or set up and synchronize that mirror when requested.

    81k GitHub stars~1.2k tokensUpdated today
    Auto-check: notes
  • Triage Codacy

    netdata/netdata

    Inspect, analyze, troubleshoot, or review Codacy findings and local analyzer/API helpers.

    81k GitHub stars~2.2k tokensUpdated today
    Auto-check: notes
  • Triage Coverity

    netdata/netdata

    Inspect or review Coverity Scan defects and saved CID bundles; fetch live findings or apply verified triage decisions when requested.

    81k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Triage Sonarqube

    netdata/netdata

    Inspect, review, or apply authorized triage decisions to SonarCloud issues and security hotspots; also review the Sonar helpers.

    81k GitHub stars~2.8k tokensUpdated today
    Auto-check: notes
  • Create, review or validate Netdata Prometheus chart profiles, exporter dashboard design, collection policy and stock semantic proofs.

    81k GitHub stars~4.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Triage Agent Events

What does Triage Agent Events do?

Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries. Triage Agent Events is an agent skill from netdata/netdata. Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries.

When should I use Triage Agent Events?

Triage Agent Events fits situations like: restart/dedup timing; structured filters; version comparisons and reviews of these investigation helpers.

How do I install Triage Agent Events in Claude Code?

Run `npx skills add netdata/netdata --skill triage-agent-events -a claude-code`. Or copy the skill folder (.agents/skills/triage-agent-events in netdata/netdata) into .claude/skills/triage-agent-events in your project. Claude Code loads it when a task matches its description.

How do I install Triage Agent Events in Codex?

Run `npx skills add netdata/netdata --skill triage-agent-events -a codex`. Or copy the skill folder (.agents/skills/triage-agent-events in netdata/netdata) into .agents/skills/triage-agent-events in your project. Codex loads it when a task matches its description.

Can I use Triage Agent Events in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add netdata/netdata --skill triage-agent-events -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage-agent-events, .gemini/skills/triage-agent-events, .github/skills/triage-agent-events and .opencode/skills/triage-agent-events in your project.

What does Triage Agent Events need to run?

Going by SKILL.md and its folder, Triage Agent Events needs a shell and Python for the scripts in its folder and credentials named NETDATA_CLOUD_TOKEN. Our summary lists: Python 3; A Bash shell; A credential in NETDATA_CLOUD_TOKEN.

Does Triage Agent Events access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triage Agent Events safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Triage Agent Events use?

Triage Agent Events is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triage Agent Events use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triage Agent Events?

Skills that share tags, products or a category with Triage Agent Events: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 35k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triage Agent Events?

netdata (a GitHub organization) maintains it in netdata/netdata, which has 80,820 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 7, 2026.

Source: netdata/netdata on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.