Monitor CI
nrwl/nx
Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.
Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries.
$ npx skills add netdata/netdata --skill triage-agent-events -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install netdata/netdata triage-agent-events --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/netdata/netdata.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/triage-agent-events .claude/skills/triage-agent-events && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "triage-agent-events" agent skill from https://github.com/netdata/netdata/tree/master/.agents/skills/triage-agent-events into .claude/skills/triage-agent-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-agent-events", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/netdata/netdata/tree/master/.agents/skills/triage-agent-eventsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add netdata/netdata --skill triage-agent-events -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install netdata/netdata triage-agent-events --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/netdata/netdata.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/triage-agent-events .agents/skills/triage-agent-events && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "triage-agent-events" agent skill from https://github.com/netdata/netdata/tree/master/.agents/skills/triage-agent-events into .agents/skills/triage-agent-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-agent-events", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add netdata/netdata --skill triage-agent-events -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install netdata/netdata triage-agent-events --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/netdata/netdata.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/triage-agent-events .cursor/skills/triage-agent-events && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "triage-agent-events" agent skill from https://github.com/netdata/netdata/tree/master/.agents/skills/triage-agent-events into .cursor/skills/triage-agent-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-agent-events", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/netdata/netdata.git --path .agents/skills/triage-agent-events--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add netdata/netdata --skill triage-agent-events -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install netdata/netdata triage-agent-events --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/netdata/netdata.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/triage-agent-events .gemini/skills/triage-agent-events && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "triage-agent-events" agent skill from https://github.com/netdata/netdata/tree/master/.agents/skills/triage-agent-events into .gemini/skills/triage-agent-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-agent-events", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install netdata/netdata triage-agent-eventsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add netdata/netdata --skill triage-agent-events -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/netdata/netdata.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/triage-agent-events .github/skills/triage-agent-events && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "triage-agent-events" agent skill from https://github.com/netdata/netdata/tree/master/.agents/skills/triage-agent-events into .github/skills/triage-agent-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-agent-events", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add netdata/netdata --skill triage-agent-events -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install netdata/netdata triage-agent-events --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/netdata/netdata.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/triage-agent-events .opencode/skills/triage-agent-events && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "triage-agent-events" agent skill from https://github.com/netdata/netdata/tree/master/.agents/skills/triage-agent-events into .opencode/skills/triage-agent-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triage-agent-events", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
triage-agent-eventsInvestigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries.
Triage Agent Events is an agent skill from netdata/netdata. Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries. Use for AE fields, restart/dedup timing, structured filters, version comparisons and reviews of these investigation helpers. Ordinary logs use the Agent/Cloud query skills.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including scripts (for example `AE_FIELDS.md`, `finding-crashes.md` and `finding-fatals.md`).
It sits in DevOps & Cloud. The repository describes itself as: The fastest path to AI-powered full stack observability, even for lean teams. The licence is GPL-3.0.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 9fe30d9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Shell and Python, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NETDATA_CLOUD_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Triage Agent Events loads about 2.4k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 990 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
and the synthetic self-test do not load `.env`.All values live in `<repo>/.env` (gitignored). SeeAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from netdata/netdata at commit 9fe30d9, republished under its GPL-3.0 licence (© netdata). 990 words, ~2,393 tokens.
.claude/skills/triage-agent-events/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.Private developer skill for triaging crashes, panics, and
fatals across the Netdata fleet. Reads the agent-events
systemd-journal namespace via the Netdata systemd-journal
Function (Cloud-proxied or direct-agent transport) and ships
scripts that bake in index-friendly query patterns.
| Task | Read or use |
|---|---|
| Analyze a supplied capture | AE_FIELDS.md, relevant crash/fatal/recipe guidance, analyze-events.sh --input PATH; no credential setup |
| Explain fields, timing or query construction | AE_FIELDS.md, update-cadence.md, query-discipline.md; transport reference only as needed |
| Fetch live evidence | Selected transport and query discipline, then configured environment and get-events.sh within existing authorization |
| Review helper or investigation changes | Affected contracts/source and existing tests; examples do not authorize live queries or bug fixes |
The no-leak self-test uses synthetic transport in a subshell, without environment loading or network access. It checks Cloud/direct dispatch success and visible request masking; arbitrary event response contents remain private evidence.
Historical observations found 40k-200k daily status events and a fleet around 1.5M agents. These are not current size guarantees; the dataset can be large and noisy (many unupdated agents report crashes that have been fixed). Naive "grep all" queries are slow and wasteful. This skill teaches the maintainer (and any AI assistant helping them) how to slice the dataset efficiently and how to interpret what comes back.
+-------------------------+ +---------------------+
| get-events.sh | --> | <run>.json |
| (cloud or agent API) | | in .local/audits |
+-------------------------+ +---------------------+
|
v
+------------------------+
| analyze-events.sh |
| --by signal|version| |
| function|... |
+------------------------+
|
v
+------------------------+
| cluster + read source |
| + report the finding |
+------------------------+The skill is a bug-investigation tool, not a generic logs
query tool. The two existing query-netdata-cloud and
query-netdata-agents skills already cover transport
mechanics; this skill EXTENDS them with the agent-events
specifics (what fields are present, what predicates are
index-friendly, what each enum value means for triage).
The dataset: fleet-scale status events, not a complete census of crashes at occurrence time. Naive full-namespace queries with bare FTS are slow.
Index-friendly queries (HARD RULE): use multi-value
field filters FIRST. The Netdata systemd-journal plugin
supports the syntax:
(FIELD1 in A, B, C) AND (FIELD2 in D, E, F) AND ...Between fields = AND. Between values = OR. This is a
facet-engine feature, NOT raw journalctl. Use FTS via
query= only as a residual narrower over the structured
slice. See query-discipline.md.
Three transports (priority order):
/api/v3/function?function=systemd-journal. Primary for
scripts.transports.md but no scripted ssh transport.After-the-fact event model: agents POST events ONLY on start (the previous session's exit reason). They commit status to disk on start, stop, and at most every 10 minutes. So the meaningful query unit is "events posted in the last 24 hours"; "the last hour" misses real crashes that haven't restarted yet.
23h client-side dedup (src/daemon/status-file-dedup.c:11):
same agent + same event-content hash within 23h ->
suppressed at the producer. So 1 record per agent per
event-signature per day is the natural unit. Different
agents posting the same crash signature -> both arrive
(server does not dedup).
Default time + version filters: 24h time window;
highest numeric stable + up to three nightlies observed
in the discovery response, not the published release catalog. This
focuses triage on bugs that still matter. Wide windows
(--since '7d ago' or longer) are reserved for rare crashes
(1-per-few-days class) and for "when did this start /
get fixed" investigations.
AE_ field naming*: every JSON path in the producer's
status document becomes an AE_-prefixed journal field
(per log2journal --prefix 'AE_' on the ingestion server).
See AE_FIELDS.md for the verified map and enum meanings.
get-events.sh fetches one page (default 500 rows), not a paginated census. Before count or absence claims, inspect
status, partial/sampling flags and matched/returned limits as in
./how-tos/trace-stack-symbol-regression-to-mutator.md#3-prove-that-the-response-is-complete. Narrow or paginate through
the transport API when needed. Client-side version regexes change rows only; facets/totals still describe the server
response. Prefer an explicit --input when analyzing a particular run; the default latest-file choice is a convenience.
| Doc | Purpose |
|---|---|
AE_FIELDS.md | Verified field map (~80 rows) + enum meanings for triage. Indispensable. |
transports.md | Cloud API + direct agent API call patterns; ssh footnote. |
update-cadence.md | After-the-fact model, dedup, push timing, disk commits, query implications. |
query-discipline.md | The multi-value filter syntax, structured-filters-first rule, anti-patterns. |
finding-crashes.md | Recipe: signal crashes (SIGSEGV / SIGBUS / SIGFPE / SIGABRT) on stable. |
finding-fatals.md | Recipe: deliberate fatals (OOM, disk full, asserts). |
recipes/INDEX.md | Live catalog of recipes (find-by-function, find-by-version, find-related-to-work). |
how-tos/INDEX.md | Catalog of reusable investigation how-tos. |
Capture timing and authorization follow AGENTS.md#knowledge-capture. Investigation recipes live in how-tos/ and
are listed in ./how-tos/INDEX.md; check the per-domain guides and recipes/ before adding another recipe.
| Script | Purpose |
|---|---|
_lib.sh | Helpers (agentevents_* prefix). Sources query-netdata-agents/scripts/_lib.sh. Token-safe; ships a no-leak self-test. |
get-events.sh | Fetch events of interest. Index-friendly defaults. Fresh private JSON output to .local/audits/query-agent-events/; explicit output paths must be new. |
analyze-events.sh | Group-by stats over a downloaded dump (signal, version, fatal_function, architecture, etc.). |
redact-events.sh | Opt-in redaction (machine_guid / claim_id / host_id / ephemeral_id -> placeholders). For sharing only. |
This skill follows
<repo>/.agents/sensitive-data-discipline.md:
<repo>/src/...).${NETDATA_REPOS_DIR}/<repo>/....src/daemon/status-file.c:988.<repo>/.local/audits/query-agent-events/<run>.json
(gitignored). Do NOT paste raw event JSON into committed
artifacts.get-events.sh requires Bash 4 or later for associative arrays; select a compatible Bash on systems with an older
default. Only live get-events.sh calls load these settings. Its current loader requires all listed values for
either transport;
offline analysis/redaction, explanation, source review and the synthetic self-test do not load .env.
| Key | Role |
|---|---|
NETDATA_CLOUD_TOKEN | Cloud REST token (long-lived). |
NETDATA_CLOUD_HOSTNAME | Cloud REST API host. |
AGENT_EVENTS_HOSTNAME | Dual-duty: ssh host AND direct-HTTP host of the ingestion node. Can be IP or DNS name. NOT the journalctl namespace (hardcoded agent-events); NOT the Cloud room name (also hardcoded agent-events). |
AGENT_EVENTS_MACHINE_GUID | Agent machine GUID for direct-agent transport. |
AGENT_EVENTS_NODE_ID | Cloud node UUID for cloud-proxy transport. |
All values live in <repo>/.env (gitignored). See
<repo>/.agents/ENV.md for setup (where each value comes
from, sample formats, common mistakes).
query-netdata-cloud -- transport: Cloud REST API.query-netdata-agents -- transport: direct agent REST + bearer auto-mint._lib.sh helpers.© netdata, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 17 other files (scripts) in .agents/skills/triage-agent-events of netdata/netdata.
Open the folder on GitHubat commit 9fe30d9
Triage Agent Events next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Triage Agent Events this skillnetdata/netdata | 81k | — | ~2.4k | Automated safety check: Notes | GPL-3.0 | |
| Monitor CInrwl/nx | 29k | 5 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Terraform and OpenTofu Guideagentscope-ai/QwenPaw | 35k | 6 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Vercel Optimize Auditvercel-labs/agent-skills | 32k | 8 repos | ~4.3k | Automated safety check: Pass | None | |
| Analyze GitHub Action Logswithastro/astro | 63k | 1 repos | ~1.3k | Automated safety check: Pass | Custom licence | |
| Iron Proxy Gateway for NanoClawnanocoai/nanoclaw | 31k | — | ~4.6k | Automated safety check: Notes | MIT |
nrwl/nx
Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.
agentscope-ai/QwenPaw
Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.
vercel-labs/agent-skills
Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.
withastro/astro
Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.
nanocoai/nanoclaw
Installs or refreshes Iron Proxy and its Iron Control web console for NanoClaw, with a local Docker setup, database, credentials and a human approval bridge.
antonbabenko/terraform-skill
A skill your agent uses when writing, reviewing, or debugging Terraform/OpenTofu modules, tests, CI, scans, or state ops - diagnoses failure mode (identity churn, secrets, blast radius, CI drift…
netdata/netdata
Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.
netdata/netdata
Inspect Netdata-org source checkouts under NETDATAREPOSDIR, or set up and synchronize that mirror when requested.
netdata/netdata
Inspect, analyze, troubleshoot, or review Codacy findings and local analyzer/API helpers.
netdata/netdata
Inspect or review Coverity Scan defects and saved CID bundles; fetch live findings or apply verified triage decisions when requested.
netdata/netdata
Inspect, review, or apply authorized triage decisions to SonarCloud issues and security hotspots; also review the Sonar helpers.
netdata/netdata
Create, review or validate Netdata Prometheus chart profiles, exporter dashboard design, collection policy and stock semantic proofs.
Categories
Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries. Triage Agent Events is an agent skill from netdata/netdata. Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries.
Triage Agent Events fits situations like: restart/dedup timing; structured filters; version comparisons and reviews of these investigation helpers.
Run `npx skills add netdata/netdata --skill triage-agent-events -a claude-code`. Or copy the skill folder (.agents/skills/triage-agent-events in netdata/netdata) into .claude/skills/triage-agent-events in your project. Claude Code loads it when a task matches its description.
Run `npx skills add netdata/netdata --skill triage-agent-events -a codex`. Or copy the skill folder (.agents/skills/triage-agent-events in netdata/netdata) into .agents/skills/triage-agent-events in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add netdata/netdata --skill triage-agent-events -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage-agent-events, .gemini/skills/triage-agent-events, .github/skills/triage-agent-events and .opencode/skills/triage-agent-events in your project.
Going by SKILL.md and its folder, Triage Agent Events needs a shell and Python for the scripts in its folder and credentials named NETDATA_CLOUD_TOKEN. Our summary lists: Python 3; A Bash shell; A credential in NETDATA_CLOUD_TOKEN.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Triage Agent Events is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Triage Agent Events: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 35k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
netdata (a GitHub organization) maintains it in netdata/netdata, which has 80,820 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 7, 2026.
Source: netdata/netdata on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.