Agent skill

Health Alert Authoring

by netdata in netdata/netdata

Author, adapt, modify, or review Netdata health alerts and alert templates.

GPL-3.0Auto-check passedDevOps & Cloud

Install Health Alert Authoring

skills CLI
$ npx skills add netdata/netdata --skill health-alert-authoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install netdata/netdata health-alert-authoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/netdata/netdata.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/health-alert-authoring .claude/skills/health-alert-authoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
health-alert-authoring
GitHub stars
81k
Token cost
~7.1k tokens
SKILL.md length
3,713 words
Files
2
Skills in repo
27
Repo updated
First seen
Licence
GPL-3.0

At a glance

Author, adapt, modify, or review Netdata health alerts and alert templates.

  • Works in 5 steps: An exact chart ID is checked first. → A chart name may match. → If the prefix is a context, every chart… → …
  • Translating alerts from another system
  • SKILL.md covers Read The Right Sources, Establish The Alert Contract…, Combine Values Across… and Model The Lifecycle Truthfully, plus 6 more sections
  • Calls git

What it does

Health Alert Authoring is an agent skill from netdata/netdata. Author, adapt, modify, or review Netdata health alerts and alert templates. Use when translating alerts from another system; changing src/health/health.d/.conf, lookup/calc/warn/crit expressions, lifecycle, timing, routing, ownership, or missing-data behavior; writing health-config tests; or selecting an alert's chart/context/label identity.

Its SKILL.md is about 7.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in DevOps & Cloud, covering Translation. The repository describes itself as: The fastest path to AI-powered full stack observability, even for lean teams. The licence is GPL-3.0.

When your agent uses it

  • Translating alerts from another system
  • Changing src/health/health.d/.conf
  • Lookup/calc/warn/crit expressions
  • Missing-data behavior

Example prompts

  • “/health-alert-authoring”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. An exact chart ID is checked first.
  2. A chart name may match.
  3. If the prefix is a context, every chart instance in that context contributes candidates for the final dimension.
  4. Every candidate is scored against the alert's chart labels by counting equal key/value labels.
  5. The candidate with the highest score wins. Equal scores are resolved by candidate traversal order, so avoid designs

What it can do on your machine

Read from SKILL.md and the folder at commit 5bcdf74. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Health Alert Authoring loads about 7.1k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 3,713 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~7.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from netdata/netdata at commit 5bcdf74, republished under its GPL-3.0 licence (© netdata). 3,713 words, ~7,146 tokens.

Download SKILL.mdSave it as .claude/skills/health-alert-authoring/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
health-alert-authoring
description
Author, adapt, modify, or review Netdata health alerts and alert templates. Use when translating alerts from another system; changing `src/health/health.d/*.conf`, lookup/calc/warn/crit expressions, lifecycle, timing, routing, ownership, or missing-data behavior; writing health-config tests; or selecting an alert's chart/context/label identity.

Author Netdata Health Alerts

Use this skill when changing or reviewing a health-alert definition. Alerts are production policy: a syntactically valid expression can still page incorrectly, manufacture a recovery, duplicate an incident owner, or make a disappearing entity look healthy.

The normal goal when translating an alert from another system is Netdata-adapted operational equivalence, not execution-engine emulation. Preserve the operator-visible incident as closely as Netdata can express it through NIDL instances, alert beats, database lookups, native gap behavior, and chart obsoletion. Treat differences from the source engine as explicit design facts to document and test, not as defects by default.

Read The Right Sources

Read the target alert, its collector/profile signal contract, and the existing incident owner. Search for duplicate names, contexts and equivalent generic alerts. Select the affected references:

QuestionSources
Identity, scope, variables and expression syntaxdocs/NIDL-Framework.md, src/health/REFERENCE.md; current variable/evaluator code when affected
Template/alarm ownership or overridessrc/health/alert-configuration-ordering.md and existing named/generic rules
Lifecycle, gaps or timingsrc/health/README.md, relevant lookup/delay sections of src/health/REFERENCE.md, and runtime owners below
Runtime semantics changed or uncertainhealth_event_loop.c, health_variable.c under src/health/; src/web/api/queries/query-execute.c, selected grouping implementation and src/libnetdata/eval/eval-evaluate.c as applicable
Collector/profile signal or identity changesMatching collector/profile skill and its affected references; this skill does not replace them

For review, use the contract and completion sections as questions against the complete assigned change and existing design/test evidence. Read an active SOW when available; do not create one or demand a new authoring note merely to review. Missing evidence matters when it leaves an affected contract unproven. Implementation follows the root SOW gate.

Never query or reconfigure a live Agent merely to validate an alert unless the user has explicitly authorized that access.

Establish The Alert Contract Before Editing

Write down the following in the active SOW before changing a non-trivial alert:

  • Incident and owner: What real incident does this alert represent? Which one source owns it? Do not create a product-named duplicate of an existing collection-failure, component-failure, or generic host alert just to change routing or severity.
  • Signal contract: Identify every source value and its meanings, including zero, non-zero, tri-state values, absent dimensions, temporary collection failure, and known entity disappearance.
  • Source intent and adaptation: Separate the source alert's operator intent from its engine syntax. Record which condition, scope, severity, persistence intent, identity, and recovery semantics Netdata preserves; record every deliberate difference. Use NETDATA-ADAPTED as the normal classification. Reserve EXACT for a proven coincidental match across timing, gaps, identity, recovery, and removal—not merely a similar expression.
  • NIDL instance map: Record the monitored component, context, one instance type, RRDSET/chart ID identity, dimensions, stable identity labels, current metadata labels, and the collector's obsoletion condition. An alert attaches to one RRDSET/chart instance; a template applies that same rule independently to matching instances in one context.
  • Identity: Choose an alarm only for a specific chart instance; choose a template for a context-wide rule. An identity label MUST identify the same monitored entity across its intended lifetime. A label whose source value can change for that entity is current metadata, not RRDSET identity: preserve the chart ID and update/promote the label. Confirm labels preserve the incident identity and do not create unbounded instances. Prove whether each promoted source label is live metadata, creation-time metadata, or source-frozen history. Do not describe a label as "current" merely because Netdata can update labels.
  • Variables and aggregation: An unqualified variable resolves in the alert's local monitored component/instance before broader candidates. Use a fully qualified chart/dimension reference only after proving the other chart is the matching instance and labels select it unambiguously. Never synthesize an infrastructure-level alert by combining multiple RRDSET instances in alert configuration; require a source-owned aggregate RRDSET at the higher component level instead. The alert component: classification field does not replace this NIDL mapping.
  • Lifecycle: Specify expected UNINITIALIZED, CLEAR, WARNING/CRITICAL, UNDEFINED, and REMOVED behavior. State whether an ordinary recovery emits a zero or whether the chart/dimension disappears instead.
  • Timing: State the source update cadence, explicit alert cadence, lookup window, source persistence intent, selected Netdata adaptation, startup behavior, partial-gap behavior, stale-chart behavior, and notification policy. Do not call a lookup window an implementation of another engine's for: state machine.
  • Validation: List the boundary, transition, missing-data, recovery, and duplicate-ownership tests that prove the contract.
  • Policy change disclosure: for a change to a shipped alert's lookup, window, threshold or recovery, you MUST show the old and the new outcome on one short source trace that includes an active event followed by quiet samples, and state what clears the alert and whether it can clear while the underlying cumulative value stays nonzero. A lookup name and window do not disclose that trade-off; the user-owned decision is not informed without it.

Pause for a user decision when the change creates a public alert contract, changes default notification policy, changes the owner of an incident, needs shared health/query framework work, or cannot preserve the approved operator-visible incident closely enough with Netdata-native behavior. A difference from Prometheus execution alone is not such a blocker.

Combine Values Across Dimensions, Charts, And Alerts

Choose the smallest pattern that can express the condition truthfully. Do not introduce cross-chart or intermediate-alert plumbing when the required values already share the alert's chart.

Pattern 1 — dimensions on the alert's chart

Use unqualified dimension variables in calc, warn, or crit when every required value is a dimension on the chart named by on::

text
template: example_ratio
      on: app.state
     calc: $total - $used

Runtime facts:

  • Matching is by dimension ID or name in the alert's own chart.
  • The default dimension value is collector.last_stored_value: the latest database-stored value after interpolation, not a database-window aggregate and not necessarily the raw last collector sample.
  • A missing dimension makes the expression fail with an unknown variable and the alert value become NaN.
  • Prefer explicit non-finite guards when UNDEFINED is the required missing-input state; do not rely on arithmetic or comparisons to preserve NaN.

Use $dimension_raw only when the contract specifically needs the last collected value before interpolation/storage presentation, and $dimension_last_collected_t only when it needs that dimension's collection timestamp.

Do not treat equal dimension timestamps as a collection identifier: PLUGINSD SET calls timestamp dimensions independently. pluginsd_set calls rrddim_set_by_pointer (or its double variant), which stamps a timeval with microsecond precision; health's variable_lookup_in_chart exposes only its tv_sec field. One complete collection can cross a second boundary. Conversely, an omitted dimension retains its previous raw value, so freshness and arithmetic coherence alone can miss a partial collection. When using raw values, also test stored-input gaps: rrdset_done_interpolate sets the normal dimension value to NaN when an omitted dimension reaches a storage beat. Validate omission after a valid zero, as well as a never-present dimension. These checks provide native freshness and gap eligibility; they do not prove atomic transaction membership, particularly when collections occur faster than the storage cadence.

Pattern 2 — dimensions from another chart or context

Use a dotted chart/context reference when values live on different charts:

text
template: example_cross_chart
      on: app.usage
     calc: $this * 100 / ${app.capacity.total}

The dotted prefix is interpreted from right to left as chart/context plus the final dimension name. Runtime resolution:

  1. An exact chart ID is checked first.
  2. A chart name may match.
  3. If the prefix is a context, every chart instance in that context contributes candidates for the final dimension.
  4. Every candidate is scored against the alert's chart labels by counting equal key/value labels.
  5. The candidate with the highest score wins. Equal scores are resolved by candidate traversal order, so avoid designs where a tie is possible.

Consequences:

  • This is the natural way to compare related chart instances, but the matching labels must make the intended instance unique. A generic shared label such as only component=ceph is usually insufficient across many cluster instances.
  • Like Pattern 1, the selected dimension value is the latest stored value, not a query over a time window.
  • Fully qualified names containing punctuation must use ${...} braces.
  • A reference that resolves to no candidate fails as an unknown variable and produces NaN.
Pattern 3 — an intermediate alert as a computed variable

Create a non-notifying helper alert when the required value itself needs a database lookup or multi-stage computation, then reference that alert by name from the consuming alert:

text
template: example_window
      on: app.work
   lookup: average -1h of requests
     calc: $this
      to: silent

template: example_consumes
      on: app.current
   lookup: average -1m of requests
     calc: $this / $example_window
     warn: $this > 1
      to: sysadmin

Runtime facts:

  • Every running alert with the referenced name is a candidate.
  • Candidates are selected by the same equal-label score used for cross-chart variables.
  • The selected candidate contributes its current alert value (rc->value), after that helper's own lookup and calc.
  • This is the only one of the three patterns that can combine database-window results such as “max of the last hour of X with the average of the last minute of Y”.
  • Health evaluates all lookup/calculation phases before warning/critical phases, but helper snapshots are published as each alert completes. A consumer with a different cadence can therefore read the helper's previous published value on its first beat or after cadence drift. Align every: or disclose this timing difference.
  • Prefer to: silent on helpers and document why they exist; a helper is instrumentation, not a second incident owner.
Choosing a pattern
  • Same chart, latest values: Pattern 1.
  • Different charts/contexts, latest values: Pattern 2.
  • Any input needs a database window or a staged calculation: Pattern 3.
  • Never use Pattern 3 merely to avoid a qualified dotted reference; its added timing and lifecycle coupling must earn its place.
  • Never combine multiple RRDSET instances into one infrastructure-level alert on the alert side. If no source-owned chart provides the required aggregate, that is a collector/profile/framework gap, not an alert-expression workaround.

Model The Lifecycle Truthfully

Source situationAlert-lifecycle consequence
Valid numeric input and both conditions are falseCLEAR
Valid numeric input and warning/critical condition is trueWARNING or CRITICAL
Some values exist in a lookup window and some are NULLThe lookup continues using numeric values; NULL is not automatically a failed condition
Collector misses a collection while the chart remains liveA runnable alert may evaluate a lookup based on stored data; no fresh sample is fabricated
Lookup runs and its selected window has no usable values$this is NaN; the final state is UNDEFINED only when the condition preserves it
A lookup's newest stored sample is too old for its runnable-history gateThe health loop skips that evaluation; it does not manufacture CLEAR or guarantee an UNDEFINED transition
Collector knows the entity is gone and obsoletes the chartThe alert enters REMOVED, not an ordinary zero/CLEAR recovery

Important rules:

  • A collection failure is not disappearance. Collectors MUST preserve a gap when they cannot measure and MUST obsolete only an entity they know is gone. Do not add a false zero to make an alert recover.
  • A chart can remain alert-eligible during a gap. The health loop skips an obsolete chart, but does not require a newly collected value for every scheduled evaluation. $last_collected_t and $update_every expose freshness when an alert is the deliberate stale-collection owner.
  • A live chart does not guarantee every lookup remains runnable forever. For a relative database lookup, the health loop eventually skips evaluation when the newest stored point is too old for the requested window plus its bounded update-interval tolerance. Skipping evaluation preserves the prior alert state; it is different from evaluating an all-null result to UNDEFINED.
  • Obsoletion ends the instance alert. Once the collector knows an entity is gone and obsoletes its RRDSET, the health rule no longer evaluates that instance. Do not attempt to emulate an infrastructure-level continuation in another alert; collect a source-owned aggregate instance if that is the required product signal.
  • A null result does not fabricate healthy state. The health loop assigns NaN to an empty lookup result. A condition that itself evaluates to NaN is UNDEFINED.
  • Comparisons can consume NaN. In the expression evaluator, NaN is false in boolean contexts and a comparison such as $this == 0 produces a finite false result. That can yield CLEAR, not UNDEFINED. If the contract requires UNDEFINED, make and test an expression path that preserves NaN; never assume a comparison does so.
  • Some NULLs are not all NULLs. Query aggregation accepts numeric points and omits non-numeric ones. Do not assume a partial collection gap invalidates a window unless the current runtime and the approved product contract say it does. A general strict-coverage rule is shared-framework work, not an alert-file shortcut.
  • Do not duplicate freshness ownership. Use $now - $last_collected_t for an explicit collection/staleness alert only when no existing generic collection-failure alert already owns the incident. A data-state alert normally owns the measured condition, not the collector outage.

For a data-state alert that MUST become UNDEFINED when its lookup is non-finite, use and test this condition shape:

text
warn: ($this == nan or $this == inf) ? (nan) : (<numeric predicate>)

Use crit: in the same way. This preserves the non-finite result only; it does not turn a partial-null window into a collection-failure alert or replace a known-obsolete chart's REMOVED lifecycle.

Adapt Timing To Netdata

Current State

Use calc when the source's current value is the full condition. Test start-up, missing input, normal recovery, and chart obsoletion separately. delay: controls notification hysteresis, not the evaluated alert state; it is not a substitute for a Prometheus for: duration.

Control Flapping With Three Combinable Layers

Stability involves signal smoothing, threshold hysteresis and notification delay. Address them in that order. State-dependent thresholds are hysteresis; delay: postpones notification execution while the evaluated alert state changes immediately. See src/health/REFERENCE.md#alert-line-delay and its conditional-operator hysteresis section.

Layer 1 — stabilize the queried value

Use lookup: to make $this a stable aggregate over an explicit observation window:

text
lookup: average -5m unaligned of latency
  • average smooths noisy utilization, rate, latency, and utilization-like signals.
  • min requires every observed numeric sample to remain active, appropriate for persisted binary fault states.
  • max preserves worst-case excursions when the incident is defined by peaks.
  • countif expresses percent-of-observed-time semantics.

This reduces noise but cannot prevent a stable aggregate from hovering near one threshold. A 5-minute average near 100 ms can still cross a 100 ms threshold repeatedly.

Show full SKILL.md (1,447 more words)Show less
Layer 2 — separate raise and clear thresholds

For policy thresholds whose signal may hover near the boundary, use a state-dependent predicate:

text
warn: $this > (($status >= $WARNING) ? (90) : (100))
crit: $this > (($status == $CRITICAL) ? (105) : (110))

This changes one threshold into two thresholds:

  • clear-to-warning raises above 100;
  • the warning predicate becomes false at or below 90 (CLEAR when the critical predicate is also false);
  • critical can independently use another raise/clear pair.

Important semantics:

  • This is threshold hysteresis, without a time delay.
  • The alert continues evaluating at its normal every: cadence.
  • It only changes the boundary used by the current status, so a genuine crossing of the recovery threshold acts immediately.
  • Preserve a non-finite guard around the complete expression when UNDEFINED must remain possible.

Prefer this when independent raise/clear boundaries are meaningful and the signal is expected to linger near one boundary. Do not use it to redefine a categorical exact condition into a policy band.

Layer 3 — delay transition notifications

Use delay: when notification execution should wait after a state change:

text
delay: down 5m multiplier 1.5 max 1h
  • up is selected when the new numeric status is greater than the old; down otherwise. This includes UNDEFINED (-1) → CLEAR (1) using up, not just warning/critical escalation. Owners: src/health/health_event_loop.c (delay selection), src/health/rrdcalc.h (status values).
  • multiplier grows the delay when the state changes during the delay.
  • max caps the accumulated delay.

This is the only layer that postpones an alert transition notification. It can suppress rapid clear/reactivate notification cycles, but it can also postpone notice of a real incident. Use it sparingly and record the expected notification delay in the alert contract. The status still changes on the evaluation beat: health_event_loop.c updates it immediately, while health_notifications.c checks the notification due time.

Selection procedure
  1. Establish whether the incident is categorical, threshold policy, or derived arithmetic.
  2. Select the smallest truthful lookup: window and aggregation first.
  3. For a noisy policy threshold, choose explicit raise/clear thresholds before adding transition delay.
  4. Add delay: only when rapid transition notifications are independently harmful and later notification is acceptable.
  5. Test each layer: input noise, boundary crossing, recovery, reactivation, non-finite input, partial gap, and the exact expected notification time.
Persistence Intent

Treat another system's for: D as an operator intent to suppress transient conditions. It is not a requirement to emulate that system's pending-state machine. Choose the closest safe Netdata-native behavior from the source's actual numeric state space, set an explicit every:, and disclose the differences:

  • For a binary source where 1 means the active fault, min -5m unaligned with an active predicate is true only when every observed numeric value in the selected window is active. This is a Netdata observation-window adaptation, not Prometheus for: 5m.
  • For a binary source where 0 means the active fault, use the complementary aggregation/predicate that requires all observed values to be zero; consult the lookup reference and prove it with state-sequence tests.
  • For a tri-state or enumerated source, do not apply a binary min/max rule by analogy. Use an exact predicate over the full state space. For example, countif(!=target) returns the percentage of observed values outside target; zero means every observed value matched.
  • min, max, and countif operate on the numeric samples they receive. Test active-to-other-state transitions in both directions and record the intended partial-gap behavior.
  • A new chart may become runnable with up to one chart update interval less history than the requested lookup window. Therefore an already-active condition can raise up to roughly one source beat before D after chart creation. Do not conceal this with a hand-authored "window complete" flag.
  • A missing sample does not reset an observation window. Values before and after a partial gap may contribute to the same lookup. Collection-failure ownership remains separate unless the alert contract explicitly owns freshness.
  • One health rule has one historical database lookup. A predicate combining multiple dimensions cannot gain exact historical persistence by looking up one dimension and combining it with the others' current values. Prefer, in order: a source-owned derived dimension when the persisted compound condition is essential; otherwise a truthful current-state adaptation; never a mixed-time formula that changes the incident meaning.
  • Do not lengthen the window mechanically by one assumed collector interval. Collector cadence is configurable, and that does not repair partial gaps or create a source-engine pending state.

For every persistence adaptation, prove: chart startup, observed active history, each recovery state, each higher/lower enumerated state, a partial source gap, a stale non-runnable lookup, collection resumption, chart obsoletion, and the configured evaluation cadence. Never invent a pending state that Netdata does not expose, and never declare fidelity from the look of an expression alone.

Keep Ownership And Identity Non-Duplicating

  • Search existing stock alerts before adding one. Reuse an existing alert only when it really owns the same logical incident; disclose routing/severity/lifecycle differences rather than silently relabeling it.
  • Keep source collection failure, component/API collection failure, source data-state failure, and client-observed failure as separate owners when they identify different operator actions.
  • Filter templates with chart labels only when they select the intended RRDSET instance without changing its identity. A current metadata label may be used as a filter only when the alert contract explicitly wants that current metadata; do not turn it into instances.by_labels merely to make filtering convenient. Exclude known named rules from generic fallbacks, including special sources whose recovery is chart removal rather than zero.
  • Preserve ordinary zero recovery. Do not convert an active-to-zero source into disappearance, and do not treat a disappearing source as a normal CLEAR.
  • Use the ordering guide for template/alarm precedence and user-versus-stock override behavior. Same-name definitions are an override mechanism; different names coexist and can therefore duplicate incidents.

Validate The Actual Contract

Run the smallest relevant tests first, then the full affected suite. A complete alert change normally needs all applicable items below:

  1. Run /usr/sbin/netdata -W healthconfigtest for the built-in health parser and lookup suite.
  2. Verify the shipped alert against an independent oracle: exercise the signal's behavior, or compare independently owned artifacts with the collecttest health checks named in .agents/skills/collectors-go-framework-v2/chart-template.md#tests. A test MUST NOT copy the definition's lookup, cadence, thresholds, expressions or recipients into expected literals; that restates the file and pins whatever it said, right or wrong.
  3. Test the signal's lifecycle through the real query/health runtime where practical. If a lower-level deterministic model is necessary, derive it directly from runtime timestamps and numeric-point selection; do not pass an arbitrary windowComplete flag or label skipped NULLs as continuous evaluation. Cover startup, active transition, normal zero recovery, a true all-null query, partial-null behavior, stale non-runnable behavior, collection resumption, known disappearance/REMOVED, and label identity.
  4. For an expression that relies on NaN, test the evaluator result directly. Check both direct NaN propagation and any comparison/conditional branch; do not infer the result from ordinary floating-point intuition.
  5. Run collector/profile validation when the alert depends on a collector or profile change. Include source absence and collision-bearing labels where relevant.
  6. Search for same-incident alerts, duplicate contexts, old names, fallback overlap, and generated artifacts. Record the result in the SOW.
  7. For an externally sourced alert pack, commit a source-pinned mapping that records the original condition, scope, severity, persistence intent, supported releases, Netdata owner, adaptation, and known differences. Tests MUST consume that mapping rather than restating the intended result independently.
  8. Validate every published configuration example through the same job-construction prerequisites users need. In particular, a job referencing vnode: is incomplete unless the example defines or clearly links the required vnode.
  9. Run git diff --check and the project-required validation/review gate before claiming completion.

healthconfigtest runs built-in health parser and lookup cases. It does not load every stock health template, prove that a template attaches to the intended chart, or prove the runtime lifecycle. Cover those separately with source-aware contract and transition tests.

Completion Check

Before requesting review, confirm all of the following:

  • The alert has exactly one logical owner and a stable scope.
  • Its NIDL instance map shows one monitored component and one instance-level alert target; any required aggregate is a source-owned higher-level RRDSET, not an alert-side merge.
  • Its source state values, gaps, absence, and recovery have been tested rather than assumed.
  • Its NaN, UNDEFINED, CLEAR, and REMOVED transitions match the recorded contract.
  • Its timing is the closest safe Netdata adaptation, with startup/gap/stale differences disclosed; delay: is not standing in for another engine's persistence state machine.
  • Notification defaults follow the approved product policy.
  • No shared health/query behavior was added or assumed without the required separate scope approval.

Authoritative References

  • Syntax, variables, lookups, and stock patterns: src/health/REFERENCE.md
  • State model and missing-data summary: src/health/README.md
  • NIDL component/instance/dimension/label model: docs/NIDL-Framework.md
  • Template/alarm and user/stock precedence: src/health/alert-configuration-ordering.md
  • Alert eligibility, lookup execution, and REMOVED: src/health/health_event_loop.c
  • $now, $last_collected_t, $update_every, dimension freshness, same-chart variables, cross-chart/context variables, alert variables, and label-score selection: src/health/health_variable.c
  • Equal-label score implementation: src/database/rrdlabels.c:rrdlabels_common_count()
  • Query gaps and grouping: src/web/api/queries/query-execute.c and the relevant grouping implementation
  • NaN expression semantics: src/libnetdata/eval/eval-evaluate.c

© netdata, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/health-alert-authoring of netdata/netdata.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 5bcdf74

Compare with similar skills

Health Alert Authoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Health Alert Authoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Health Alert Authoring this skillnetdata/netdata81k—~7.1kAutomated safety check: PassGPL-3.0
Migrate Vstest To Mtprunceel/ReactiveProperty944—~4.3kAutomated safety check: PassMIT
Doca Telemetry UtilsNVIDIA/skills3.6k—~4.2kAutomated safety check: PassApache-2.0
Lokalise Deploy Integrationjeremylongshore/tons-of-skills-marketplace2.8k—~3.4kAutomated safety check: PassMIT
Azure TranslatorMicrosoftDocs/Agent-Skills776—~4.4kAutomated safety check: PassCC-BY-4.0
Build Openshell Mxc WindowsNVIDIA/OpenShell16k—~4.9kAutomated safety check: PassApache-2.0

Similar skills

  • Migrate Vstest To Mtp

    runceel/ReactiveProperty

    Migrates .NET test projects from VSTest to Microsoft.Testing.Platform (MTP).

    944 GitHub stars~4.3k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Official

    A skill your agent uses when the user is invoking docatelemetryutils on a host with DOCA installed — discovering the diagnostic-counter schema, translating counter names to binary Data IDs…

    3.6k GitHub stars~4.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Lokalise Deploy Integration

    jeremylongshore/tons-of-skills-marketplace

    Deploy Lokalise integrations to Vercel, Netlify, and Cloud Run platforms.

    2.8k GitHub stars~3.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Azure Translator

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Translator development including troubleshooting, best practices, decision making, limits & quotas, security, configuration, integrations & coding patterns, and deployment.

    776 GitHub stars~4.4k tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • Official

    Maintain and validate OpenShell's build-only Windows MSVC lane for x64 and ARM64.

    16k GitHub stars~4.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Pi Runbook Writer

    Pr1p/pi-runbook

    A skill your agent uses when creating, editing, or polishing pi-runbook content: README/index pages, docs, journal notes, experiments, bilingual documentation, source-reading summaries…

    143 GitHub stars~796 tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from netdata/netdata

All 27 skills in this repo
  • Docs Learn PR Preview

    netdata/netdata

    Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.

    81k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Repo Mirror Sources

    netdata/netdata

    Inspect Netdata-org source checkouts under NETDATAREPOSDIR, or set up and synchronize that mirror when requested.

    81k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check: notes
  • Triage Agent Events

    netdata/netdata

    Investigate Netdata crashes, panics and fatals from agent-events captures or authorized fleet queries.

    81k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check: notes
  • Triage Codacy

    netdata/netdata

    Inspect, analyze, troubleshoot, or review Codacy findings and local analyzer/API helpers.

    81k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Triage Coverity

    netdata/netdata

    Inspect or review Coverity Scan defects and saved CID bundles; fetch live findings or apply verified triage decisions when requested.

    81k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Triage Sonarqube

    netdata/netdata

    Inspect, review, or apply authorized triage decisions to SonarCloud issues and security hotspots; also review the Sonar helpers.

    81k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check: notes

Questions about Health Alert Authoring

What does Health Alert Authoring do?

Author, adapt, modify, or review Netdata health alerts and alert templates. Health Alert Authoring is an agent skill from netdata/netdata. Author, adapt, modify, or review Netdata health alerts and alert templates.

When should I use Health Alert Authoring?

Health Alert Authoring fits situations like: translating alerts from another system; changing src/health/health.d/.conf; lookup/calc/warn/crit expressions; missing-data behavior.

How do I install Health Alert Authoring in Claude Code?

Run `npx skills add netdata/netdata --skill health-alert-authoring -a claude-code`. Or copy the skill folder (.agents/skills/health-alert-authoring in netdata/netdata) into .claude/skills/health-alert-authoring in your project. Claude Code loads it when a task matches its description.

How do I install Health Alert Authoring in Codex?

Run `npx skills add netdata/netdata --skill health-alert-authoring -a codex`. Or copy the skill folder (.agents/skills/health-alert-authoring in netdata/netdata) into .agents/skills/health-alert-authoring in your project. Codex loads it when a task matches its description.

Can I use Health Alert Authoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add netdata/netdata --skill health-alert-authoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/health-alert-authoring, .gemini/skills/health-alert-authoring, .github/skills/health-alert-authoring and .opencode/skills/health-alert-authoring in your project.

What does Health Alert Authoring need to run?

Going by SKILL.md and its folder, Health Alert Authoring needs the command-line tools its instructions call (git).

Does Health Alert Authoring access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Health Alert Authoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Health Alert Authoring use?

Health Alert Authoring is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Health Alert Authoring use?

About 7.1k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Health Alert Authoring?

Skills that share tags, products or a category with Health Alert Authoring: Migrate Vstest To Mtp (runceel/ReactiveProperty, 944 stars), Doca Telemetry Utils (NVIDIA/skills, 3.6k stars), Lokalise Deploy Integration (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Azure Translator (MicrosoftDocs/Agent-Skills, 776 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Health Alert Authoring?

netdata (a GitHub organization) maintains it in netdata/netdata, which has 80,863 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 10, 2026.

Source: netdata/netdata on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.