Agent skill

Grafana Observability

by automateyournetwork in automateyournetwork/netclaw

Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools).

Apache-2.0Auto-check: notesDevOps & Cloud

Install Grafana Observability

skills CLI
$ npx skills add automateyournetwork/netclaw --skill grafana-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install automateyournetwork/netclaw grafana-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/automateyournetwork/netclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workspace/skills/grafana-observability .claude/skills/grafana-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
grafana-observability
GitHub stars
676
Token cost
~2.7k tokens
SKILL.md length
966 words
Files
1
Skills in repo
120
Repo updated
First seen
Licence
Apache-2.0

At a glance

Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools).

  • Works in 7 steps: Find dashboards: search_dashboards with… → Dashboard overview:… → Query metrics: query_prometheus with… → …
  • Querying Grafana dashboards
  • SKILL.md covers MCP Server, How to Run, Environment Variables and Key Tool Categories, plus 8 more sections
  • Calls uvx; needs GRAFANA_SERVICE_ACCOUNT_TOKEN and GRAFANA_PASSWORD

What it does

Grafana Observability is an agent skill from automateyournetwork/netclaw. Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools). Use when querying Grafana dashboards, searching Loki logs for syslog events, investigating firing alerts, or checking who is on call. For a direct PromQL query with no Grafana dashboard/panel/Loki/OnCall context needed, or when Grafana isn't configured, use prometheus-monitoring instead.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Monitoring and alerting and Incident response. It works with Grafana, Prometheus and Model Context Protocol. The repository describes itself as: An AI agent that claws through your network. The licence is Apache-2.0.

When your agent uses it

  • Querying Grafana dashboards
  • Searching Loki logs for syslog events
  • Investigating firing alerts
  • Checking who is on call

Example prompts

  • “/grafana-observability”

Requirements

  • A credential in GRAFANA_SERVICE_ACCOUNT_TOKEN

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Find dashboards: search_dashboards with keyword (e.g., "network", "interface", "BGP")
  2. Dashboard overview: get_dashboard_summary for panel list without full JSON
  3. Query metrics: query_prometheus with PromQL for specific metrics
  4. Check alerts: list_alert_rules to see active alerting thresholds
  5. Search logs: query_loki_logs for syslog or SNMP trap data
  6. Report: Metrics summary with alert status and log correlation
  7. GAIT: Record all queries in audit trail

What it can do on your machine

Read from SKILL.md and the folder at commit aa90e7d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uvx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GRAFANA_SERVICE_ACCOUNT_TOKEN
    • GRAFANA_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Grafana Observability loads about 2.7k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 966 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:232
    A_SERVICE_ACCOUNT_TOKEN` in `~/.openclaw/.env`. Verify service account has Editor role or required RBAC permissions.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from automateyournetwork/netclaw at commit aa90e7d, republished under its Apache-2.0 licence (© automateyournetwork). 966 words, ~2,714 tokens.

Download SKILL.mdSave it as .claude/skills/grafana-observability/SKILL.md (or your agent's skills folder).
name
grafana-observability
description
Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools). Use when querying Grafana dashboards, searching Loki logs for syslog events, investigating firing alerts, or checking who is on call. For a direct PromQL query with no Grafana dashboard/panel/Loki/OnCall context needed, or when Grafana isn't configured, use `prometheus-monitoring` instead.
license
Apache-2.0
user-invocable
true

Grafana Observability

MCP Server

PropertyValue
Sourcegrafana/mcp-grafana
Transportstdio (default), SSE, or streamable-http
LanguageGo (runs via uvx mcp-grafana)
Tools75+ (dashboards, Prometheus, Loki, alerting, incidents, OnCall, annotations, admin)
AuthService account token (preferred) or username/password
RequiresGrafana 9.0+, service account with Editor role or granular RBAC

How to Run

bash
# stdio mode (default — used by NetClaw)
uvx mcp-grafana

# Read-only mode (prevents dashboard/alert modifications)
uvx mcp-grafana --disable-write

Environment Variables

VariableRequiredExampleDescription
GRAFANA_URLYeshttp://grafana.example.com:3000Grafana instance URL
GRAFANA_SERVICE_ACCOUNT_TOKENYes*glsa_abc123...Service account token (preferred auth)
GRAFANA_USERNAMEAltadminBasic auth username (alternative to token)
GRAFANA_PASSWORDAltchangemeBasic auth password
GRAFANA_ORG_IDNo1Organization ID for multi-org setups

*Either service account token or username/password required.

Key Tool Categories

Dashboard Operations
ToolWhat It Does
search_dashboardsFind dashboards by title or metadata
get_dashboard_summaryLightweight overview (context-efficient — use this first)
get_dashboard_by_uidFull dashboard JSON (large — use sparingly)
get_dashboard_propertyExtract specific fields via JSONPath
get_dashboard_panel_queriesExtract panel query details
update_dashboardCreate or modify dashboards
patch_dashboardTargeted modifications without full JSON replacement
Prometheus (PromQL)
ToolWhat It Does
query_prometheusExecute instant or range PromQL queries
list_prometheus_metric_namesDiscover available metrics
list_prometheus_label_namesList labels matching selectors
list_prometheus_label_valuesRetrieve values for a specific label
query_prometheus_histogramCalculate percentiles (p50, p90, p95, p99)
list_prometheus_metric_metadataMetric type, help text, unit
Loki (LogQL)
ToolWhat It Does
query_loki_logsExecute LogQL queries against log streams
list_loki_label_namesDiscover available log labels
list_loki_label_valuesList values for a specific log label
query_loki_statsStream statistics (volume, rate)
query_loki_patternsDetect log structure patterns
Alerting
ToolWhat It Does
list_alert_rulesView all Grafana and datasource-managed alert rules
get_alert_rule_by_uidRetrieve specific alert rule details
create_alert_ruleCreate new alert rule
update_alert_ruleModify existing alert rule
delete_alert_ruleRemove alert rule
list_contact_pointsView notification endpoints (email, Slack, PagerDuty, etc.)
Incident Management
ToolWhat It Does
list_incidentsView Grafana Incidents with filtering
get_incidentSingle incident details
create_incidentCreate a new incident
add_activity_to_incidentAdd timeline entry to incident
OnCall
ToolWhat It Does
list_oncall_schedulesView on-call rotation schedules
get_oncall_shiftShift details
get_current_oncall_usersWho is on call right now
list_alert_groupsOnCall alert groups with filtering
Annotations & Rendering
ToolWhat It Does
get_annotationsQuery annotations with time/tag filters
create_annotationAdd annotation to dashboard/panel
get_panel_imageRender a panel or dashboard as PNG image
generate_deeplinkCreate accurate Grafana URLs for sharing
Investigation (Sift)
ToolWhat It Does
list_sift_investigationsList automated investigations
get_sift_investigationInvestigation details
find_error_pattern_logsDetect elevated error patterns in logs
find_slow_requestsIdentify slow requests via Tempo traces

Workflow: Network Infrastructure Monitoring

When checking network device metrics in Grafana:

  1. Find dashboards: search_dashboards with keyword (e.g., "network", "interface", "BGP")
  2. Dashboard overview: get_dashboard_summary for panel list without full JSON
  3. Query metrics: query_prometheus with PromQL for specific metrics:
    • Interface traffic: rate(ifHCInOctets{instance="router1"}[5m]) * 8
    • BGP peer state: bgp_peer_state{peer="10.1.1.2"}
    • CPU utilization: device_cpu_utilization{device="core-rtr-01"}
    • Interface errors: increase(ifInErrors{device=~".*"}[1h])
  4. Check alerts: list_alert_rules to see active alerting thresholds
  5. Search logs: query_loki_logs for syslog or SNMP trap data
  6. Report: Metrics summary with alert status and log correlation
  7. GAIT: Record all queries in audit trail
Example: Interface Utilization Check
search_dashboards(title="Network Interfaces")
get_dashboard_summary(uid="abc123")
query_prometheus(expr="rate(ifHCInOctets{device='core-rtr-01'}[5m]) * 8", time_range="1h")
query_prometheus(expr="rate(ifHCOutOctets{device='core-rtr-01'}[5m]) * 8", time_range="1h")
list_alert_rules(folder="Network")

Workflow: Alert Investigation

When investigating Grafana alerts:

  1. List alerts: list_alert_rules — find firing or pending rules
  2. Alert details: get_alert_rule_by_uid — thresholds, conditions, datasource
  3. Query metrics: query_prometheus — check the metric that triggered the alert
  4. Search logs: query_loki_logs — correlate with log events around alert time
  5. Check incidents: list_incidents — is this already tracked?
  6. Contact points: list_contact_points — verify notification routes
  7. Report: Alert analysis with root cause and metric evidence
Show full SKILL.md (415 more words)Show less

Workflow: Incident Response

When responding to a Grafana incident:

  1. List incidents: list_incidents — find open incidents
  2. Incident details: get_incident — timeline, severity, labels
  3. OnCall: get_current_oncall_users — who should be notified
  4. Correlate metrics: query_prometheus — check affected service metrics
  5. Correlate logs: query_loki_logs — find error patterns around incident time
  6. Investigate: find_error_pattern_logs — automated error pattern detection
  7. Update incident: add_activity_to_incident — add findings to timeline
  8. Annotate: create_annotation — mark event on relevant dashboards

Workflow: Log Analysis

When investigating network logs stored in Loki:

  1. Discover labels: list_loki_label_names — find available labels (host, severity, facility)
  2. Label values: list_loki_label_values — enumerate hosts, severity levels
  3. Query logs: query_loki_logs with LogQL:
    • By device: {host="core-rtr-01"}
    • By severity: {host="core-rtr-01"} |= "error"
    • Pattern match: {job="syslog"} |~ "BGP|OSPF"
  4. Patterns: query_loki_patterns — detect recurring log structures
  5. Stats: query_loki_stats — log volume and rate analysis

Integration with Other Skills

SkillIntegration
pyats-health-checkCross-reference pyATS health data with Grafana metrics and dashboards
pyats-routingCorrelate OSPF/BGP state changes with Grafana metric timelines
gait-session-trackingRecord all Grafana queries and findings in GAIT audit trail
slack-network-alertsGrafana alerts fed through Slack + NetClaw for automated investigation
servicenow-change-workflowAnnotate Grafana dashboards during change windows; correlate incidents with CRs
te-network-monitoringPair ThousandEyes path data with Grafana infrastructure metrics
aws-cloud-monitoringCompare Grafana dashboards with CloudWatch data for hybrid visibility
markmap-vizVisualize Grafana alert rule hierarchies as mind maps

Context Window Management

Grafana dashboards can be large JSON documents. Use these strategies:

  1. Always start with get_dashboard_summary — lightweight overview, not full JSON
  2. Use get_dashboard_property with JSONPath for specific fields
  3. Avoid get_dashboard_by_uid unless you need the complete dashboard definition
  4. Use get_dashboard_panel_queries to extract just the query definitions

Important Rules

  • Prefer read-only operations — use search_dashboards, get_dashboard_summary, query_prometheus, query_loki_logs, list_alert_rules before any write operations
  • Dashboard modifications require ServiceNow CR — unless in lab/dev Grafana instance
  • Alert rule changes require approval — creating/updating/deleting alert rules affects production monitoring
  • Token-efficient queries — use get_dashboard_summary over get_dashboard_by_uid, use time ranges to limit Prometheus/Loki result size
  • GAIT audit mandatory — record all Grafana queries, dashboard modifications, alert changes, and incident updates
  • No secrets in queries — never embed credentials or sensitive data in PromQL/LogQL expressions

Error Handling

  • Auth fails (401/403): Check GRAFANA_URL and GRAFANA_SERVICE_ACCOUNT_TOKEN in ~/.openclaw/.env. Verify service account has Editor role or required RBAC permissions.
  • Datasource not found: Use list_datasources to discover available datasource UIDs and names.
  • PromQL/LogQL errors: Use list_prometheus_metric_names or list_loki_label_names to discover valid metric/label names before querying.
  • Dashboard not found: Use search_dashboards to find dashboards by title before using UID-based tools.
  • Rate limiting: Grafana may rate-limit API requests; space out large query batches.

© automateyournetwork, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in workspace/skills/grafana-observability of automateyournetwork/netclaw.

Open the folder on GitHubat commit aa90e7d

Compare with similar skills

Grafana Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Grafana Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Grafana Observability this skillautomateyournetwork/netclaw676—~2.7kAutomated safety check: NotesApache-2.0
Syncmetapawurb/hotpath-rs1.9k—~1.2kAutomated safety check: NotesMIT
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence
Frontmcp Observabilityagentfront/frontmcp146—~4.6kAutomated safety check: PassApache-2.0
Live Debugmacro-inc/macro4.6k—~2.4kAutomated safety check: NotesAGPL-3.0

Similar skills

  • Syncmeta

    pawurb/hotpath-rs

    Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).

    1.9k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Live Debug

    macro-inc/macro

    Debug the running local stack with traces, logs, and a shared headless browser.

    4.6k GitHub stars~2.4k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Analyze the experiment precompute result-consistency canary across prod-US and prod-EU, deep-dive any issues, and produce an actionable report.

    40k GitHub stars~3.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from automateyournetwork/netclaw

All 120 skills in this repo
  • EVE-NG Lab Topology Design

    automateyournetwork/netclaw

    Entry point for designing EVE-NG network labs: classifies the request, gathers missing requirements, proposes options and validates the resulting topology.

    677 GitHub stars~612 tokensUpdated today
    Auto-check passed
  • ACI Policy Change Deployment

    automateyournetwork/netclaw

    Deploys Cisco ACI policy changes only behind an approved ServiceNow Change Request, capturing pre and post-change fault baselines and rolling back automatically on a fault delta.

    677 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Cisco ACI Fabric Health Audit

    automateyournetwork/netclaw

    Runs a phased health audit of a Cisco ACI fabric through MCP tools: node status, links, tenant and policy review, faults and endpoint learning.

    677 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Anta Validation

    automateyournetwork/netclaw

    Validate Arista EOS network state against ANTA's pre-built 208-test catalogue, with structured pass/fail verdicts.

    677 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Arista Cvp

    automateyournetwork/netclaw

    Arista CloudVision Portal (CVP) automation via REST API — device inventory, events, connectivity monitoring, tag management (4 tools).

    677 GitHub stars~2.2k tokensUpdated today
    Auto-check: notes
  • AWS Cloud Monitoring

    automateyournetwork/netclaw

    AWS CloudWatch monitoring — metrics, alarms, log queries, VPC flow log analysis, network performance.

    677 GitHub stars~1k tokensUpdated today
    Auto-check passed

Categories

Questions about Grafana Observability

What does Grafana Observability do?

Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools). Grafana Observability is an agent skill from automateyournetwork/netclaw. Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools).

When should I use Grafana Observability?

Grafana Observability fits situations like: querying Grafana dashboards; searching Loki logs for syslog events; investigating firing alerts; checking who is on call.

How do I install Grafana Observability in Claude Code?

Run `npx skills add automateyournetwork/netclaw --skill grafana-observability -a claude-code`. Or copy the skill folder (workspace/skills/grafana-observability in automateyournetwork/netclaw) into .claude/skills/grafana-observability in your project. Claude Code loads it when a task matches its description.

How do I install Grafana Observability in Codex?

Run `npx skills add automateyournetwork/netclaw --skill grafana-observability -a codex`. Or copy the skill folder (workspace/skills/grafana-observability in automateyournetwork/netclaw) into .agents/skills/grafana-observability in your project. Codex loads it when a task matches its description.

Can I use Grafana Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add automateyournetwork/netclaw --skill grafana-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grafana-observability, .gemini/skills/grafana-observability, .github/skills/grafana-observability and .opencode/skills/grafana-observability in your project.

What does Grafana Observability need to run?

Going by SKILL.md and its folder, Grafana Observability needs the command-line tools its instructions call (uvx) and credentials named GRAFANA_SERVICE_ACCOUNT_TOKEN and GRAFANA_PASSWORD. Our summary lists: A credential in GRAFANA_SERVICE_ACCOUNT_TOKEN.

Does Grafana Observability access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Grafana Observability safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Grafana Observability use?

Grafana Observability is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Grafana Observability use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Grafana Observability?

Skills that share tags, products or a category with Grafana Observability: Syncmeta (pawurb/hotpath-rs, 1.9k stars), Alerting Irm (grafana/skills, 282 stars), Archestra Dev Observability (archestra-ai/archestra, 4.4k stars) and Frontmcp Observability (agentfront/frontmcp, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Grafana Observability?

automateyournetwork (a GitHub user) maintains it in automateyournetwork/netclaw, which has 676 GitHub stars. The repository holds 120 skills in this directory. The repository was last updated on October 9, 2026.

Source: automateyournetwork/netclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.