Agent skill

Sre Dashboards

by sickn33 in sickn33/agentic-awesome-skills

Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services.

MITAuto-check passedDevOps & Cloud

Install Sre Dashboards

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill sre-dashboards -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills sre-dashboards --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sre-dashboards .claude/skills/sre-dashboards && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sre-dashboards
GitHub stars
47k
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
365 words
Files
1
Skills in repo
1,493
Repo updated
First seen
Licence
MIT

At a glance

Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services.

  • Works in 3 steps: Executive Reliability View: SLO… → Service Health View: RED/USE metrics,… → Deep-Dive View: Per-endpoint latency,…
  • Tasks that involve Site reliability engineering
  • SKILL.md covers When to Use This Skill, Prerequisites, Dashboard Architecture and Core SRE Panels, plus 5 more sections
  • Calls git and kubectl

What it does

Sre Dashboards is an agent skill from sickn33/agentic-awesome-skills. Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and…

It sits in DevOps & Cloud, covering Site reliability engineering. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Site reliability engineering

Example prompts

  • “/sre-dashboards”

Requirements

  • Compatibility (from SKILL.md): Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Executive Reliability View: SLO attainment, incident counts, MTTR trends.
  2. Service Health View: RED/USE metrics, dependency health, release markers.
  3. Deep-Dive View: Per-endpoint latency, resource saturation, error categories.

What it can do on your machine

Read from SKILL.md and the folder at commit 680176d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

Sre Dashboards loads about 1.1k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 365 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 680176d, republished under its MIT licence (© sickn33). 365 words, ~1,150 tokens.

Download SKILL.mdSave it as .claude/skills/sre-dashboards/SKILL.md (or your agent's skills folder).
name
sre-dashboards
description
Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services.
compatibility
Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

SRE Dashboards

Build dashboards that help teams detect, triage, and prevent reliability incidents.

When to Use This Skill

Use this skill when:

  • Defining service-level dashboards for production systems
  • Tracking SLO health and error-budget burn
  • Creating incident command-center views
  • Standardizing dashboard patterns across teams

Prerequisites

  • Metrics pipeline (Prometheus, OpenTelemetry, or vendor equivalent)
  • Logs/traces linked to services and environments
  • Agreed service taxonomy (team, service, tier, environment)

Dashboard Architecture

Structure dashboards in layers:

  1. Executive Reliability View: SLO attainment, incident counts, MTTR trends.
  2. Service Health View: RED/USE metrics, dependency health, release markers.
  3. Deep-Dive View: Per-endpoint latency, resource saturation, error categories.

Keep each view answer-oriented:

  • Are customers impacted?
  • What changed?
  • Where is the bottleneck?

Core SRE Panels

Golden Signals
  • Latency: p50/p95/p99 request duration by endpoint
  • Traffic: request throughput and queue depth
  • Errors: 5xx rate, failed jobs, timeout ratio
  • Saturation: CPU, memory, disk I/O, thread/connection pool exhaustion
SLO Panels
  • Current SLI value (rolling windows: 5m, 1h, 24h, 30d)
  • Error-budget remaining (%)
  • Burn-rate panels (fast and slow windows)
  • Multi-window burn alert status
Change Correlation
  • Deployment markers and config-change annotations
  • Feature flag state overlays
  • Upstream/downstream dependency error rates

Example PromQL Snippets

promql
# API error rate (%)
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))
promql
# p95 latency by route
histogram_quantile(0.95,
  sum by (le, route) (rate(http_request_duration_seconds_bucket[5m]))
)
promql
# Fast burn rate (5m / 1h)
(
  sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))
)
/
(
  sum(rate(http_requests_total{status=~"5.."}[1h]))
  / sum(rate(http_requests_total[1h]))
)

Operational Guidelines

  • Use consistent color semantics (green=healthy, yellow=degrading, red=breach)
  • Label units explicitly (ms, req/s, %, cores)
  • Default time windows to incident-friendly ranges (15m, 1h, 6h, 24h)
  • Minimize panel count per dashboard to reduce cognitive load
  • Add runbook links directly in panel descriptions
Show full SKILL.md (139 more words)Show less

Troubleshooting

Panel appears flat or empty
  • Verify label cardinality and filters (service, env, region)
  • Confirm scrape/ingest latency is within expected range
  • Check metric rename regressions after instrumentation updates
High cardinality slows dashboards
  • Aggregate by stable dimensions (service, route_group) instead of raw IDs
  • Use recording rules for expensive percentile and ratio queries
  • Split deep-dive dashboards from NOC summary dashboards
  • prometheus-grafana (prometheus-grafana) - Dashboard implementation and PromQL
  • opentelemetry (opentelemetry) - Standardized telemetry instrumentation
  • alerting-oncall (alerting-oncall) - Reliability alert routing and escalation
  • agent-observability (agent-observability) - AI workload reliability telemetry

Limitations

  • Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
  • Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
Example
bash
git status && git diff --stat
kubectl diff -f manifest.yaml

Adapted from BagelHole/DevOps-Security-Agent-Skills (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/sre-dashboards of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit 680176d

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sre Dashboards next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sre Dashboards compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sre Dashboards this skillsickn33/agentic-awesome-skills47k1 repos~1.1kAutomated safety check: PassMIT
Inference Autopilotrednote-machine-learning/Inference-autopilot144—~4.5kAutomated safety check: PassApache-2.0
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Alerting Irmgrafana/skills2811 repos~1.9kAutomated safety check: PassApache-2.0
Slo Implementationwshobson/agents40k11 repos~1.7kAutomated safety check: PassMIT
Agentforce D360 Analyzeforcedotcom/sf-skills1.1k—~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Inference Autopilot

    rednote-machine-learning/Inference-autopilot

    Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

    144 GitHub stars~4.5k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    281 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Slo Implementation

    wshobson/agents

    Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.

    40k GitHub starsUsed in 11 repos~1.7k tokens
    DevOps & CloudAuto-check passed
  • Agentforce D360 Analyze

    forcedotcom/sf-skills

    Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills.

    1.1k GitHub stars~3.4k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    281 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,493 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Categories

Questions about Sre Dashboards

What does Sre Dashboards do?

Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services. Sre Dashboards is an agent skill from sickn33/agentic-awesome-skills. Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services.

When should I use Sre Dashboards?

Sre Dashboards fits situations like: tasks that involve Site reliability engineering.

How do I install Sre Dashboards in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill sre-dashboards -a claude-code`. Or copy the skill folder (skills/sre-dashboards in sickn33/agentic-awesome-skills) into .claude/skills/sre-dashboards in your project. Claude Code loads it when a task matches its description.

How do I install Sre Dashboards in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill sre-dashboards -a codex`. Or copy the skill folder (skills/sre-dashboards in sickn33/agentic-awesome-skills) into .agents/skills/sre-dashboards in your project. Codex loads it when a task matches its description.

Can I use Sre Dashboards in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill sre-dashboards -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sre-dashboards, .gemini/skills/sre-dashboards, .github/skills/sre-dashboards and .opencode/skills/sre-dashboards in your project.

What does Sre Dashboards need to run?

Going by SKILL.md and its folder, Sre Dashboards needs the command-line tools its instructions call (git and kubectl). Compatibility (from SKILL.md): Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled..

Does Sre Dashboards access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Sre Dashboards safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sre Dashboards use?

Sre Dashboards is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sre Dashboards use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sre Dashboards?

Skills that share tags, products or a category with Sre Dashboards: Inference Autopilot (rednote-machine-learning/Inference-autopilot, 144 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Alerting Irm (grafana/skills, 281 stars) and Slo Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sre Dashboards?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,379 GitHub stars. The repository holds 1,493 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.