Axiom Dashboard Builder
openclaw/clawhub
Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana.
Write a monitoring setup guide for a service — defining what to measure, how to alert on it, and how to build the observability stack covering the four golden signals, business metrics, log…
$ npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mohitagw15856/pm-claude-skills monitoring-setup-guide --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/monitoring-setup-guide .claude/skills/monitoring-setup-guide && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "monitoring-setup-guide" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/monitoring-setup-guide into .claude/skills/monitoring-setup-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-setup-guide", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/monitoring-setup-guideType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mohitagw15856/pm-claude-skills monitoring-setup-guide --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/monitoring-setup-guide .agents/skills/monitoring-setup-guide && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "monitoring-setup-guide" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/monitoring-setup-guide into .agents/skills/monitoring-setup-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-setup-guide", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mohitagw15856/pm-claude-skills monitoring-setup-guide --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/monitoring-setup-guide .cursor/skills/monitoring-setup-guide && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "monitoring-setup-guide" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/monitoring-setup-guide into .cursor/skills/monitoring-setup-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-setup-guide", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mohitagw15856/pm-claude-skills.git --path skills/monitoring-setup-guide--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mohitagw15856/pm-claude-skills monitoring-setup-guide --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/monitoring-setup-guide .gemini/skills/monitoring-setup-guide && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "monitoring-setup-guide" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/monitoring-setup-guide into .gemini/skills/monitoring-setup-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-setup-guide", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mohitagw15856/pm-claude-skills monitoring-setup-guideInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/monitoring-setup-guide .github/skills/monitoring-setup-guide && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "monitoring-setup-guide" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/monitoring-setup-guide into .github/skills/monitoring-setup-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-setup-guide", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mohitagw15856/pm-claude-skills monitoring-setup-guide --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/monitoring-setup-guide .opencode/skills/monitoring-setup-guide && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "monitoring-setup-guide" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/monitoring-setup-guide into .opencode/skills/monitoring-setup-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-setup-guide", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
monitoring-setup-guideWrite a monitoring setup guide for a service — defining what to measure, how to alert on it, and how to build the observability stack covering the four golden signals, business metrics, log…
Monitoring Setup Guide is an agent skill from mohitagw15856/pm-claude-skills. Write a monitoring setup guide for a service — defining what to measure, how to alert on it, and how to build the observability stack covering the four golden signals, business metrics, log strategy, distributed tracing, alerting rules, dashboard layout, and observability debt. Use when asked to set up monitoring for a service, define alerting strategy, write an observability plan, create a dashboard specification, or document logging standards for a team. Produces a metric definitions table, alert rules…
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Monitoring and alerting, Observability and UI design. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are json, python and yaml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Monitoring Setup Guide loads about 5.4k tokens when it runs. Until then it costs about 160 tokens; SKILL.md has 1,652 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 1,652 words, ~5,439 tokens.
.claude/skills/monitoring-setup-guide/SKILL.md (or your agent's skills folder).Produce a complete monitoring setup guide for a service — defining exactly what to measure, how to structure logs, how to configure alerts with actionable thresholds, and how to build dashboards that answer real operational questions. A good monitoring guide eliminates "we don't know what's happening in production" as a root cause category, and gives on-call engineers a single source of truth for what healthy looks like.
Ask for these if not already provided:
Team: [Team name] | Tech lead: [Name] Stack: [Language/Framework] on [Infrastructure] Monitoring platform: [Datadog / Prometheus+Grafana / CloudWatch / etc.] Date: [Date] | Review cycle: Quarterly
Good monitoring answers three questions:
This guide defines the answers for [Service Name]. Every alert must be actionable — if an on-call engineer cannot take a specific action in response to the alert, the alert should not exist.
Key user journeys monitored:
Apply the four golden signals specifically to [Service Name]:
Latency measures how long requests take to complete. Track it separately for successful and failed requests — slow failures hide behind fast errors if you only measure aggregate latency.
| Metric | Description | Source | Dimensions |
|---|---|---|---|
[service].request.duration_ms | End-to-end request latency | Application instrumentation | endpoint, method, status_code |
[service].db.query_duration_ms | Database query latency | ORM / query instrumentation | query_name, table |
[service].external.request_duration_ms | Outbound call latency to dependencies | HTTP client instrumentation | target_service, endpoint |
[service].queue.processing_duration_ms | Time to process one message (if applicable) | Consumer instrumentation | queue_name, message_type |
Latency SLO targets:
| Endpoint / operation | p50 target | p95 target | p99 target |
|---|---|---|---|
GET /api/v1/[resource] | < [50] ms | < [200] ms | < [500] ms |
POST /api/v1/[resource] | < [100] ms | < [400] ms | < [1000] ms |
GET /health | < [10] ms | < [20] ms | < [50] ms |
| [Background job name] | < [5] sec | < [15] sec | < [60] sec |
Traffic measures demand on the system. Use it to detect unexpected spikes, traffic drops (which can indicate upstream failures), and to capacity-plan.
| Metric | Description | Source |
|---|---|---|
[service].request.count | Requests per second | Application / load balancer |
[service].request.count_by_endpoint | RPS broken down by endpoint | Application |
[service].queue.messages_consumed_per_second | Consumer throughput | Queue consumer |
[service].queue.depth | Messages waiting in queue | Queue metrics |
Traffic baselines (update after observing production for 2+ weeks):
| Time period | Expected RPS | Low-traffic floor | Spike ceiling |
|---|---|---|---|
| Peak (weekday business hours) | [N] RPS | [N × 0.5] RPS | [N × 5] RPS |
| Off-peak (nights/weekends) | [N × 0.2] RPS | [N × 0.05] RPS | [N] RPS |
Errors measure the fraction of requests that fail. Distinguish between client errors (4xx — caller is doing something wrong) and server errors (5xx — the service is broken).
| Metric | Description | Alert on? |
|---|---|---|
[service].request.error_rate | 5xx errors / total requests | Yes — see alert rules |
[service].request.client_error_rate | 4xx errors / total requests | Threshold alert — sudden spike may indicate API misuse |
[service].dependency.error_rate | Errors calling downstream dependencies | Yes — upstream health signal |
[service].queue.dlq_depth | Messages in dead-letter queue | Yes — indicates processing failures |
Saturation measures how "full" the service is — how close to maximum capacity are the constrained resources.
| Resource | Metric | Alert threshold | Source |
|---|---|---|---|
| CPU | [service].cpu.utilisation_pct | >80% sustained 5 min | Container / VM metrics |
| Memory | [service].memory.utilisation_pct | >85% sustained 5 min | Container / VM metrics |
| DB connections | [service].db.connection_pool.utilisation_pct | >75% | Application / DB metrics |
| Thread pool / goroutines | [service].runtime.goroutine_count / thread_count | >N (establish baseline) | Runtime metrics |
| Disk (if applicable) | [service].disk.utilisation_pct | >75% | Infrastructure |
| Queue depth (if applicable) | [service].queue.depth | >[backlog threshold] | Queue metrics |
Beyond the golden signals, track metrics that measure whether the service is delivering business value. These matter for SLO reporting and product dashboards.
| Metric | Description | Source | Alert? |
|---|---|---|---|
[service].[primary_action].success_rate | [e.g. "Payment success rate"] | Application | Yes — if drops >5% vs 1h average |
[service].[primary_action].count | [e.g. "Payments processed per minute"] | Application | Yes — sudden drop (traffic anomaly) |
[service].[resource].created_per_hour | [e.g. "New accounts created"] | Application / DB | No — informational |
[service].cache.hit_rate | Fraction of requests served from cache | Cache instrumentation | Yes — if drops below [60]% |
[service].job.[name].success_rate | [Background job success rate] | Job framework | Yes — if drops below [99]% |
All logs must be structured JSON. Do not emit unstructured text logs in production. Every log line must include the mandatory fields.
Mandatory fields (every log line):
{
"timestamp": "2024-01-15T10:23:45.123Z",
"level": "info",
"service": "[service-name]",
"version": "[git-sha-short]",
"trace_id": "[uuid-from-request-context]",
"span_id": "[span-uuid]",
"request_id": "[uuid-per-request]",
"message": "[human readable description]"
}Request log (emit for every HTTP request):
{
"timestamp": "...",
"level": "info",
"service": "[service-name]",
"event": "http_request",
"method": "POST",
"path": "/api/v1/[resource]",
"status_code": 201,
"duration_ms": 45,
"user_id": "[uuid — DO NOT log PII directly]",
"request_id": "[uuid]",
"trace_id": "[uuid]"
}Error log (emit for every error with context):
{
"timestamp": "...",
"level": "error",
"service": "[service-name]",
"event": "error",
"error_code": "[application-error-code]",
"error_message": "[description — no sensitive data]",
"stack_trace": "[stack trace]",
"request_id": "[uuid]",
"trace_id": "[uuid]",
"context": {
"[key]": "[relevant context without PII]"
}
}| Level | Use when | Example |
|---|---|---|
error | Something failed that requires attention — this should page on-call eventually | Database query failed, external API returned 5xx, required config missing |
warn | Something unexpected happened but service is still functioning | Retry succeeded after failure, cache miss on expected hit, rate limit approaching |
info | Significant business events and request lifecycle | Request received, payment processed, user authenticated, job started/completed |
debug | Detailed diagnostic information — off in production by default | Query parameters, intermediate computation results, cache key lookups |
Never log:
GET /health from access logs)Distributed tracing is mandatory for any service that calls other services. It enables root-cause analysis across service boundaries.
[ ] Tracing library installed:
- Go: go.opentelemetry.io/otel
- Python: opentelemetry-sdk, opentelemetry-instrumentation
- Node: @opentelemetry/sdk-node
- Java: opentelemetry-java-instrumentation
[ ] Tracer initialized at service startup with service name and version
[ ] Trace context propagated via W3C Trace Context headers:
traceparent: 00-[trace-id]-[span-id]-01
tracestate: [optional vendor-specific]
[ ] Automatic instrumentation enabled for:
[ ] Inbound HTTP/gRPC requests (creates root span)
[ ] Outbound HTTP/gRPC calls (creates child spans)
[ ] Database queries (creates child spans with sanitized query)
[ ] Cache operations (Redis, Memcached)
[ ] Message queue produce/consume
[ ] Custom spans added for:
[ ] Key business operations ([e.g. payment processing, user lookup])
[ ] Background jobs (each job execution = root span)
[ ] Third-party API calls with custom attributes
[ ] Span attributes to capture on all spans:
- user.id (if authenticated — no PII)
- deployment.environment (production/staging)
- service.version (git SHA)
- [service-specific key attributes]
[ ] Trace exporter configured to: [Datadog / Jaeger / Tempo / OTLP endpoint]
[ ] Sampling rate configured:
- Production: [1–10]% of requests (adjust based on volume and cost)
- Always sample: errors, slow requests (>p99 threshold), and 100% of [critical endpoint]# Python — OpenTelemetry example
from opentelemetry import trace
tracer = trace.get_tracer("[service-name]")
def process_payment(payment_data):
with tracer.start_as_current_span("process_payment") as span:
span.set_attribute("payment.amount_cents", payment_data["amount"])
span.set_attribute("payment.currency", payment_data["currency"])
# Never: span.set_attribute("payment.card_number", ...)
try:
result = _do_process(payment_data)
span.set_status(trace.StatusCode.OK)
return result
except PaymentError as e:
span.set_status(trace.StatusCode.ERROR, str(e))
span.record_exception(e)
raiseEvery alert must have: a name, a condition, a threshold, a severity, and a clear on-call action. Alerts without a clear action should not exist.
| Alert name | Condition | Threshold | Severity | On-call action |
|---|---|---|---|---|
[Service]HighErrorRate | 5xx error rate, 5-min rolling window | >1% for 2 consecutive windows | P1 | Check recent deploys; inspect error logs; see runbook [link] |
[Service]CriticalErrorRate | 5xx error rate, 2-min rolling window | >5% | P1 — immediate | Same as above — page immediately, do not wait |
[Service]HighP99Latency | p99 latency on key endpoints | >2× SLO target for 3 min | P2 | Check DB latency, cache hit rate, and upstream dependencies |
[Service]LatencySLOBreach | p99 latency | >SLO target for 5 consecutive minutes | P1 | SLO burn — page on-call, escalate if not resolved in 20 min |
[Service]HighCPU | CPU utilisation | >80% sustained for 5 min | P2 | Check for traffic spike; scale up if needed; check for runaway processes |
[Service]HighMemory | Memory utilisation | >85% sustained for 5 min | P2 | Check for memory leak (especially after deploys); restart pod if OOM imminent |
[Service]DBConnectionPoolHigh | DB connection pool utilisation | >75% | P2 | Check for long-running queries; consider scaling service or increasing pool size |
[Service]DLQDepthHigh | Dead-letter queue depth | >10 messages | P2 | Inspect DLQ messages for error pattern; fix bug and replay if safe |
[Service]TrafficDropAnomaly | RPS, compared to same hour yesterday | >50% drop sustained 5 min | P1 | Upstream may be down; check caller health; check load balancer |
[Service]PrimaryActionSuccessRateDrop | [Business metric success rate] | <[95]% over 10 min | P1 | [Service-specific action — e.g. "Check payment provider status"] |
[Service]DownstreamDependencyErrors | Error rate calling [dependency] | >5% over 5 min | P2 | Check [dependency] status page; enable fallback if available |
# Prometheus / Grafana alerting rules (adapt for your platform)
groups:
- name: [service-name]-alerts
rules:
- alert: [Service]HighErrorRate
expr: |
(
sum(rate([service]_http_requests_total{status=~"5.."}[5m]))
/
sum(rate([service]_http_requests_total[5m]))
) > 0.01
for: 2m
labels:
severity: critical
team: [team-name]
annotations:
summary: "High error rate on [Service Name]"
description: "Error rate is {{ $value | humanizePercentage }} (threshold: 1%)"
runbook_url: "[runbook link]"
- alert: [Service]HighP99Latency
expr: |
histogram_quantile(0.99,
sum(rate([service]_http_request_duration_seconds_bucket[5m])) by (le, endpoint)
) > [0.5]
for: 3m
labels:
severity: warning
team: [team-name]
annotations:
summary: "p99 latency elevated on [Service Name]"
description: "p99 latency on {{ $labels.endpoint }} is {{ $value | humanizeDuration }}"
runbook_url: "[runbook link]"# Datadog monitor configuration (Python SDK or Terraform)
import datadog
datadog.initialize(api_key="[key]", app_key="[key]")
datadog.api.Monitor.create(
type="metric alert",
query=f"sum(last_5m):sum:{{service}}.http.errors{{service:[service-name]}} / sum:{{service}}.http.requests{{service:[service-name]}} > 0.01",
name="[Service] High Error Rate",
message="Error rate exceeded 1%. @pagerduty-[service-oncall]\n\nRunbook: [link]",
tags=["service:[service-name]", "team:[team-name]"],
options={
"thresholds": {"critical": 0.01, "warning": 0.005},
"notify_no_data": False,
"evaluation_delay": 60,
}
)The primary service dashboard must answer "is the service healthy right now?" at a glance. Use this layout:
┌─────────────────────────────────────────────────────────────────────┐
│ [SERVICE NAME] — Service Health Dashboard [Time range ▼] │
├───────────────┬───────────────┬───────────────┬─────────────────────┤
│ Error rate │ p99 Latency │ RPS (current)│ SLO budget remaining│
│ [BIG NUMBER] │ [BIG NUMBER] │ [BIG NUMBER] │ [BIG NUMBER / days] │
│ vs SLO: 0.1% │ vs SLO: 500ms│ vs avg: [N] │ [Error budget gauge]│
├───────────────┴───────────────┴───────────────┴─────────────────────┤
│ Error rate over time (24h) │
│ [Time series: 5xx rate line, SLO threshold line] │
├─────────────────────────────────┬───────────────────────────────────┤
│ Latency percentiles over time │ Request throughput over time │
│ [Lines: p50, p95, p99, p999] │ [Bars: RPS by endpoint] │
│ [SLO threshold horizontal line]│ │
├─────────────────────────────────┴───────────────────────────────────┤
│ Latency heatmap (all requests — shows distribution shape) │
├─────────────────────────────────┬───────────────────────────────────┤
│ CPU utilisation over time │ Memory utilisation over time │
│ [All instances/pods — lines] │ [All instances/pods — lines] │
│ [Alert threshold: 80%] │ [Alert threshold: 85%] │
├─────────────────────────────────┴───────────────────────────────────┤
│ DB: connection pool utilisation│ DB: query latency (p99 per query)│
├─────────────────────────────────┴───────────────────────────────────┤
│ [Business metric 1 over time] │ [Business metric 2 over time] │
│ e.g. Payment success rate │ e.g. Orders created/min │
└─────────────────────────────────┴───────────────────────────────────┘Second dashboard — Dependency Health:
┌─────────────────────────────────────────────────────────────────────┐
│ [SERVICE NAME] — Dependency Health │
├─────────────────────────────────────────────────────────────────────┤
│ For each dependency: error rate | latency | current status │
│ [Database] [N]% errors | [N]ms p99 | ● Healthy / ⚠ Degraded │
│ [Redis] [N]% errors | [N]ms p99 | ● Healthy │
│ [External API][N]% errors | [N]ms p99 | ● Healthy │
├─────────────────────────────────────────────────────────────────────┤
│ Outbound call latency over time (one line per dependency) │
├─────────────────────────────────────────────────────────────────────┤
│ Circuit breaker / fallback state (if implemented) │
└─────────────────────────────────────────────────────────────────────┘Honest assessment of what is missing today and what the priority to add it is:
| Gap | Impact | Priority | Effort | Owner | Target date |
|---|---|---|---|---|---|
| [e.g. No distributed tracing — can't see cross-service latency] | High — blind to dependency issues | P1 | [2 days] | [Name] | [Date] |
| [e.g. No business metric alerts — only infra alerts] | High — silent business failures | P1 | [1 day] | [Name] | [Date] |
| [e.g. Logs are unstructured text — not searchable] | Medium — slow incident investigation | P2 | [3 days] | [Name] | [Date] |
| [e.g. No dead-letter queue monitoring] | Medium — failed messages go unnoticed | P2 | [4 hours] | [Name] | [Date] |
| [e.g. Alert thresholds not calibrated to production baseline] | Medium — alert fatigue or missed alerts | P2 | [1 day] | [Name] | [Date] |
| [e.g. No latency heatmap — outliers invisible in averages] | Low — harder to spot tail latency issues | P3 | [2 hours] | [Name] | [Date] |
Total observability debt: [N] items | Estimated effort: [N days]
© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/monitoring-setup-guide of mohitagw15856/pm-claude-skills.
Open the folder on GitHubat commit 1cbf1f0
Monitoring Setup Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Monitoring Setup Guide this skillmohitagw15856/pm-claude-skills | 1.4k | — | ~5.4k | Automated safety check: Pass | MIT | |
| Axiom Dashboard Builderopenclaw/clawhub | 9.5k | — | ~4.9k | Automated safety check: Pass | MIT | |
| Happy Infra Metrics and Grafanaslopus/happy | 24k | — | ~2k | Automated safety check: Notes | MIT | |
| Axiom Cost Controlopenclaw/clawhub | 9.5k | — | ~1.7k | Automated safety check: Pass | MIT | |
| WizTelemetry Platform Servicekubesphere/kubesphere | 17k | — | ~1.8k | Automated safety check: Pass | Custom licence | |
| Sentry Miniapp SDKlizhiyao/sentry-miniapp | 686 | — | ~5k | Automated safety check: Pass | MIT |
openclaw/clawhub
Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana.
slopus/happy
Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.
openclaw/clawhub
Finds unused data in Axiom by analyzing query patterns, then deploys a cost dashboard and ingest monitors to keep spend under the contract limit.
kubesphere/kubesphere
Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions.
lizhiyao/sentry-miniapp
Full Sentry SDK setup for Mini Programs — error monitoring, tracing, offline cache, source maps.
openclaw/clawhub
Explores and queries OpenTelemetry metrics in Axiom MetricsDB, listing datasets, metrics and tags first and picking the right aggregation for each metric's type.
mohitagw15856/pm-claude-skills
Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.
mohitagw15856/pm-claude-skills
Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.
mohitagw15856/pm-claude-skills
Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.
mohitagw15856/pm-claude-skills
Compare two or more job offers as total-comp curves over four years — vesting cliffs, bonuses, 401(k) match, and the crossover year computed, not vibed.
mohitagw15856/pm-claude-skills
Compute the month a refinance actually starts saving money — payment delta, breakeven month, and total interest on both paths including the term-reset trap.
mohitagw15856/pm-claude-skills
Model rent-vs-buy honestly — year-by-year net position for both paths including the assumption everyone drops (the renter invests the difference), with a breakeven horizon instead of a verdict.
Categories
Write a monitoring setup guide for a service — defining what to measure, how to alert on it, and how to build the observability stack covering the four golden signals, business metrics, log…. Monitoring Setup Guide is an agent skill from mohitagw15856/pm-claude-skills. Write a monitoring setup guide for a service — defining what to measure, how to alert on it, and how to build the observability stack covering the four golden signals, business metrics, log strategy, distributed tracing, alerting rules, dashboard layout, and observability debt.
Monitoring Setup Guide fits situations like: asked to set up monitoring for a service; define alerting strategy; write an observability plan; create a dashboard specification.
Run `npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a claude-code`. Or copy the skill folder (skills/monitoring-setup-guide in mohitagw15856/pm-claude-skills) into .claude/skills/monitoring-setup-guide in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a codex`. Or copy the skill folder (skills/monitoring-setup-guide in mohitagw15856/pm-claude-skills) into .agents/skills/monitoring-setup-guide in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill monitoring-setup-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitoring-setup-guide, .gemini/skills/monitoring-setup-guide, .github/skills/monitoring-setup-guide and .opencode/skills/monitoring-setup-guide in your project.
SKILL.md names no scripts, command-line tools or credentials: Monitoring Setup Guide is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Monitoring Setup Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Monitoring Setup Guide: Axiom Dashboard Builder (openclaw/clawhub, 9.5k stars), Happy Infra Metrics and Grafana (slopus/happy, 24k stars), Axiom Cost Control (openclaw/clawhub, 9.5k stars) and WizTelemetry Platform Service (kubesphere/kubesphere, 17k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,431 GitHub stars. The repository holds 1,322 skills in this directory. The repository was last updated on October 7, 2026.
Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.