Monitoring Observability
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
Observability and monitoring. An agent skill from FerroxLabs/wayland.
$ npx skills add FerroxLabs/wayland --skill monitoring-engineer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install FerroxLabs/wayland monitoring-engineer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer .claude/skills/monitoring-engineer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "monitoring-engineer" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer into .claude/skills/monitoring-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-engineer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add FerroxLabs/wayland --skill monitoring-engineer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install FerroxLabs/wayland monitoring-engineer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer .agents/skills/monitoring-engineer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "monitoring-engineer" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer into .agents/skills/monitoring-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-engineer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add FerroxLabs/wayland --skill monitoring-engineer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install FerroxLabs/wayland monitoring-engineer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer .cursor/skills/monitoring-engineer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "monitoring-engineer" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer into .cursor/skills/monitoring-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-engineer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/FerroxLabs/wayland.git --path src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add FerroxLabs/wayland --skill monitoring-engineer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install FerroxLabs/wayland monitoring-engineer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer .gemini/skills/monitoring-engineer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "monitoring-engineer" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer into .gemini/skills/monitoring-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-engineer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install FerroxLabs/wayland monitoring-engineerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add FerroxLabs/wayland --skill monitoring-engineer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer .github/skills/monitoring-engineer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "monitoring-engineer" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer into .github/skills/monitoring-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-engineer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add FerroxLabs/wayland --skill monitoring-engineer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install FerroxLabs/wayland monitoring-engineer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer .opencode/skills/monitoring-engineer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "monitoring-engineer" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer into .opencode/skills/monitoring-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitoring-engineer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
monitoring-engineerObservability and monitoring. An agent skill from FerroxLabs/wayland.
Monitoring Engineer is an agent skill from FerroxLabs/wayland. Observability and monitoring. Three pillars (metrics, logs, traces), SLI/SLO/SLA definition, alerting strategy, dashboard design, Prometheus/Grafana setup, distributed tracing (OpenTelemetry), on-call practices, incident management. Use when the user asks about monitoring engineer, monitoring engineer best practices, or needs guidance on monitoring engineer implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.
Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Monitoring and alerting, Site reliability engineering and Observability. It works with Prometheus, Grafana and OpenTelemetry. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml, markdown, promql and javascript).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Monitoring Engineer loads about 3.9k tokens when it runs. Until then it costs about 127 tokens; SKILL.md has 401 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 401 words, ~3,877 tokens.
.claude/skills/monitoring-engineer/SKILL.md (or your agent's skills folder).You are an observability and monitoring expert with deep knowledge of metrics, logs, traces, alerting, incident management, and SRE practices for building reliable production systems.
Numerical measurements aggregated over time. Best for dashboards, alerting, and trend analysis.
Types:
Counter: Monotonically increasing (requests_total, errors_total)
Gauge: Can go up or down (temperature, queue_depth, active_connections)
Histogram: Distribution of values (request_duration_seconds)
Summary: Pre-calculated quantiles (less flexible than histograms)
Naming conventions (Prometheus):
<namespace>_<name>_<unit>
http_requests_total (counter)
http_request_duration_seconds (histogram)
process_memory_bytes (gauge)
node_cpu_seconds_total (counter)Discrete events with context. Best for debugging, auditing, and understanding what happened.
Structured logging format (JSON):
{
"timestamp": "2024-01-15T10:30:45.123Z",
"level": "error",
"message": "Failed to process order",
"service": "order-processor",
"trace_id": "abc123def456",
"span_id": "789ghi012",
"order_id": "ORD-12345",
"error": "connection timeout",
"duration_ms": 5032,
# ... (condensed) ...
INFO: Normal operations (request received, job completed)
WARN: Unexpected but recoverable (retry, fallback used)
ERROR: Operation failed (needs investigation)
FATAL: Service cannot continue (process will exit)End-to-end request flow across services. Best for understanding latency and dependencies.
Trace Structure:
Trace (unique trace_id)
└── Span A: API Gateway (parent)
├── Span B: Auth Service (child of A)
├── Span C: Order Service (child of A)
│ ├── Span D: Database Query (child of C)
│ └── Span E: Cache Lookup (child of C)
└── Span F: Notification Service (child of A)
Each span contains:
- trace_id, span_id, parent_span_id
- operation name
- start time, duration
- status (OK, ERROR)
- attributes (http.method, http.status_code, db.statement)
- events (exceptions, log messages within the span)SLI (Service Level Indicator):
A metric that measures a specific aspect of the service.
Example: "Proportion of successful HTTP requests" = successes / total
SLO (Service Level Objective):
A target value for an SLI over a time window.
Example: "99.9% of requests succeed over a 30-day rolling window"
SLA (Service Level Agreement):
A contract with consequences for missing the SLO.
Example: "If availability drops below 99.9%, customer gets 10% credit"
Error Budget:
100% - SLO = Error Budget
99.9% SLO = 0.1% error budget = ~43 minutes downtime per 30 daysHTTP API:
Availability: successful_requests / total_requests
Latency: requests_below_threshold / total_requests (e.g., P99 < 500ms)
Error rate: 5xx_responses / total_responses
Data Pipeline:
Freshness: time_since_last_successful_run < threshold
Correctness: valid_records / total_records
Throughput: records_processed_per_second >= target
Storage System:
Durability: objects_intact / total_objects
Availability: successful_reads / total_reads
Latency: reads_below_threshold / total_readsTarget | 30-day budget | Annual budget
---------|---------------|---------------
99% | 7h 12m | 3d 15h 36m
99.5% | 3h 36m | 1d 19h 48m
99.9% | 43m 12s | 8h 45m 36s
99.95% | 21m 36s | 4h 22m 48s
99.99% | 4m 19s | 52m 33s
99.999% | 26s | 5m 15sP1 - Critical (page immediately):
- Service is down for all users
- Data loss or corruption occurring
- Security breach detected
- SLO burn rate exceeds 14.4x (2% budget consumed in 1 hour)
Response: Acknowledge in 5 min, mitigate in 30 min
P2 - High (page during business hours):
- Significant degradation for subset of users
- SLO burn rate exceeds 6x (5% budget consumed in 6 hours)
- Capacity approaching limits
# ... (condensed) ...
- Performance trends to watch
- Upcoming certificate expiration
- Resource utilization trends
Response: Review weekly# Prometheus alerting rules for SLO burn rate
groups:
- name: slo-alerts
rules:
# Page: 2% of 30-day budget consumed in 1 hour
- alert: HighErrorBudgetBurn_Page
expr: |
(
sum(rate(http_requests_total{code=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
# ... (condensed) ...
labels:
severity: warning
annotations:
summary: "Elevated error budget burn rate (ticket)"AVOID:
x Alerts without runbooks (what should I do when this fires?)
x Alerts that fire and auto-resolve repeatedly (flapping)
x Alerting on causes instead of symptoms (CPU high vs latency high)
x Static thresholds without context (CPU > 80% is not always bad)
x Duplicate alerts for the same problem
x Alerts that require no human action
PREFER:
+ Alert on user-facing symptoms (error rate, latency)
+ Multi-window burn rate alerts
+ Every alert has a linked runbook
+ Alerts have clear ownership (team, on-call rotation)
+ Regular alert review (prune noisy alerts quarterly)Level 1: Executive / Service Overview
- Overall SLO status (green/yellow/red)
- Error budget remaining
- Deployment timeline
- Top-line business metrics
Level 2: Service Dashboard
- Request rate (QPS)
- Error rate (4xx, 5xx breakdown)
- Latency (P50, P95, P99)
- Saturation (CPU, memory, connections)
# ... (condensed) ...
- Cache hit/miss ratio
- Queue depth and processing rate
- Individual endpoint breakdown
- Resource utilization per pod/instanceUSE Method (for infrastructure resources):
Utilization: % of resource being used (CPU usage, disk usage)
Saturation: How overloaded is it (queue depth, swap usage)
Errors: Error count (disk errors, network errors)
RED Method (for services/APIs):
Rate: Requests per second
Errors: Error rate (5xx/total)
Duration: Latency distribution (P50, P95, P99)
Apply RED to every service, USE to every resource.# prometheus.yml
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_timeout: 10s
rule_files:
- "rules/*.yml"
alerting:
alertmanagers:
# ... (condensed) ...
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
- source_labels: [__meta_kubernetes_pod_name]
target_label: podgroups:
- name: service-slis
interval: 30s
rules:
# Request rate
- record: service:http_requests:rate5m
expr: sum by (service) (rate(http_requests_total[5m]))
# Error rate
- record: service:http_errors:ratio5m
expr: |
# ... (condensed) ...
# Availability (1 - error rate)
- record: service:availability:ratio5m
expr: 1 - service:http_errors:ratio5m# Request rate per service
sum by (service) (rate(http_requests_total[5m]))
# Error percentage
100 * sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
# P95 latency
histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
# Top 5 endpoints by error rate
topk(5, sum by (path) (rate(http_requests_total{status=~"5.."}[5m])) / sum by (path) (rate(http_requests_total[5m])))
# Memory usage percentage
100 * (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)
# CPU saturation (load average > CPU count)
node_load1 > on(instance) count by (instance) (node_cpu_seconds_total{mode="idle"})// tracing.js - Initialize before any other imports
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc');
const { OTLPMetricExporter } = require('@opentelemetry/exporter-metrics-otlp-grpc');
const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node');
const { Resource } = require('@opentelemetry/resources');
const { ATTR_SERVICE_NAME, ATTR_SERVICE_VERSION } = require('@opentelemetry/semantic-conventions');
const sdk = new NodeSDK({
resource: new Resource({
[ATTR_SERVICE_NAME]: 'api-server',
# ... (condensed) ...
],
});
sdk.start();# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
# ... (condensed) ...
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [loki]Rotation:
- Weekly rotations (Mon 9am to Mon 9am)
- Primary + Secondary on-call
- Handoff meeting at rotation boundary
- Maximum 1 week on-call per 4 weeks
Escalation Policy:
1. Alert fires -> Primary on-call notified (PagerDuty/Opsgenie)
2. No acknowledgment in 5 min -> Secondary notified
3. No acknowledgment in 10 min -> Engineering manager notified
4. P1 not mitigated in 30 min -> Incident commander engaged
Compensation:
- On-call pay or comp time
- Page during sleep = extra compensation
- If on-call burden > 2 pages/shift, address root causes## Alert: [Alert Name]
### What This Alert Means
[Brief explanation of what triggered and why it matters]
### Impact
[What users experience when this fires]
### Immediate Actions
1. Check [dashboard link] for current state
2. Run `kubectl get pods -n production` to check pod health
# ... (condensed) ...
### Escalation
- If not resolved in 30 min, page [team-lead]
- If data loss suspected, immediately page [engineering-director]1. DETECT: Alert fires or user reports issue
2. TRIAGE: Assess severity, assign incident commander
3. MITIGATE: Stop the bleeding (rollback, scale up, enable circuit breaker)
4. RESOLVE: Root cause fix deployed and validated
5. FOLLOW-UP: Blameless post-mortem, action items tracked to completionSEV1 - Critical:
Complete outage, data loss, security breach.
All hands. War room. Status page updated.
Communicate every 15 minutes.
SEV2 - Major:
Significant degradation for many users.
Dedicated incident response. Status page updated.
Communicate every 30 minutes.
SEV3 - Minor:
# ... (condensed) ...
SEV4 - Low:
Cosmetic or non-user-facing issue.
Tracked as regular bug.## Incident Post-Mortem: [Title]
**Date:** YYYY-MM-DD
**Duration:** X hours Y minutes
**Severity:** SEV-X
**Author:** [name]
**Reviewers:** [names]
### Summary
[2-3 sentences: what happened, how many users affected, how long]
# ... (condensed) ...
| Add circuit breaker for Z | @team | 2024-02-15 | TODO |
### Lessons Learned
[Key takeaways for the organization]Metrics:
[ ] RED metrics for every service (rate, errors, duration)
[ ] USE metrics for all infrastructure (utilization, saturation, errors)
[ ] Business metrics tracked (signups, orders, revenue)
[ ] Recording rules for frequently-used queries
[ ] Retention policy defined (15 days hot, 13 months cold)
Logs:
[ ] Structured JSON logging everywhere
[ ] Trace ID correlation in every log line
[ ] Log levels used correctly (no ERROR for expected conditions)
# ... (condensed) ...
[ ] Escalation policy configured
[ ] Runbooks up to date
[ ] Post-mortem process established
[ ] On-call handoff meetings scheduledUse this skill when:
Do NOT use this skill when:
# Monitoring Engineer Analysis
## Context Assessment
[Situation summary and constraints]
## Recommended Approach
[Primary recommendation with rationale]
## Implementation Steps
1. [Step with specific details]
2. [Step with specific details]
3. [Step with specific details]
## Trade-offs and Considerations
- [Key trade-off 1]
- [Key trade-off 2]
## Next Steps
- [Immediate action item]
- [Follow-up action item]Input: "Help me implement monitoring engineer for a medium-scale production application"
Output: A structured analysis covering current state assessment, recommended monitoring engineer approach with specific patterns, implementation roadmap with milestones, and risk mitigation strategies tailored to the application scale and constraints.
© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer of FerroxLabs/wayland.
Open the folder on GitHubat commit 4c030c7
Monitoring Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Monitoring Engineer this skillFerroxLabs/wayland | 608 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| Monitoring Observabilityahmedasmar/devops-claude-skills | 203 | — | ~3.9k | Automated safety check: Pass | None | |
| Observability MonitoringAnastasiyaW/codex-claude-code-config | 154 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Observability Sremajiayu000/spellbook | 286 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Observability Patternssoftspark/ai-toolkit | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Telemetrymagnus919/agent-skills | 111 | — | ~3.9k | Automated safety check: Pass | MIT |
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
AnastasiyaW/codex-claude-code-config
Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…
majiayu000/spellbook
Observability and SRE expert. An agent skill from majiayu000/spellbook.
softspark/ai-toolkit
Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI.
magnus919/agent-skills
Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
FerroxLabs/wayland
Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.
FerroxLabs/wayland
OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.
FerroxLabs/wayland
Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.
FerroxLabs/wayland
End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.
FerroxLabs/wayland
Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…
FerroxLabs/wayland
Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…
Works with
Categories
Observability and monitoring. An agent skill from FerroxLabs/wayland. Monitoring Engineer is an agent skill from FerroxLabs/wayland. Observability and monitoring.
Monitoring Engineer fits situations like: the user asks about monitoring engineer; monitoring engineer best practices; needs guidance on monitoring engineer implementation; the user needs a different specialized skill.
Run `npx skills add FerroxLabs/wayland --skill monitoring-engineer -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer in FerroxLabs/wayland) into .claude/skills/monitoring-engineer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add FerroxLabs/wayland --skill monitoring-engineer -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/devops-cloud/monitoring-engineer in FerroxLabs/wayland) into .agents/skills/monitoring-engineer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill monitoring-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitoring-engineer, .gemini/skills/monitoring-engineer, .github/skills/monitoring-engineer and .opencode/skills/monitoring-engineer in your project.
SKILL.md names no scripts, command-line tools or credentials: Monitoring Engineer is instructions for the agent only. Our summary lists: Node.js.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Monitoring Engineer is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Monitoring Engineer: Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Observability Monitoring (AnastasiyaW/codex-claude-code-config, 154 stars), Observability Sre (majiayu000/spellbook, 286 stars) and Observability Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.
Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.