Service Mesh Observability
wshobson/agents
Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.
Scheduled probes that run CONTINUOUSLY after release. An agent skill from petrkindlmann/qa-skills.
$ npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install petrkindlmann/qa-skills synthetic-monitoring --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/synthetic-monitoring .claude/skills/synthetic-monitoring && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "synthetic-monitoring" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/synthetic-monitoring into .claude/skills/synthetic-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-monitoring", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/petrkindlmann/qa-skills/tree/main/skills/synthetic-monitoringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install petrkindlmann/qa-skills synthetic-monitoring --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/synthetic-monitoring .agents/skills/synthetic-monitoring && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "synthetic-monitoring" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/synthetic-monitoring into .agents/skills/synthetic-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-monitoring", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install petrkindlmann/qa-skills synthetic-monitoring --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/synthetic-monitoring .cursor/skills/synthetic-monitoring && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "synthetic-monitoring" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/synthetic-monitoring into .cursor/skills/synthetic-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-monitoring", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/petrkindlmann/qa-skills.git --path skills/synthetic-monitoring--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install petrkindlmann/qa-skills synthetic-monitoring --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/synthetic-monitoring .gemini/skills/synthetic-monitoring && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "synthetic-monitoring" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/synthetic-monitoring into .gemini/skills/synthetic-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-monitoring", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install petrkindlmann/qa-skills synthetic-monitoringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/synthetic-monitoring .github/skills/synthetic-monitoring && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "synthetic-monitoring" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/synthetic-monitoring into .github/skills/synthetic-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-monitoring", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install petrkindlmann/qa-skills synthetic-monitoring --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/synthetic-monitoring .opencode/skills/synthetic-monitoring && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "synthetic-monitoring" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/synthetic-monitoring into .opencode/skills/synthetic-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-monitoring", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
synthetic-monitoringScheduled probes that run CONTINUOUSLY after release. An agent skill from petrkindlmann/qa-skills.
Synthetic Monitoring is an agent skill from petrkindlmann/qa-skills. Scheduled probes that run CONTINUOUSLY after release. Covers probe design for critical user journeys, alerting integration, SLA validation, multi-region monitoring, and the boundary between QA and SRE. Use when: "synthetic monitoring," "uptime testing," "scheduled probes," "SLA validation," "availability monitoring," "post-deploy checks." Not for: safe-release techniques during rollout — use testing-in-production. Not for: designing tests from prod telemetry — use observability-driven-testing. Not for: a one-shot…
Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/platforms-and-ci.md` and `references/probe-implementations.md`).
It sits in DevOps & Cloud, covering Monitoring and alerting, Feature launches and release readiness and Site reliability engineering. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
sre.googleopenslo.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Synthetic Monitoring loads about 5.8k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 172 tokens; SKILL.md has 2,506 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,506 words, ~5,762 tokens.
.claude/skills/synthetic-monitoring/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.<objective>
Synthetic monitoring runs scripted tests against production on a schedule, 24/7. It catches outages, performance degradation, and broken flows before real users report them — at 3 AM when traffic is zero, probes are the only thing checking your app works. A login page that always returns 200 but never authenticates passes a naive uptime check; a synthetic probe that asserts the dashboard loads catches it. This skill covers probe design, alerting integration, SLA validation, multi-region execution, and the runbook discipline that keeps a 3 AM page actionable.
</objective>
| Situation | Go to |
|---|---|
| Picking a platform (Checkly, Datadog, Grafana, CloudWatch…) | Platform Options |
| Writing a probe (login, API health, search) | Probe Design → references/probe-implementations.md |
| Probe runs but users still report outages | Failure Modes |
| Alerts too noisy or missing real outages | Alerting Integration |
| Calculating downtime budget for an SLA | SLA Validation |
| Probe failed at 3 AM and on-call is lost | runbook template in references/platforms-and-ci.md |
Check .agents/qa-project-context.md first. If it exists, use it as context and skip questions already answered there. Each answer changes the probe set, the alerting config, or the SLA math.
Critical flows (decides which probes you write):
Current monitoring (decides where the gaps are):
Infrastructure (decides regions and CDN assertions):
x-cache/origin headers.Test accounts (decides data safety):
RUM tells you what happened; synthetic tells you what is happening right now, whether or not real users are active. At 3 AM when traffic is zero, synthetic probes are the only thing checking the app works. Complement RUM, never replace it: synthetic covers known paths with predictable inputs, RUM discovers the creative ways real users break things.
A probe that takes 2 minutes and touches 15 pages is not a probe — it is an E2E suite running in production. Probes are short (per-probe wall-clock budget under 30 seconds), focused (one critical path each), and stable (retries: 0, zero flakiness tolerance). The 30-second ceiling is a hard budget: a slow probe that "still passes" is masking a degradation users feel.
A single probe failure is noise — network blips and DNS hiccups cause them constantly. Two consecutive failures are a signal; failures from 2+ regions confirm it is not local. Configure consecutive-failure and multi-region thresholds, or alert fatigue teaches the team to ignore the pager.
Probes run every few minutes, 24/7. Even small side effects (a created record, an incremented counter) compound. Probes must be non-destructive, created data cleaned up immediately, and synthetic traffic excluded from analytics, billing, and — critically — from SLO/error-budget math itself, or probes inflate your own reliability numbers.
A login page that returns 200 but never authenticates is broken. A search page that returns 200 with zero results is broken. Probes assert that the user can accomplish their goal — data loads, auth succeeds, results appear — not merely that the page returns a 2xx.
Design probes around critical user journeys, not infrastructure components. Users do not care if your load balancer is healthy — they care if they can log in and use the product.
| Probe | What It Validates | Frequency | Timeout |
|---|---|---|---|
| Homepage load | DNS, CDN, server, basic rendering | 1 min | 10s |
| Login flow | Authentication service, session management | 5 min | 15s |
| Core workflow | Primary value-delivering action (create document, run report) | 5 min | 20s |
| API health | Backend services, database connectivity | 1 min | 5s |
| Search | Search index, query processing, result rendering | 5 min | 15s |
| Checkout (if applicable) | Payment integration (sandbox mode), cart, order creation | 10 min | 25s |
| Third-party integrations | OAuth providers, email delivery, file storage | 10 min | 15s |
Three probe shapes cover most needs: a login flow (Playwright browser), an API health check (status + auth + latency budget), and a search probe (fill query → submit → assert on results, not just status). Keep each to one critical path, a tight timeout, and retries: 0. See references/probe-implementations.md for runnable login-flow, API-health, and search probes plus the environment-aware config.
Safe patterns:
- Read-only operations: GET requests, page loads, searches
- Sandbox transactions: payment in test mode, email to internal addresses
- Create-then-delete: create a draft, verify, delete immediately
- Synthetic flag: operations tagged as synthetic, excluded from processing
Dangerous patterns (avoid):
- Creating real orders, tickets, or user-facing records
- Triggering real notifications (email, SMS, push)
- Modifying shared resources (config, permissions, settings)
- Operations that cannot be automatically cleaned upAccount requirements:
- Clearly identifiable: email contains "synthetic" or "monitor"
- Flagged in database with an is_synthetic flag (is_synthetic = true)
- Excluded from analytics, billing, email campaigns, support queues
- Excluded from RUM and SLO/error-budget pipelines (synthetic traffic is not real traffic)
- Pre-populated with stable test data that probes can rely on
- Credentials in a secrets manager (process.env), rotated quarterly
- Separate accounts per concurrent probe (avoid state conflicts)| Platform | Strengths | When to Use |
|---|---|---|
| Checkly | Playwright-native, code-first, Git integration; Rocky AI agent (GA 2026) now does automated root-cause analysis across Playwright/API/Multistep/TCP/DNS/ICMP checks; CLI access from any AI agent; MCP server | Teams already using Playwright for E2E |
| Datadog Synthetic | Deep APM integration, browser and API tests | Teams on the Datadog platform |
| Grafana Synthetic Monitoring | Open source, integrates with Grafana dashboards; pairs with k6 2.x; pin a version channel (v1.x/v2.x) for reproducibility | Teams using the Grafana stack |
| AWS CloudWatch Synthetics | Blueprints (heartbeat, API, broken-link, visual diff); Python or Node Puppeteer canaries | Teams already on AWS |
| New Relic Synthetics | Full-stack observability integration | Teams on the New Relic platform |
| Better Stack | Lightweight uptime + status pages + on-call | SMB-friendly, fast setup |
| Uptime Kuma | OSS, self-hosted, lightweight | Self-hosting requirement, small surface area |
| Custom (Playwright / k6 / Puppeteer + cron) | Full control, no vendor lock-in | Budget-constrained or custom requirements |
Probes can be authored in Playwright (TS/JS), Puppeteer, k6 (JS — k6 2.0 shipped May 2026 with AI-assisted test authoring and a clearer Assertions API; first-class for synthetic), or Python (CloudWatch Synthetics, Checkly). Pick what your team already maintains. Avoid: Pingdom for new setups — it is a legacy uptime tool; prefer code-first alternatives (Checkly, Grafana, custom Playwright) that version-control probes alongside your app.
Self-managed: schedule Playwright probes with a GitHub Actions schedule cron (every 5 minutes), inject prod credentials via secrets, and report results to a monitoring webhook. Managed: Checkly is Playwright-native and runs from multiple locations on a fixed frequency. Either way, make probes environment-aware so the same code runs against staging and production with different base URLs and thresholds. See references/platforms-and-ci.md for the GitHub Actions workflow, the Checkly config, the alert-routing rules, and the runbook template.
Not every probe failure is an incident. Configure rules that cut noise while catching real problems.
Alerting rules:
- Single failure: log, do not alert (transient network issue)
- 2 consecutive failures from same region: warning (possible issue)
- 2 consecutive failures from 2+ regions: alert on-call (confirmed outage)
- Latency >2x baseline for 10 minutes: warning (performance degradation)
- Latency >3x baseline for 5 minutes: alert on-call (severe degradation)
- Any probe timeout: alert if 3 consecutive (service unresponsive)Routing. Critical failures on revenue paths (login, checkout, api-health) page on-call (PagerDuty + Slack incidents) with a short repeat interval; warnings on secondary probes go to a Slack monitoring channel; info-level events use a long repeat interval. The probe must tag each result with a severity and probe label for these routes to match — see the routing note in references/platforms-and-ci.md for the full alerting-rules.yaml and the tagging step.
Suppress synthetic alerts during planned maintenance. A maintenance window should silence synthetic paging (the probes will fail by design) and exclude that window from error-budget math, or scheduled work burns budget and pages on-call for nothing.
Alert message format — include enough context to start investigating immediately:
Alert template:
Title: [SYNTHETIC] {probe_name} failing from {region}
Severity: {critical|warning|info}
Consecutive failures: {count}
Last success: {timestamp}
Error: {error_message}
Duration: {last_response_time_ms}ms (threshold: {threshold}ms)
Dashboard: {link_to_dashboard}
Runbook: {link_to_runbook}
Regions affected: {list_of_failing_regions}The {link_to_runbook} points at a per-probe runbook (six lines: what it tests, first checks, manual repro, escalation, dashboard, owner). See the template in references/platforms-and-ci.md.
Availability = (total_minutes - downtime_minutes) / total_minutes × 100
Where:
- total_minutes = calendar month in minutes (43,200 for a 30-day month)
- downtime_minutes = minutes where synthetic probes detected failure
SLA tiers (downtime/month, monthly basis on 43,200 min):
99.0% = 432 min = 7h 12min (basic web app)
99.9% = 43.2 min = 43min 12s (business application)
99.95% = 21.6 min = 21min 36s (critical SaaS)
99.99% = 4.32 min = 4min 19s (infrastructure/platform)These are common availability targets, not prescriptive tiers. Pick targets from a user-impact analysis, not by tier name. Modern practice (Google SRE Workbook, OpenSLO) favors explicit SLO + error-budget policies over labelled tiers — define what user-visible failure looks like, set the budget user impact tolerates, and let the SLO follow. References: https://sre.google/workbook/ ; https://openslo.com/
Track percentiles, not averages — averages hide the worst experiences.
SLA response time targets (example):
Homepage load: P50 < 1s, P95 < 3s, P99 < 5s
API response: P50 < 200ms, P95 < 500ms, P99 < 1s
Search results: P50 < 500ms, P95 < 2s, P99 < 4s
Login flow: P50 < 2s, P95 < 5s, P99 < 8sError budget connects SLA targets to engineering decisions.
Error budget calculation:
SLO: 99.9% availability on a 43,200-min month
Budget: 0.1% of total time = 43.2 minutes/month
Budget consumed this month: 12 minutes (28%)
Budget remaining: 31.2 minutes (72%)
Actions by budget status:
>50% remaining: normal operations, ship features
25-50% remaining: caution, review recent changes
<25% remaining: freeze non-critical deploys, focus on reliability
Budget exhausted: incident mode, every deploy needs extra scrutinyExclude synthetic-probe downtime caused by your own maintenance windows from this calculation, or planned work shows as budget burn.
Run probes from regions where your users are. A service that works from us-east-1 but is broken from ap-southeast-1 is broken for APAC users.
Region selection strategy:
- Minimum 3 regions for global services
- Always include: closest to primary infrastructure, largest user base, farthest from primary
- Example for US-primary service: us-east-1, eu-west-1, ap-southeast-1
- Example for EU-primary service: eu-west-1, us-east-1, ap-northeast-1Track regional latency separately — a global average hides regional degradation.
Regional latency dashboard:
Region | P50 | P95 | Status
us-east-1 | 120ms | 340ms | healthy
eu-west-1 | 280ms | 620ms | healthy
ap-southeast | 450ms | 1200ms | warning (P95 above threshold)
Alert when:
- Any region's P95 exceeds its regional threshold
- Latency difference between regions exceeds 5x (CDN or routing issue)
- A region that was healthy becomes consistently degradedCDN validation. Synthetic probes verify CDN caching by checking response headers (x-cache, cf-cache-status) for HIT and confirming the server header matches the expected provider. This catches CDN misconfigurations — and origin failures masked by a stale cache — before users hit slow uncached responses.
A probe that navigates 10 pages, fills 5 forms, and asserts on 20 elements is an E2E test, not a synthetic probe. When it breaks, you cannot tell if the app is down or the probe is flaky. Fix: one critical path per probe, under 30 seconds, under 5 assertions. A failure should make it immediately clear what is broken.
Network blips, DNS hiccups, and transient cloud issues cause occasional failures. Alerting on every one produces noise that teaches the team to ignore alerts. Fix: require 2-3 consecutive failures and failures from 2+ regions before paging. Escalating severity: first failure logs, second warns, third pages.
Probes sharing one account interfere — one probe changes a setting, another fails because it expected the default. Fix: one dedicated synthetic account per concurrent probe, flagged synthetic, excluded from analytics and billing.
Probes that only check "page loads, returns 200" miss broken functionality behind a loading page. A login that returns 200 but never authenticates is not working. Fix: assert on meaningful content — data loads, auth succeeds, the core action completes, search returns results. One "can the user accomplish their goal" probe is worth ten "does the page return 200" probes.
An alert fires at 3 AM. On-call sees "Login probe failing" but has no idea what to check first, what the probe does, or how to tell a real outage from a probe issue. Fix: every probe links a runbook (what it tests; first checks — third-party status, recent deploys, app telemetry; manual repro; escalation; dashboard link). See the template in references/platforms-and-ci.md.
Emerging pattern (2026): agent-driven first response. Tools like Checkly Rocky (GA 2026, automated root-cause analysis across check types) and Honeycomb Canvas Skills (Agent Observability, launched May 2026; open-source honeycombio/agent-skill repo ships a honeycomb-investigator and instrumentation-advisor for Claude Code and Cursor) read probe context + linked runbook + telemetry and post a candidate diagnosis to chat. Treat this as triage assist — it shortens MTTR for routine failures, but it does not replace on-call judgment for novel incidents.
| Symptom | Likely cause | Fix or check |
|---|---|---|
| Probe green but users report an outage | Probe asserts only HTTP 200, not the goal | Add content/state assertions (dashboard heading, search results count, body.status) |
| Flapping alerts (fire/resolve repeatedly) | No consecutive-failure or multi-region rule | Require 2+ consecutive failures and 2+ regions before paging |
| Probe passes locally, fails in CI/region | Region-specific outage or CDN routing | Compare per-region results; check latency-difference and x-cache assertions |
| Probe failures with no real outage | Synthetic-account state drift (shared account) | One isolated account per concurrent probe; reset/seed stable data |
| Origin is down but probe stays green | CDN serving stale cache | Assert cf-cache-status/origin server header, not just 200 |
| Error budget burns with no incident | Maintenance windows counted as downtime | Suppress synthetic paging during maintenance; exclude window from budget math |
| On-call paged but can't act | Missing/empty runbook on the alert | Populate the six-line runbook; verify {link_to_runbook} resolves |
| Reliability numbers look too good | Synthetic traffic counted as real in SLO/RUM | Exclude is_synthetic traffic from RUM, billing, and error-budget pipelines |
Prove the monitoring works before trusting it — smallest check first.
npx playwright test probes/ --reporter=list against staging and confirm every probe passes within its declared timeout (no probe exceeds the 30s ceiling).503 or a wrong path), let it run the configured consecutive-failure count, and confirm the alert routes to on-call within the SLA detection time — and resolves when you revert.is_synthetic traffic and confirm it is filtered out.retries: 0.severity + probe so routes match.references/)© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/synthetic-monitoring of petrkindlmann/qa-skills.
Open the folder on GitHubat commit b3bb61b
Synthetic Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Synthetic Monitoring this skillpetrkindlmann/qa-skills | 170 | — | ~5.8k | Automated safety check: Pass | MIT | |
| Service Mesh Observabilitywshobson/agents | 40k | 9 repos | ~607 | Automated safety check: Pass | MIT | |
| Monitoring Observabilityahmedasmar/devops-claude-skills | 203 | — | ~3.9k | Automated safety check: Pass | None | |
| Prometheus Error Rate Investigatorprometheus/prometheus-mcp | 121 | — | ~592 | Automated safety check: Pass | Apache-2.0 | |
| Monitoring ExpertJeffallan/claude-skills | 12k | — | ~1.6k | Automated safety check: Pass | MIT | |
| Oma Observabilityfirst-fluke/oh-my-agent | 1.3k | — | ~4.9k | Automated safety check: Pass | MIT |
wshobson/agents
Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
prometheus/prometheus-mcp
Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.
Jeffallan/claude-skills
Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery.
first-fluke/oh-my-agent
Intent-based observability + traceability router across layers, boundaries, and signals.
elastic/agent-skills
Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health…
petrkindlmann/qa-skills
Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
petrkindlmann/qa-skills
Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.
petrkindlmann/qa-skills
Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Categories
Scheduled probes that run CONTINUOUSLY after release. An agent skill from petrkindlmann/qa-skills. Synthetic Monitoring is an agent skill from petrkindlmann/qa-skills. Scheduled probes that run CONTINUOUSLY after release.
Synthetic Monitoring fits situations like: : synthetic monitoring; scheduled probes; availability monitoring; post-deploy checks. Not for: safe-release techniques during rollout — use testing-in-production.
Run `npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a claude-code`. Or copy the skill folder (skills/synthetic-monitoring in petrkindlmann/qa-skills) into .claude/skills/synthetic-monitoring in your project. Claude Code loads it when a task matches its description.
Run `npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a codex`. Or copy the skill folder (skills/synthetic-monitoring in petrkindlmann/qa-skills) into .agents/skills/synthetic-monitoring in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill synthetic-monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/synthetic-monitoring, .gemini/skills/synthetic-monitoring, .github/skills/synthetic-monitoring and .opencode/skills/synthetic-monitoring in your project.
Going by SKILL.md and its folder, Synthetic Monitoring needs the command-line tools its instructions call (npx). Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: sre.google and openslo.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Synthetic Monitoring is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Synthetic Monitoring: Service Mesh Observability (wshobson/agents, 40k stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars), Prometheus Error Rate Investigator (prometheus/prometheus-mcp, 121 stars) and Monitoring Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 170 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.
Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.