Service Mesh Observability
wshobson/agents
Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.
Build a cloud, SLO, and incident-readiness register after intake.
$ npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sickn33/agentic-awesome-skills observability-cloud-planning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/observability-cloud-planning .claude/skills/observability-cloud-planning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "observability-cloud-planning" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/observability-cloud-planning into .claude/skills/observability-cloud-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-cloud-planning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/observability-cloud-planningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sickn33/agentic-awesome-skills observability-cloud-planning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/observability-cloud-planning .agents/skills/observability-cloud-planning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "observability-cloud-planning" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/observability-cloud-planning into .agents/skills/observability-cloud-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-cloud-planning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sickn33/agentic-awesome-skills observability-cloud-planning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/observability-cloud-planning .cursor/skills/observability-cloud-planning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "observability-cloud-planning" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/observability-cloud-planning into .cursor/skills/observability-cloud-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-cloud-planning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sickn33/agentic-awesome-skills.git --path skills/observability-cloud-planning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sickn33/agentic-awesome-skills observability-cloud-planning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/observability-cloud-planning .gemini/skills/observability-cloud-planning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "observability-cloud-planning" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/observability-cloud-planning into .gemini/skills/observability-cloud-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-cloud-planning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sickn33/agentic-awesome-skills observability-cloud-planningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/observability-cloud-planning .github/skills/observability-cloud-planning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "observability-cloud-planning" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/observability-cloud-planning into .github/skills/observability-cloud-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-cloud-planning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sickn33/agentic-awesome-skills observability-cloud-planning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/observability-cloud-planning .opencode/skills/observability-cloud-planning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "observability-cloud-planning" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/observability-cloud-planning into .opencode/skills/observability-cloud-planning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-cloud-planning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
observability-cloud-planningBuild a cloud, SLO, and incident-readiness register after intake.
Observability Cloud Planning is an agent skill from sickn33/agentic-awesome-skills. Build a cloud, SLO, and incident-readiness register after intake. Use when an SME needs monitoring scope, alert ownership, cost limits, or service planning.
Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/common-pitfalls.md`, `references/related-skills.md` and `references/reusable-prompt.md`).
It sits in DevOps & Cloud, covering Site reliability engineering and Observability. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 680176d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml, csv, sql, json and markdown).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
json-schema.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Observability Cloud Planning loads about 6.5k tokens when it runs, and up to ~8.1k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 2,666 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sickn33/agentic-awesome-skills at commit 680176d, republished under its MIT licence (© sickn33). 2,666 words, ~6,516 tokens.
.claude/skills/observability-cloud-planning/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.What it is: the plan for what the business runs on, what it watches, what wakes someone up, and what that costs - written before any of it is bought.
Works out the smallest monitoring plan that would actually catch this business's worst failure, then builds it only when asked. The default output is a short recommendation, not a service register. The register - CSV, SQL DDL, JSON Schema, Notion mapping - is produced on request, from one field list so the four cannot drift apart.
Layer: Layer 8: Operate. Fits: Growth stage. Table code: n/a.
The rule this table exists to enforce: an alert is a human cost, so Alert Channel and
Severity decide whether a person is woken at 3am. Everything else in this table exists to
make that decision possible: what the service does (SLI Definition), what good looks like
(SLO Target), what is measured (Dashboard URL), what it costs (Monthly Cost Estimate),
and - the field most tables leave out - what happens when the alert fires (Runbook URL).
An alert with no runbook converts a technical problem into a panicked one.
The second rule: Phase is the field that keeps this affordable. A ten-service stack
with four signals each, on a plan tier, ordered by what would hurt the business most. A
business with no developer should be running a managed platform and three alerts, not a
stack of collectors.
Do not use it for: application code, infrastructure as code, or a deployment pipeline
(devops-pipeline-designer, ci-cd-pipeline-builder in the engineering pack); a security
architecture review (security-and-privacy); or a real incident, which is a live
situation, not a plan.
Follow the shared execution contract. The module-specific rules below define only domain fields, decisions, calculations, and safety constraints.
Read the request and pick the intent before asking anything.
Ask only if this is the highest-value missing fact; otherwise proceed without an opener:
Q: What breaks for the business if the site or app is down for one hour, and who would notice first - a customer, a member of staff, or an automated system?
Treat ambiguous replies as unanswered and ask which explicit option the user means. Record unknown values as Unknown; Unknown is not zero. A record must not be Done when a required check fails.
Skip anything already answered. Ask the rest one at a time, and stop as soon as the remaining answers would not change the plan.
Never invent an answer. Traffic numbers, costs, hostnames, provider names, service
identifiers, error rates, uptime figures and recovery times the user has not supplied are
Unknown. A monitoring plan built on invented traffic or invented spend is worthless, and
invented SLOs are worse - they create an alert nobody can satisfy.
module: observability-cloud-planning
intent: null # set up | review | report | import
scale: null # Starter | Growth | Scale, only if the answer changes it
areas:
"Failure": null
"Stack": null
"Traffic": null
"People": null
"Data and money": null
requested_outputs: []
confirmed_facts: []
open_questions: []Build an already requested artifact without asking again. For advice-only requests, give a short recommendation and offer the relevant artifact.
Recommended approach: Start with the one thing whose failure stops the business earning - usually the checkout or the booking form - and watch exactly that from outside the infrastructure, so the check does not fail with the thing it is checking. Three signals and no more: is it reachable, is it answering within a target time, and is the thing that transacts money or data actually succeeding. One alert channel that reaches a human, a runbook for each alert, and a named owner. Everything else - traces, log aggregation, five custom dashboards, a status page - goes in a later phase once the first phase has proven someone reads it.
Why this one: The failure mode of observability is not too little monitoring, it is too much of the wrong. A business that installs fifteen alerts stops reading them within a week and returns to finding out from customers. And external monitoring matters more than internal: an uptime check that runs on the same box as the site reports the site is fine right up to the moment it is not.
Workflow: Worst failure named → What must be measured to catch it, from outside → Baseline measured, not assumed → Three signals defined with targets → Alerts with a channel and a runbook each → Costs estimated against a ceiling → Owner and review date set → Runbook rehearsed → Phase 2 only after phase 1 is being read → Quarterly review of alerts kept, deleted and added
Once the user asks for it, derive the fields from the confirmed context and emit the requested artifacts. For machine-readable text, keep prose outside the data; for files, provide a usable link. Report material validation failures or limitations separately.
A selected Notion output is rendered by notion-manual-import, so route the
Notion step there. When the user selects Notion, hand that step to
](https://github.com/sickn33/agentic-awesome-skills/blob/main/skills/notion-manual-import/SKILL.md): it holds the CSV, the property
mapping, the import steps and the verification checklist, and it renders the Field
Reference below instead of defining a table of its own. Do not restate the mapping
here and do not improvise the import steps. Manual CSV and mapping outputs need no
connection. For requested workspace changes, follow the shared contract: verify actual
tool access and the target before writing. A user saying "connected" is not tool evidence.
Never ask for a Notion password or token.
Service ID,Service Name,Purpose,Criticality,Owner,Environment,Monitored From,SLI Definition,SLO Target,Dashboard URL,Log Source,Alert Channel,Severity,Alert Threshold,Runbook URL,Escalation Path,Monthly Cost Estimate,Data Handled,Deployment Method,Dependencies,Phase,Review Date,Status,Notes
,Example Booking Service,Takes customer bookings and confirms by email,High,Example Owner,Production,External region outside the hosting provider,Uptime of the public booking page,99.5% monthly,Not yet created,Application and web server logs,Email and SMS to on-call,Critical,3 consecutive failed checks or 5 minutes above target,Not yet created,On-call then business owner,Unknown,Payment card details,Managed platform,None,1,2026-10-27,Not started,Example row - replace every value before use.-- Engine assumption: PostgreSQL. For another engine use the engine's auto-increment
-- equivalent and keep the rest portable.
CREATE TABLE obs_service (
service_id BIGINT PRIMARY KEY,
service_name VARCHAR(100) NOT NULL,
purpose TEXT NOT NULL,
criticality VARCHAR(50) NOT NULL,
owner VARCHAR(255) NOT NULL,
environment VARCHAR(50) NOT NULL,
monitored_from VARCHAR(100) NOT NULL,
sli_definition TEXT NOT NULL,
slo_target VARCHAR(50) NOT NULL,
dashboard_url VARCHAR(255),
log_source VARCHAR(255) NOT NULL,
alert_channel VARCHAR(255) NOT NULL,
severity VARCHAR(50) NOT NULL,
alert_threshold VARCHAR(255) NOT NULL,
runbook_url VARCHAR(255),
escalation_path TEXT,
monthly_cost_estimate NUMERIC(10,2),
data_handled VARCHAR(255) NOT NULL,
deployment_method VARCHAR(100) NOT NULL,
dependencies TEXT,
phase VARCHAR(50) NOT NULL,
review_date DATE,
status VARCHAR(50) NOT NULL,
notes TEXT,
created_at TIMESTAMP DEFAULT NOW(),
updated_at TIMESTAMP DEFAULT NOW(),
CONSTRAINT obs_service_phase CHECK (phase IN ('1 - core','2 - deepen','3 - optional','Sunset')),
CONSTRAINT obs_service_status CHECK (status IN ('Not started','In progress','Blocked','Done','Cancelled')),
CONSTRAINT obs_service_severity CHECK (severity IN ('Info','Warning','Critical')),
CONSTRAINT obs_service_criticality CHECK (criticality IN ('Low','Medium','High','Critical')),
CONSTRAINT obs_service_cost_non_negative CHECK (monthly_cost_estimate IS NULL OR monthly_cost_estimate >= 0)
);
CREATE INDEX idx_obs_service_status ON obs_service (status);
CREATE INDEX idx_obs_service_phase ON obs_service (phase);
CREATE INDEX idx_obs_service_review ON obs_service (review_date);{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "Cloud and Observability Plan",
"type": "object",
"additionalProperties": false,
"properties": {
"Service ID": { "type": "integer" },
"Service Name": { "type": "string" },
"Purpose": { "type": "string" },
"Criticality": { "type": "string" },
"Owner": { "type": "string" },
"Environment": { "type": "string" },
"Monitored From": { "type": "string" },
"SLI Definition": { "type": "string" },
"SLO Target": { "type": "string" },
"Dashboard URL": { "type": "string" },
"Log Source": { "type": "string" },
"Alert Channel": { "type": "string" },
"Severity": { "type": "string" },
"Alert Threshold": { "type": "string" },
"Runbook URL": { "type": "string" },
"Escalation Path": { "type": "string" },
"Monthly Cost Estimate": { "type": "number" },
"Data Handled": { "type": "string" },
"Deployment Method": { "type": "string" },
"Dependencies": { "type": "string" },
"Phase": { "type": "string" },
"Review Date": { "type": "string", "format": "date" },
"Status": { "type": "string" },
"Notes": { "type": "string" }
},
"required": [
"Service Name",
"Purpose",
"Criticality",
"Owner",
"Environment",
"Monitored From",
"SLI Definition",
"SLO Target",
"Log Source",
"Alert Channel",
"Severity",
"Alert Threshold",
"Data Handled",
"Deployment Method",
"Phase",
"Status"
]
}| CSV column | Notion property | Set after import |
|---|---|---|
| Service ID | Text (preserve source ID) | Keep imported IDs as Text; optionally add a separate Unique ID property |
| Service Name | Title | Use as the database title |
| Purpose | Text | Leave as Text. What the business loses if this stops. Not a technical description |
| Criticality | Select (add options after import) | Convert to Select, add options: "Low", "Medium", "High", "Critical". Critical means the business stops earning |
| Owner | Text | Leave as Text. A named person. A service with no owner is unmonitored in practice |
| Environment | Select (add options after import) | Convert to Select, add options: "Production", "Staging", "Development". Monitoring a staging URL and calling it uptime is a common and expensive mistake |
| Monitored From | Text | Leave as Text. Where the check runs. It must not run on the thing it is checking |
| SLI Definition | Text | Leave as Text. The measurement in plain words - "public booking page answers in under 2 seconds for 95% of requests" |
| SLO Target | Text | Leave as Text. The number and the window, e.g. "99.5% monthly". A target nobody has ever met is a target that gets deleted |
| Dashboard URL | Text | Leave as Text. Blank until a dashboard exists, which is the correct state in phase 1 |
| Log Source | Text | Leave as Text. Where the logs are today, including "none - not collected" |
| Alert Channel | Text | Leave as Text. The specific route: which email, which SMS, which app. "Alerts" is not a channel |
| Severity | Select (add options after import) | Convert to Select, add options: "Info", "Warning", "Critical". Only Critical should page a human outside working hours |
| Alert Threshold | Text | Leave as Text. The condition that fires, written so it is unambiguous and testable |
| Runbook URL | Text | Leave as Text. Blank is a finding, not a formatting gap. An alert with no runbook must not ship |
| Escalation Path | Text | Leave as Text. Who is woken first, then who, then who gives up and calls the host |
| Monthly Cost Estimate | Number | Convert to Number, two decimal places. An estimate against a ceiling, in the business's own currency |
| Data Handled | Select (add options after import) | Convert to Select, add options: "None", "Personal data", "Payment card details", "Health data", "Credentials or secrets", "Business confidential", "Public only". Drives the log retention and access rules |
| Deployment Method | Text | Leave as Text. How a change reaches this service, and by whom |
| Dependencies | Text | Leave as Text. What this needs to work. "None" is a valid and worth recording answer |
| Phase | Select (add options after import) | Convert to Select, add options: "1 - core", "2 - deepen", "3 - optional", "Sunset". The affordability control |
| Review Date | Date | Convert to Date. When the alerts, the target and the cost are re-examined |
| Status | Select (add options after import) | Convert to Select, add options: "Not started", "In progress", "Blocked", "Done", "Cancelled" |
| Notes | Text | Leave as Text |The rows above are documentation examples only. Emit empty templates unless the user explicitly requests examples. Dashboard URL and Runbook URL read Not yet created rather than a plausible-looking link, because a dead or wrong link in a runbook
column is worse than an obvious gap.
| # | Field | Type | SQL | JSON Schema | Notion | CSV example |
|---|---|---|---|---|---|---|
| 1 | Service ID | id | BIGINT PRIMARY KEY | integer | Text or Notion auto-ID | (blank) |
| 2 | Service Name | text | VARCHAR(100) | string | Text | (blank) |
| 3 | Purpose | long_text | TEXT | string | Text | (blank) |
| 4 | Criticality | select | VARCHAR(50) | string | Select | High |
| 5 | Owner | text | VARCHAR(255) | string | Text | (blank) |
| 6 | Environment | select | VARCHAR(50) | string | Select | Production |
| 7 | Monitored From | text | VARCHAR(100) | string | Text | (blank) |
| 8 | SLI Definition | long_text | TEXT | string | Text | (blank) |
| 9 | SLO Target | text | VARCHAR(50) | string | Text | 99.5% monthly |
| 10 | Dashboard URL | text | VARCHAR(255) | string | Text | (blank) |
| 11 | Log Source | text | VARCHAR(255) | string | Text | (blank) |
| 12 | Alert Channel | text | VARCHAR(255) | string | Text | (blank) |
| 13 | Severity | select | VARCHAR(50) | string | Select | Critical |
| 14 | Alert Threshold | text | VARCHAR(255) | string | Text | (blank) |
| 15 | Runbook URL | text | VARCHAR(255) | string | Text | (blank) |
| 16 | Escalation Path | long_text | TEXT | string | Text | (blank) |
| 17 | Monthly Cost Estimate | number | NUMERIC(10,2) | number | Number | Unknown |
| 18 | Data Handled | select | VARCHAR(255) | string | Select | Personal data |
| 19 | Deployment Method | text | VARCHAR(100) | string | Text | (blank) |
| 20 | Dependencies | long_text | TEXT | string | Text | None |
| 21 | Phase | select | VARCHAR(50) | string | Select | 1 - core |
| 22 | Review Date | date | DATE | string, format: date | Date | (blank) |
| 23 | Status | select | VARCHAR(50) | string | Select | Not started |
| 24 | Notes | long_text | TEXT | string | Text | (blank) |
Criticality - what the business loses. Critical means the business stops earning, not
that the service is technically important. A search index that goes down is Low; a
checkout that goes down is Critical.
Low | Medium | High | CriticalEnvironment - and the mistake this prevents: monitoring a staging URL and reporting it as uptime. Production is the only environment that earns an alert.
Production | Staging | DevelopmentSeverity - the difference between a notification and a woken person. Only Critical
should reach a human outside working hours, because a Warning that pages at 3am trains
the on-call to ignore the channel.
Info | Warning | CriticalData Handled - this is not documentation, it is an instruction. Payment card details,
Health data and Credentials or secrets change the log retention period, who may read
the logs, and whether the logs may be shipped to a third-party service at all.
None | Personal data | Payment card details | Health data | Credentials or secrets | Business confidential | Public onlyPhase - the field that keeps this affordable. Phase 1 is what catches the worst failure and is read. Nothing enters phase 2 until phase 1 is being read and acted on; a plan that ships all phases at once is a plan nobody maintains.
1 - core | 2 - deepen | 3 - optional | SunsetStatus - Blocked is the honest value whenever something depends on a person, a
contract or a decision that has not happened.
Not started | In progress | Blocked | Done | CancelledDependencies names other rows in this table. A cycle here is a design smell, not an
error, and it is worth finding deliberately.Log Source and Data Handled point at the data and access rules in
data-privacy-controls (me-ops-system-builder). Confirm the join before making it
live; a third-party log destination is a data transfer.Runbook URL points at runbook-playbook in the engineering pack. Blank there and blank
here is the same finding.Prompt
We run a booking site on a VPS. Last month it was down for a whole Sunday and we only
found out on Monday from a customer. We want to know if it goes down, and we do not know
what we should be paying for all this.Context first - one question per message, nothing already answered:
Q: What breaks for the business if the site or app is down for one hour, and who would notice first - a customer, a member of staff, or an automated system? A: The booking form. A customer, obviously.
Q: What is actually running today, and who set it up? A: A VPS with the site on it, set up by an agency we no longer deal with.
Q: Roughly how many visitors or bookings a day? A: Maybe 20 bookings. It is much busier in December.
Q: Is there anyone who could be on call, and at what hours? A: Me, evenings and weekends.
Q: What is the current monthly spend, roughly, and the ceiling? A: Forty a month for the server. I would not want to spend much more than that.
Recommended next step - offered, not built:
One external uptime check on the booking page, run from outside your provider, plus a check that a test booking actually completes. That catches the Sunday problem. Two alerts, both to your phone: the page unreachable for three checks running, and a test booking that fails. No dashboard, no log shipping, no tracing - the site does not earn that yet. A one-page runbook for each: what to check, how to restart it, and the agency number if it does not come back. December is your peak, so test it properly in November. Keep it near 40 a month, or you will pay more for storage than you do for the server.
Workflow: Worst failure named → What must be measured, from outside → Baseline measured → Three signals with targets → Alerts with a channel and a runbook each → Costs estimated → Owner and review date set → Runbook rehearsed → Phase 2 only if phase 1 is being read
Want the CSV, SQL DDL, JSON Schema and Notion mapping for this?
Critical pages a human. Everything else is a notification for working hours, or it
is deleted. Alert fatigue is the normal end state of an unmaintained alerting setup.Runbook URL is a finding that blocks going live.Data Handled exists to make that decision explicit before the logs are shipped anywhere.Monthly Cost Estimate as a planning figure, never as a
quote, and never as a provider price list.security-and-privacy.See the Security & Safety Notes reference for the full guidance.
© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/observability-cloud-planning of sickn33/agentic-awesome-skills.
Open the folder on GitHubat commit 680176d
We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.
Observability Cloud Planning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Observability Cloud Planning this skillsickn33/agentic-awesome-skills | 47k | 1 repos | ~6.5k | Automated safety check: Pass | MIT | |
| Service Mesh Observabilitywshobson/agents | 40k | 9 repos | ~607 | Automated safety check: Pass | MIT | |
| Observability2SSK/dot-files | 247 | — | ~676 | Automated safety check: Pass | MIT | |
| Monitoring Observabilityahmedasmar/devops-claude-skills | 203 | — | ~3.9k | Automated safety check: Pass | None | |
| Prometheus Error Rate Investigatorprometheus/prometheus-mcp | 120 | — | ~592 | Automated safety check: Pass | Apache-2.0 | |
| Error HandlerEliasOulkadi/shokunin | 114 | — | ~3.6k | Automated safety check: Notes | MIT |
wshobson/agents
Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality.
2SSK/dot-files
Observability best practices. An agent skill from 2SSK/dot-files.
ahmedasmar/devops-claude-skills
Monitoring and observability strategy, implementation, and troubleshooting.
prometheus/prometheus-mcp
Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.
EliasOulkadi/shokunin
Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…
davila7/claude-code-templates
Build production-ready monitoring, logging, and tracing systems.
sickn33/agentic-awesome-skills
Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.
sickn33/agentic-awesome-skills
Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.
sickn33/agentic-awesome-skills
Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.
sickn33/agentic-awesome-skills
Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.
sickn33/agentic-awesome-skills
Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.
sickn33/agentic-awesome-skills
Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.
Categories
Build a cloud, SLO, and incident-readiness register after intake. Observability Cloud Planning is an agent skill from sickn33/agentic-awesome-skills. Build a cloud, SLO, and incident-readiness register after intake.
Observability Cloud Planning fits situations like: an SME needs monitoring scope; alert ownership; service planning.
Run `npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a claude-code`. Or copy the skill folder (skills/observability-cloud-planning in sickn33/agentic-awesome-skills) into .claude/skills/observability-cloud-planning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a codex`. Or copy the skill folder (skills/observability-cloud-planning in sickn33/agentic-awesome-skills) into .agents/skills/observability-cloud-planning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill observability-cloud-planning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-cloud-planning, .gemini/skills/observability-cloud-planning, .github/skills/observability-cloud-planning and .opencode/skills/observability-cloud-planning in your project.
SKILL.md names no scripts, command-line tools or credentials: Observability Cloud Planning is instructions for the agent only.
SKILL.md names 1 domain. In commands or code: json-schema.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Observability Cloud Planning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Observability Cloud Planning: Service Mesh Observability (wshobson/agents, 40k stars), Observability (2SSK/dot-files, 247 stars), Monitoring Observability (ahmedasmar/devops-claude-skills, 203 stars) and Prometheus Error Rate Investigator (prometheus/prometheus-mcp, 120 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,379 GitHub stars. The repository holds 1,493 skills in this directory. The repository was last updated on October 9, 2026.
Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.