Agent skill

Tracking Service Reliability

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Define and track SLAs, SLIs, and SLOs for service reliability including availability, latency, and error rates.

MITAuto-check passedDevOps & Cloud

Install Tracking Service Reliability

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill tracking-service-reliability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace tracking-service-reliability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/tracking-service-reliability .claude/skills/tracking-service-reliability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tracking-service-reliability
GitHub stars
2.8k
Token cost
~1k tokens
SKILL.md length
465 words
Files
5 (incl. scripts, references, assets)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Define and track SLAs, SLIs, and SLOs for service reliability including availability, latency, and error rates.

  • Works in 3 steps: SLI Definition: The skill guides the… → SLO Target Setting: The skill assists in… → SLA Establishment: The skill helps in…
  • Establishing reliability targets
  • SKILL.md covers Overview, How It Works, When to Use This Skill and Examples, plus 7 more sections
  • Runs Python scripts from its folder

What it does

Tracking Service Reliability is an agent skill from jeremylongshore/tons-of-skills-marketplace. Define and track SLAs, SLIs, and SLOs for service reliability including availability, latency, and error rates. Use when establishing reliability targets or monitoring service health. Trigger with phrases like "define SLOs", "track SLI metrics", or "calculate error budget".

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/README.md`, `references/README.md` and `scripts/README.md`). Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering Site reliability engineering. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Establishing reliability targets
  • Monitoring service health
  • With phrases like define SLOs
  • Track SLI metrics

Example prompts

  • “define SLOs”
  • “track SLI metrics”
  • “calculate error budget”
  • “/tracking-service-reliability”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Glob, Bash(monitoring:*), Bash(metrics:*)

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. SLI Definition: The skill guides the user to define Service Level Indicators (SLIs) such as availability, latency, error rate, and…
  2. SLO Target Setting: The skill assists in setting Service Level Objectives (SLOs) by establishing target values for the defined SLIs (e.g…
  3. SLA Establishment: The skill helps in formalizing Service Level Agreements (SLAs), which are customer-facing commitments based on the…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Bash(monitoring:*)
    • Bash(metrics:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Tracking Service Reliability loads about 1k tokens when it runs, and up to ~1k if it reads all its reference files. Until then it costs about 76 tokens; SKILL.md has 465 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 465 words, ~1,012 tokens.

Download SKILL.mdSave it as .claude/skills/tracking-service-reliability/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
tracking-service-reliability
description
Define and track SLAs, SLIs, and SLOs for service reliability including availability, latency, and error rates. Use when establishing reliability targets or monitoring service health. Trigger with phrases like "define SLOs", "track SLI metrics", or "calculate error budget".
allowed-tools
Read, Write, Edit, Grep, Glob, Bash(monitoring:*), Bash(metrics:*)
compatibility
Designed for Claude Code
version
1.22.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
performance, monitoring, tracking-service

Sla Sli Tracker

Define and track SLAs, SLIs, and SLOs for service reliability including availability targets, latency budgets, error rate thresholds, and error budget burn rates.

Overview

This skill provides a structured approach to defining and tracking SLAs, SLIs, and SLOs, which are essential for ensuring service reliability. It automates the process of setting performance targets and monitoring actual performance, enabling proactive identification and resolution of potential issues.

How It Works

  1. SLI Definition: The skill guides the user to define Service Level Indicators (SLIs) such as availability, latency, error rate, and throughput.
  2. SLO Target Setting: The skill assists in setting Service Level Objectives (SLOs) by establishing target values for the defined SLIs (e.g., 99.9% availability).
  3. SLA Establishment: The skill helps in formalizing Service Level Agreements (SLAs), which are customer-facing commitments based on the defined SLOs.

When to Use This Skill

This skill activates when you need to:

  • Define SLAs, SLIs, and SLOs for a service.
  • Track service performance against defined objectives.
  • Calculate error budgets based on SLOs.

Examples

Example 1: Defining SLOs for a New Service

User request: "Create SLOs for our new payment processing service."

The skill will:

  1. Prompt the user to define SLIs (e.g., latency, error rate).
  2. Assist in setting target values for each SLI (e.g., p99 latency < 100ms, error rate < 0.01%).
Example 2: Tracking Availability

User request: "Track the availability SLI for the database service."

The skill will:

  1. Guide the user in setting up the tracking of the availability SLI.
  2. Visualize availability performance against the defined SLO.

Best Practices

  • Granularity: Define SLIs that are specific and measurable.
  • Realism: Set SLOs that are challenging but achievable.
  • Alignment: Ensure SLAs align with the defined SLOs and business requirements.
Show full SKILL.md (179 more words)Show less

Integration

This skill can be integrated with monitoring tools to automatically collect SLI data and track performance against SLOs. It can also be used in conjunction with alerting systems to trigger notifications when SLO violations occur.

Prerequisites

  • SLI definitions stored in ${CLAUDE_SKILL_DIR}/slos/sli-definitions.yaml
  • Access to monitoring and metrics systems
  • Historical performance data for baseline
  • Business requirements for service reliability

Instructions

  1. Define Service Level Indicators (availability, latency, error rate, throughput)
  2. Set Service Level Objectives with target values (e.g., 99.9% availability)
  3. Formalize Service Level Agreements with customer commitments
  4. Configure automated SLI data collection
  5. Calculate error budgets based on SLOs
  6. Track performance and alert on SLO violations

Output

  • SLI/SLO/SLA definition documents
  • Real-time SLI metric dashboards
  • Error budget calculations and burn rate
  • SLO compliance reports
  • Alerting configurations for violations

Error Handling

If SLI/SLO tracking fails:

  • Verify SLI definition completeness
  • Check metric collection infrastructure
  • Validate data accuracy and granularity
  • Ensure alerting system connectivity
  • Review error budget calculation logic

Resources

  • Google SRE book on SLIs and SLOs
  • Error budget implementation guides
  • Service reliability engineering practices
  • SLO definition templates and examples

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/.curated/tracking-service-reliability of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • assets/README.md
  • references/README.md
  • scripts/README.md
  • scripts/generate_sla_report.py

Open the folder on GitHubat commit cfae287

Compare with similar skills

Tracking Service Reliability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tracking Service Reliability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tracking Service Reliability this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1kAutomated safety check: PassMIT
Inference Autopilotrednote-machine-learning/Inference-autopilot144—~4.5kAutomated safety check: PassApache-2.0
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0
Slo Implementationwshobson/agents40k11 repos~1.7kAutomated safety check: PassMIT
Agentforce D360 Analyzeforcedotcom/sf-skills1.1k—~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Inference Autopilot

    rednote-machine-learning/Inference-autopilot

    Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

    144 GitHub stars~4.5k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Slo Implementation

    wshobson/agents

    Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.

    40k GitHub starsUsed in 11 repos~1.7k tokens
    DevOps & CloudAuto-check passed
  • Agentforce D360 Analyze

    forcedotcom/sf-skills

    Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills.

    1.1k GitHub stars~3.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    282 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Tracking Service Reliability

What does Tracking Service Reliability do?

Define and track SLAs, SLIs, and SLOs for service reliability including availability, latency, and error rates. Tracking Service Reliability is an agent skill from jeremylongshore/tons-of-skills-marketplace. Define and track SLAs, SLIs, and SLOs for service reliability including availability, latency, and error rates.

When should I use Tracking Service Reliability?

Tracking Service Reliability fits situations like: establishing reliability targets; monitoring service health; with phrases like define SLOs; track SLI metrics.

How do I install Tracking Service Reliability in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill tracking-service-reliability -a claude-code`. Or copy the skill folder (skills/.curated/tracking-service-reliability in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/tracking-service-reliability in your project. Claude Code loads it when a task matches its description.

How do I install Tracking Service Reliability in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill tracking-service-reliability -a codex`. Or copy the skill folder (skills/.curated/tracking-service-reliability in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/tracking-service-reliability in your project. Codex loads it when a task matches its description.

Can I use Tracking Service Reliability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill tracking-service-reliability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tracking-service-reliability, .gemini/skills/tracking-service-reliability, .github/skills/tracking-service-reliability and .opencode/skills/tracking-service-reliability in your project.

What does Tracking Service Reliability need to run?

Going by SKILL.md and its folder, Tracking Service Reliability needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Glob, Bash(monitoring:*), Bash(metrics:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Tracking Service Reliability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tracking Service Reliability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Tracking Service Reliability use?

Tracking Service Reliability is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tracking Service Reliability use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15 tokens, read only when the agent opens those files.

What are the alternatives to Tracking Service Reliability?

Skills that share tags, products or a category with Tracking Service Reliability: Inference Autopilot (rednote-machine-learning/Inference-autopilot, 144 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Alerting Irm (grafana/skills, 282 stars) and Slo Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tracking Service Reliability?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.