Agent skill

It Operations

by davila7 in davila7/claude-code-templates

Manages IT infrastructure, monitoring, incident response, and service reliability.

MITAuto-check passedDevOps & Cloud

Install It Operations

skills CLI
$ npx skills add davila7/claude-code-templates --skill it-operations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates it-operations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/development/it-operations .claude/skills/it-operations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
it-operations
GitHub stars
32k
Used in
1 other repo
Token cost
~3.7k tokens
SKILL.md length
748 words
Files
7
Skills in repo
477
Repo updated
First seen
Licence
MIT

At a glance

Manages IT infrastructure, monitoring, incident response, and service reliability.

  • Works in 8 steps: Service Reliability First → Automation Over Manual Processes → ITIL Service Management → …
  • Tasks that involve Site reliability engineering
  • SKILL.md covers Core Principles, Core Workflow, Decision Frameworks and Common Operational Challenges, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

It Operations is an agent skill from davila7/claude-code-templates. Manages IT infrastructure, monitoring, incident response, and service reliability. Provides frameworks for ITIL service management, observability strategies, automation, backup/recovery, capacity planning, and operational excellence practices.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `README.md`, `reference/automation.md` and `reference/backup-recovery.md`).

It sits in DevOps & Cloud, covering Site reliability engineering, Incident response and Observability. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Tasks that involve Site reliability engineering
  • Tasks that involve Incident response
  • Tasks that involve Observability

Example prompts

  • “Use the it-operations skill to manage IT infrastructure, monitoring, incident response, and service reliability”
  • “/it-operations”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Service Reliability First
  2. Automation Over Manual Processes
  3. ITIL Service Management
  4. Operational Excellence
  5. Blameless Post-Mortems
  6. Runbook Standards
  7. On-Call Best Practices
  8. Change Management Discipline

What it can do on your machine

Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

It Operations loads about 3.7k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 748 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 748 words, ~3,736 tokens.

Download SKILL.mdSave it as .claude/skills/it-operations/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
it-operations
description
Manages IT infrastructure, monitoring, incident response, and service reliability. Provides frameworks for ITIL service management, observability strategies, automation, backup/recovery, capacity planning, and operational excellence practices.

IT Operations Expert

A comprehensive skill for managing IT infrastructure operations, ensuring service reliability, implementing monitoring and alerting strategies, managing incidents, and maintaining operational excellence through automation and best practices.

Core Principles

1. Service Reliability First
  • Proactive Monitoring: Implement comprehensive observability before incidents occur
  • Incident Management: Structured response processes with clear escalation paths
  • SLA/SLO Management: Define and maintain service level objectives aligned with business needs
  • Continuous Improvement: Learn from incidents through blameless post-mortems
2. Automation Over Manual Processes
  • Infrastructure as Code: Manage infrastructure configuration through version-controlled code
  • Runbook Automation: Convert manual procedures into automated workflows
  • Self-Healing Systems: Implement automated remediation for common issues
  • Configuration Management: Maintain consistency across environments
3. ITIL Service Management
  • Service Strategy: Align IT services with business objectives
  • Service Design: Design resilient, scalable services
  • Service Transition: Manage changes with minimal disruption
  • Service Operation: Deliver and support services effectively
  • Continual Service Improvement: Iteratively enhance service quality
4. Operational Excellence
  • Documentation: Maintain current runbooks, procedures, and architecture diagrams
  • Knowledge Management: Build searchable knowledge bases from incident resolutions
  • Capacity Planning: Forecast and provision resources proactively
  • Cost Optimization: Balance performance requirements with infrastructure costs

Core Workflow

Infrastructure Operations Workflow
1. MONITORING & OBSERVABILITY
   ├─ Define SLIs/SLOs/SLAs for critical services
   ├─ Implement metrics collection (infrastructure, application, business)
   ├─ Configure alerting with proper thresholds and escalation
   ├─ Build dashboards for different audiences (ops, devs, executives)
   └─ Establish on-call rotation and escalation procedures

2. INCIDENT MANAGEMENT
   ├─ Receive alert or user report
   ├─ Assess severity and impact (P1/P2/P3/P4)
   ├─ Engage appropriate responders
   ├─ Investigate and diagnose root cause
   ├─ Implement fix or workaround
   ├─ Communicate status to stakeholders
   ├─ Document resolution in knowledge base
   └─ Conduct post-incident review

3. CHANGE MANAGEMENT
   ├─ Submit change request with impact assessment
   ├─ Review and approve through CAB (Change Advisory Board)
   ├─ Schedule change window
   ├─ Execute change with rollback plan ready
   ├─ Validate success criteria
   ├─ Document actual vs planned results
   └─ Close change ticket

4. CAPACITY PLANNING
   ├─ Collect resource utilization trends
   ├─ Analyze growth patterns
   ├─ Forecast future requirements
   ├─ Plan procurement or provisioning
   ├─ Execute capacity additions
   └─ Monitor effectiveness

5. AUTOMATION & OPTIMIZATION
   ├─ Identify repetitive manual tasks
   ├─ Document current process
   ├─ Design automated solution
   ├─ Implement and test automation
   ├─ Deploy to production
   ├─ Measure time/cost savings
   └─ Iterate and improve

Decision Frameworks

Alert Configuration Decision Matrix
ScenarioAlert TypeThresholdResponse TimeEscalation
Service completely downPageImmediate< 5 minImmediate to on-call
Service degradedPage2-3 failures< 15 minAfter 15 min to on-call
High resource usageWarning> 80% sustained< 1 hourAfter 2 hours to team lead
Approaching capacityInfo> 70% trend< 24 hoursWeekly capacity review
Configuration driftTicketAny deviation< 7 daysMonthly review
Incident Severity Classification

Priority 1 (Critical)

  • Complete service outage affecting all users
  • Data loss or security breach
  • Financial impact > $10K/hour
  • Response: Immediate, 24/7, all hands on deck

Priority 2 (High)

  • Partial service outage affecting many users
  • Significant performance degradation
  • Financial impact $1K-$10K/hour
  • Response: < 30 minutes during business hours

Priority 3 (Medium)

  • Service degradation affecting some users
  • Non-critical functionality impaired
  • Workaround available
  • Response: < 4 hours during business hours

Priority 4 (Low)

  • Minor issues with minimal impact
  • Cosmetic problems
  • Enhancement requests
  • Response: Next business day
Change Management Risk Assessment
Risk Level = Impact × Likelihood × Complexity

Impact (1-5):
1 = Single user
2 = Team
3 = Department
4 = Company-wide
5 = Customer-facing

Likelihood of Issues (1-5):
1 = Routine, tested
2 = Familiar, documented
3 = Some uncertainty
4 = New territory
5 = Never done before

Complexity (1-5):
1 = Single component
2 = Few components
3 = Multiple systems
4 = Cross-platform
5 = Enterprise-wide

Risk Score Interpretation:
1-20: Standard change (pre-approved)
21-50: Normal change (CAB review)
51-75: High-risk change (extensive testing, senior approval)
76-125: Emergency change only (executive approval)
Monitoring Tool Selection
RequirementPrometheus + GrafanaDatadogNew RelicELK StackSplunk
CostFree (self-hosted)$$$$$$$$Free-$$$$$$$
MetricsExcellentExcellentExcellentGoodGood
LogsVia LokiExcellentExcellentExcellentExcellent
TracesVia TempoExcellentExcellentLimitedGood
Learning CurveSteepModerateModerateSteepSteep
Cloud-NativeExcellentExcellentExcellentGoodGood
On-PremisesExcellentGoodGoodExcellentExcellent
APMVia exportersExcellentExcellentLimitedGood

Common Operational Challenges

Challenge 1: Alert Fatigue

Problem: Too many false positive alerts causing team burnout

Solution:

yaml
Alert Tuning Process:
1. Measure baseline alert volume and false positive rate
2. Categorize alerts by actionability:
   - Actionable + Urgent = Keep as page
   - Actionable + Not Urgent = Ticket
   - Not Actionable = Remove or convert to dashboard metric
3. Implement alert aggregation (group similar alerts)
4. Add context to alerts (runbook links, relevant metrics)
5. Regular review meetings (weekly) to tune thresholds
6. Track metrics:
   - MTTA (Mean Time to Acknowledge): < 5 min target
   - False Positive Rate: < 20% target
   - Alert Volume per Week: Trending down
Challenge 2: Incident Documentation During Crisis

Problem: Teams skip documentation during high-pressure incidents

Solution:

  • Assign dedicated scribe role (not the incident commander)
  • Use incident management tools (PagerDuty, Opsgenie) with automatic timeline
  • Template-based incident reports with required fields
  • Post-incident review scheduled automatically (within 48 hours)
  • Gamify documentation (track and recognize thorough documentation)
Show full SKILL.md (275 more words)Show less
Challenge 3: Knowledge Silos

Problem: Critical knowledge trapped in individual team members' heads

Solution:

yaml
Knowledge Transfer Strategy:
- Pair Programming/Shadowing: 20% of sprint capacity
- Runbook Requirements: Every system must have runbook
- Lunch & Learn Sessions: Weekly 30-min knowledge sharing
- Cross-Training Matrix: Track who knows what, identify gaps
- On-Call Rotation: Everyone rotates to spread knowledge
- Post-Incident Reviews: Mandatory team sharing
- Documentation Sprints: Quarterly focus on doc completion
Challenge 4: Balancing Stability vs Innovation

Problem: Operations team resists change to maintain stability

Solution:

  • Implement change windows (planned maintenance periods)
  • Use blue-green or canary deployments for lower risk
  • Establish "innovation time" (Google 20% time model)
  • Create sandbox environments for experimentation
  • Measure and reward both stability AND improvement metrics
  • Include "toil reduction" as OKR target

Key Metrics & KPIs

Service Reliability Metrics
yaml
Availability:
  Formula: (Total Time - Downtime) / Total Time × 100
  Target: 99.9% (43.8 min/month downtime)
  Measurement: Per service, monthly

MTTR (Mean Time to Recovery):
  Formula: Sum of recovery times / Number of incidents
  Target: < 30 minutes for P1, < 4 hours for P2
  Measurement: Per severity level, monthly

MTBF (Mean Time Between Failures):
  Formula: Total operational time / Number of failures
  Target: > 720 hours (30 days)
  Measurement: Per service, quarterly

MTTA (Mean Time to Acknowledge):
  Formula: Sum of acknowledgment times / Number of alerts
  Target: < 5 minutes for pages
  Measurement: Per on-call engineer, weekly

Change Success Rate:
  Formula: Successful changes / Total changes × 100
  Target: > 95%
  Measurement: Monthly

Incident Recurrence Rate:
  Formula: Repeat incidents / Total incidents × 100
  Target: < 10%
  Measurement: Quarterly (same root cause within 90 days)
Operational Efficiency Metrics
yaml
Toil Percentage:
  Definition: Time spent on manual, repetitive tasks
  Target: < 30% of team capacity
  Measurement: Weekly time tracking

Automation Coverage:
  Formula: Automated tasks / Total repetitive tasks × 100
  Target: > 70%
  Measurement: Quarterly audit

On-Call Load:
  Formula: Alerts per on-call shift
  Target: < 5 actionable alerts per shift
  Measurement: Per engineer, weekly

Runbook Coverage:
  Formula: Services with runbooks / Total services × 100
  Target: 100%
  Measurement: Monthly audit

Knowledge Base Utilization:
  Formula: Incidents resolved via KB / Total incidents × 100
  Target: > 40%
  Measurement: Monthly

Integration Points

With Development Teams
  • Participate in design reviews for operational requirements
  • Provide deployment automation and CI/CD pipeline support
  • Share monitoring and logging requirements
  • Collaborate on incident response and post-mortems
  • Joint ownership of SLOs and error budgets
With Security Teams
  • Implement security monitoring and alerting
  • Manage access controls and authentication systems
  • Coordinate vulnerability patching and remediation
  • Conduct security incident response
  • Maintain compliance with security policies
With Business Stakeholders
  • Report on service availability and performance
  • Communicate planned maintenance windows
  • Provide capacity planning forecasts
  • Translate technical metrics to business impact
  • Participate in business continuity planning

Best Practices

1. Blameless Post-Mortems
markdown
Post-Incident Review Template:
- Incident Summary (what happened, when, impact)
- Timeline of Events (detailed chronology)
- Root Cause Analysis (5 Whys or Fishbone)
- What Went Well (strengths during response)
- What Could Be Improved (opportunities)
- Action Items (with owners and due dates)
- Lessons Learned (shareable insights)

Rules:
- No blame or punishment
- Focus on systems and processes, not people
- Everyone can speak freely
- Action items must be tracked to completion
2. Runbook Standards
yaml
Runbook Contents:
  - Service Overview: Purpose, dependencies, architecture
  - SLIs/SLOs/SLAs: Defined thresholds and targets
  - Common Issues: Symptoms, causes, solutions
  - Troubleshooting Steps: Step-by-step procedures
  - Escalation Paths: Who to contact and when
  - Useful Commands: Copy-paste ready commands
  - Dashboard Links: Direct links to relevant dashboards
  - Recent Changes: Link to change log
  - Contact Information: Team, product owner, SMEs

Maintenance:
  - Review quarterly or after major incidents
  - Test procedures during low-traffic periods
  - Update after every significant change
  - Track usage metrics (page views, helpfulness ratings)
3. On-Call Best Practices
yaml
On-Call Preparation:
  - Laptop with VPN access
  - Mobile device with notification apps
  - Contact list (escalation paths)
  - Access to all critical systems
  - Runbooks bookmarked
  - Backup on-call identified

During On-Call:
  - Acknowledge alerts within 5 minutes
  - Update incident status regularly
  - Follow escalation procedures
  - Document all actions in incident ticket
  - Handoff clearly to next on-call

Post On-Call:
  - Complete incident reports
  - Submit toil reduction tickets
  - Provide feedback on runbooks
  - Update on-call documentation
4. Change Management Discipline
yaml
Standard Change Process:
  1. Create change request (RFC)
  2. Document:
     - What: Specific changes being made
     - Why: Business justification
     - When: Proposed date/time
     - Who: Change implementer and approver
     - How: Step-by-step procedure
     - Risk: Assessment and mitigation
     - Rollback: Detailed rollback plan
     - Testing: Validation steps
  3. Submit for CAB review (7 days advance notice)
  4. Implement during approved window
  5. Validate success criteria
  6. Close change with actual results
  7. Post-implementation review if issues occurred

Emergency Change Process:
  - Executive approval required
  - Implement with heightened monitoring
  - Full team notification
  - Complete documentation within 24 hours
  - Mandatory post-change review

Reference Files

For detailed technical guidance, see:

Getting Started

  1. For New Infrastructure: Start with reference/infrastructure.md for setup guidance
  2. For Monitoring Setup: Review reference/monitoring.md for observability strategy
  3. For Incident Response: See reference/incident-management.md for procedures
  4. For Automation Projects: Check reference/automation.md for tooling recommendations
  5. For DR Planning: Consult reference/backup-recovery.md for recovery strategies

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files in cli-tool/components/skills/development/it-operations of davila7/claude-code-templates.

  • SKILL.md
  • README.md
  • reference/automation.md
  • reference/backup-recovery.md
  • reference/incident-management.md
  • reference/infrastructure.md
  • reference/monitoring.md

Open the folder on GitHubat commit 14680ec

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

It Operations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

It Operations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
It Operations this skilldavila7/claude-code-templates32k1 repos~3.7kAutomated safety check: PassMIT
Structured Logging Litemajiayu000/spellbook286—~1.7kAutomated safety check: PassMIT
Release Itwondelai/skills2.4k—~4kAutomated safety check: PassMIT
Guidewire Observability And Incident Responsejeremylongshore/tons-of-skills-marketplace2.8k—~3.1kAutomated safety check: PassMIT
Observability And Reliabilitycbrock84/headcount2k—~931Automated safety check: PassMIT
Monitoringericrisco/rsc-harness167—~3.1kAutomated safety check: PassMIT

Similar skills

  • Structured Logging Lite

    majiayu000/spellbook

    Design, audit, or implement application structured logging architecture from repository evidence.

    286 GitHub stars~1.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Release It

    wondelai/skills

    Build production-ready systems with stability patterns: circuit breakers, bulkheads, timeouts, and retry logic.

    2.4k GitHub stars~4k tokensUpdated 27 days ago
    DevOps & CloudAuto-check passed
  • Guidewire Observability And Incident Response

    jeremylongshore/tons-of-skills-marketplace

    Operate a Guidewire Cloud API integration in production — define SLIs/SLOs for token availability, bind success rate, FNOL p99 latency; route alerts so the on-call gets paged for real outages and…

    2.8k GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure.

    2k GitHub stars~931 tokensUpdated 20 days ago
    DevOps & CloudAuto-check passed
  • Monitoring

    ericrisco/rsc-harness

    A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…

    167 GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Categories

Questions about It Operations

What does It Operations do?

Manages IT infrastructure, monitoring, incident response, and service reliability. It Operations is an agent skill from davila7/claude-code-templates. Manages IT infrastructure, monitoring, incident response, and service reliability.

When should I use It Operations?

It Operations fits situations like: tasks that involve Site reliability engineering; tasks that involve Incident response; tasks that involve Observability.

How do I install It Operations in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill it-operations -a claude-code`. Or copy the skill folder (cli-tool/components/skills/development/it-operations in davila7/claude-code-templates) into .claude/skills/it-operations in your project. Claude Code loads it when a task matches its description.

How do I install It Operations in Codex?

Run `npx skills add davila7/claude-code-templates --skill it-operations -a codex`. Or copy the skill folder (cli-tool/components/skills/development/it-operations in davila7/claude-code-templates) into .agents/skills/it-operations in your project. Codex loads it when a task matches its description.

Can I use It Operations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill it-operations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/it-operations, .gemini/skills/it-operations, .github/skills/it-operations and .opencode/skills/it-operations in your project.

What does It Operations need to run?

SKILL.md names no scripts, command-line tools or credentials: It Operations is instructions for the agent only.

Does It Operations access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is It Operations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does It Operations use?

It Operations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does It Operations use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to It Operations?

Skills that share tags, products or a category with It Operations: Structured Logging Lite (majiayu000/spellbook, 286 stars), Release It (wondelai/skills, 2.4k stars), Guidewire Observability And Incident Response (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Observability And Reliability (cbrock84/headcount, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains It Operations?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.