Structured Logging Lite
majiayu000/spellbook
Design, audit, or implement application structured logging architecture from repository evidence.
Manages IT infrastructure, monitoring, incident response, and service reliability.
$ npx skills add davila7/claude-code-templates --skill it-operations -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install davila7/claude-code-templates it-operations --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/development/it-operations .claude/skills/it-operations && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "it-operations" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/it-operations into .claude/skills/it-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "it-operations", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/it-operationsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add davila7/claude-code-templates --skill it-operations -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install davila7/claude-code-templates it-operations --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .agents/skills && cp -r skills-src/cli-tool/components/skills/development/it-operations .agents/skills/it-operations && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "it-operations" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/it-operations into .agents/skills/it-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "it-operations", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davila7/claude-code-templates --skill it-operations -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install davila7/claude-code-templates it-operations --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/cli-tool/components/skills/development/it-operations .cursor/skills/it-operations && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "it-operations" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/it-operations into .cursor/skills/it-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "it-operations", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/davila7/claude-code-templates.git --path cli-tool/components/skills/development/it-operations--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add davila7/claude-code-templates --skill it-operations -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install davila7/claude-code-templates it-operations --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/cli-tool/components/skills/development/it-operations .gemini/skills/it-operations && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "it-operations" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/it-operations into .gemini/skills/it-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "it-operations", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install davila7/claude-code-templates it-operationsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add davila7/claude-code-templates --skill it-operations -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .github/skills && cp -r skills-src/cli-tool/components/skills/development/it-operations .github/skills/it-operations && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "it-operations" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/it-operations into .github/skills/it-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "it-operations", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davila7/claude-code-templates --skill it-operations -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install davila7/claude-code-templates it-operations --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/cli-tool/components/skills/development/it-operations .opencode/skills/it-operations && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "it-operations" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/development/it-operations into .opencode/skills/it-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "it-operations", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
it-operationsManages IT infrastructure, monitoring, incident response, and service reliability.
It Operations is an agent skill from davila7/claude-code-templates. Manages IT infrastructure, monitoring, incident response, and service reliability. Provides frameworks for ITIL service management, observability strategies, automation, backup/recovery, capacity planning, and operational excellence practices.
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `README.md`, `reference/automation.md` and `reference/backup-recovery.md`).
It sits in DevOps & Cloud, covering Site reliability engineering, Incident response and Observability. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
It Operations loads about 3.7k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 748 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 748 words, ~3,736 tokens.
.claude/skills/it-operations/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.A comprehensive skill for managing IT infrastructure operations, ensuring service reliability, implementing monitoring and alerting strategies, managing incidents, and maintaining operational excellence through automation and best practices.
1. MONITORING & OBSERVABILITY
├─ Define SLIs/SLOs/SLAs for critical services
├─ Implement metrics collection (infrastructure, application, business)
├─ Configure alerting with proper thresholds and escalation
├─ Build dashboards for different audiences (ops, devs, executives)
└─ Establish on-call rotation and escalation procedures
2. INCIDENT MANAGEMENT
├─ Receive alert or user report
├─ Assess severity and impact (P1/P2/P3/P4)
├─ Engage appropriate responders
├─ Investigate and diagnose root cause
├─ Implement fix or workaround
├─ Communicate status to stakeholders
├─ Document resolution in knowledge base
└─ Conduct post-incident review
3. CHANGE MANAGEMENT
├─ Submit change request with impact assessment
├─ Review and approve through CAB (Change Advisory Board)
├─ Schedule change window
├─ Execute change with rollback plan ready
├─ Validate success criteria
├─ Document actual vs planned results
└─ Close change ticket
4. CAPACITY PLANNING
├─ Collect resource utilization trends
├─ Analyze growth patterns
├─ Forecast future requirements
├─ Plan procurement or provisioning
├─ Execute capacity additions
└─ Monitor effectiveness
5. AUTOMATION & OPTIMIZATION
├─ Identify repetitive manual tasks
├─ Document current process
├─ Design automated solution
├─ Implement and test automation
├─ Deploy to production
├─ Measure time/cost savings
└─ Iterate and improve| Scenario | Alert Type | Threshold | Response Time | Escalation |
|---|---|---|---|---|
| Service completely down | Page | Immediate | < 5 min | Immediate to on-call |
| Service degraded | Page | 2-3 failures | < 15 min | After 15 min to on-call |
| High resource usage | Warning | > 80% sustained | < 1 hour | After 2 hours to team lead |
| Approaching capacity | Info | > 70% trend | < 24 hours | Weekly capacity review |
| Configuration drift | Ticket | Any deviation | < 7 days | Monthly review |
Priority 1 (Critical)
Priority 2 (High)
Priority 3 (Medium)
Priority 4 (Low)
Risk Level = Impact × Likelihood × Complexity
Impact (1-5):
1 = Single user
2 = Team
3 = Department
4 = Company-wide
5 = Customer-facing
Likelihood of Issues (1-5):
1 = Routine, tested
2 = Familiar, documented
3 = Some uncertainty
4 = New territory
5 = Never done before
Complexity (1-5):
1 = Single component
2 = Few components
3 = Multiple systems
4 = Cross-platform
5 = Enterprise-wide
Risk Score Interpretation:
1-20: Standard change (pre-approved)
21-50: Normal change (CAB review)
51-75: High-risk change (extensive testing, senior approval)
76-125: Emergency change only (executive approval)| Requirement | Prometheus + Grafana | Datadog | New Relic | ELK Stack | Splunk |
|---|---|---|---|---|---|
| Cost | Free (self-hosted) | $$$$ | $$$$ | Free-$$ | $$$$$ |
| Metrics | Excellent | Excellent | Excellent | Good | Good |
| Logs | Via Loki | Excellent | Excellent | Excellent | Excellent |
| Traces | Via Tempo | Excellent | Excellent | Limited | Good |
| Learning Curve | Steep | Moderate | Moderate | Steep | Steep |
| Cloud-Native | Excellent | Excellent | Excellent | Good | Good |
| On-Premises | Excellent | Good | Good | Excellent | Excellent |
| APM | Via exporters | Excellent | Excellent | Limited | Good |
Problem: Too many false positive alerts causing team burnout
Solution:
Alert Tuning Process:
1. Measure baseline alert volume and false positive rate
2. Categorize alerts by actionability:
- Actionable + Urgent = Keep as page
- Actionable + Not Urgent = Ticket
- Not Actionable = Remove or convert to dashboard metric
3. Implement alert aggregation (group similar alerts)
4. Add context to alerts (runbook links, relevant metrics)
5. Regular review meetings (weekly) to tune thresholds
6. Track metrics:
- MTTA (Mean Time to Acknowledge): < 5 min target
- False Positive Rate: < 20% target
- Alert Volume per Week: Trending downProblem: Teams skip documentation during high-pressure incidents
Solution:
Problem: Critical knowledge trapped in individual team members' heads
Solution:
Knowledge Transfer Strategy:
- Pair Programming/Shadowing: 20% of sprint capacity
- Runbook Requirements: Every system must have runbook
- Lunch & Learn Sessions: Weekly 30-min knowledge sharing
- Cross-Training Matrix: Track who knows what, identify gaps
- On-Call Rotation: Everyone rotates to spread knowledge
- Post-Incident Reviews: Mandatory team sharing
- Documentation Sprints: Quarterly focus on doc completionProblem: Operations team resists change to maintain stability
Solution:
Availability:
Formula: (Total Time - Downtime) / Total Time × 100
Target: 99.9% (43.8 min/month downtime)
Measurement: Per service, monthly
MTTR (Mean Time to Recovery):
Formula: Sum of recovery times / Number of incidents
Target: < 30 minutes for P1, < 4 hours for P2
Measurement: Per severity level, monthly
MTBF (Mean Time Between Failures):
Formula: Total operational time / Number of failures
Target: > 720 hours (30 days)
Measurement: Per service, quarterly
MTTA (Mean Time to Acknowledge):
Formula: Sum of acknowledgment times / Number of alerts
Target: < 5 minutes for pages
Measurement: Per on-call engineer, weekly
Change Success Rate:
Formula: Successful changes / Total changes × 100
Target: > 95%
Measurement: Monthly
Incident Recurrence Rate:
Formula: Repeat incidents / Total incidents × 100
Target: < 10%
Measurement: Quarterly (same root cause within 90 days)Toil Percentage:
Definition: Time spent on manual, repetitive tasks
Target: < 30% of team capacity
Measurement: Weekly time tracking
Automation Coverage:
Formula: Automated tasks / Total repetitive tasks × 100
Target: > 70%
Measurement: Quarterly audit
On-Call Load:
Formula: Alerts per on-call shift
Target: < 5 actionable alerts per shift
Measurement: Per engineer, weekly
Runbook Coverage:
Formula: Services with runbooks / Total services × 100
Target: 100%
Measurement: Monthly audit
Knowledge Base Utilization:
Formula: Incidents resolved via KB / Total incidents × 100
Target: > 40%
Measurement: MonthlyPost-Incident Review Template:
- Incident Summary (what happened, when, impact)
- Timeline of Events (detailed chronology)
- Root Cause Analysis (5 Whys or Fishbone)
- What Went Well (strengths during response)
- What Could Be Improved (opportunities)
- Action Items (with owners and due dates)
- Lessons Learned (shareable insights)
Rules:
- No blame or punishment
- Focus on systems and processes, not people
- Everyone can speak freely
- Action items must be tracked to completionRunbook Contents:
- Service Overview: Purpose, dependencies, architecture
- SLIs/SLOs/SLAs: Defined thresholds and targets
- Common Issues: Symptoms, causes, solutions
- Troubleshooting Steps: Step-by-step procedures
- Escalation Paths: Who to contact and when
- Useful Commands: Copy-paste ready commands
- Dashboard Links: Direct links to relevant dashboards
- Recent Changes: Link to change log
- Contact Information: Team, product owner, SMEs
Maintenance:
- Review quarterly or after major incidents
- Test procedures during low-traffic periods
- Update after every significant change
- Track usage metrics (page views, helpfulness ratings)On-Call Preparation:
- Laptop with VPN access
- Mobile device with notification apps
- Contact list (escalation paths)
- Access to all critical systems
- Runbooks bookmarked
- Backup on-call identified
During On-Call:
- Acknowledge alerts within 5 minutes
- Update incident status regularly
- Follow escalation procedures
- Document all actions in incident ticket
- Handoff clearly to next on-call
Post On-Call:
- Complete incident reports
- Submit toil reduction tickets
- Provide feedback on runbooks
- Update on-call documentationStandard Change Process:
1. Create change request (RFC)
2. Document:
- What: Specific changes being made
- Why: Business justification
- When: Proposed date/time
- Who: Change implementer and approver
- How: Step-by-step procedure
- Risk: Assessment and mitigation
- Rollback: Detailed rollback plan
- Testing: Validation steps
3. Submit for CAB review (7 days advance notice)
4. Implement during approved window
5. Validate success criteria
6. Close change with actual results
7. Post-implementation review if issues occurred
Emergency Change Process:
- Executive approval required
- Implement with heightened monitoring
- Full team notification
- Complete documentation within 24 hours
- Mandatory post-change reviewFor detailed technical guidance, see:
© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files in cli-tool/components/skills/development/it-operations of davila7/claude-code-templates.
Open the folder on GitHubat commit 14680ec
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.
It Operations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| It Operations this skilldavila7/claude-code-templates | 32k | 1 repos | ~3.7k | Automated safety check: Pass | MIT | |
| Structured Logging Litemajiayu000/spellbook | 286 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Release Itwondelai/skills | 2.4k | — | ~4k | Automated safety check: Pass | MIT | |
| Guidewire Observability And Incident Responsejeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~3.1k | Automated safety check: Pass | MIT | |
| Observability And Reliabilitycbrock84/headcount | 2k | — | ~931 | Automated safety check: Pass | MIT | |
| Monitoringericrisco/rsc-harness | 167 | — | ~3.1k | Automated safety check: Pass | MIT |
majiayu000/spellbook
Design, audit, or implement application structured logging architecture from repository evidence.
wondelai/skills
Build production-ready systems with stability patterns: circuit breakers, bulkheads, timeouts, and retry logic.
jeremylongshore/tons-of-skills-marketplace
Operate a Guidewire Cloud API integration in production — define SLIs/SLOs for token availability, bind success rate, FNOL p99 latency; route alerts so the on-call gets paged for real outages and…
cbrock84/headcount
Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure.
ericrisco/rsc-harness
A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and…
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
davila7/claude-code-templates
Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.
davila7/claude-code-templates
Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.
davila7/claude-code-templates
Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.
davila7/claude-code-templates
Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.
davila7/claude-code-templates
Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.
davila7/claude-code-templates
Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.
Categories
Manages IT infrastructure, monitoring, incident response, and service reliability. It Operations is an agent skill from davila7/claude-code-templates. Manages IT infrastructure, monitoring, incident response, and service reliability.
It Operations fits situations like: tasks that involve Site reliability engineering; tasks that involve Incident response; tasks that involve Observability.
Run `npx skills add davila7/claude-code-templates --skill it-operations -a claude-code`. Or copy the skill folder (cli-tool/components/skills/development/it-operations in davila7/claude-code-templates) into .claude/skills/it-operations in your project. Claude Code loads it when a task matches its description.
Run `npx skills add davila7/claude-code-templates --skill it-operations -a codex`. Or copy the skill folder (cli-tool/components/skills/development/it-operations in davila7/claude-code-templates) into .agents/skills/it-operations in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill it-operations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/it-operations, .gemini/skills/it-operations, .github/skills/it-operations and .opencode/skills/it-operations in your project.
SKILL.md names no scripts, command-line tools or credentials: It Operations is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
It Operations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with It Operations: Structured Logging Lite (majiayu000/spellbook, 286 stars), Release It (wondelai/skills, 2.4k stars), Guidewire Observability And Incident Response (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Observability And Reliability (cbrock84/headcount, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.
Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.