Agent skill

Maintenance And Reliability

by cbrock84 in cbrock84/headcount

Keeps production assets available — ranking equipment by consequence of failure, setting preventive and predictive intervals, sizing spares, and moving a plant off reactive maintenance.

MITAuto-check passed

Install Maintenance And Reliability

skills CLI
$ npx skills add cbrock84/headcount --skill maintenance-and-reliability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cbrock84/headcount maintenance-and-reliability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .claude/skills && cp -r skills-src/verticals/industrial/skills/operations/maintenance-and-reliability .claude/skills/maintenance-and-reliability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
maintenance-and-reliability
GitHub stars
2k
Token cost
~1.4k tokens
SKILL.md length
857 words
Files
1
Skills in repo
178
Repo updated
First seen
Licence
MIT

At a glance

Keeps production assets available — ranking equipment by consequence of failure, setting preventive and predictive intervals, sizing spares, and moving a plant off reactive maintenance.

  • Works in 4 steps: Wears predictably → interval-based… → Degrades observably → condition-based.… → Fails randomly, consequence low → run to… → …
  • SKILL.md covers Rank by consequence, not by…, Not everything wears out, The backlog is the leading… and Spares are an availability…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Maintenance And Reliability is an agent skill from cbrock84/headcount. Keeps production assets available — ranking equipment by consequence of failure, setting preventive and predictive intervals, sizing spares, and moving a plant off reactive maintenance. Use this to build or fix a maintenance program, decide what to put on a PM schedule, diagnose repeat failures on one asset, justify spares inventory, or work out why a plant that maintains everything still has unplanned downtime.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: An agent organization structured as a company — 15+ departments, 125+ skills, each independently installable, citing the standards and regulators that settle the question. Runs… The licence is MIT.

Example prompts

  • “Use the maintenance-and-reliability skill to keep production assets available — ranking equipment by consequence of failure, setting preventive and…”
  • “/maintenance-and-reliability”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Wears predictably → interval-based replacement, set from observed life rather than the
  2. Degrades observably → condition-based. Vibration, thermal, oil analysis, current draw. Replace
  3. Fails randomly, consequence low → run to failure, deliberately, with the decision recorded so
  4. Fails randomly, consequence high → redundancy or detection, because no interval helps.

What it can do on your machine

Read from SKILL.md and the folder at commit 98d1c17. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Maintenance And Reliability loads about 1.4k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 857 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cbrock84/headcount at commit 98d1c17, republished under its MIT licence (© cbrock84). 857 words, ~1,444 tokens.

Download SKILL.mdSave it as .claude/skills/maintenance-and-reliability/SKILL.md (or your agent's skills folder).
name
maintenance-and-reliability
description
Keeps production assets available — ranking equipment by consequence of failure, setting preventive and predictive intervals, sizing spares, and moving a plant off reactive maintenance. Use this to build or fix a maintenance program, decide what to put on a PM schedule, diagnose repeat failures on one asset, justify spares inventory, or work out why a plant that maintains everything still has unplanned downtime.

Maintenance and reliability

A plant that maintains everything equally maintains the wrong things. Maintenance effort is finite, and spending it evenly across assets means the machine whose failure stops the line gets the same attention as the one with three spares in the cabinet.

Rank by consequence, not by age or cost

The question is not how likely a machine is to fail. It is what happens when it does.

Rank every asset on what its failure costs: production stopped, product scrapped, a safety event, a regulatory exposure, or nothing anyone notices before the next shift. That ranking, not the capital value, decides where preventive effort goes.

Two things this exposes immediately in most plants:

  • Assets with no redundancy and no spare — the ones that will stop the plant and keep it stopped while a part ships. These are the whole list worth arguing about.
  • Assets on a PM schedule for no reason — inherited from a manual, consuming hours, preventing nothing that would have mattered.

Not everything wears out

The intuition behind interval-based maintenance is that failure probability rises with age, so replacing on a schedule beats waiting. That holds for things that genuinely wear — belts, bearings, seals, tooling, anything in contact.

It does not hold for most electronics, instrumentation and control systems, where failure rate is roughly flat with age and intervention is itself a source of failure. Scheduled replacement of a component that does not wear out converts a stable asset into one that gets disturbed on a cycle, and introduces the installation errors that follow every disturbance.

So the interval has to be chosen by failure mode:

  1. Wears predictably → interval-based replacement, set from observed life rather than the manual's default.
  2. Degrades observably → condition-based. Vibration, thermal, oil analysis, current draw. Replace on the signal rather than the calendar.
  3. Fails randomly, consequence low → run to failure, deliberately, with the decision recorded so it does not read as neglect.
  4. Fails randomly, consequence high → redundancy or detection, because no interval helps.

Option 3 is a legitimate strategy and it is the one nobody writes down, which is why it gets mistaken for the program failing.

The backlog is the leading indicator

Downtime is a lagging measure. By the time it moves, the condition that produced it has been present for months.

The honest leading indicator is the state of the planned-work backlog: how much identified work is waiting, how old the oldest item is, and what fraction of the week's hours went to work that was planned before the week started. A plant doing 80% planned work is in a different regime from one doing 30%, and the second cannot schedule production reliably no matter how good its scheduler is.

The reactive trap is self-sustaining. Unplanned failures consume the hours that would have prevented the next ones, so the ratio degrades on its own. Breaking it costs a deliberate, temporary over-allocation to planned work while the backlog is worked down, and that is a decision someone has to fund rather than a habit the team can adopt.

Show full SKILL.md (347 more words)Show less

Spares are an availability decision priced as inventory

A spare is bought against a downtime cost, not against a usage rate, which is why usage-based reorder logic gets critical spares wrong in both directions. The part used twice a year sits at zero and the plant waits six weeks for it.

Size the critical few on lead time and consequence: what it costs per day stopped, multiplied by the realistic days to obtain. Everything else can run on consumption. Say which list each part is on, because the finance conversation about slow-moving inventory will otherwise delete the critical ones first — they are, by design, the parts that never move.

Tooling

A computerized maintenance management system is the asset register, the work order history and the PM schedule in one place. Fiix, Limble, UpKeep, eMaint, IBM Maximo and similar differ mostly in how much configuration they demand; the failure mode is identical across all of them, which is a system populated with assets and never with completed work orders. History is the entire value — without it there is no observed failure interval and every PM stays at the manual's default.

Condition monitoring is a sensor and analysis decision before it is a software one. Portable vibration and thermal instruments on a route cover most plants. Permanently installed monitoring on a machine ranked at the top of the consequence list pays; installed everywhere it becomes a data stream nobody reads.

Where the ERP holds the asset master — SAP PM, Oracle, Infor and similar — keep one register rather than two. A maintenance system with its own asset list diverges from the financial one within a year, and then neither is trusted.

Never

  • Put an asset on a preventive schedule without naming the failure mode it prevents.
  • Treat repeat failures on one asset as bad luck rather than as an unfixed cause.
  • Let unplanned work consume the hours allocated to planned work and call the ratio a result.
  • Judge a maintenance program on downtime alone, which moves months after the cause does.
  • Cut a critical spare because it has not moved.

© cbrock84, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in verticals/industrial/skills/operations/maintenance-and-reliability of cbrock84/headcount.

Open the folder on GitHubat commit 98d1c17

Compare with similar skills

Maintenance And Reliability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Maintenance And Reliability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Maintenance And Reliability this skillcbrock84/headcount2k—~1.4kAutomated safety check: PassMIT
Equipment Maintenance Logaipoch/medical-research-skills2k—~1.4kAutomated safety check: PassMIT
Google Cloud Waf Reliabilitygoogle/skills21k—~2kAutomated safety check: PassApache-2.0
Doc Maintenancepaperclipai/paperclip99k—~1.1kAutomated safety check: PassMIT
Google Cloud Waf Reliabilitydavila7/claude-code-templates32k—~1.8kAutomated safety check: PassMIT
AvailabilityBuilderIO/agent-native7.1k—~387Automated safety check: PassNone

Similar skills

  • Equipment Maintenance Log

    aipoch/medical-research-skills

    Track lab equipment calibration dates and send maintenance reminders for pipettes, balances, centrifuges, and other instruments.

    2k GitHub stars~1.4k tokensUpdated 22 days ago
    Business, Finance & HRAuto-check passed
  • Official

    Generates guidance for reliability, resilience, availability, redundancy, fault-tolerance, and disaster recovery (DR) for Google Cloud workloads based on the design principles and recommendations in…

    21k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Doc Maintenance

    paperclipai/paperclip

    Keep project docs aligned with recent code and feature changes — detect drift, update affected pages, and add release-relevant notes without rewriting unchanged sections.

    99k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Google Cloud Waf Reliability

    davila7/claude-code-templates

    Generates reliability-focused guidance for Google Cloud workloads based on the Google Cloud Well-Architected Framework.

    32k GitHub stars~1.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Availability

    BuilderIO/agent-native

    How schedules, weekly rules, date overrides, travel schedules, and out-of-office entries combine to determine when someone is bookable.

    7.1k GitHub stars~387 tokensUpdated today
    Auto-check passed
  • Gke Reliability

    google/skills

    Official

    Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints.

    21k GitHub stars~1.8k tokensUpdated today
    DevOps & CloudAuto-check passed

More from cbrock84/headcount

All 178 skills in this repo
  • Agent Hierarchy

    cbrock84/headcount

    Designs orchestrator-and-subagent hierarchies for a repository — splitting agents by exclusive write surface, pairing every producer with an independent auditor, and enforcing the split with a…

    2k GitHub stars~1.2k tokensUpdated 22 days ago
    Auto-check passed
  • Access And Identity

    cbrock84/headcount

    Designs and audits who can reach what — authentication, authorization models, privileged access, service credentials, and joiner-mover-leaver process.

    2k GitHub stars~1.1k tokensUpdated 22 days ago
    Auto-check passed
  • Account Based Marketing

    cbrock84/headcount

    Concentrates marketing and sales effort on a named set of accounts rather than on volume — qualifying whether the model fits your economics at all, building the account list and the buying group…

    2k GitHub stars~1.2k tokensUpdated 22 days ago
    Auto-check passed
  • Activation

    cbrock84/headcount

    Gets new users from signup to first real value — signup flow, onboarding, time-to-value, and the early experience that determines whether someone becomes a user or a lapsed account.

    2k GitHub stars~865 tokensUpdated 22 days ago
    Auto-check passed
  • AI ML Governance

    cbrock84/headcount

    Governs models and AI systems in production — intended use, evaluation, monitoring, human oversight, documentation, and the decision to deploy or retire.

    2k GitHub stars~1k tokensUpdated 22 days ago
    Auto-check passed
  • AI Research Analyst

    cbrock84/headcount

    Produces executive-level research — market sizing, competitor mapping, trend analysis, and strategic intelligence — grounded in cited sources with the confidence in each claim made explicit.

    2k GitHub stars~916 tokensUpdated 22 days ago
    Auto-check passed

Questions about Maintenance And Reliability

What does Maintenance And Reliability do?

Keeps production assets available — ranking equipment by consequence of failure, setting preventive and predictive intervals, sizing spares, and moving a plant off reactive maintenance. Maintenance And Reliability is an agent skill from cbrock84/headcount. Keeps production assets available — ranking equipment by consequence of failure, setting preventive and predictive intervals, sizing spares, and moving a plant off reactive maintenance.

How do I install Maintenance And Reliability in Claude Code?

Run `npx skills add cbrock84/headcount --skill maintenance-and-reliability -a claude-code`. Or copy the skill folder (verticals/industrial/skills/operations/maintenance-and-reliability in cbrock84/headcount) into .claude/skills/maintenance-and-reliability in your project. Claude Code loads it when a task matches its description.

How do I install Maintenance And Reliability in Codex?

Run `npx skills add cbrock84/headcount --skill maintenance-and-reliability -a codex`. Or copy the skill folder (verticals/industrial/skills/operations/maintenance-and-reliability in cbrock84/headcount) into .agents/skills/maintenance-and-reliability in your project. Codex loads it when a task matches its description.

Can I use Maintenance And Reliability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cbrock84/headcount --skill maintenance-and-reliability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/maintenance-and-reliability, .gemini/skills/maintenance-and-reliability, .github/skills/maintenance-and-reliability and .opencode/skills/maintenance-and-reliability in your project.

What does Maintenance And Reliability need to run?

SKILL.md names no scripts, command-line tools or credentials: Maintenance And Reliability is instructions for the agent only.

Does Maintenance And Reliability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Maintenance And Reliability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Maintenance And Reliability use?

Maintenance And Reliability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Maintenance And Reliability use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Maintenance And Reliability?

Skills that share tags, products or a category with Maintenance And Reliability: Equipment Maintenance Log (aipoch/medical-research-skills, 2k stars), Google Cloud Waf Reliability (google/skills, 21k stars), Doc Maintenance (paperclipai/paperclip, 99k stars) and Google Cloud Waf Reliability (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Maintenance And Reliability?

cbrock84 (a GitHub user) maintains it in cbrock84/headcount, which has 2,016 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on September 17, 2026.

Source: cbrock84/headcount on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.