Agent skill

Reliability

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Run one analytics task through several fresh trials to measure what repeats and what varies.

MITAuto-check passed

Install Reliability

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill reliability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst reliability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/reliability .claude/skills/reliability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
reliability
GitHub stars
304
Token cost
~929 tokens
SKILL.md length
512 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Run one analytics task through several fresh trials to measure what repeats and what varies.

  • Works in 7 steps: Record the task, model, active data,… → Use five fresh trials unless the user… → Preserve completed, failed, blocked,… → …
  • Repeated-run requests
  • SKILL.md covers What this answers, Run the evaluation and Report
  • Calls python3

What it does

Reliability is an agent skill from ai-analyst-lab/ai-analyst. Run one analytics task through several fresh trials to measure what repeats and what varies. Use for reliability, repeatability, variance, or repeated-run requests. This measures stability, not correctness.

Its SKILL.md is about 930 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • Repeated-run requests

Example prompts

  • “/reliability”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Record the task, model, active data, available tools, system fingerprint, and the tolerance that matters for the decision.
  2. Use five fresh trials unless the user chooses another count. The trials must not share answers.
  3. Preserve completed, failed, blocked, error, and unparseable trials separately.
  4. Save the raw output from every trial.
  5. Use the repository's deterministic reliability runner. It creates five independent workspaces, uses the active data source, preserves…
  6. Report exact agreement and agreement within the named tolerance separately.
  7. State what changed if this is a comparison with an earlier run.

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reliability loads about 929 tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 512 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~929

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 512 words, ~929 tokens.

Download SKILL.mdSave it as .claude/skills/reliability/SKILL.md (or your agent's skills folder).
name
reliability
description
Run one analytics task through several fresh trials to measure what repeats and what varies. Use for reliability, repeatability, variance, or repeated-run requests. This measures stability, not correctness.

Reliability

What this answers

Does this named system configuration behave consistently on this task?

It does not answer whether the result is correct. A wrong analysis can repeat perfectly.

Run the evaluation

  1. Record the task, model, active data, available tools, system fingerprint, and the tolerance that matters for the decision.
  2. Use five fresh trials unless the user chooses another count. The trials must not share answers.
  3. Preserve completed, failed, blocked, error, and unparseable trials separately.
  4. Save the raw output from every trial.
  5. Use the repository's deterministic reliability runner. It creates five independent workspaces, uses the active data source, preserves every result, and calculates the comparison.
  6. Report exact agreement and agreement within the named tolerance separately.
  7. State what changed if this is a comparison with an earlier run.

The deterministic CLI is available at python3 -m helpers.evals.cli run-reliability. Use it when the task can be executed noninteractively. For the standard run, explicitly pass --trials 5 --parallelism 5 --model claude-opus-4-6. Add --allow-code when the task requires a query. Do not lower parallelism preemptively or infer an account limit. Lower it only after the five-way run returns an actual concurrency or rate-limit error, and tell the user what failed before retrying.

Run the reliability command as one foreground task. Read its progress messages as trials finish. Do not create separate sleep commands or polling shells merely to wait for it.

The runner automatically gives each trial a clean view of the current analytical system. It excludes old outputs, test fixtures, future course examples, and inactive dataset packages. Connection credentials remain outside the workspace and are passed through the process environment. Do not ask the user to manage excluded paths or trial workspaces.

If the current system already contains a reviewed definition for the question, report that fact. Do not hide current context merely to manufacture variation.

Show full SKILL.md (202 more words)Show less

For a separate local context store, pass its path with --context-store. The runner saves a snapshot under the run's context-snapshot/ directory and installs the same contents in each trial. This does not change the project's active context configuration. Without an override, a configured source: path store is used automatically; legacy Git-cache mode must first be connected as a visible local path. Do not silently substitute local definitions if that source cannot be read.

Each trial is instructed to inspect the guide catalog and load relevant reviewed guides. Inspect trials/<number>/trace/context_loads_<analysis_id>.jsonl to see the full guide text loaded, and compare it with the query logs to check application. The snapshot proves which files were available, not that a guide was used. input-inventory.json records trial inputs and response.json preserves the response and captured tool events. Definition labels alone do not establish whether two attempts used the same calculation.

Report

Lead with:

  • how many trials succeeded;
  • every normalized result beside its raw form;
  • exact agreement;
  • tolerance-based agreement;
  • what definitions or methods changed; and
  • what the evidence does not establish.

If the trials vary, locate the source before proposing a fix. If one definition or context change is made, freeze everything else and rerun the same task.

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/reliability of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Reliability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reliability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reliability this skillai-analyst-lab/ai-analyst304—~929Automated safety check: PassMIT
Content Freshness Signalsthedaviddias/Front-End-Checklist74k—~741Automated safety check: PassMIT
Google Cloud Waf Reliabilitydavila7/claude-code-templates32k—~1.8kAutomated safety check: PassMIT
Relsa Severity AssessmentK-Dense-AI/scientific-agent-skills48k1 repos~5.2kAutomated safety check: NotesMIT
Google Cloud Waf Reliabilitygoogle/skills21k—~2kAutomated safety check: PassApache-2.0
Gke Reliabilitygoogle/skills21k—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Content Freshness Signals

    thedaviddias/Front-End-Checklist

    Audits article pages for freshness signals, covering the Last-Modified header, Article JSON-LD dateModified and a visible last-updated date, and fixes mismatches.

    74k GitHub stars~741 tokensUpdated yesterday
    Marketing & SEOAuto-check passed
  • Google Cloud Waf Reliability

    davila7/claude-code-templates

    Generates reliability-focused guidance for Google Cloud workloads based on the Google Cloud Well-Architected Framework.

    32k GitHub stars~1.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Relsa Severity Assessment

    K-Dense-AI/scientific-agent-skills

    Supports multivariate severity assessment and exploratory endpoint-time score forecasting for laboratory animal studies using the RELSA (RELative Severity Assessment) score and ARIMA-based foRcast…

    48k GitHub starsUsed in 1 repo~5.2k tokens
    Data & AnalyticsAuto-check: notes
  • Official

    Generates guidance for reliability, resilience, availability, redundancy, fault-tolerance, and disaster recovery (DR) for Google Cloud workloads based on the design principles and recommendations in…

    21k GitHub stars~2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Gke Reliability

    google/skills

    Official

    Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints.

    21k GitHub stars~1.8k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Windows Shell Reliability

    sickn33/agentic-awesome-skills

    Reliable command execution on Windows: paths, encoding, and common binary pitfalls.

    47k GitHub starsUsed in 2 repos~1k tokens
    DevOps & CloudAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Questions about Reliability

What does Reliability do?

Run one analytics task through several fresh trials to measure what repeats and what varies. Reliability is an agent skill from ai-analyst-lab/ai-analyst. Run one analytics task through several fresh trials to measure what repeats and what varies.

When should I use Reliability?

Reliability fits situations like: repeated-run requests.

How do I install Reliability in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill reliability -a claude-code`. Or copy the skill folder (.claude/skills/reliability in ai-analyst-lab/ai-analyst) into .claude/skills/reliability in your project. Claude Code loads it when a task matches its description.

How do I install Reliability in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill reliability -a codex`. Or copy the skill folder (.claude/skills/reliability in ai-analyst-lab/ai-analyst) into .agents/skills/reliability in your project. Codex loads it when a task matches its description.

Can I use Reliability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill reliability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/reliability, .gemini/skills/reliability, .github/skills/reliability and .opencode/skills/reliability in your project.

What does Reliability need to run?

Going by SKILL.md and its folder, Reliability needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Reliability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Reliability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Reliability use?

Reliability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reliability use?

About 929 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Reliability?

Skills that share tags, products or a category with Reliability: Content Freshness Signals (thedaviddias/Front-End-Checklist, 74k stars), Google Cloud Waf Reliability (davila7/claude-code-templates, 32k stars), Relsa Severity Assessment (K-Dense-AI/scientific-agent-skills, 48k stars) and Google Cloud Waf Reliability (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reliability?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.