Agent skill

Anomaly Investigation

by gaasher in gaasher/Agent-Loop-Skills

A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.

MITAuto-check passedData & Analytics

Install Anomaly Investigation

skills CLI
$ npx skills add gaasher/Agent-Loop-Skills --skill anomaly-investigation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gaasher/Agent-Loop-Skills anomaly-investigation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/loops/anomaly-investigation .claude/skills/anomaly-investigation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anomaly-investigation
GitHub stars
174
Token cost
~2.1k tokens
SKILL.md length
1,047 words
Files
2
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.

  • Works in 4 steps: Pick a candidate to test. Ideally the… → Test it. Write /iter/test.py that… → Eliminate or advance. → …
  • The user has a known
  • SKILL.md covers When to use, Setup, The loop and Ledger, plus 2 more sections
  • Calls uv

What it does

Anomaly Investigation is an agent skill from gaasher/Agent-Loop-Skills. Use when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed. Forms a slate of candidate causes, tests each against the data, and eliminates the ones the data refutes, narrowing the live candidates until exactly one survives refutation and passes a positive confirming test. The result is an investigation log with the confirmed root cause and the evidence that ruled out the alternatives. Not for…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `examples/run.example.yaml`). Compatibility notes: Requires Python 3.9+

It sits in Data & Analytics, covering Root cause analysis and Data analysis. The repository describes itself as: Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills… The licence is MIT.

When your agent uses it

  • The user has a known
  • Already-observed anomaly in their data — a metric spike
  • An unexpected number — and wants its root cause diagnosed

Example prompts

  • “/anomaly-investigation”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.9+

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Pick a candidate to test. Ideally the one whose test most cleanly splits the remaining field, so
  2. Test it. Write /iter/test.py that computes the thing that would **refute or
  3. Eliminate or advance.
  4. Log one ledger row and continue, narrowing the live set.

What it can do on your machine

Read from SKILL.md and the folder at commit f1169e6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.9+

    From compatibility in the SKILL.md frontmatter.

Context cost

Anomaly Investigation loads about 2.1k tokens when it runs. Until then it costs about 194 tokens; SKILL.md has 1,047 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~194
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from gaasher/Agent-Loop-Skills at commit f1169e6, republished under its MIT licence (© gaasher). 1,047 words, ~2,119 tokens.

Download SKILL.mdSave it as .claude/skills/anomaly-investigation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
anomaly-investigation
description
Use when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed. Forms a slate of candidate causes, tests each against the data, and eliminates the ones the data refutes, narrowing the live candidates until exactly one survives refutation and passes a positive confirming test. The result is an investigation log with the confirmed root cause and the evidence that ruled out the alternatives. Not for open-ended discovery over a dataset with no specific anomaly in hand (that is data-analysis), and not for checking an external claim against sources (that is claim-verify) — this is reactive diagnosis of one anomaly you already know about.
compatibility
Requires Python 3.9+
metadata.version
0.1.0

Anomaly Investigation Loop

A form → test → eliminate → confirm loop — root-cause analysis as a search. The artifact is an investigation log; the feedback signal is the count of live candidate explanations, driven down toward a single cause that is confirmed, not merely consistent. Each iteration you test one candidate against the data and drop the ones the data refutes, narrowing the field until one survives.

The discipline this enforces: a cause is "root" only when it both survives an honest attempt to refute it and makes a positive prediction that checks out (e.g. "if this is the cause, removing it restores normal" — and it does). A story that merely could explain the anomaly is a hypothesis, not a finding.

When to use

Use this when an anomaly is already in hand — you know roughly what looks wrong and want the cause diagnosed by elimination against the data. Default to a broad initial slate of mutually distinguishable causes, then test the one that splits the field fastest; if the anomaly is vague, your first job is to make it precise (iteration 0). Not for open-ended exploration of a dataset with no anomaly to chase (use data-analysis), and not for verifying an external claim against the literature (use claim-verify).

Setup

Resolve bindings interactively. If loop.run.yaml exists in the working dir, load it, confirm the values in one line, and skip to the loop. Otherwise: on Claude Code (the AskUserQuestion tool is available) infer a likely value for each binding and present it as the recommended option; on other hosts ask each as a quoted plain-text prompt. Then write loop.run.yaml (format: examples/run.example.yaml) and confirm the values before creating any other files.

bindingmeaningdefaulthow to infer
<dataset>data (or logs) to investigate; read-only ground truth—scan the working dir for a data/log file
<anomaly>what looks wrong: the metric, where/when, and how big the deviation is—ask the user; make precise in iter 0
<analysis_cmd>interpreter that runs analysis snippets in the user's envpython3pyproject.toml/.venv/uv in the working dir
<log>output investigation log<sandbox_root>/investigation.md—
<sandbox_root>where snippets + ledger live./sandbox—
<budget>max iterations8—

Analysis snippets run in the user's environment via <analysis_cmd>, so they may use whatever the user has installed. Keep helper code stdlib-first (csv, statistics): if a snippet needs pandas/numpy, probe with try/except ImportError and degrade to a stdlib path, or offer a consented uv pip install "pandas==<ver>" — never assume the package is installed.

The loop

Copy this checklist and tick items off:

  • Iteration 0 — characterize the anomaly precisely; form an initial slate of candidate causes in <log>.
  • Pick a candidate to test (the one whose test most cleanly splits the remaining field).
  • Test it: write <sandbox_root>/iter<N>/test.py, run with <analysis_cmd>, redirect to out.txt.
  • Eliminate (data refutes it → drop from the live set) or advance (data supports it → keep it live).
  • Confirm the survivor: when one candidate leads, run a positive test only it predicts, after a refutation attempt.
  • Append a ledger row; stop when one cause is confirmed or at <budget>.

Iteration 0 — characterize. Quantify the anomaly precisely: write and run a snippet that pins down what deviated, where/when, and how big the deviation is against the normal baseline (the same metric on surrounding periods/segments). Then form an initial slate of candidate causes — mutually distinguishable explanations, broad enough to contain the truth (a real change, a composition/mix shift, a data-quality bug, a measurement change, seasonality, an outlier segment). List them in <log> as the live candidates. Record nothing as confirmed yet.

Then, until stop (one confirmed cause, or budget):

  1. Pick a candidate to test. Ideally the one whose test most cleanly splits the remaining field, so each iteration removes as many live candidates as possible.
  2. Test it. Write <sandbox_root>/iter<N>/test.py that computes the thing that would refute or support it (slice by segment/source/time, recompute the metric, compare distributions). Run it with <analysis_cmd>, redirecting output to <sandbox_root>/iter<N>/out.txt (never flood your context).
  3. Eliminate or advance.
    • Data refutes it → mark it eliminated in <log> with the evidence; drop it from the live set.
    • Data supports it → keep it live, and if it is now the leading candidate run a confirming test: a positive prediction it uniquely makes (e.g. "remove / seasonally-adjust the suspected factor → the anomaly disappears"). Also try to refute it — a leading candidate that survives a genuine refutation attempt and passes its confirming test is the root cause.
  4. Log one ledger row and continue, narrowing the live set.
Show full SKILL.md (307 more words)Show less

Observational equivalence. Two mechanistically different candidates can make identical predictions in the data you have (e.g. a bot flood and a pipeline double-count both look like "sessions spike, conversions flat" in daily aggregates). When that happens you cannot separate them here — do not pick one arbitrarily. Report them as a single confirmed cause at the resolution of the available data, and name the additional data that would distinguish them (finer-grained logs, raw event records, an upstream check). Distinguish, too, the mechanism (how the metric moved) from the root cause (why the inputs were wrong) — confirming the mechanism is progress, but is not the cause.

Ledger

<sandbox_root>/ledger.tsv, tab-separated, never commas in the text. Header:

iter	candidate_tested	verdict	live_candidates

verdict ∈ {characterize, refuted, supported, confirmed}. Example:

iter	candidate_tested	verdict	live_candidates
0	characterize anomaly + slate	characterize	5
1	real drop across all segments	refuted	4
2	one segment's conversions fell	refuted	3
3	one source's sessions inflated	supported	2
4	removing that source restores normal	confirmed	1

Report the confirmed root cause with its confirming evidence, the alternatives and how each was ruled out, and — if you stop without a single confirmed cause — the remaining live candidates and the test that would separate them.

Constraints

  • Confirm, don't just fit. The root cause must survive an honest refutation attempt and pass a positive confirming test; "consistent with the data" is not enough, since several stories usually are.
  • Test against the data, not intuition — every elimination and the final confirmation is backed by a computation you ran, recorded in <log>.
  • Keep candidates distinguishable and prefer the test that splits the field fastest, so the live count falls; do not chase one pet theory while leaving alternatives untested.
  • Only read <dataset> — never modify it, because it is the ground truth every test is checked against. The sandbox is self-contained (no ../ escapes).
  • Do not pause the loop to ask whether to continue; run until a cause is confirmed or the budget is hit.

Stops

  • Confirmed — exactly one candidate survived refutation and passed a positive confirming test.
  • Budget — <budget> iterations reached without a single confirmed cause; report the live set.

© gaasher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in loops/anomaly-investigation of gaasher/Agent-Loop-Skills.

  • SKILL.md
  • examples/run.example.yaml

Open the folder on GitHubat commit f1169e6

Compare with similar skills

Anomaly Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anomaly Investigation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anomaly Investigation this skillgaasher/Agent-Loop-Skills174—~2.1kAutomated safety check: PassMIT
Amazon Opensearch Serviceaws/agent-toolkit-for-aws2.8k—~2.4kAutomated safety check: PassApache-2.0
Data Analysis Standardmohitagw15856/pm-claude-skills1.4k—~1.7kAutomated safety check: PassMIT
Exploratory Data Analysisspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow84k4 repos~2.2kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2122 repos~3.4kAutomated safety check: NotesMIT

Similar skills

  • Amazon Opensearch Service

    aws/agent-toolkit-for-aws

    Official

    Guides migration, provisioning, search, log-analytics, trace-analytics, and Agentic AI Assistant workflows for Amazon OpenSearch Service and Serverless across six capabilities — migration…

    2.8k GitHub stars~2.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Data Analysis Standard

    mohitagw15856/pm-claude-skills

    Structure a product data analysis, metric deep-dive, funnel analysis, or cohort study.

    1.4k GitHub stars~1.7k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    84k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    212 GitHub starsUsed in 2 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed

More from gaasher/Agent-Loop-Skills

All 20 skills in this repo
  • Alpha Evolve

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to evolve an ML model/program through population-based search rather than a single sequential refine loop — a generational evolution where parallel…

    174 GitHub stars~3.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Blue Team

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…

    174 GitHub stars~3.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Data Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants an iterative, self-checking exploratory analysis of a dataset — surfacing findings that are each verified by re-running the computation, not asserted.

    174 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Hypothesis Gen

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to generate and literature-vet a pool of novel, testable research hypotheses for a question or domain.

    174 GitHub stars~2.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.

    174 GitHub stars~2.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Dueling Autoresearch

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g.

    174 GitHub stars~2.6k tokensUpdated 3 mo ago
    Auto-check: warnings

Questions about Anomaly Investigation

What does Anomaly Investigation do?

A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed. Anomaly Investigation is an agent skill from gaasher/Agent-Loop-Skills. Use when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.

When should I use Anomaly Investigation?

Anomaly Investigation fits situations like: the user has a known; already-observed anomaly in their data — a metric spike; an unexpected number — and wants its root cause diagnosed.

How do I install Anomaly Investigation in Claude Code?

Run `npx skills add gaasher/Agent-Loop-Skills --skill anomaly-investigation -a claude-code`. Or copy the skill folder (loops/anomaly-investigation in gaasher/Agent-Loop-Skills) into .claude/skills/anomaly-investigation in your project. Claude Code loads it when a task matches its description.

How do I install Anomaly Investigation in Codex?

Run `npx skills add gaasher/Agent-Loop-Skills --skill anomaly-investigation -a codex`. Or copy the skill folder (loops/anomaly-investigation in gaasher/Agent-Loop-Skills) into .agents/skills/anomaly-investigation in your project. Codex loads it when a task matches its description.

Can I use Anomaly Investigation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gaasher/Agent-Loop-Skills --skill anomaly-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anomaly-investigation, .gemini/skills/anomaly-investigation, .github/skills/anomaly-investigation and .opencode/skills/anomaly-investigation in your project.

What does Anomaly Investigation need to run?

Going by SKILL.md and its folder, Anomaly Investigation needs the command-line tools its instructions call (uv). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.9+.

Does Anomaly Investigation access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Anomaly Investigation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Anomaly Investigation use?

Anomaly Investigation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anomaly Investigation use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Anomaly Investigation?

Skills that share tags, products or a category with Anomaly Investigation: Amazon Opensearch Service (aws/agent-toolkit-for-aws, 2.8k stars), Data Analysis Standard (mohitagw15856/pm-claude-skills, 1.4k stars), Exploratory Data Analysis (spacering-net/codeg, 3.9k stars) and Excel and CSV Data Analysis (bytedance/deer-flow, 84k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anomaly Investigation?

gaasher (a GitHub user) maintains it in gaasher/Agent-Loop-Skills, which has 174 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on June 30, 2026.

Source: gaasher/Agent-Loop-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.