Agent skill

Detecting Data Anomalies

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Process identify anomalies and outliers in datasets using machine learning algorithms.

MITAuto-check passedData & Analytics

Install Detecting Data Anomalies

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill detecting-data-anomalies -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace detecting-data-anomalies --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/detecting-data-anomalies .claude/skills/detecting-data-anomalies && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
detecting-data-anomalies
GitHub stars
2.8k
Token cost
~1.4k tokens
SKILL.md length
600 words
Files
7 (incl. scripts, references, assets)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Process identify anomalies and outliers in datasets using machine learning algorithms.

  • Works in 10 steps: Load the dataset using the Read tool and… → Profile feature distributions using… → Handle missing values via imputation… → …
  • Analyzing data for unusual patterns
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Calls pip

What it does

Detecting Data Anomalies is an agent skill from jeremylongshore/tons-of-skills-marketplace. Process identify anomalies and outliers in datasets using machine learning algorithms. Use when analyzing data for unusual patterns, outliers, or unexpected deviations from normal behavior. Trigger with phrases like "detect anomalies", "find outliers", or "identify unusual patterns".

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts, reference files and assets (for example `assets/README.md`, `references/README.md` and `references/errors.md`). Compatibility notes: Designed for Claude Code

It sits in Data & Analytics, covering Anomaly detection. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Analyzing data for unusual patterns
  • Unexpected deviations from normal behavior
  • With phrases like detect anomalies
  • Identify unusual patterns

Example prompts

  • “detect anomalies”
  • “find outliers”
  • “identify unusual patterns”
  • “/detecting-data-anomalies”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Bash(python:*), Grep, Glob

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Load the dataset using the Read tool and verify schema, column types, and row count
  2. Profile feature distributions using descriptive statistics to understand baseline behavior
  3. Handle missing values via imputation (median for numeric, mode for categorical) or row exclusion
  4. Apply StandardScaler or MinMaxScaler to numeric features to normalize magnitude differences
  5. Select the detection algorithm based on data characteristics
  6. Set the contamination parameter to the expected anomaly proportion (start with 0.01-0.05)
  7. Fit the model on the training partition and generate anomaly scores for each data point
  8. Apply the decision threshold to classify points as normal (-1) or anomalous (1)
  9. Analyze flagged anomalies for common characteristics, temporal clusters, or feature correlations
  10. Generate a summary report with detection counts, score distributions, and visualization plots

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash(python:*)
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • scikit-learn.org
    • pyod.readthedocs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Detecting Data Anomalies loads about 1.4k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 600 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 600 words, ~1,391 tokens.

Download SKILL.mdSave it as .claude/skills/detecting-data-anomalies/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
detecting-data-anomalies
description
Process identify anomalies and outliers in datasets using machine learning algorithms. Use when analyzing data for unusual patterns, outliers, or unexpected deviations from normal behavior. Trigger with phrases like "detect anomalies", "find outliers", or "identify unusual patterns".
allowed-tools
Read, Bash(python:*), Grep, Glob
compatibility
Designed for Claude Code
version
1.26.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
ai, ml, detecting-data

Detecting Data Anomalies

Overview

Identify anomalies and outliers in datasets using statistical and machine learning algorithms including Isolation Forest, One-Class SVM, Local Outlier Factor, and autoencoders. This skill handles the full detection pipeline from data ingestion and feature scaling through algorithm selection, threshold tuning, and result interpretation with anomaly scoring.

Prerequisites

  • Python 3.9+ with scikit-learn >= 1.3 (pip install scikit-learn)
  • pandas and NumPy for data manipulation (pip install pandas numpy)
  • matplotlib or seaborn for anomaly visualizations (pip install matplotlib seaborn)
  • Dataset in CSV, JSON, Parquet, or database-queryable format
  • Minimum 500 data points for statistical significance (1000+ recommended)
  • Optional: PyTorch or TensorFlow for autoencoder-based detection on complex patterns

Instructions

  1. Load the dataset using the Read tool and verify schema, column types, and row count
  2. Profile feature distributions using descriptive statistics to understand baseline behavior
  3. Handle missing values via imputation (median for numeric, mode for categorical) or row exclusion
  4. Apply StandardScaler or MinMaxScaler to numeric features to normalize magnitude differences
  5. Select the detection algorithm based on data characteristics:
    • Isolation Forest: high-dimensional data, no assumptions on distribution
    • One-Class SVM: well-defined normal class with clear decision boundary
    • Local Outlier Factor: density-varying data with local anomaly patterns
    • Autoencoder: complex temporal or image data with non-linear relationships
  6. Set the contamination parameter to the expected anomaly proportion (start with 0.01-0.05)
  7. Fit the model on the training partition and generate anomaly scores for each data point
  8. Apply the decision threshold to classify points as normal (-1) or anomalous (1)
  9. Analyze flagged anomalies for common characteristics, temporal clusters, or feature correlations
  10. Generate a summary report with detection counts, score distributions, and visualization plots

See ${CLAUDE_SKILL_DIR}/references/implementation.md for the detailed implementation guide.

Output

  • Anomaly detection summary: total points, anomaly count, contamination rate
  • Per-record anomaly scores with classification labels
  • Algorithm configuration: model type, contamination, distance metric, threshold
  • Feature importance ranking showing which dimensions drive anomaly flags
  • Visualization: scatter plot of anomaly scores, distribution histogram, t-SNE cluster plot
  • CSV export of flagged records with anomaly scores and contributing features
Show full SKILL.md (265 more words)Show less

Error Handling

ErrorCauseSolution
Insufficient data volumeFewer than 100 data points for model fittingCollect additional data or switch to simple statistical methods (z-score, IQR)
High false positive rateContamination parameter set too high or features not scaledLower contamination to 0.01; verify StandardScaler applied; refine feature selection
Algorithm OOM on large datasetIsolation Forest or LOF exceeds available memorySubsample data for training; use max_samples parameter; switch to streaming approach
Feature scaling mismatchMixed numeric and categorical features without proper encodingOne-hot encode categoricals separately; scale numeric features independently
No ground truth for validationUnlabeled dataset prevents accuracy measurementUse domain expert review on top-N anomalies; implement feedback loop to refine threshold

See ${CLAUDE_SKILL_DIR}/references/errors.md for the full error reference.

Examples

Scenario 1: Network Intrusion Detection -- Apply Isolation Forest to 50K network flow records with features: packet count, byte volume, duration, protocol type. Expected contamination: 2%. Target: flag port-scan and DDoS patterns with precision above 0.85.

Scenario 2: Manufacturing Quality Control -- Run LOF on sensor readings (temperature, vibration, pressure) from 10K production cycles. Detect equipment degradation anomalies. Visualize flagged cycles on a time-series plot with normal operating bands.

Scenario 3: Financial Transaction Monitoring -- Train an autoencoder on 100K legitimate transactions. Reconstruct test transactions and flag those with reconstruction error above the 99th percentile. Report flagged transactions with amount, merchant category, and time-of-day features.

Resources

  • scikit-learn Anomaly Detection -- Isolation Forest, LOF, One-Class SVM
  • PyOD Library -- 40+ outlier detection algorithms with unified API
  • Autoencoder anomaly detection: Keras/PyTorch reconstruction-error approach
  • Feature scaling: StandardScaler, RobustScaler, MinMaxScaler selection guide
  • Evaluation without labels: silhouette analysis, domain expert review protocols

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references, assets) in skills/.curated/detecting-data-anomalies of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • assets/README.md
  • references/README.md
  • references/errors.md
  • references/examples.md
  • references/implementation.md
  • scripts/README.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Detecting Data Anomalies next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Detecting Data Anomalies compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Detecting Data Anomalies this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMIT
TimesFM Forecastinggoogle-research/timesfm34k—~4.7kAutomated safety check: PassApache-2.0
Anomalib Adding A Modelopen-edge-platform/anomalib6.2k—~1.9kAutomated safety check: PassApache-2.0
Anomalib Tiled Ensembleopen-edge-platform/anomalib6.2k—~1.4kAutomated safety check: PassApache-2.0
Kqlmicrosoft/fabric-rti-mcp131—~6.2kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries171—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • TimesFM Forecasting

    google-research/timesfm

    Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.

    34k GitHub stars~4.7k tokensUpdated 11 days ago
    Data & AnalyticsAuto-check passed
  • Anomalib Adding A Model

    open-edge-platform/anomalib

    Adds a new anomaly-detection model to anomalib under src/anomalib/models/.

    6.2k GitHub stars~1.9k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Anomalib Tiled Ensemble

    open-edge-platform/anomalib

    Runs and configures the anomalib tiled-ensemble pipeline, which trains/evaluates one model per image tile and merges results (with optional seam smoothing) for high-resolution anomaly detection.

    6.2k GitHub stars~1.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Kql

    microsoft/fabric-rti-mcp

    Official

    KQL language expertise for writing correct, efficient Kusto queries using the Fabric RTI MCP tools.

    131 GitHub stars~6.2k tokensUpdated 10 days ago
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    171 GitHub stars~3.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Dt Obs Analytics

    Dynatrace/dynatrace-for-ai

    Analyze dashboards and notebooks using Davis analyzers — anomaly detection, novelty scoring, and correlation.

    163 GitHub stars~3.9k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Detecting Data Anomalies

What does Detecting Data Anomalies do?

Process identify anomalies and outliers in datasets using machine learning algorithms. Detecting Data Anomalies is an agent skill from jeremylongshore/tons-of-skills-marketplace. Process identify anomalies and outliers in datasets using machine learning algorithms.

When should I use Detecting Data Anomalies?

Detecting Data Anomalies fits situations like: analyzing data for unusual patterns; unexpected deviations from normal behavior; with phrases like detect anomalies; identify unusual patterns.

How do I install Detecting Data Anomalies in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill detecting-data-anomalies -a claude-code`. Or copy the skill folder (skills/.curated/detecting-data-anomalies in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/detecting-data-anomalies in your project. Claude Code loads it when a task matches its description.

How do I install Detecting Data Anomalies in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill detecting-data-anomalies -a codex`. Or copy the skill folder (skills/.curated/detecting-data-anomalies in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/detecting-data-anomalies in your project. Codex loads it when a task matches its description.

Can I use Detecting Data Anomalies in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill detecting-data-anomalies -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/detecting-data-anomalies, .gemini/skills/detecting-data-anomalies, .github/skills/detecting-data-anomalies and .opencode/skills/detecting-data-anomalies in your project.

What does Detecting Data Anomalies need to run?

Going by SKILL.md and its folder, Detecting Data Anomalies needs the command-line tools its instructions call (pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Bash(python:*), Grep, Glob. Compatibility (from SKILL.md): Designed for Claude Code.

Does Detecting Data Anomalies access the network?

SKILL.md names 2 domains. As links in the text: scikit-learn.org and pyod.readthedocs.io. This is read from the text; nothing was executed.

Is Detecting Data Anomalies safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Detecting Data Anomalies use?

Detecting Data Anomalies is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Detecting Data Anomalies use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Detecting Data Anomalies?

Skills that share tags, products or a category with Detecting Data Anomalies: TimesFM Forecasting (google-research/timesfm, 34k stars), Anomalib Adding A Model (open-edge-platform/anomalib, 6.2k stars), Anomalib Tiled Ensemble (open-edge-platform/anomalib, 6.2k stars) and Kql (microsoft/fabric-rti-mcp, 131 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Detecting Data Anomalies?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.