Agent skill

Data Scrubber

by FerroxLabs in FerroxLabs/wayland

Data cleaning automation expertise covering missing value strategies, outlier detection methods, duplicate detection and deduplication, data type correction, text normalization, date parsing across…

Apache-2.0Auto-check passedData & Analytics

Install Data Scrubber

skills CLI
$ npx skills add FerroxLabs/wayland --skill data-scrubber -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FerroxLabs/wayland data-scrubber --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/software-engineering/data-scrubber .claude/skills/data-scrubber && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-scrubber
GitHub stars
608
Token cost
~3.6k tokens
SKILL.md length
399 words
Files
1
Skills in repo
1,194
Repo updated
First seen
Licence
Apache-2.0

At a glance

Data cleaning automation expertise covering missing value strategies, outlier detection methods, duplicate detection and deduplication, data type correction, text normalization, date parsing across…

  • Works in 10 steps: Never modify raw data in place: Always… → Document every cleaning decision: Why… → Make cleaning reproducible: Scripts, not… → …
  • The user asks about data scrubber
  • SKILL.md covers Core Philosophy, Data Cleaning Pipeline Design, Missing Value Strategies and Outlier Detection, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Scrubber is an agent skill from FerroxLabs/wayland. Data cleaning automation expertise covering missing value strategies, outlier detection methods, duplicate detection and deduplication, data type correction, text normalization, date parsing across formats, encoding fixes, validation rules, pipeline design patterns, and data quality reporting. Use when the user asks about data scrubber, data scrubber best practices, or needs guidance on data scrubber implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.

When your agent uses it

  • The user asks about data scrubber
  • Data scrubber best practices
  • Needs guidance on data scrubber implementation
  • The user needs a different specialized skill

Example prompts

  • “/data-scrubber”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Never modify raw data in place: Always work on copies, keep originals
  2. Document every cleaning decision: Why was this value removed/changed?
  3. Make cleaning reproducible: Scripts, not manual edits
  4. Generate quality reports before AND after cleaning: Measure improvement
  5. Handle edge cases explicitly: Empty strings, whitespace-only, special characters
  6. Validate after cleaning: Ensure constraints are satisfied
  7. Use appropriate methods per data type: Median for numeric, mode for categorical
  8. Be conservative with outlier removal: Flag first, remove only when justified
  9. Test cleaning pipeline on sample data: Verify behavior before full run
  10. Version control cleaning scripts: Track changes to cleaning logic

What it can do on your machine

Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Scrubber loads about 3.6k tokens when it runs. Until then it costs about 136 tokens; SKILL.md has 399 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~136
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 399 words, ~3,574 tokens.

Download SKILL.mdSave it as .claude/skills/data-scrubber/SKILL.md (or your agent's skills folder).
name
data-scrubber
description
Data cleaning automation expertise covering missing value strategies, outlier detection methods, duplicate detection and deduplication, data type correction, text normalization, date parsing across formats, encoding fixes, validation rules, pipeline design patterns, and data quality reporting. Use when the user asks about data scrubber, data scrubber best practices, or needs guidance on data scrubber implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.
license
Apache-2.0
metadata.author
foundry-skills
metadata.version
1.0.0
metadata.tags
automation shell-scripting data-science
metadata.category
software-engineering
metadata.subcategory
developer-tools
metadata.disclaimer
none
metadata.difficulty
intermediate

Data Scrubber

Core Philosophy

Data cleaning is the unglamorous but critical foundation of any data-driven system. Raw data is messy: missing values, inconsistent formats, duplicates, encoding errors, and outliers. A systematic data cleaning pipeline transforms raw chaos into reliable, analysis-ready data. The goal is not perfection -- it is fitness for purpose. Every cleaning decision should be documented, reversible, and auditable.

Data Cleaning Pipeline Design

Pipeline Architecture
python
from abc import ABC, abstractmethod
from dataclasses import dataclass, field
from typing import Any
import pandas as pd

@dataclass
class CleaningReport:
    """Track all changes made during cleaning."""
    total_rows_in: int = 0
    total_rows_out: int = 0
    steps: list[dict] = field(default_factory=list)

    def add_step(self, name: str, rows_affected: int, details: str = ""):
        self.steps.append({
            "step": name,
            "rows_affected": rows_affected,
            "details": details,
        })
# ... (condensed) ...
pipeline.add_step(NormalizeText(columns=['name', 'city']))
pipeline.add_step(ParseDates(columns=['created_at', 'updated_at']))
pipeline.add_step(DetectOutliers(column='amount', method='iqr'))
pipeline.add_step(ValidateConstraints())

clean_df, report = pipeline.run(raw_df)
print(report.summary())

Missing Value Strategies

Detection
python
import pandas as pd
import numpy as np

def analyze_missing_values(df: pd.DataFrame) -> pd.DataFrame:
    """Generate a missing value report for each column."""
    missing = df.isnull().sum()
    percent = (missing / len(df)) * 100
    dtypes = df.dtypes

    report = pd.DataFrame({
        'column': missing.index,
        'missing_count': missing.values,
        'missing_pct': percent.values.round(2),
        'dtype': dtypes.values,
    }).sort_values('missing_pct', ascending=False)

    return report[report['missing_count'] > 0]

# Example output:
#       column  missing_count  missing_pct  dtype
# phone          2340          23.40        object
# address         890           8.90        object
# age             120           1.20        float64
Strategies by Data Type
python
class HandleMissingValues(CleaningStep):
    def name(self) -> str:
        return "Handle Missing Values"

    def execute(self, df: pd.DataFrame, report: CleaningReport) -> pd.DataFrame:
        total_fixed = 0

        for col in df.columns:
            missing = df[col].isnull().sum()
            if missing == 0:
                continue

            pct_missing = missing / len(df) * 100

            if pct_missing > 50:
                # Drop column if >50% missing
                df = df.drop(columns=[col])
                report.add_step(self.name(), missing, f"Dropped column '{col}' ({pct_missing:.1f}% missing)")
                # ... (condensed) ...
                if remaining > 0:
                    df = df.dropna(subset=[col])

            total_fixed += missing

        report.add_step(self.name(), total_fixed, "Total missing values handled")
        return df
Advanced Imputation
python
from sklearn.impute import KNNImputer
from sklearn.experimental import enable_iterative_imputer
from sklearn.impute import IterativeImputer

def impute_numeric_columns(df: pd.DataFrame, method: str = 'knn') -> pd.DataFrame:
    """Impute missing numeric values using ML-based methods."""
    numeric_cols = df.select_dtypes(include=[np.number]).columns

    if method == 'knn':
        imputer = KNNImputer(n_neighbors=5, weights='distance')
    elif method == 'iterative':
        imputer = IterativeImputer(max_iter=10, random_state=42)
    else:
        raise ValueError(f"Unknown method: {method}")

    df[numeric_cols] = imputer.fit_transform(df[numeric_cols])
    return df

Outlier Detection

Statistical Methods
python
class DetectOutliers(CleaningStep):
    def __init__(self, column: str, method: str = 'iqr', action: str = 'flag'):
        self.column = column
        self.method = method
        self.action = action  # 'flag', 'remove', 'cap'

    def name(self) -> str:
        return f"Outlier Detection ({self.column})"

    def execute(self, df: pd.DataFrame, report: CleaningReport) -> pd.DataFrame:
        if self.method == 'iqr':
            outlier_mask = self._iqr_method(df)
        elif self.method == 'zscore':
            outlier_mask = self._zscore_method(df)
        elif self.method == 'modified_zscore':
            outlier_mask = self._modified_zscore_method(df)
        else:
            raise ValueError(f"Unknown method: {self.method}")
# ... (condensed) ...
        Q1 = df[self.column].quantile(0.25)
        Q3 = df[self.column].quantile(0.75)
        IQR = Q3 - Q1
        lower = Q1 - 1.5 * IQR
        upper = Q3 + 1.5 * IQR
        df[self.column] = df[self.column].clip(lower=lower, upper=upper)
        return df

Duplicate Detection

python
class RemoveDuplicates(CleaningStep):
    def __init__(self, subset: list[str] | None = None, strategy: str = 'exact'):
        self.subset = subset
        self.strategy = strategy

    def name(self) -> str:
        return "Remove Duplicates"

    def execute(self, df: pd.DataFrame, report: CleaningReport) -> pd.DataFrame:
        before = len(df)

        if self.strategy == 'exact':
            df = df.drop_duplicates(subset=self.subset, keep='first')
        elif self.strategy == 'fuzzy':
            df = self._fuzzy_dedup(df)

        removed = before - len(df)
        report.add_step(self.name(), removed,
                       # ... (condensed) ...
            for j in range(i + 1, len(values)):
                if j in to_remove:
                    continue
                if fuzz.ratio(str(values[i]).lower(), str(values[j]).lower()) > 90:
                    to_remove.add(j)

        return df.drop(index=list(to_remove)).reset_index(drop=True)

Text Normalization

python
import re
import unicodedata

class NormalizeText(CleaningStep):
    def __init__(self, columns: list[str]):
        self.columns = columns

    def name(self) -> str:
        return "Normalize Text"

    def execute(self, df: pd.DataFrame, report: CleaningReport) -> pd.DataFrame:
        total_modified = 0
        for col in self.columns:
            if col not in df.columns:
                continue
            original = df[col].copy()
            df[col] = df[col].apply(self._normalize)
            modified = (original != df[col]).sum()
            # ... (condensed) ...
        return None
    local, domain = email.rsplit('@', 1)
    # Remove dots in Gmail local part
    if domain in ('gmail.com', 'googlemail.com'):
        local = local.replace('.', '').split('+')[0]
        domain = 'gmail.com'
    return f"{local}@{domain}"

Date Parsing

python
from dateutil import parser as dateparser

class ParseDates(CleaningStep):
    COMMON_FORMATS = [
        '%Y-%m-%d',
        '%Y-%m-%dT%H:%M:%S',
        '%Y-%m-%dT%H:%M:%SZ',
        '%Y-%m-%dT%H:%M:%S%z',
        '%m/%d/%Y',
        '%d/%m/%Y',
        '%m-%d-%Y',
        '%d-%m-%Y',
        '%B %d, %Y',
        '%b %d, %Y',
        '%d %B %Y',
        '%Y%m%d',
    ]

    # ... (condensed) ...
            except (ValueError, TypeError):
                continue
        # Fallback to dateutil parser (slower but handles more formats)
        try:
            return pd.Timestamp(dateparser.parse(value, dayfirst=self.dayfirst))
        except (ValueError, TypeError):
            return None

Encoding Fixes

python
import chardet

def detect_and_fix_encoding(file_path: str) -> pd.DataFrame:
    """Detect file encoding and read with correct encoding."""
    # Detect encoding
    with open(file_path, 'rb') as f:
        raw_data = f.read(100000)  # Read first 100KB for detection
    detected = chardet.detect(raw_data)
    encoding = detected['encoding']
    confidence = detected['confidence']

    print(f"Detected encoding: {encoding} (confidence: {confidence:.0%})")

    # Try detected encoding, fall back to common alternatives
    encodings_to_try = [encoding, 'utf-8', 'latin-1', 'cp1252', 'iso-8859-1']
    for enc in encodings_to_try:
        try:
            df = pd.read_csv(file_path, encoding=enc)
            # ... (condensed) ...
        'ö': 'o',  'ü': 'u',  'ñ': 'n',  'ç': 'c',
        '’': "'",  '“': '"',  'â€\x9d': '"',  'â€"': '-',
        'â€"': '--', '…': '...',
    }
    for bad, good in replacements.items():
        text = text.replace(bad, good)
    return text

Validation Rules

python
class ValidateConstraints(CleaningStep):
    def name(self) -> str:
        return "Validate Constraints"

    def execute(self, df: pd.DataFrame, report: CleaningReport) -> pd.DataFrame:
        violations = []

        # Email format
        if 'email' in df.columns:
            email_pattern = r'^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$'
            invalid_emails = ~df['email'].str.match(email_pattern, na=False)
            count = invalid_emails.sum()
            if count > 0:
                violations.append(f"Invalid emails: {count}")
                df.loc[invalid_emails, 'email'] = None

        # Numeric ranges
        if 'age' in df.columns:
            # ... (condensed) ...
                missing = df[col].isnull().sum()
                if missing > 0:
                    violations.append(f"Missing required '{col}': {missing}")

        report.add_step(self.name(), len(violations),
                       "; ".join(violations) if violations else "All constraints satisfied")
        return df

Data Quality Reporting

python
def generate_quality_report(df: pd.DataFrame) -> dict:
    """Generate a comprehensive data quality report."""
    return {
        "overview": {
            "total_rows": len(df),
            "total_columns": len(df.columns),
            "total_cells": len(df) * len(df.columns),
            "total_missing": df.isnull().sum().sum(),
            "completeness_pct": round((1 - df.isnull().sum().sum() / (len(df) * len(df.columns))) * 100, 2),
        },
        "columns": {
            col: {
                "dtype": str(df[col].dtype),
                "non_null": int(df[col].notna().sum()),
                "null_count": int(df[col].isnull().sum()),
                "null_pct": round(df[col].isnull().sum() / len(df) * 100, 2),
                "unique_count": int(df[col].nunique()),
                "unique_pct": round(df[col].nunique() / max(df[col].notna().sum(), 1) * 100, 2),
                "sample_values": df[col].dropna().head(3).tolist(),
            }
            for col in df.columns
        },
    }

Best Practices

  1. Never modify raw data in place: Always work on copies, keep originals
  2. Document every cleaning decision: Why was this value removed/changed?
  3. Make cleaning reproducible: Scripts, not manual edits
  4. Generate quality reports before AND after cleaning: Measure improvement
  5. Handle edge cases explicitly: Empty strings, whitespace-only, special characters
  6. Validate after cleaning: Ensure constraints are satisfied
  7. Use appropriate methods per data type: Median for numeric, mode for categorical
  8. Be conservative with outlier removal: Flag first, remove only when justified
  9. Test cleaning pipeline on sample data: Verify behavior before full run
  10. Version control cleaning scripts: Track changes to cleaning logic

When to Use

Use this skill when:

  • Designing or implementing data scrubber solutions
  • Reviewing or improving existing data scrubber approaches
  • Making architectural or implementation decisions about data scrubber
  • Learning data scrubber patterns and best practices
  • Troubleshooting data scrubber-related issues

Do NOT use this skill when:

  • The question is about a fundamentally different technology domain
  • A more specific sibling skill covers the exact topic needed
  • The user needs a complete hands-on tutorial rather than expert guidance
Show full SKILL.md (123 more words)Show less

Output Format

markdown
# Data Scrubber Analysis

## Context Assessment
[Situation summary and constraints]

## Recommended Approach
[Primary recommendation with rationale]

## Implementation Steps
1. [Step with specific details]
2. [Step with specific details]
3. [Step with specific details]

## Trade-offs and Considerations
- [Key trade-off 1]
- [Key trade-off 2]

## Next Steps
- [Immediate action item]
- [Follow-up action item]

Example

Input: "Help me implement data scrubber for a medium-scale production application"

Output: A structured analysis covering current state assessment, recommended data scrubber approach with specific patterns, implementation roadmap with milestones, and risk mitigation strategies tailored to the application scale and constraints.

Edge Cases

  • Legacy system integration: When data scrubber must coexist with legacy approaches, provide a gradual migration path rather than a complete rewrite
  • Scale mismatch: When the solution complexity exceeds the project scale, recommend a simpler approach and note when to revisit
  • Team skill gaps: When the team lacks experience with the recommended approach, include learning resources and simpler alternatives
  • Conflicting requirements: When constraints conflict (e.g., performance vs. maintainability), explicitly state the trade-off and recommend based on stated priorities

© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/process/resources/skills-library/bodies/skills/software-engineering/data-scrubber of FerroxLabs/wayland.

Open the folder on GitHubat commit 4c030c7

Compare with similar skills

Data Scrubber next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Scrubber compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Scrubber this skillFerroxLabs/wayland608—~3.6kAutomated safety check: PassApache-2.0
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 9 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from FerroxLabs/wayland

All 1,194 skills in this repo
  • Star Office Helper

    FerroxLabs/wayland

    Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.

    608 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Openclaw Setup

    FerroxLabs/wayland

    OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.

    608 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Tvcontrol Setup

    FerroxLabs/wayland

    Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.

    608 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Ab Testing Specialist

    FerroxLabs/wayland

    End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

    608 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Academic Writer

    FerroxLabs/wayland

    Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…

    608 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Accessibility Auditor

    FerroxLabs/wayland

    Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…

    608 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed

Questions about Data Scrubber

What does Data Scrubber do?

Data cleaning automation expertise covering missing value strategies, outlier detection methods, duplicate detection and deduplication, data type correction, text normalization, date parsing across…. Data Scrubber is an agent skill from FerroxLabs/wayland. Data cleaning automation expertise covering missing value strategies, outlier detection methods, duplicate detection and deduplication, data type correction, text normalization, date parsing across formats, encoding fixes, validation rules, pipeline design patterns, and data quality reporting.

When should I use Data Scrubber?

Data Scrubber fits situations like: the user asks about data scrubber; data scrubber best practices; needs guidance on data scrubber implementation; the user needs a different specialized skill.

How do I install Data Scrubber in Claude Code?

Run `npx skills add FerroxLabs/wayland --skill data-scrubber -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/software-engineering/data-scrubber in FerroxLabs/wayland) into .claude/skills/data-scrubber in your project. Claude Code loads it when a task matches its description.

How do I install Data Scrubber in Codex?

Run `npx skills add FerroxLabs/wayland --skill data-scrubber -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/software-engineering/data-scrubber in FerroxLabs/wayland) into .agents/skills/data-scrubber in your project. Codex loads it when a task matches its description.

Can I use Data Scrubber in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill data-scrubber -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-scrubber, .gemini/skills/data-scrubber, .github/skills/data-scrubber and .opencode/skills/data-scrubber in your project.

What does Data Scrubber need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Scrubber is instructions for the agent only. Our summary lists: Python 3.

Does Data Scrubber access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Scrubber safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Scrubber use?

Data Scrubber is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Scrubber use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Scrubber?

Skills that share tags, products or a category with Data Scrubber: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Issues Deduplication (JetBrains/ideavim, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Scrubber?

FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.

Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.