Agent skill

Data Cog Guide

by wentorai in wentorai/research-plugins

Upload messy CSVs with minimal prompting for deep automated analysis

MITAuto-check passedDocuments & Office

Install Data Cog Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill data-cog-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins data-cog-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/wrangling/data-cog-guide .claude/skills/data-cog-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-cog-guide
GitHub stars
298
Used in
1 other repo
Token cost
~1.8k tokens
SKILL.md length
456 words
Files
1
Skills in repo
428
Repo updated
First seen
Licence
MIT

At a glance

Upload messy CSVs with minimal prompting for deep automated analysis

  • Works in 6 steps: Schema overview: Column names, inferred… → Univariate statistics: Mean, median,… → Missing data matrix: Heatmap-style… → …
  • Tasks that involve CSV and tabular files
  • SKILL.md covers Overview, Automated Ingestion Pipeline, Deep Automated Profiling and Interactive Analysis Workflow, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Cog Guide is an agent skill from wentorai/research-plugins. Upload messy CSVs with minimal prompting for deep automated analysis

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering CSV and tabular files. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve CSV and tabular files

Example prompts

  • “/data-cog-guide”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Schema overview: Column names, inferred types, semantic roles (ID, feature, target, timestamp).
  2. Univariate statistics: Mean, median, mode, std, skewness, kurtosis for numeric columns; frequency tables for categoricals.
  3. Missing data matrix: Heatmap-style report of missingness patterns across all columns.
  4. Correlation analysis: Pairwise Pearson, Spearman, and Cramér's V correlations.
  5. Distribution flags: Columns that are heavily skewed, zero-inflated, or constant.
  6. Duplicate detection: Exact row duplicates and near-duplicate clusters.

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pandas.pydata.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Cog Guide loads about 1.8k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 456 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 456 words, ~1,782 tokens.

Download SKILL.mdSave it as .claude/skills/data-cog-guide/SKILL.md (or your agent's skills folder).
name
data-cog-guide
description
Upload messy CSVs with minimal prompting for deep automated analysis

Data Cog Guide

An intelligent data analysis assistant that accepts messy, poorly documented CSV files and automatically infers structure, cleans anomalies, and produces deep analytical reports with minimal user prompting. Designed for researchers who need quick insights from unfamiliar or inherited datasets without spending hours on manual data preparation.

Overview

Researchers frequently receive datasets from collaborators, public repositories, or legacy systems that lack documentation, use inconsistent formatting, and contain mixed data quality. Traditional analysis requires significant upfront effort to understand and prepare such data. Data Cog automates this process by applying heuristic inference, pattern recognition, and iterative cleaning to produce analysis-ready data along with a comprehensive profile report.

The skill implements a "zero-configuration" philosophy: provide the CSV file path and an optional research question, and it handles encoding detection, delimiter inference, type casting, missingness assessment, and initial exploratory statistics automatically.

Automated Ingestion Pipeline

Smart Loading
python
import pandas as pd
import chardet
import io

def smart_load_csv(filepath: str) -> tuple:
    """
    Intelligently load a CSV file, auto-detecting encoding,
    delimiter, header row, and comment lines.
    """
    # Step 1: Detect encoding
    with open(filepath, 'rb') as f:
        raw = f.read(100000)
    encoding = chardet.detect(raw)['encoding']

    # Step 2: Detect delimiter
    import csv
    with open(filepath, 'r', encoding=encoding, errors='replace') as f:
        sample = f.read(8192)
    sniffer = csv.Sniffer()
    try:
        dialect = sniffer.sniff(sample)
        delimiter = dialect.delimiter
    except csv.Error:
        delimiter = ','

    # Step 3: Detect header row (skip comment lines)
    skip_rows = 0
    with open(filepath, 'r', encoding=encoding, errors='replace') as f:
        for line in f:
            if line.startswith('#') or line.startswith('//') or line.strip() == '':
                skip_rows += 1
            else:
                break

    # Step 4: Load with inferred parameters
    df = pd.read_csv(
        filepath, encoding=encoding, delimiter=delimiter,
        skiprows=skip_rows, low_memory=False
    )

    metadata = {
        'encoding': encoding,
        'delimiter': repr(delimiter),
        'skipped_rows': skip_rows,
        'shape': df.shape
    }
    return df, metadata
Automatic Type Inference
python
def auto_cast_columns(df: pd.DataFrame) -> pd.DataFrame:
    """
    Automatically cast columns to their most appropriate types.
    Handles dates, numerics stored as strings, booleans, and categories.
    """
    for col in df.columns:
        # Try numeric conversion
        numeric = pd.to_numeric(df[col], errors='coerce')
        if numeric.notna().mean() > 0.85:
            df[col] = numeric
            continue

        # Try datetime conversion
        datetime = pd.to_datetime(df[col], errors='coerce', infer_datetime_format=True)
        if datetime.notna().mean() > 0.85:
            df[col] = datetime
            continue

        # Try boolean detection
        unique_lower = df[col].dropna().astype(str).str.lower().unique()
        if set(unique_lower).issubset({'true', 'false', 'yes', 'no', '1', '0', 'y', 'n'}):
            df[col] = df[col].astype(str).str.lower().map(
                {'true': True, 'false': False, 'yes': True, 'no': False,
                 '1': True, '0': False, 'y': True, 'n': False}
            )
            continue

        # Convert low-cardinality strings to category
        if df[col].nunique() / len(df) < 0.05 and df[col].nunique() < 50:
            df[col] = df[col].astype('category')

    return df

Deep Automated Profiling

Profile Report Generation

The profiling stage produces a structured report covering:

  1. Schema overview: Column names, inferred types, semantic roles (ID, feature, target, timestamp).
  2. Univariate statistics: Mean, median, mode, std, skewness, kurtosis for numeric columns; frequency tables for categoricals.
  3. Missing data matrix: Heatmap-style report of missingness patterns across all columns.
  4. Correlation analysis: Pairwise Pearson, Spearman, and Cramér's V correlations.
  5. Distribution flags: Columns that are heavily skewed, zero-inflated, or constant.
  6. Duplicate detection: Exact row duplicates and near-duplicate clusters.
MetricNumeric ColumnsCategorical Columns
Central tendencyMean, median, modeMode, frequency
DispersionStd, IQR, range, CVUnique count, entropy
ShapeSkewness, kurtosisImbalance ratio
QualityMissing %, zero %, outlier %Missing %, rare labels %

Interactive Analysis Workflow

Show full SKILL.md (188 more words)Show less
Minimal-Prompt Usage Pattern

The recommended workflow requires only three inputs:

  1. File path: The CSV to analyze.
  2. Research question (optional): A one-sentence description of what you want to learn.
  3. Output format: "summary", "full_report", or "cleaned_csv".
User: Analyze /data/survey_results_2025.csv
      Question: What factors predict participant satisfaction?
      Output: full_report

Data Cog will:
  1. Load and profile the dataset (auto-detect everything)
  2. Clean and transform (handle missing data, encode categoricals)
  3. Run correlation analysis focused on satisfaction-related columns
  4. Generate regression models predicting satisfaction
  5. Produce a structured report with findings and visualizations
Iterative Refinement

After the initial automated analysis, you can refine by asking targeted follow-up questions:

  • "Focus only on respondents from Group A"
  • "Exclude the first 50 rows (pilot data)"
  • "Treat column X as ordinal with levels: low < medium < high"
  • "Run the same analysis but with log-transformed income"

Best Practices

  • Always review the auto-generated profile before trusting downstream results.
  • Verify that automatic type inference made sensible choices, especially for ambiguous columns.
  • Provide a research question when possible to guide feature selection and analysis focus.
  • Save the cleaning audit log alongside your results for reproducibility.
  • For datasets over 1 million rows, consider sampling for the initial profile to save time.

References

  • Breck, E., et al. (2019). Data Validation for Machine Learning. MLSys 2019.
  • Hynes, N., et al. (2017). The Data Linter: Lightweight, Automated Sanity Checking for ML Data Sets. NIPS MLSys Workshop.
  • Pandas Development Team (2024). pandas: Powerful Python Data Analysis Toolkit. https://pandas.pydata.org/

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/analysis/wrangling/data-cog-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Data Cog Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Cog Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Cog Guide this skillwentorai/research-plugins2981 repos~1.8kAutomated safety check: PassMIT
Data Cleanupsgharlow/claude-code-recipes388—~566Automated safety check: PassCustom licence
Sn Da Image CaptionMichaelYang-lyx/AIDABench1111 repos~2kAutomated safety check: PassNone
Douban Skilldaymade/claude-code-skills1.4k—~1.1kAutomated safety check: PassMIT
Champion Trackergooseworks-ai/goose-skills1.2k1 repos~1.1kAutomated safety check: NotesMIT
Tabular Cleanupgaasher/Agent-Loop-Skills174—~4kAutomated safety check: PassMIT

Similar skills

  • Data Cleanup

    sgharlow/claude-code-recipes

    Clean and standardize messy tabular data (CSV, spreadsheet paste, system exports) into an analysis-ready dataset — consistent dates and names, typed columns, duplicates identified, missing values…

    388 GitHub stars~566 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Sn Da Image Caption

    MichaelYang-lyx/AIDABench

    图片理解与数据提取 skill。当图片文件(.png/.jpg/.jpeg/.gif/.webp/.bmp)是主要输入且用户需要理解、提取数据或分析图片内容时使用。提供预配置的 caption 脚本(scripts/caption.py),通过 vision 模型将图片转为文本描述,无需额外配置 API Key。覆盖:(1) 通过 scripts/caption.py…

    111 GitHub starsUsed in 1 repo~2k tokens
    Documents & OfficeAuto-check passed
  • Douban Skill

    daymade/claude-code-skills

    Export and sync Douban (豆瓣) book/movie/music/game collections to local CSV files via Frodo API.

    1.4k GitHub stars~1.1k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Champion Tracker

    gooseworks-ai/goose-skills

    Track product champions for job changes and qualify their new companies against ICP.

    1.2k GitHub starsUsed in 1 repo~1.1k tokens
    Documents & OfficeAuto-check: notes
  • Tabular Cleanup

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a messy tabular data dump (CSV/TSV/parquet/Excel/JSON) and wants it iteratively cleaned to an inferred data contract — a checklist of deterministic…

    174 GitHub stars~4k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Meta Sensitivity Plot

    aipoch/medical-research-skills

    Generate leave-one-out sensitivity analysis plots for meta-analysis.

    2k GitHub stars~1.6k tokensUpdated 20 days ago
    Documents & OfficeAuto-check passed

More from wentorai/research-plugins

All 428 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Data Cog Guide

What does Data Cog Guide do?

Upload messy CSVs with minimal prompting for deep automated analysis. Data Cog Guide is an agent skill from wentorai/research-plugins.

When should I use Data Cog Guide?

Data Cog Guide fits situations like: tasks that involve CSV and tabular files.

How do I install Data Cog Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill data-cog-guide -a claude-code`. Or copy the skill folder (skills/analysis/wrangling/data-cog-guide in wentorai/research-plugins) into .claude/skills/data-cog-guide in your project. Claude Code loads it when a task matches its description.

How do I install Data Cog Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill data-cog-guide -a codex`. Or copy the skill folder (skills/analysis/wrangling/data-cog-guide in wentorai/research-plugins) into .agents/skills/data-cog-guide in your project. Codex loads it when a task matches its description.

Can I use Data Cog Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill data-cog-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-cog-guide, .gemini/skills/data-cog-guide, .github/skills/data-cog-guide and .opencode/skills/data-cog-guide in your project.

What does Data Cog Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Cog Guide is instructions for the agent only. Our summary lists: Python 3.

Does Data Cog Guide access the network?

SKILL.md names 1 domain. As links in the text: pandas.pydata.org. This is read from the text; nothing was executed.

Is Data Cog Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Cog Guide use?

Data Cog Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Cog Guide use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Cog Guide?

Skills that share tags, products or a category with Data Cog Guide: Data Cleanup (sgharlow/claude-code-recipes, 388 stars), Sn Da Image Caption (MichaelYang-lyx/AIDABench, 111 stars), Douban Skill (daymade/claude-code-skills, 1.4k stars) and Champion Tracker (gooseworks-ai/goose-skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Cog Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 428 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.