Agent skill

Data Analysis

by damianvtran in damianvtran/local-operator

Analysing data rigorously: profile before trust (keys, missingness), reproducible queries, provenance per number, units and sanity checks.

MITAuto-check passedData & Analytics

Install Data Analysis

skills CLI
$ npx skills add damianvtran/local-operator --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install damianvtran/local-operator data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/damianvtran/local-operator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/local_operator/skills/builtin/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
217
Token cost
~709 tokens
SKILL.md length
389 words
Files
1
Repo updated
First seen
Licence
MIT

At a glance

Analysing data rigorously: profile before trust (keys, missingness), reproducible queries, provenance per number, units and sanity checks.

  • Exploring datasets
  • SKILL.md covers Profile before trust, Reproducible queries, Guard silent joins and filters and Provenance and checks, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Computing metrics

What it does

Data Analysis is an agent skill from damianvtran/local-operator. Analysing data rigorously: profile before trust (keys, missingness), reproducible queries, provenance per number, units and sanity checks. Use when exploring datasets, computing metrics, or answering data questions.

Its SKILL.md is about 710 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data analysis. The repository describes itself as: An open-source AI agent hub for your own machine: build organizations of collaborating agents that message each other, run in the background around the clock, and share your… The licence is MIT.

When your agent uses it

  • Exploring datasets
  • Computing metrics
  • Answering data questions

Example prompts

  • “/data-analysis”

What it can do on your machine

Read from SKILL.md and the folder at commit 05e4c97. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Analysis loads about 709 tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 389 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~709

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from damianvtran/local-operator at commit 05e4c97, republished under its MIT licence (© damianvtran). 389 words, ~709 tokens.

Download SKILL.mdSave it as .claude/skills/data-analysis/SKILL.md (or your agent's skills folder).
name
data-analysis
description
Analysing data rigorously: profile before trust (keys, missingness), reproducible queries, provenance per number, units and sanity checks. Use when exploring datasets, computing metrics, or answering data questions.

Data analysis

Numbers that survive scrutiny: profile first, keep queries reproducible, carry provenance, sanity-check the result.

Profile before trust

CheckWhy it matters
Shape: rows, columns, time rangeKnow what you have before asking questions of it
Types and units per columnA text date, or cents read as dollars, corrupts every downstream number
Keys and duplicatesDuplicate keys silently change every aggregate
Missingness per columnNull-heavy columns change what a rate means
Distributions and rangesOutliers and impossible values (negative ages, future dates) surface here
Timezone and calendar of timestampsDay boundaries move with zones

Keep the profile as a small set of queries you can rerun, not ad-hoc typing.

Reproducible queries

  • Keep the exact query behind every number in the report. A number nobody can re-derive is a rumour.
  • Deterministic: pin the snapshot or time window, make ordering explicit before any limit, and avoid "latest row" ambiguity.
  • Work in steps; save intermediates. After changing an input, rerun the step rather than mentally adjusting its output.

Guard silent joins and filters

  • Count rows before and after every join. An inner join on an incomplete key drops rows without an error.
  • Check filter semantics: inclusive or exclusive bounds, null membership, timezone edges. Apply each filter once.
  • State the population: "of the 4,210 rows matching X" beats "of the data".
Show full SKILL.md (170 more words)Show less

Provenance and checks

  • Every reported number carries its source (table, snapshot), filter, and derivation. Derive it, or say where it came from.
  • Reconcile: totals equal the sum of their parts; an independent cut of the same data reproduces the number.
  • Units on everything: counts, rates, currency. Percentages name their denominator. Magnitudes are checked against expectation; a 1000x jump is a units or join error until proven otherwise.
  • Small n: below your stated threshold, report the count, not a rate. "60% (n=5)" misleads; "3 of 5" does not.

Output

  • Separate raw from summary: show the summary and keep the raw one step away so the reader can drill down.
  • State assumptions in the report: time window, definitions, exclusions, null handling.

Red flags

  • Answering from the first query without looking at the data's shape.
  • A join with no row count before and after.
  • A percentage with no denominator; a total with no source.
  • Dropping rows to "clean" data without recording the rule.
  • A distribution treated as stable when it trends over time.

© damianvtran, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in local_operator/skills/builtin/data-analysis of damianvtran/local-operator.

Open the folder on GitHubat commit 05e4c97

Compare with similar skills

Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Analysis this skilldamianvtran/local-operator217—~709Automated safety check: PassMIT
Exploratory Data Analysisspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow84k4 repos~2.2kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2122 repos~3.4kAutomated safety check: NotesMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    84k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    212 GitHub starsUsed in 2 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed

More from damianvtran/local-operator

All 8 skills in this repo
  • Design QA

    damianvtran/local-operator

    Deterministic design review for UI/UX on web and terminal screens: measure contrast, spacing, overlap, clipping and copy defects, then judge the rest.

    217 GitHub stars~729 tokensUpdated today
    Auto-check passed
  • Skill Authoring

    damianvtran/local-operator

    Writing and improving agent skills: description as routing signal, progressive disclosure, token budgets, testing before shipping, and improving from repeated mistakes.

    217 GitHub stars~748 tokensUpdated today
    Auto-check passed
  • Code Review

    damianvtran/local-operator

    Reviewing code for defects: diff-first reading, severity classification, evidence over inference, finding caps, and terminal verdicts.

    217 GitHub stars~672 tokensUpdated today
    Auto-check passed
  • Financial Analysis

    damianvtran/local-operator

    Financial figures done carefully: periods and fiscal calendars, units and currencies, reconciliation, magnitude sanity checks, sourced assumptions.

    217 GitHub stars~657 tokensUpdated today
    Auto-check passed
  • Frontend Design

    damianvtran/local-operator

    Designing interfaces with a deliberate point of view: plan before pixels, avoid the generic-default look, typography and colour discipline, a quality floor that is not negotiable.

    217 GitHub stars~773 tokensUpdated today
    Auto-check passed
  • Responding To Review

    damianvtran/local-operator

    Answering review findings: verify each claim before implementing, no performative agreement, per-item fix/reject/defer with evidence, one batched remediation round.

    217 GitHub stars~616 tokensUpdated today
    Auto-check passed

Questions about Data Analysis

What does Data Analysis do?

Analysing data rigorously: profile before trust (keys, missingness), reproducible queries, provenance per number, units and sanity checks. Data Analysis is an agent skill from damianvtran/local-operator. Analysing data rigorously: profile before trust (keys, missingness), reproducible queries, provenance per number, units and sanity checks.

When should I use Data Analysis?

Data Analysis fits situations like: exploring datasets; computing metrics; answering data questions.

How do I install Data Analysis in Claude Code?

Run `npx skills add damianvtran/local-operator --skill data-analysis -a claude-code`. Or copy the skill folder (local_operator/skills/builtin/data-analysis in damianvtran/local-operator) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Data Analysis in Codex?

Run `npx skills add damianvtran/local-operator --skill data-analysis -a codex`. Or copy the skill folder (local_operator/skills/builtin/data-analysis in damianvtran/local-operator) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add damianvtran/local-operator --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Data Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Analysis is instructions for the agent only.

Does Data Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Analysis use?

Data Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Analysis use?

About 709 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Analysis?

Skills that share tags, products or a category with Data Analysis: Exploratory Data Analysis (spacering-net/codeg, 3.9k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 84k stars), Exploratory Data Analysis (Oleafly/Oleafly, 212 stars) and Pandas Pro (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Analysis?

damianvtran (a GitHub user) maintains it in damianvtran/local-operator, which has 217 GitHub stars. The repository was last updated on October 11, 2026.

Source: damianvtran/local-operator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.