Agent skill

Data

by notque in notque/vexjoy-agent

“Data analysis and reference enrichment.”

— description from SKILL.md by notque
MITAuto-check: notesData & Analytics

Install Data

skills CLI
$ npx skills add notque/vexjoy-agent --skill data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install notque/vexjoy-agent data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/data .claude/skills/data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data
GitHub stars
435
Used in
1 other repo
Token cost
~2.4k tokens
SKILL.md length
1,044 words
Files
9 (incl. scripts, references)
Skills in repo
61
Repo updated
First seen
Licence
MIT

At a glance

  • Works in 11 steps: FRAME → DEFINE → EXTRACT → …
  • SKILL.md covers A. Data Analysis, B. Reference Enrichment and Deep References
  • Runs Python scripts from its folder; calls python3

About this skill

Data is a skill in notque/vexjoy-agent (435 stars). Its SKILL.md is about 2.4k tokens, with 8 other files in the folder (scripts, references), and copies of it appear in 1 other owners' repositories. Licence: MIT.

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Bash, Grep, Glob, Edit, Task, Agent

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. FRAME
  2. DEFINE
  3. EXTRACT
  4. ANALYZE
  5. CONCLUDE
  6. DECOMPOSE (when --decompose or "extract references")
  7. DISCOVER
  8. RESEARCH
  9. COMPILE
  10. VALIDATE
  11. INTEGRATE

What it can do on your machine

Read from SKILL.md and the folder at commit 5218674. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Bash
    • Grep
    • Glob
    • Edit
    • Task
    • Agent

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data loads about 2.4k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 11 tokens; SKILL.md has 1,044 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~11
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Bash, Grep, Glob, Edit, Task, Agent

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from notque/vexjoy-agent at commit 5218674, republished under its MIT licence (© notque). 1,044 words, ~2,360 tokens.

Download SKILL.mdSave it as .claude/skills/data/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
data
description
Data analysis and reference enrichment.
allowed-tools
Read, Write, Bash, Grep, Glob, Edit, Task, Agent
user-invocable
true
argument-hint
<dataset-or-component-name> [--decompose]
routing.triggers
analyze data, data analysis, CSV, dataset, metrics, trend, cohort, A/B test, statistical, distribution, correlation, KPI, funnel, experiment results, data…
routing.not_for
database schema (agents handle directly), code review (use review)
routing.pairs_with
workflow, assessment
routing.complexity
medium
routing.category
analysis

Data Skill

Two modes. Match the request to a section.

SignalMode
Analyze data, CSV, metrics, A/B test, trend, KPI, funnel, distributionA. Data Analysis
Enrich references, generate references, decompose skill, improve depthB. Reference Enrichment

A. Data Analysis

Every analysis starts with the decision it supports, works backward to evidence required, then touches the data. Analysis without a decision is arithmetic.

Phase 1: FRAME

Establish what decision this analysis supports.

  1. Identify the decision, decision-maker, options, and default action if no analysis is done.
  2. If the user cannot articulate a decision, ask: "What will you do differently based on this analysis?" If exploratory, switch to Exploratory Mode (apply rigor gates, make no causal claims).
  3. Define evidence requirements: what evidence favors each option, minimum threshold for changing the default, deal-breakers.
  4. Save analysis-frame.md.

Gate: Decision identified, options enumerated, evidence requirements saved.

Phase 2: DEFINE

Lock metric definitions before loading data. Defining after seeing data enables cherry-picking.

For each metric: name, exact formula (numerator/denominator), population (included/excluded), time window, segments. For comparisons: define groups and verify fairness.

Save metric-definitions.md. Definitions are locked once Phase 3 starts. If data reveals a definition is unworkable, return here, update, and document the change.

Gate: All metrics defined with formulas and populations.

Phase 3: EXTRACT

Load data. Assess quality. No interpretation.

  1. Detect tools: try import pandas; fall back to csv.DictReader + statistics.
  2. Profile: row count, column types, missing values, date range, distribution stats.
  3. Quality checks (load references/rigor-gates.md Gate 1):
CheckMinimumIf failed
Sample fractionReport N of MWarn if <5% coverage
Time windowNo gaps >10%Adjust or note limitation
Segment size30+ per segmentMerge small segments or exclude
Missing rate<20% per critical columnImpute with disclosure or exclude
  1. Save data-quality-report.md.

Gate: Data loaded, quality assessed, failures documented as limitations.

Phase 4: ANALYZE

Compute metrics per Phase 2 definitions. Report confidence intervals, not point estimates.

  1. Compute using exact formulas. Wilson score CI for proportions.
  2. Fairness gate (comparisons): same time window, same population, confounders documented, survivorship checked (load references/rigor-gates.md Gate 2).
  3. Multiple testing (6+ comparisons): apply Bonferroni (threshold = 0.05/N). Report all segments tested (Gate 3).
  4. Practical significance: report effect size alongside statistical significance. Base-rate context ("from 2.1% to 2.3%", not "+10% lift") (Gate 4).
  5. Save analysis-results.md.

Gate: All metrics computed. Rigor gates applied.

Phase 5: CONCLUDE

Lead with insights. Return to the decision.

  1. Headline finding: one sentence addressing the Phase 1 decision.
  2. Supporting evidence: primary metric with CI, secondary metrics, segment breakdowns.
  3. Limitations: wide CIs are the finding, not a formatting problem.
  4. Decision mapping: does evidence meet threshold? Deal-breakers triggered? Recommended action? Additional data needed?
  5. Save analysis-report.md (load references/output-templates.md for analysis-type templates).

Gate: Report saved with headline, limitations, recommendation tied to decision.

Error Handling (Data Analysis)
ErrorRecovery
No decision contextAsk "What will you do differently?" Switch to Exploratory if none.
Parse failureTry utf-8, latin-1, utf-8-sig. Detect delimiter. Max 3 attempts.
Insufficient segment data (<30)Merge small segments, remove segmentation, or accept with disclosure.
Metrics changed after seeing dataReturn to Phase 2, document changes. Max 2 revisions.
Wide CI on primary metricState: "Data does not support a confident decision." Suggest more data.

B. Reference Enrichment

Enrich agent/skill reference files from Level 0-2 to Level 3+, or decompose bloated body files by extracting domain content into references.

Phase 0: DECOMPOSE (when --decompose or "extract references")

Extract domain-heavy content from a bloated SKILL.md into reference files.

  1. Run python3 scripts/detect-decomposition-targets.py --skill {name} (or --agent).
  2. If no extractable blocks, report "nothing to decompose" and stop.
  3. Snapshot: cp {path} /tmp/decomp-before-{name}.md.
  4. For each block: create reference file, remove from body (MOVE, not copy), add loading table entry.
  5. Retain in body: frontmatter, overview, phase workflow, loading table, error handling.
  6. Validate: python3 scripts/validate-decomposition.py --before /tmp/decomp-before-{name}.md --after {path} --refs {refs_dir}/.
  7. If fails: restore from snapshot. If passes: python3 scripts/validate-references.py --skill {name}.

Load references/decomposition-prompt.md for the autonomous decomposition prompts.

Gate: Validation passes. Body reduced. All extracted content in references.

Show full SKILL.md (372 more words)Show less
Phase 1: DISCOVER
  1. Run python3 scripts/gap-analyzer.py --agent {name} (or --skill).
  2. Read the component's .md and existing references. Map coverage.
  3. Compare stated domains against covered domains. Output gap report.

Gate: At least one gap identified. If Level 3 already, stop.

Phase 2: RESEARCH

For each gap: identify version-specific patterns, failure modes with detection commands (grep -rn "pattern"), error-fix mappings, project conventions. Dispatch up to 5 parallel research agents per sub-domain.

Gate: Each gap has 10+ concrete findings (version numbers, function names, grep patterns). Generic advice does not count.

Phase 3: COMPILE

Create one reference file per sub-domain (max 500 lines) following references/reference-file-template.md. Include: overview, pattern table with version ranges, failure mode table with detection commands, error-fix mappings.

Do-pairing rule: every failure mode needs a "Do instead" counterpart. No bare negative blocks.

Validate: python3 scripts/validate-references.py --agent {name} and --check-do-framing. Both must exit 0. Then run condense on each file.

Gate: Each file 80-500 lines. Both validations pass.

Phase 4: VALIDATE

Tier 1: python3 scripts/audit-reference-depth.py --agent {name} --json. Level must be 3. Tier 2: Apply references/quality-rubric.md. For each pattern: detection command present? Would a reviewer using only this file produce Level 3 output?

Gate: Both tiers pass. Max 2 loops per gap before flagging for manual review.

Phase 5: INTEGRATE
  1. Add/update loading table in the component body.
  2. Validate: python3 scripts/validate-references.py --agent {name} and python3 -m pytest scripts/tests/test_reference_loading.py -k {name} -v.
  3. Stage changes.

Gate: Validation passes. Report level change (was N, now M) and new file list.

Error Handling (Reference Enrichment)
ErrorRecovery
Gap analyzer failsCheck both agents/ and skills/ directories.
Phase 2 gate fails (<10 findings)Domain may be narrow. Flag for manual enrichment.
Phase 4 still below Level 3Files too generic. Target Phase 2 at weakest section.
Decomposition validation failsRestore from snapshot. Check for partial extractions.

Deep References

All references are >100 lines of domain-specific content. Load as directed by sections above.

SignalReferenceLines
Phase 3-4: statistical gates, sample adequacy, fairnessreferences/rigor-gates.md378
Phase 5: report templates (A/B, trend, distribution, cohort)references/output-templates.md489
Failure mode recognition (p-hacking, survivorship, Simpson's)references/preferred-patterns.md240
Classifying reference depth Level 0-3references/quality-rubric.md173
Writing new reference filesreferences/reference-file-template.md166
Running headless decompositionreferences/decomposition-prompt.md205
Running headless enrichmentreferences/enrichment-prompt.md117

© notque, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/analysis/data of notque/vexjoy-agent.

  • SKILL.md
  • references/decomposition-prompt.md
  • references/enrichment-prompt.md
  • references/output-templates.md
  • references/preferred-patterns.md
  • references/quality-rubric.md
  • references/reference-file-template.md
  • references/rigor-gates.md
  • scripts/gap-analyzer.py

Open the folder on GitHubat commit 5218674

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in notque/vexjoy-agent, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data this skillnotque/vexjoy-agent4351 repos~2.4kAutomated safety check: NotesMIT
Exploratory Data Analysisspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2052 repos~3.4kAutomated safety check: NotesMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT

Similar skills

  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    205 GitHub starsUsed in 2 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Yichen Wecom Local Vault

    mcncarl/yichen-skills

    Read, decrypt, query, search, and export local WeCom/企业微信 5.x desktop databases on macOS into a private read-only vault.

    4.3k GitHub stars~1.3k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed

More from notque/vexjoy-agent

All 61 skills in this repo
  • Game Asset Generator

    notque/vexjoy-agent

    Deterministic palette/matrix pixel art (not AI). An agent skill from notque/vexjoy-agent.

    435 GitHub stars~2.3k tokensUpdated 4 days ago
    Auto-check: notes
  • PR Workflow

    notque/vexjoy-agent

    Pull request lifecycle: commit, codex review, sync, review, fix, status, cleanup, and PR mining.

    435 GitHub stars~2.8k tokensUpdated 4 days ago
    Auto-check: notes
  • Architecture Deepening

    notque/vexjoy-agent

    Improve architecture across modules by deepening interfaces.

    435 GitHub stars~3.3k tokensUpdated 4 days ago
    Auto-check: notes
  • Code Quality

    notque/vexjoy-agent

    Code quality: cleanup, linting, formatting, quality gates. An agent skill from notque/vexjoy-agent.

    435 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check: notes
  • Codebase Analyzer

    notque/vexjoy-agent

    Statistical rule discovery from Go codebase patterns. An agent skill from notque/vexjoy-agent.

    435 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check: notes
  • Comment Quality

    notque/vexjoy-agent

    Review and fix temporal references in code comments. An agent skill from notque/vexjoy-agent.

    435 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check: notes

Questions about Data

How do I install Data in Claude Code?

Run `npx skills add notque/vexjoy-agent --skill data -a claude-code`. Or copy the skill folder (skills/analysis/data in notque/vexjoy-agent) into .claude/skills/data in your project. Claude Code loads it when a task matches its description.

How do I install Data in Codex?

Run `npx skills add notque/vexjoy-agent --skill data -a codex`. Or copy the skill folder (skills/analysis/data in notque/vexjoy-agent) into .agents/skills/data in your project. Codex loads it when a task matches its description.

Can I use Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notque/vexjoy-agent --skill data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data, .gemini/skills/data, .github/skills/data and .opencode/skills/data in your project.

What does Data need to run?

Going by SKILL.md and its folder, Data needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Its frontmatter pre-approves these tools: Read, Write, Bash, Grep, Glob, Edit, Task, Agent.

Does Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data use?

Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.

What are the alternatives to Data?

Skills that share tags, products or a category with Data: Exploratory Data Analysis (spacering-net/codeg, 3.8k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 205 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data?

notque (a GitHub user) maintains it in notque/vexjoy-agent, which has 435 GitHub stars. The repository holds 61 skills in this directory. The repository was last updated on October 3, 2026.

Source: notque/vexjoy-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.