Agent skill

Exploratory Data Analysis

by jaechang-hits in jaechang-hits/SciAgent-Skills

Methodology for exploratory data analysis on scientific files.

CC-BY-4.0Auto-check passedData & Analytics

Install Exploratory Data Analysis

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill exploratory-data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills exploratory-data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scientific-computing/exploratory-data-analysis .claude/skills/exploratory-data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
exploratory-data-analysis
GitHub stars
370
Token cost
~3.1k tokens
SKILL.md length
1,294 words
Files
2 (incl. references)
Skills in repo
163
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Methodology for exploratory data analysis on scientific files.

  • Works in 6 steps: File Identification → Initial Assessment → Data Loading → …
  • Given a data file for initial exploration
  • SKILL.md covers Overview, Key Concepts, Decision Framework and Best Practices, plus 5 more sections
  • Calls pip

What it does

Exploratory Data Analysis is an agent skill from jaechang-hits/SciAgent-Skills. Methodology for exploratory data analysis on scientific files. Decision frameworks by data type (tabular, sequence, image, spectral, structural, omics), quality assessment, report generation, format detection across 200+ formats. Use when given a data file for initial exploration or to pick an analysis before a pipeline.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/file_format_reference.md`).

It sits in Data & Analytics, covering Data analysis. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.

When your agent uses it

  • Given a data file for initial exploration
  • Pick an analysis before a pipeline

Example prompts

  • “/exploratory-data-analysis”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. File Identification
  2. Initial Assessment
  3. Data Loading
  4. Quality Assessment
  5. Statistical Analysis
  6. Report Generation

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • openmicroscopy.org
    • psidev.info

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Exploratory Data Analysis loads about 3.1k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 1,294 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 1,294 words, ~3,088 tokens.

Download SKILL.mdSave it as .claude/skills/exploratory-data-analysis/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
exploratory-data-analysis
description
Methodology for exploratory data analysis on scientific files. Decision frameworks by data type (tabular, sequence, image, spectral, structural, omics), quality assessment, report generation, format detection across 200+ formats. Use when given a data file for initial exploration or to pick an analysis before a pipeline.
license
CC-BY-4.0

Exploratory Data Analysis for Scientific Data

Overview

Exploratory data analysis (EDA) is the systematic examination of scientific data files to understand their structure, content, quality, and characteristics before formal analysis. This knowhow covers methodology for detecting file types, selecting appropriate analysis approaches, assessing data quality, and generating comprehensive reports across all major scientific data domains.

Key Concepts

Scientific Data Type Categories
CategoryCommon FormatsTypical AnalysisKey Libraries
TabularCSV, TSV, XLSX, ParquetSummary statistics, distributions, correlations, missing valuespandas, polars
SequenceFASTA, FASTQ, SAM/BAMLength distribution, quality scores, GC content, alignment statsBioPython, pysam
Image/MicroscopyTIFF, ND2, CZI, DICOMDimensions (XYZCT), intensity stats, metadata, calibrationtifffile, aicsimageio, nd2reader
SpectralmzML, SPC, JCAMP, FIDPeak detection, baseline, S/N ratio, resolutionpymzml, nmrglue, pyteomics
StructuralPDB, CIF, MOL, SDFAtom counts, bond validation, B-factors, completenessBioPython, RDKit, MDAnalysis
Array/TensorNPY, HDF5, Zarr, NetCDFShape, dtype, value range, NaN/Inf check, chunk structurenumpy, h5py, zarr, xarray
OmicsH5AD, MTX, VCF, BEDFeature/sample counts, sparsity, annotation completenessscanpy, pyranges, cyvcf2
Format Detection Strategy
  1. Extension-based: Map file extension to category (primary method)
  2. Magic bytes: Check file header for binary format identification (HDF5: \x89HDF, GZIP: \x1f\x8b)
  3. Content sniffing: For ambiguous extensions (.txt, .dat, .csv), inspect first lines for delimiters, headers, or format markers
  4. Compound extensions: Handle .ome.tiff, .nii.gz, .tar.gz by checking from the rightmost extension inward
Data Quality Dimensions
  • Completeness: Missing values, empty fields, truncated records
  • Consistency: Matching dimensions, compatible dtypes, cross-field validation
  • Validity: Values within expected ranges, proper encoding, format compliance
  • Provenance: Instrument metadata, software versions, processing history
  • Uniqueness: Duplicate records, redundant features

Decision Framework

Data file received
├── What is the file type?
│   ├── Known extension → Look up in format reference
│   ├── Unknown extension → Magic bytes / content sniffing
│   └── Directory (e.g., .d, .zarr) → Check internal structure
│
├── What category does it belong to?
│   ├── Tabular → Summary stats, distributions, correlations
│   ├── Sequence → Length/quality distributions, composition
│   ├── Image → Dimensions, channels, intensity, metadata
│   ├── Spectral → Peaks, baseline, resolution, S/N
│   ├── Structural → Atom/bond validation, geometry checks
│   ├── Array → Shape, dtype, value range, sparsity
│   └── Omics → Feature counts, sample QC, annotation check
│
├── How large is the file?
│   ├── Small (<100 MB) → Load fully, comprehensive analysis
│   ├── Medium (100 MB–1 GB) → Sample or lazy evaluation
│   └── Large (>1 GB) → Stream/chunk, representative sampling
│
└── What is the analysis goal?
    ├── Pre-pipeline QC → Focus on completeness, format compliance
    ├── Data understanding → Statistics, distributions, patterns
    ├── Troubleshooting → Compare against expected format/values
    └── Documentation → Full report with recommendations
Quick Reference: Analysis Approach by Data Type
Data TypeFirst CheckCore AnalysisVisualization
Tabulardtypes, shape, nullsdescribe(), correlations, outliershistograms, scatter, heatmap
Sequencerecord count, formatlength dist., quality, compositionquality plots, length histogram
Imagedimensions, bit depthintensity stats, channel infothumbnail, histogram
Spectralscan count, m/z rangepeak detection, TIC, baselinespectrum plot, TIC chromatogram
Structuralatom/residue countB-factors, missing residuesRamachandran, contact map
Arrayshape, dtypestatistics, NaN checkslice visualization
Omicsgenes × cells matrixsparsity, QC metricsviolin plots, PCA

Best Practices

  1. Always check file integrity first — verify file is complete (not truncated) and readable before deep analysis. Check file size against expectations
  2. Sample large files before full analysis — for files with millions of records, analyze a representative sample (first N records, random sample, or stratified sample) to get quick feedback
  3. Use lazy/streaming readers when available — pl.scan_parquet(), h5py dataset slicing, pysam indexed access prevent memory overflows
  4. Validate metadata against data — cross-check stated dimensions vs actual data, verify column count matches header, confirm timestamps are monotonic
  5. Report data quality quantitatively — "5.2% missing values in column X" is more useful than "some missing values". Include completeness percentages, outlier counts, and format compliance scores
  6. Consider data provenance — note instrument type, software version, processing steps, and any preprocessing already applied. This context affects downstream analysis choices
  7. Generate actionable recommendations — don't just describe the data; suggest specific preprocessing steps (normalization method, imputation strategy), appropriate analyses, and potential issues to watch for
  8. Preserve analysis reproducibility — include library versions, parameter choices, and random seeds in reports so EDA can be reproduced
  9. Check for batch effects early — in multi-sample datasets, compare distributions across batches, plates, or runs before combining
  10. Handle vendor-specific formats carefully — many instruments produce proprietary formats (.nd2, .czi, .raw). Document which reader library and version was used, as format support varies

Common Pitfalls

  1. Assuming CSV means clean tabular data — CSV files can have inconsistent delimiters, mixed encodings, embedded newlines, or malformed quoting. How to avoid: Use pd.read_csv(engine='python') for robustness; check encoding with chardet

  2. Ignoring missing value encoding — scientific data uses diverse null representations: NA, NaN, -999, empty string, #N/A, .. How to avoid: Specify na_values parameter; check for sentinel values in numeric columns

  3. Drawing conclusions from truncated files — large file transfers can fail silently. How to avoid: Check file size, verify record counts against expected values, check for EOF markers

  4. Applying wrong reader to file — some extensions are ambiguous (.raw = Thermo MS, XRD, or image; .d = Agilent directory or generic data). How to avoid: Use magic bytes and context (source instrument) to disambiguate

  5. Memory overflow on large datasets — loading a 10 GB CSV into a pandas DataFrame will fail. How to avoid: Check file size first; use chunked reading, lazy evaluation, or sampling for files >100 MB

  6. Ignoring coordinate systems and units — microscopy data may use pixels vs microns; spectroscopy data may use wavelength vs wavenumber vs energy. How to avoid: Extract and report units from metadata; verify calibration information

  7. Treating all columns as independent — scientific tabular data often has hierarchical structure (replicates nested within conditions). How to avoid: Identify experimental design from column names and metadata before computing correlations

  8. Skipping format-specific quality metrics — generic statistics miss domain-specific issues (e.g., Phred quality scores in FASTQ, R-factors in crystallography, mass accuracy in MS). How to avoid: Consult the format reference for domain-specific QC metrics

  9. Overinterpreting small samples — EDA on first 1000 rows may not represent the full dataset's distribution. How to avoid: Sample from multiple positions in the file; report sample size and sampling method

  10. Not checking for duplicates — duplicate records are common in merged datasets and database exports. How to avoid: Check for exact and near-duplicates early; report the duplication rate

Show full SKILL.md (399 more words)Show less

Workflow

Step 1: File Identification
  • Extract file extension (handle compound extensions: .ome.tiff, .nii.gz)
  • Look up format in reference catalog
  • If unknown: check magic bytes, inspect first lines, ask user about source instrument
  • Record: format name, category, typical content, recommended libraries
Step 2: Initial Assessment
  • File size, creation date, permissions
  • For binary formats: verify file integrity (header/footer checks)
  • For text formats: detect encoding, delimiter, line endings
  • For directories: enumerate contents and check expected structure
Step 3: Data Loading
  • Install required library if missing (provide pip install command)
  • Use appropriate reader with explicit parameters (dtype, encoding, columns)
  • For large files: load metadata first, then sample data
  • Record: dimensions (rows × columns, or X×Y×Z×C×T), dtypes, memory footprint
Step 4: Quality Assessment
  • Completeness: count nulls per column/field, check for truncation
  • Validity: range checks, dtype verification, format compliance
  • Consistency: cross-field validation, duplicate detection
  • Domain-specific: Phred scores (FASTQ), B-factors (PDB), mass accuracy (MS), intensity range (microscopy)
Step 5: Statistical Analysis
  • Summary statistics (mean, median, std, min, max, quartiles)
  • Distribution characteristics (skewness, modality, outliers)
  • Correlation structure (for multi-variable data)
  • Domain-specific metrics (GC content, coverage, resolution, S/N ratio)
Step 6: Report Generation

Generate a structured markdown report containing:

  1. Header: filename, timestamp, file size
  2. Format Information: type, description, typical use cases
  3. Data Structure: dimensions, dtypes, memory usage
  4. Quality Assessment: completeness scores, validity checks, issues found
  5. Statistical Summary: key metrics and distributions
  6. Key Findings: notable patterns, potential issues, anomalies
  7. Recommendations: preprocessing steps, appropriate analyses, tools to use

Save as {original_filename}_eda_report.md.

Bundled Resources

  • references/file_format_reference.md — Quick-reference catalog of the most common scientific file formats across all 6 categories (bioinformatics, chemistry, microscopy, spectroscopy, proteomics/metabolomics, general), with extension, description, Python library, and key EDA approach for each format

Not migrated from original: The 6 category-specific format catalog files (3,616 lines total) contained detailed entries for 200+ formats. The bundled reference consolidates the ~50 most commonly encountered formats. For rare or vendor-specific formats, consult official library documentation.

Further Reading

  • McKinney, W. (2017). Python for Data Analysis (2nd ed.) — pandas-based EDA methodology
  • VanderPlas, J. (2016). Python Data Science Handbook — EDA with matplotlib and scikit-learn
  • Tukey, J. (1977). Exploratory Data Analysis — foundational EDA philosophy
  • Bio-Formats documentation: https://www.openmicroscopy.org/bio-formats/ — microscopy format reference
  • PSI standard formats: https://www.psidev.info/ — proteomics/MS data standards
  • matplotlib-scientific-plotting — visualization for EDA reports
  • polars-dataframes — efficient tabular data loading and analysis
  • scanpy-scrna-seq — single-cell omics EDA workflows
  • zarr-python — chunked array data access for large datasets
  • pysam-genomic-files — indexed access to BAM/CRAM/VCF files

© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/scientific-computing/exploratory-data-analysis of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/file_format_reference.md

Open the folder on GitHubat commit 82c862c

Compare with similar skills

Exploratory Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Exploratory Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Exploratory Data Analysis this skilljaechang-hits/SciAgent-Skills370—~3.1kAutomated safety check: PassCC-BY-4.0
Exploratory Data Analysisspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2063 repos~3.4kAutomated safety check: NotesMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT

Similar skills

  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    206 GitHub starsUsed in 3 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Yichen Wecom Local Vault

    mcncarl/yichen-skills

    Read, decrypt, query, search, and export local WeCom/企业微信 5.x desktop databases on macOS into a private read-only vault.

    4.3k GitHub stars~1.3k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 163 skills in this repo
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    370 GitHub stars~3.2k tokensUpdated 9 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    370 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    370 GitHub stars~6.9k tokensUpdated 9 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub stars~2.3k tokensUpdated 9 days ago
    Auto-check passed
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 2 repos~5.8k tokens
    Auto-check passed

Questions about Exploratory Data Analysis

What does Exploratory Data Analysis do?

Methodology for exploratory data analysis on scientific files. Exploratory Data Analysis is an agent skill from jaechang-hits/SciAgent-Skills. Methodology for exploratory data analysis on scientific files.

When should I use Exploratory Data Analysis?

Exploratory Data Analysis fits situations like: given a data file for initial exploration; pick an analysis before a pipeline.

How do I install Exploratory Data Analysis in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill exploratory-data-analysis -a claude-code`. Or copy the skill folder (skills/scientific-computing/exploratory-data-analysis in jaechang-hits/SciAgent-Skills) into .claude/skills/exploratory-data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Exploratory Data Analysis in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill exploratory-data-analysis -a codex`. Or copy the skill folder (skills/scientific-computing/exploratory-data-analysis in jaechang-hits/SciAgent-Skills) into .agents/skills/exploratory-data-analysis in your project. Codex loads it when a task matches its description.

Can I use Exploratory Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill exploratory-data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/exploratory-data-analysis, .gemini/skills/exploratory-data-analysis, .github/skills/exploratory-data-analysis and .opencode/skills/exploratory-data-analysis in your project.

What does Exploratory Data Analysis need to run?

Going by SKILL.md and its folder, Exploratory Data Analysis needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Exploratory Data Analysis access the network?

SKILL.md names 2 domains. As links in the text: openmicroscopy.org and psidev.info. This is read from the text; nothing was executed.

Is Exploratory Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Exploratory Data Analysis use?

Exploratory Data Analysis is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Exploratory Data Analysis use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Exploratory Data Analysis?

Skills that share tags, products or a category with Exploratory Data Analysis: Exploratory Data Analysis (spacering-net/codeg, 3.8k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 206 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Exploratory Data Analysis?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 163 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.