Agent skill

Sample Group Sankey Plot

by aipoch in aipoch/medical-research-skills

A skill your agent uses when generating Sankey or alluvial plots from sample annotation tables where rows are samples and selected columns are categorical stages such as risk group, response status…

MITAuto-check passedData & Analytics

Install Sample Group Sankey Plot

skills CLI
$ npx skills add aipoch/medical-research-skills --skill sample-group-sankey-plot -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills sample-group-sankey-plot --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'awesome-med-research-skills/Data Analysis/sample-group-sankey-plot' .claude/skills/sample-group-sankey-plot && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sample-group-sankey-plot
GitHub stars
2k
Token cost
~2.4k tokens
SKILL.md length
952 words
Files
12 (incl. scripts, references)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when generating Sankey or alluvial plots from sample annotation tables where rows are samples and selected columns are categorical stages such as risk group, response status…

  • Works in 4 steps: Validate Input → Prepare Sankey Data → Generate Visualization → …
  • Generating Sankey
  • SKILL.md covers Input Validation, Agent Response Contract, When to Read External Files and Usage, plus 8 more sections
  • Runs R scripts from its folder

What it does

Sample Group Sankey Plot is an agent skill from aipoch/medical-research-skills. Use when generating Sankey or alluvial plots from sample annotation tables where rows are samples and selected columns are categorical stages such as risk group, response status, subtype, or cohort labels. NOT for: gene network flow analysis, continuous-value trajectories, or graph-structured pathway visualization.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts and reference files (for example `eval_report_sample-group-sankey-plot_result.json`, `references/algorithm.md` and `references/cli-guide.md`).

It sits in Data & Analytics. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Generating Sankey
  • Alluvial plots from sample annotation tables where rows are samples and selected columns are categorical stages such as risk group
  • Response status

Example prompts

  • “/sample-group-sankey-plot”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Validate Input
  2. Prepare Sankey Data
  3. Generate Visualization
  4. Record Outputs

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sample Group Sankey Plot loads about 2.4k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 85 tokens; SKILL.md has 952 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 952 words, ~2,422 tokens.

Download SKILL.mdSave it as .claude/skills/sample-group-sankey-plot/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
sample-group-sankey-plot
description
Use when generating Sankey or alluvial plots from sample annotation tables where rows are samples and selected columns are categorical stages such as risk group, response status, subtype, or cohort labels. NOT for: gene network flow analysis, continuous-value trajectories, or graph-structured pathway visualization.
license
MIT
skill-author
AIPOCH

Sample Group Sankey Plot

Builds a reproducible Sankey/alluvial visualization from a tabular sample annotation file and exports the selected annotations, lodes-format table, plot PDF, and session metadata.

Input Validation

This skill accepts: a sample annotation table in CSV or TSV format where rows are samples and selected columns are categorical stages (e.g., risk group, response status, subtype, cohort label). At least 2 stage columns are required.

If the user's request does not involve generating a Sankey or alluvial flow diagram from categorical sample annotations — for example, asking to visualize a gene regulatory network, plot continuous-value trajectories, analyze pathway flow, or process non-tabular data — do not proceed with the workflow. Instead respond:

"sample-group-sankey-plot is designed to generate Sankey/alluvial plots from categorical sample annotation tables. Your request appears to be outside this scope. Please provide a sample annotation table with at least 2 categorical stage columns, or use a more appropriate tool for gene network visualization or pathway analysis."

Readability guidance: Sankey plots are recommended for fewer than 8 unique values per stage and fewer than 5 stages total. For larger inputs, filter or aggregate categories before plotting to ensure readable output.

Agent Response Contract

After a successful run, report to the caller:

Sankey plot generated successfully.
Stages plotted : <comma-separated stage column names>
Samples        : <row count>
Output prefix  : <output_prefix>
Outputs:
  table/selected_annotations.csv
  table/sankey_lodes.csv
  plot/<output_prefix>.pdf
  data/session_info.txt
Readability warnings (if any): <advisory messages or "none">

If the script exits with a non-zero status, surface the SKILL_* error code and message verbatim. Do not attempt to continue or retry silently.

When to Read External Files

SituationFile to ReadPurpose
Need to run analysisscripts/main.RExecute: Rscript scripts/main.R --input_file ... --output_dir ...
Need algorithm detailsreferences/algorithm.mdAlluvial transformation logic, assumptions, and plotting choices
Encounter errorsreferences/troubleshooting.mdCommon errors and solutions
Need CLI examplesreferences/cli-guide.mdDetailed CLI examples
Need test datatests/data/Sample annotation tables for smoke tests and regression checks

→ Reference files algorithm.md, troubleshooting.md, and cli-guide.md are in references/. If absent, rely on the Error Handling table below for common issues.


Usage

Environment Setup

Install required R packages before the first run:

bash
Rscript scripts/install_dependencies.R

Note: install_dependencies.R installs from CRAN without version pinning. Tested with ggalluvial >= 0.12.5 and ggplot2 >= 3.4.0. For reproducible CI environments, consider using remotes::install_version().

Basic Command
bash
Rscript scripts/main.R \
  --input_file ./annotations.csv \
  --output_dir ./output \
  --columns risk,Responder \
  --seed 42

Arguments

ShortLongTypeDefaultDescription
-i--input_filecharacterrequiredInput CSV/TSV annotation table
-o--output_dircharacter./output/Output directory
-c--columnscharacterall columnsComma-separated stage columns to include in the plot. When omitted, all columns in the file are used as stages.
-p--output_prefixcharactersankey_plotPrefix for generated output files (alphanumeric, dot, underscore, or hyphen characters only)
--widthnumeric7Plot width in inches
--heightnumeric5Plot height in inches
--alphanumeric0.5Flow transparency between 0 and 1
--label_sizenumeric4.5Stratum label size
--missing_labelcharacterMissingReplacement label for blank or NA strata
--titlecharacteremptyOptional plot title
-s--seedinteger42Random seed recorded for reproducibility
--timeoutinteger3600Maximum allowed elapsed runtime in seconds; use 0 to disable

Input Format

Annotation Table (input_file)

Delimited text file where rows represent samples and each selected column is a categorical stage shown in the Sankey plot.

csv
SampleID,risk,Responder,Subtype
S1,High,Yes,Basal
S2,Low,No,LumA
S3,High,Yes,Basal
S4,Low,No,LumB

Requirements:

  • The file must contain at least 2 columns if --columns is omitted.
  • Each selected column must exist in the file header.
  • Selected columns are interpreted as categorical stages and will be converted to character values.
  • Blank strings and NA values are replaced with --missing_label.
  • CSV and TSV inputs are supported.
  • Recommended: fewer than 8 unique values per stage and fewer than 5 stages total for readable plots.

Show full SKILL.md (393 more words)Show less

Output Files

FileDescription
table/selected_annotations.csvFiltered table containing only the plotted stage columns
table/sankey_lodes.csvLong-format lodes table used to build the Sankey plot
plot/{output_prefix}.pdfSankey/alluvial plot as a PDF file; default filename is sankey_plot.pdf
data/session_info.txtR session information and runtime parameters
selected_annotations.csv
ColumnTypeDescription
stage columnscharacterOne column per plotted stage in original order
sankey_lodes.csv
ColumnTypeDescription
sample_idcharacterSynthetic row identifier used as the alluvium key
xcharacterStage name
stratumcharacterCategory label for the stage

Workflow

Step 1: Validate Input
  • Check that the input file exists and is readable.
  • Detect CSV vs TSV input.
  • Validate that at least 2 stage columns are available.
  • Validate that user-specified columns exist.
Step 2: Prepare Sankey Data
  • Subset the selected stage columns.
  • Replace missing or blank labels.
  • Add a row-level sample_id identifier.
  • Convert the table with ggalluvial::to_lodes_form().
Step 2a: Readability Advisories

After column selection, the script emits log_warn advisories when:

  • More than 5 stages are selected: "More than 5 stages selected; plot may be hard to read. Consider filtering."
  • A stage has more than 8 unique values: "Stage <name> has <n> unique values; consider aggregating for readability."

These are advisory only — the script continues and produces the plot.

Step 3: Generate Visualization
  • Build the Sankey/alluvial plot with geom_flow() and geom_stratum().
  • Render stratum labels.
  • Save the plot as PDF.
Step 4: Record Outputs
  • Save the selected annotations and lodes-format table as CSV files.
  • Save sessionInfo() and runtime arguments for reproducibility.

Examples

Reproduce the Original Two-Column Plot
bash
Rscript scripts/main.R \
  -i tests/data/sample_annotations.csv \
  -o tests/output \
  -c risk,Responder
Plot Three Annotation Stages
bash
Rscript scripts/main.R \
  -i tests/data/sample_annotations.csv \
  -o tests/output_three_stage \
  -c risk,Responder,Subtype \
  --title "Risk to response transitions"
Use All Columns Automatically
bash
Rscript scripts/main.R \
  -i tests/data/minimal_annotations.csv \
  -o tests/output_all_columns

Error Handling

Common Errors
ErrorCauseSolution
SKILL_FILE_NOT_FOUNDInput file does not existCheck --input_file
SKILL_EMPTY_DATAThe input file has zero rows or fewer than 2 usable columnsProvide a non-empty table with at least 2 stage columns
SKILL_MISSING_COLUMNSA requested stage column is absentCorrect --columns or fix the input header
SKILL_INVALID_PARAMETERWidth, height, alpha, label size, or output prefix is invalidProvide valid arguments per the Arguments table
SKILL_DEPENDENCY_MISSINGA required R package is unavailableRun Rscript scripts/install_dependencies.R
SKILL_IO_ERROROutput directory cannot be created or writtenCheck permissions on --output_dir

IF error persists, READ: references/troubleshooting.md


Testing

Test with Sample Data
bash
Rscript scripts/install_dependencies.R

Rscript scripts/main.R --help

Rscript scripts/main.R \
  -i tests/data/sample_annotations.csv \
  -o tests/output \
  -c risk,Responder,Subtype

Rscript tests/test_skill.R

Rscript tests/run_smoke_test.R
Validation Commands
bash
ls -la tests/output/table
ls -la tests/output/plot
ls -la tests/output/data

References

  1. Brunson JC (2020) ggalluvial: Layered Grammar for Alluvial Plots. Journal of Open Source Software. doi:10.21105/joss.02017
  2. Wickham H (2016) ggplot2: Elegant Graphics for Data Analysis. Springer. doi:10.1007/978-3-319-24277-4

For detailed algorithm, READ: references/algorithm.md

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references) in awesome-med-research-skills/Data Analysis/sample-group-sankey-plot of aipoch/medical-research-skills.

  • SKILL.md
  • eval_report_sample-group-sankey-plot_result.json
  • references/algorithm.md
  • references/cli-guide.md
  • references/troubleshooting.md
  • scripts/functions.R
  • scripts/install_dependencies.R
  • scripts/main.R
  • scripts/run_analysis.R
  • scripts/utils.R
  • tests/data/sample_annotations.csv
  • tests/run_smoke_test.R

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Sample Group Sankey Plot next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sample Group Sankey Plot compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sample Group Sankey Plot this skillaipoch/medical-research-skills2k—~2.4kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Exploratory Data AnalysisOleafly/Oleafly2063 repos~3.4kAutomated safety check: NotesMIT
CSV Data Summarizercoffeefuelbump/csv-data-summarizer-claude-skill4682 repos~1.4kAutomated safety check: PassNone
Ieee Figure TableCloudWave818/ieee-skills355—~1kAutomated safety check: PassMIT
Raccoon DataanalysisSenseTime-Copilot/raccoon-dataanalysis-skill137—~1.9kAutomated safety check: PassNone

Similar skills

  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Perform bounded, local exploratory analysis of explicitly supported scientific files.

    206 GitHub starsUsed in 3 repos~3.4k tokens
    Data & AnalyticsAuto-check: notes
  • CSV Data Summarizer

    coffeefuelbump/csv-data-summarizer-claude-skill

    Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.

    468 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Ieee Figure Table

    CloudWave818/ieee-skills

    Audit, redesign, generate, and improve IEEE manuscript figures, tables, captions, result presentation, plotting scripts, visual polish, hybrid Python/R plus vector-editor workflows…

    355 GitHub stars~1k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Raccoon Dataanalysis

    SenseTime-Copilot/raccoon-dataanalysis-skill

    Raccoon (小浣熊) Data Analysis - Remote code interpreter and data visualization service powered by SenseTime.

    137 GitHub stars~1.9k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • Football Match Report

    ricardoherediaj/football-analytics-tutorials

    Build post-match team + player reports (24-chart dashboard, per-player dashboards, stats CSV) from a WhoScored URL.

    119 GitHub stars~2.7k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 21 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 21 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 21 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 21 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed

Questions about Sample Group Sankey Plot

What does Sample Group Sankey Plot do?

A skill your agent uses when generating Sankey or alluvial plots from sample annotation tables where rows are samples and selected columns are categorical stages such as risk group, response status…. Sample Group Sankey Plot is an agent skill from aipoch/medical-research-skills. Use when generating Sankey or alluvial plots from sample annotation tables where rows are samples and selected columns are categorical stages such as risk group, response status, subtype, or cohort labels.

When should I use Sample Group Sankey Plot?

Sample Group Sankey Plot fits situations like: generating Sankey; alluvial plots from sample annotation tables where rows are samples and selected columns are categorical stages such as risk group; response status.

How do I install Sample Group Sankey Plot in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill sample-group-sankey-plot -a claude-code`. Or copy the skill folder (awesome-med-research-skills/Data Analysis/sample-group-sankey-plot in aipoch/medical-research-skills) into .claude/skills/sample-group-sankey-plot in your project. Claude Code loads it when a task matches its description.

How do I install Sample Group Sankey Plot in Codex?

Run `npx skills add aipoch/medical-research-skills --skill sample-group-sankey-plot -a codex`. Or copy the skill folder (awesome-med-research-skills/Data Analysis/sample-group-sankey-plot in aipoch/medical-research-skills) into .agents/skills/sample-group-sankey-plot in your project. Codex loads it when a task matches its description.

Can I use Sample Group Sankey Plot in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill sample-group-sankey-plot -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sample-group-sankey-plot, .gemini/skills/sample-group-sankey-plot, .github/skills/sample-group-sankey-plot and .opencode/skills/sample-group-sankey-plot in your project.

What does Sample Group Sankey Plot need to run?

Going by SKILL.md and its folder, Sample Group Sankey Plot needs R for the scripts in its folder.

Does Sample Group Sankey Plot access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sample Group Sankey Plot safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Sample Group Sankey Plot use?

Sample Group Sankey Plot is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sample Group Sankey Plot use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Sample Group Sankey Plot?

Skills that share tags, products or a category with Sample Group Sankey Plot: Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Exploratory Data Analysis (Oleafly/Oleafly, 206 stars), CSV Data Summarizer (coffeefuelbump/csv-data-summarizer-claude-skill, 468 stars) and Ieee Figure Table (CloudWave818/ieee-skills, 355 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sample Group Sankey Plot?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,974 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.