Agent skill

Bio Reporting Jupyter Reports

by GPTomics in GPTomics/bioSkills

Runs parameterized Jupyter notebooks as reproducible batch report generators with papermill, renders them to HTML/PDF with nbconvert, aggregates results across samples, and makes notebook outputs…

MITAuto-check passedData & Analytics

Install Bio Reporting Jupyter Reports

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-reporting-jupyter-reports -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-reporting-jupyter-reports --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/reporting/jupyter-reports .claude/skills/bio-reporting-jupyter-reports && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-reporting-jupyter-reports
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
1,447 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Runs parameterized Jupyter notebooks as reproducible batch report generators with papermill, renders them to HTML/PDF with nbconvert, aggregates results across samples, and makes notebook outputs…

  • Generating per-sample analysis reports
  • SKILL.md covers Version Compatibility, The Load-Bearing Idea: A…, Two Artifacts People Conflate and Parameterizing with papermill, plus 9 more sections
  • Runs Python scripts from its folder; calls jupyter, pip and git
  • Executing a notebook template across many datasets

What it does

Bio Reporting Jupyter Reports is an agent skill from GPTomics/bioSkills. Runs parameterized Jupyter notebooks as reproducible batch report generators with papermill, renders them to HTML/PDF with nbconvert, aggregates results across samples, and makes notebook outputs trustworthy. Use when generating per-sample analysis reports, executing a notebook template across many datasets, or fixing notebooks that do not reproduce.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/papermill_batch.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Jupyter notebooks. It works with Jupyter. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Generating per-sample analysis reports
  • Executing a notebook template across many datasets
  • Fixing notebooks that do not reproduce

Example prompts

  • “/bio-reporting-jupyter-reports”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • jupyter
    • pip
    • git
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Reporting Jupyter Reports loads about 2.9k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 1,447 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,447 words, ~2,941 tokens.

Download SKILL.mdSave it as .claude/skills/bio-reporting-jupyter-reports/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-reporting-jupyter-reports
description
Runs parameterized Jupyter notebooks as reproducible batch report generators with papermill, renders them to HTML/PDF with nbconvert, aggregates results across samples, and makes notebook outputs trustworthy. Use when generating per-sample analysis reports, executing a notebook template across many datasets, or fixing notebooks that do not reproduce.
tool_type
python
primary_tool
papermill
goal_approach_exempt
true

Version Compatibility

Reference examples tested with: papermill 2.6+, nbconvert 7.16+, nbclient 0.10+, jupyter-client 8+, scrapbook 0.5+, jupytext 1.16+, nbstripout 0.7+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show papermill then help(papermill.execute_notebook)
  • CLI: papermill --help, jupyter nbconvert --help

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Jupyter Reports with Papermill

"Generate a reproducible analysis report" -> Execute a parameterized notebook in a clean kernel, producing an executed-notebook artifact, then render it to a code-hidden HTML/PDF report.

  • Python: papermill.execute_notebook(input, output, parameters={...})
  • CLI: jupyter nbconvert --execute --to html notebook.ipynb

The Load-Bearing Idea: A Notebook Has Hidden State; the Report Is the Executed Artifact

A .ipynb is JSON holding cell source, saved outputs, and per-cell execution_count integers - and those three are decoupled. The kernel is a long-lived process with a mutable namespace, so a user can run cell 5, edit cell 2, rerun it, delete cell 3, run cell 7, and save. The stored outputs then reflect a kernel state that NO top-to-bottom rerun reproduces. Non-monotonic execution counts are the forensic tell.

This is intrinsic to the REPL-on-a-document model, not a bug. The discipline answer is mechanical: Restart Kernel and Run All before sharing, so saved outputs equal one clean linear run. papermill and nbconvert --execute are that discipline automated - they always spin up a FRESH kernel and run cells strictly in document order, so the output is by construction the record of one clean run. The value is not "trust the saved outputs," it is "regenerate them from a known-empty state."

The empirical case (Pimentel et al.): of ~1.4M GitHub notebooks, only ~24% of those executed finished without error and only ~4% reproduced their stored outputs; ~36% had out-of-order cells, and the dominant failure was ImportError - environment, not logic. (It is a public-GitHub corpus, so the absolute rates carry selection bias, but the failure mode is the lesson.) Two lessons: saved outputs are not reproducible (always re-execute), and the #1 fix is pinning the environment (see below).

Two Artifacts People Conflate

  • Executed notebook - papermill's output, or nbconvert --execute --to notebook. A record of the run: same cells, freshly computed outputs, an injected-parameters cell. For audit/debug/aggregation, not for humans to read as prose.
  • Human report - nbconvert --to html/pdf, optionally --no-input to hide code. The rendered deliverable.

A papermill pipeline does both: execute the parameterized notebook -> output .ipynb (the evidence) -> convert to HTML/PDF (the report).

Parameterizing with papermill

Tag ONE cell parameters holding defaults. At execution papermill inserts a NEW cell tagged injected-parameters immediately AFTER it, containing only the overrides; because Python runs top-to-bottom, the injected cell shadows the defaults.

python
import papermill as pm
pm.execute_notebook('template.ipynb', 'out/sampleA.ipynb',
                    parameters={'sample_id': 'sampleA', 'fdr_threshold': 0.05})
  • If NO cell is tagged parameters, the injected cell goes at the TOP and downstream references to a param raise NameError. Keep all parameters in one tagged cell.
  • Type coercion is the #1 CLI gotcha. -p name value YAML-parses the value (-p n 5 -> int, -p flag true -> bool); -r name value keeps it a raw string (-r chrom 1 -> "1", needed for sample IDs like 007 or chromosome "1"). -y/--parameters_yaml and -f/--parameters_file pass lists/dicts. The Python API passes native objects directly with no YAML round-trip - prefer it in pipelines to avoid coercion surprises.

Execution Controls That Matter in Pipelines

  • execution_timeout (CLI --execution-timeout) - seconds per cell; default is forever (None). A long bioinformatics cell hangs a pipeline silently without this. Set it.
  • kernel_name / --kernel - must be a REGISTERED kernelspec; a mismatch is a top failure cause. The kernel must point at the pinned environment.
  • log_output=True / --log-output - stream each cell's stdout/stderr for CI visibility.
  • On a cell exception papermill raises PapermillExecutionError, STILL writes the output notebook with the traceback captured (evidence preserved), then exits non-zero. Tag a cell raises-exception to allow it to fail without aborting. Fail loud, keep the artifact.
  • papermill reads/writes notebooks from s3://, gs://, adl:///abs://, and http(s):// out of the box (cloud connectors are extras: pip install papermill[s3]), so serverless per-sample execution works.
  • The fresh-kernel guarantee holds when papermill owns the kernel lifecycle (the default execute_notebook path). Advanced patterns that reuse a kernel across notebooks forgo the from-empty-state guarantee - let papermill create the kernel.

Rendering to a Report

nbconvert --execute runs the notebook top-to-bottom in a fresh kernel and writes regenerated outputs - that re-execution, not the saved outputs, is what makes a converted report reproducible.

bash
jupyter nbconvert --to html --no-input out/sampleA.ipynb     # code-hidden stakeholder report
jupyter nbconvert --execute --to html template.ipynb         # execute then render in one step
jupyter nbconvert --to webpdf out/sampleA.ipynb              # PDF without LaTeX
  • --to pdf needs a LaTeX toolchain (xelatex + pandoc) - the classic CI pain. --to webpdf renders HTML then prints via headless Chromium (pip install nbconvert[webpdf], --allow-chromium-download) and avoids TeX entirely. The tradeoff: webpdf uses Chromium page geometry, so wide tables and long code lines can clip at the page edge, whereas LaTeX --to pdf paginates and wraps better for table-heavy reports.
  • --no-input hides all code; --no-prompt drops the In[ ]:/Out[ ]: prompts; TagRemovePreprocessor.remove_cell_tags strips cells by tag (tag setup cells to remove them). Custom branded layouts use the nbconvert 6+ directory-template system (--template <name>).

Aggregating Results Across Samples

papermill.record is deprecated (since papermill 1.0). Use scrapbook: in the template, import scrapbook as sb; sb.glue('auc', 0.91) records named scraps (and sb.glue('fig', obj, display=True) for figures). Downstream, sb.read_notebooks('reports/').papermill_dataframe aggregates every executed notebook's scraps into one tidy table - the canonical pattern for looping papermill over a sample sheet then collecting per-sample QC into a cohort summary.

Show full SKILL.md (574 more words)Show less

Version Control: Never Commit Outputs

.ipynb is JSON with embedded base64 outputs and execution_count, so committing it raw gives giant unreviewable diffs, brutal merge conflicts, and leaked data. Three complementary fixes:

  • nbstripout - a git clean filter (*.ipynb filter=nbstripout in .gitattributes) that strips outputs and execution counts on git add; the working copy keeps its outputs. Highest-leverage single fix.
  • jupytext - pairs .ipynb with a text twin (py:percent is runnable and diff-friendly, or Markdown/.qmd); commit the text file as source of truth, regenerate outputs. Lets a notebook be code-reviewed as a normal PR.
  • jupyter-cache - caches executed outputs keyed by a code-cell hash; the engine Jupyter Book/MyST-NB and Quarto use to skip re-running expensive cells. Right tool when notebooks are doc sources with long compute.

Reproducible Execution Is Not Reproducible Science

State this plainly: papermill and nbconvert --execute guarantee EXECUTION ORDER and a clean kernel - they do NOT guarantee the ENVIRONMENT or numerical determinism. The same parameterized notebook rerun against newer numpy/scanpy, a different BLAS, or an unseeded RNG runs flawlessly and produces DIFFERENT numbers. Order-reproducibility is necessary, not sufficient. Pair papermill with: a pinned environment (conda lockfile / requirements.txt with exact versions, ideally a container; the registered kernel must point at it), seeded RNGs for any stochastic step (clustering, UMAP, splits, bootstraps), and pinned reference/DB versions. To catch silent drift, nbval (pytest --nbval) re-executes and compares new outputs against stored ones in CI.

When a Notebook Report Is the Wrong Tool

The notebook is the REPORT, not the PIPELINE. Heavy compute (alignment, variant calling, large scanpy integration, anything multi-hour or needing a scheduler, retries, or parallel fan-out) belongs in a workflow manager (Snakemake/Nextflow/WDL). papermill notebooks shine as the final per-sample or per-cohort summary plus figures over already-computed results. Rule of thumb: if --retry, --cluster, or a DAG is wanted, it is a pipeline; if a parameterized HTML/PDF of results is wanted, it is a papermill report. Note Quarto can consume a .ipynb directly, so papermill (parameterize/execute) and Quarto (render) interoperate.

Common Errors

SymptomCauseFix
Saved outputs do not match a rerunOut-of-order interactive execution / hidden stateRestart-and-run-all; let papermill/nbconvert --execute regenerate
Parameter is a bool when a string was wanted-p YAML-parses valuesuse -r name value for raw strings, or the Python API
NameError on a parameterNo parameters-tagged celltag one cell parameters; keep all params there
Pipeline hangs on one cellexecution_timeout defaults to foreverset execution_timeout
--to pdf fails in CIno LaTeX toolchainuse --to webpdf (headless Chromium)
Giant notebook diffs / leaked data in gitcommitting outputsnbstripout filter or jupytext pairing
Reruns months later give different numbersenvironment/seed not pinnedpin env + container, seed RNGs, nbval in CI
  • reporting/quarto-reports - Document-first reporting that can render the same notebooks
  • reporting/rmarkdown-reports - R-based literate reports
  • workflows/scrnaseq-pipeline - Heavy compute that belongs in a pipeline, with the notebook as the summary report
  • single-cell/preprocessing - Analysis embedded inside per-sample notebook templates

References

  • Pimentel JF, Murta L, Braganholo V, Freire J. A large-scale study about quality and reproducibility of Jupyter notebooks. MSR '19, IEEE/ACM. 2019:507-517. doi:10.1109/MSR.2019.00077
  • Pimentel JF, Murta L, Braganholo V, Freire J. Understanding and improving the quality and reproducibility of Jupyter notebooks. Empir Softw Eng. 2021;26:65. doi:10.1007/s10664-021-09961-9
  • Rule A, Birmingham A, Zuniga C, et al. Ten simple rules for writing and sharing computational analyses in Jupyter Notebooks. PLoS Comput Biol. 2019;15(7):e1007007. doi:10.1371/journal.pcbi.1007007
  • Kluyver T, Ragan-Kelley B, Pérez F, et al. Jupyter Notebooks - a publishing format for reproducible computational workflows. ELPUB 2016, IOS Press. 2016:87-90. doi:10.3233/978-1-61499-649-1-87

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in reporting/jupyter-reports of GPTomics/bioSkills.

  • SKILL.md
  • examples/papermill_batch.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Reporting Jupyter Reports next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Reporting Jupyter Reports compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Reporting Jupyter Reports this skillGPTomics/bioSkills1.2k1 repos~2.9kAutomated safety check: PassMIT
Jupyter Notebookmicrosoft/ai-agents-for-beginners77k8 repos~1kAutomated safety check: PassApache-2.0
Notebook For ExperimentJetBrains/intellij-community21k—~4.4kAutomated safety check: WarnCustom licence
Jupyter To Marimoericmjl/llamabot1832 repos~475Automated safety check: PassNone
Nbreviewawdeorio/dotfiles102—~2.2kAutomated safety check: PassNone
Save Research Notebooknapjon/krisk117—~702Automated safety check: PassBSD-3-Clause

Similar skills

  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    A skill your agent uses when the user asks to create, scaffold, or edit Jupyter notebooks (.ipynb) for experiments, explorations, or tutorials; prefer the bundled templates and run the helper script…

    77k GitHub starsUsed in 8 repos~1k tokens
    Data & AnalyticsAuto-check passed
  • Notebook For Experiment

    JetBrains/intellij-community

    Official

    Create reproducible Jupyter notebooks for performance experiments.

    21k GitHub stars~4.4k tokensUpdated today
    Data & AnalyticsAuto-check: warnings
  • Jupyter To Marimo

    ericmjl/llamabot

    Convert a Jupyter notebook (.ipynb) to a marimo notebook (.py).

    183 GitHub starsUsed in 2 repos~475 tokens
    Data & AnalyticsAuto-check passed
  • Nbreview

    awdeorio/dotfiles

    This skill should be used when the user asks to "review this notebook", "nbreview", "check my notebook", or otherwise requests a Jupyter notebook review.

    102 GitHub stars~2.2k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed
  • Convert a completed data-analysis conversation into evidence-backed, reproducible living research through the Krisk MCP server.

    117 GitHub stars~702 tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    Kasuta, kui kasutaja palub luua, üles ehitada või redigeerida Jupyteri märkmikke (.ipynb) katsetuste, uurimiste või juhendite jaoks; eelista kaasasolevaid malle ja käivita abiskript newnotebook.py…

    77k GitHub stars~1.2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Reporting Jupyter Reports

What does Bio Reporting Jupyter Reports do?

Runs parameterized Jupyter notebooks as reproducible batch report generators with papermill, renders them to HTML/PDF with nbconvert, aggregates results across samples, and makes notebook outputs…. Bio Reporting Jupyter Reports is an agent skill from GPTomics/bioSkills. Runs parameterized Jupyter notebooks as reproducible batch report generators with papermill, renders them to HTML/PDF with nbconvert, aggregates results across samples, and makes notebook outputs trustworthy.

When should I use Bio Reporting Jupyter Reports?

Bio Reporting Jupyter Reports fits situations like: generating per-sample analysis reports; executing a notebook template across many datasets; fixing notebooks that do not reproduce.

How do I install Bio Reporting Jupyter Reports in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-reporting-jupyter-reports -a claude-code`. Or copy the skill folder (reporting/jupyter-reports in GPTomics/bioSkills) into .claude/skills/bio-reporting-jupyter-reports in your project. Claude Code loads it when a task matches its description.

How do I install Bio Reporting Jupyter Reports in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-reporting-jupyter-reports -a codex`. Or copy the skill folder (reporting/jupyter-reports in GPTomics/bioSkills) into .agents/skills/bio-reporting-jupyter-reports in your project. Codex loads it when a task matches its description.

Can I use Bio Reporting Jupyter Reports in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-reporting-jupyter-reports -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-reporting-jupyter-reports, .gemini/skills/bio-reporting-jupyter-reports, .github/skills/bio-reporting-jupyter-reports and .opencode/skills/bio-reporting-jupyter-reports in your project.

What does Bio Reporting Jupyter Reports need to run?

Going by SKILL.md and its folder, Bio Reporting Jupyter Reports needs Python for the scripts in its folder and the command-line tools its instructions call (jupyter, pip, git and pytest). Our summary lists: Python 3.

Does Bio Reporting Jupyter Reports access the network?

SKILL.md contains no URLs. Its commands use pip and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Reporting Jupyter Reports safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Reporting Jupyter Reports use?

Bio Reporting Jupyter Reports is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Reporting Jupyter Reports use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Reporting Jupyter Reports?

Skills that share tags, products or a category with Bio Reporting Jupyter Reports: Jupyter Notebook (microsoft/ai-agents-for-beginners, 77k stars), Notebook For Experiment (JetBrains/intellij-community, 21k stars), Jupyter To Marimo (ericmjl/llamabot, 183 stars) and Nbreview (awdeorio/dotfiles, 102 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Reporting Jupyter Reports?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.