Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Clusters temporally variable genes by expression-profile SHAPE (not significance) using Mfuzz fuzzy c-means, TCseq, DEGreport degPatterns, and tslearn DTW/soft-DTW.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-temporal-clustering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/temporal-genomics/temporal-clustering .claude/skills/bio-temporal-genomics-temporal-clustering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-temporal-genomics-temporal-clustering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/temporal-clustering into .claude/skills/bio-temporal-genomics-temporal-clustering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-temporal-clustering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/temporal-clusteringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-temporal-clustering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/temporal-genomics/temporal-clustering .agents/skills/bio-temporal-genomics-temporal-clustering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-temporal-genomics-temporal-clustering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/temporal-clustering into .agents/skills/bio-temporal-genomics-temporal-clustering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-temporal-clustering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-temporal-clustering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/temporal-genomics/temporal-clustering .cursor/skills/bio-temporal-genomics-temporal-clustering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-temporal-genomics-temporal-clustering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/temporal-clustering into .cursor/skills/bio-temporal-genomics-temporal-clustering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-temporal-clustering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path temporal-genomics/temporal-clustering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-temporal-clustering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/temporal-genomics/temporal-clustering .gemini/skills/bio-temporal-genomics-temporal-clustering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-temporal-genomics-temporal-clustering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/temporal-clustering into .gemini/skills/bio-temporal-genomics-temporal-clustering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-temporal-clustering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-temporal-clusteringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/temporal-genomics/temporal-clustering .github/skills/bio-temporal-genomics-temporal-clustering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-temporal-genomics-temporal-clustering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/temporal-clustering into .github/skills/bio-temporal-genomics-temporal-clustering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-temporal-clustering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-temporal-clustering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/temporal-genomics/temporal-clustering .opencode/skills/bio-temporal-genomics-temporal-clustering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-temporal-genomics-temporal-clustering" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/temporal-clustering into .opencode/skills/bio-temporal-genomics-temporal-clustering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-temporal-clustering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-temporal-genomics-temporal-clusteringClusters temporally variable genes by expression-profile SHAPE (not significance) using Mfuzz fuzzy c-means, TCseq, DEGreport degPatterns, and tslearn DTW/soft-DTW.
Bio Temporal Genomics Temporal Clustering is an agent skill from GPTomics/bioSkills. Clusters temporally variable genes by expression-profile SHAPE (not significance) using Mfuzz fuzzy c-means, TCseq, DEGreport degPatterns, and tslearn DTW/soft-DTW. Use when grouping pre-selected time-course genes into shared trajectory programs (co-expression modules), choosing between soft vs hard clustering, picking k, selecting a distance metric (Euclidean/correlation/DTW), or interpreting clusters with per-cluster enrichment. Requires temporally variable genes selected FIRST…
Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/tslearn_clustering.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (R and Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Temporal Genomics Temporal Clustering loads about 5.2k tokens when it runs. Until then it costs about 171 tokens; SKILL.md has 1,891 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,891 words, ~5,243 tokens.
.claude/skills/bio-temporal-genomics-temporal-clustering/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: Mfuzz 2.64+, TCseq 1.14+, DEGreport 1.30+ (R/Bioconductor); tslearn 0.8+, scikit-learn 1.4+ (Python).
Before using code patterns, verify installed versions match. If versions differ:
pip show tslearn scikit-learn then help(module.function) to check signaturespackageVersion('Mfuzz') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Group my time-course genes by expression pattern shape" -> Partition PRE-SELECTED temporally variable genes into co-expression modules by trajectory shape (fuzzy c-means, hierarchical, or DTW), producing candidate temporal programs.
Mfuzz::mfuzz() (fuzzy/soft), TCseq::timeclust(), DEGreport::degPatterns()tslearn.clustering.TimeSeriesKMeans (Euclidean / DTW / soft-DTW)Clustering answers only "which genes share a temporal SHAPE." It is DESCRIPTIVE and UNSUPERVISED: it has no null model, no p-value, and no notion of a "true" cluster count, so it ALWAYS returns clusters from whatever it is handed. It is strictly DOWNSTREAM of gene selection.
If a user asks "cluster my RNA-seq time course," the first question is always: have these genes already been selected for temporal change, and how? If the answer is "no, it is all 20,000 genes," stop and prefilter.
Soft (fuzzy) clustering is preferred for expression. Genes participate in multiple regulatory programs, so forcing one gene into one cluster (hard k-means) is biologically false at boundaries and brittle: a gene between two centroids flips clusters under trivial noise. Futschik & Carlisle (2005) established fuzzy c-means as noise-ROBUST for expression time courses - low-membership (ambiguous, likely-noise) genes are down-weighted in centroid estimation, so centroids track the high-confidence core of each program, and ambiguity is exposed as a continuous membership score to threshold rather than hidden inside a hard label.
Z-score per gene is mandatory (Mfuzz standardise(), TCseq standardize=TRUE, tslearn TimeSeriesScalerMeanVariance()). Without it, MAGNITUDE dominates SHAPE: a high-abundance housekeeping gene sits far (Euclidean) from a low-abundance gene of identical shape, while two high-abundance genes co-cluster on abundance alone. Clustering-by-shape requires removing each gene's mean and scaling to unit variance across timepoints.
Goal: Group temporally variable genes into soft co-expression clusters by trajectory shape.
Approach: Build an ExpressionSet, gate out flat genes (filter.std), z-score (standardise), estimate then VALIDATE the fuzzifier, run fuzzy c-means, and filter genes by membership. Mfuzz wraps e1071::cmeans (it does not implement its own optimizer); distance is Euclidean on z-scored profiles.
library(Mfuzz)
library(Biobase)
# Rows = genes (already selected as temporally variable), columns = timepoints (mean across replicates)
expr_mat <- as.matrix(read.csv('temporal_expression.csv', row.names = 1))
eset <- ExpressionSet(assayData = expr_mat)
# filter.std: flat-gene GATE (keeps the governing principle true). min.std=0.5 is a starting
# point; inspect the SD distribution and set it above the flat-gene noise floor for your data.
eset <- filter.std(eset, min.std = 0.5)
# Per-gene mean 0, sd 1 across timepoints (British spelling; no 'standardize' alias)
eset <- standardise(eset)Goal: Pick a fuzzifier m that keeps clusters informative for THIS number of timepoints.
Approach: mestimate() implements Schwaemmle & Jensen (2010): it returns the smallest m that stops fuzzy c-means from finding tight clusters in RANDOMIZED data. The estimate is dominated by D (number of timepoints) via a D^-2 term, so it can go degenerate at the extremes - inspect the returned m AND the membership distribution rather than trusting either the estimate or the historical m=2 default.
# With FEW timepoints (small D), mestimate pushes m HIGH -> over-fuzzy: memberships flatten
# toward 1/c and an acore(0.5) filter can discard nearly everything.
# With MANY timepoints (large D), m falls toward ~1.05-1.2 -> near-hard, soft advantage evaporates.
m <- mestimate(eset)
cat(sprintf('Estimated fuzzifier m: %.2f\n', m))
cl <- mfuzz(eset, c = 8, m = m) # c=8: starting point for 6-12 timepoints; refine below
# VALIDATE m: what fraction of genes clears the alpha-core cutoff? If very few do, m is too high.
max_mem <- apply(cl$membership, 1, max)
cat(sprintf('Genes with max membership >= 0.5: %.0f%%\n', 100 * mean(max_mem >= 0.5)))
# Sanity check the estimate's own criterion: cluster a permuted copy; it should NOT form tight clusters.# acore returns, per cluster, genes with MAX membership >= min.acore ("alpha cores").
# 0.5 is a convention; it discards a data-dependent fraction (larger m -> more discarded).
# Relaxing to 0.3 is legitimate for exploratory work but admits more noise. Always report the retained fraction.
core_genes <- acore(eset, cl, min.acore = 0.5)
# Minimum centroid distance vs k: as k grows the closest centroid pair collapses; a knee hints at
# over-splitting. This is a WEAK, monotone-ish signal, not an oracle -- triangulate with stability (below).
min_dist <- sapply(4:20, function(k) {
d <- as.matrix(dist(mfuzz(eset, c = k, m = m)$centers))
diag(d) <- Inf
min(d)
})
plot(4:20, min_dist, type = 'b', xlab = 'k', ylab = 'Min centroid distance')mfuzz.plot2(eset, cl, mfrow = c(2, 4), time.labels = colnames(expr_mat), centre = TRUE, x11 = FALSE)
overlap.plot(cl, over = overlap(cl), thres = 0.05) # centroid-overlap view; merges hint at over-clusteringTCseq was built for time-course SEQUENCING (RNA-seq/ATAC-seq); upstream DE/peak steps live in the same package, and timeclust clusters the summarized (per-gene, per-timepoint) matrix.
library(TCseq)
# algo='cm': fuzzy c-means (soft, Mfuzz-like). Also 'km' (hard k-means), 'pam', 'hc' (hierarchical).
# standardize=TRUE does the mandatory per-gene z-score.
tc <- timeclust(expr_mat, algo = 'cm', k = 6, standardize = TRUE)
timeclustplot(tc, value = 'z-score', cols = 3)
tc_km <- timeclust(expr_mat, algo = 'km', k = 6, standardize = TRUE) # hard alternativeGoal: Hierarchical clustering with automatic k and design-aware grouping.
Approach: degPatterns takes replicate-level data plus metadata, collapses samples within each (time, col) group to a MEAN internally, then clusters on correlation distance and cuts the tree. Convenient, but "auto k" is really "cut + merge under minc," a heuristic - not an optimum.
library(DEGreport)
# time, col: COLUMN NAMES in metadata (col defaults to NULL). minc=15: minimum cluster size;
# clusters smaller than minc are DROPPED -- this both blocks singletons AND silently discards genes,
# so it can yield fewer clusters than the tree suggested. Set deliberately.
patterns <- degPatterns(expr_mat, metadata = sample_info, time = 'timepoint', col = 'condition', minc = 15)
cluster_df <- patterns$df # gene -> cluster assignments
degPlotCluster(patterns$normalized, time = 'timepoint', color = 'condition') # note: 'color', not 'col'Goal: Cluster time-series profiles, optionally warping the time axis for phase-shifted genes.
Approach: Z-score, then TimeSeriesKMeans. The DISTANCE METRIC matters more than the algorithm - default to Euclidean-on-zscore (which, after standardization, is monotone in Pearson correlation and captures "same shape, different amplitude"). Escalate to DTW ONLY for real, expected phase shifts, and ALWAYS constrain it.
import numpy as np
from tslearn.clustering import TimeSeriesKMeans, silhouette_score
from tslearn.preprocessing import TimeSeriesScalerMeanVariance
# expr_mat: (n_genes, n_timepoints) of PRE-SELECTED temporally variable genes
expr_scaled = TimeSeriesScalerMeanVariance().fit_transform(expr_mat[:, :, np.newaxis])
# Default, safe choice: Euclidean on z-scored profiles (phase-SENSITIVE, cheap, no fabricated structure)
model = TimeSeriesKMeans(n_clusters=8, metric='euclidean', max_iter=50, random_state=42)
labels = model.fit_predict(expr_scaled)DTW (Sakoe & Chiba 1978) warps the time axis so a profile peaking one timepoint later can still match - the ONLY reason to reach for it (signaling cascades, developmental heterochrony, unequal sampling). Its default failure mode is the SINGULARITY: unconstrained DTW maps one point of series A onto a long run of points of series B, manufacturing apparent co-regulation from noise. tslearn's default global_constraint=None is exactly this singularity-prone configuration.
# The Sakoe-Chiba BAND caps how far in time a point may be matched -- kills most singularities AND
# cuts cost. This constraint is mandatory, not optional, for DTW clustering.
# sakoe_chiba_radius: warping-window half-width in timepoints; small (1-2) for tight sampling.
model = TimeSeriesKMeans(
n_clusters=8, metric='dtw',
metric_params={'global_constraint': 'sakoe_chiba', 'sakoe_chiba_radius': 2},
max_iter=50, random_state=42)
labels = model.fit_predict(expr_scaled)
# Soft-DTW: replaces DTW's hard min with a soft-min -> DIFFERENTIABLE loss, enabling proper
# soft-DTW barycenters (cluster centers). It is NOT "faster" -- still quadratic; use it for smooth,
# well-defined averaging, not speed. gamma via metric_params (NOT the deprecated gamma_sdtw kwarg).
soft = TimeSeriesKMeans(n_clusters=8, metric='softdtw', metric_params={'gamma': 0.5},
max_iter=50, random_state=42)When DTW is worth it: only when phase shift is real and expected, the band is set, AND DTW has been checked against fabricating structure. On data with NO phase shifts, DTW should not beat Euclidean - if it "finds more clusters" there, that is invented structure, not signal.
# Scoring DTW clusters with a EUCLIDEAN silhouette is geometrically inconsistent: clusters were
# formed under DTW geometry but ranked under Euclidean, which can pick a DIFFERENT (wrong) k.
# tslearn.clustering.silhouette_score takes metric='dtw'/'softdtw' and precomputes the matching
# distances internally -- score under the SAME geometry that formed the clusters.
dtw_params = {'global_constraint': 'sakoe_chiba', 'sakoe_chiba_radius': 2}
scores = {}
for k in range(3, 11):
km = TimeSeriesKMeans(n_clusters=k, metric='dtw', metric_params=dtw_params, max_iter=30, random_state=42)
labels_k = km.fit_predict(expr_scaled)
scores[k] = silhouette_score(expr_scaled, labels_k, metric='dtw', metric_params=dtw_params)
best_k = max(scores, key=scores.get)Under a pure-Euclidean pipeline, sklearn.metrics.silhouette_score(expr_scaled.squeeze(), labels) is consistent and fast. It is only the DTW/Euclidean MISMATCH that mis-ranks k.
No index is authoritative; triangulate and let biology and stability decide.
| Signal | What it says | Caveat |
|---|---|---|
| Min centroid distance / Dmin | knee where centroids start collapsing = over-splitting | weak, monotone-ish |
| Silhouette | within- vs nearest-other-cluster separation | must match the clustering metric (DTW vs Euclidean) |
| Within-cluster dispersion / elbow / gap | dispersion drop-off | elbow subjective; gap assumes a null reference, expensive |
| Biology heuristic | does +1 cluster split a coherent program or resolve two real shapes? | the honest arbiter |
| Stability (bootstrap/consensus) | do the same genes co-cluster under resampling? | the real validation, not a lone index |
Over-clustering FRAGMENTS one real program across centroids (the same GO terms then reappear in three clusters); under-clustering MERGES distinct programs into an averaged centroid matching no gene. Report a stable partition, not a single silhouette peak.
| Metric | Captures | Phase shifts | Cost | Use when |
|---|---|---|---|---|
| Euclidean on z-score | shape + amplitude (monotone in Pearson after z-score) | NO | cheap | default for aligned timepoints |
| Correlation (DEGreport) | shape, amplitude-invariant | NO | cheap | shape-only focus |
| DTW (constrained) | shape with time warping | YES | O(n·T^2)/pair, worse for clustering | genuine, expected phase shifts only |
Selecting genes by a temporal criterion, clustering them, then TESTING those clusters for the same temporal signal is circular and inflates everything. If genes were selected for temporal variability, a follow-up test asking "are these clusters temporally structured / rhythmic?" is guaranteed to say yes - the signal was baked in at selection (Kriegeskorte-style non-independence). Interpreting per-cluster centroid p-values after DE selection is the same error: the genes are already significant by construction. Selection -> clustering is fine as a DESCRIPTIVE pipeline; what is not permissible is a test on the same data whose null was already violated by selection. Test clusters only against INDEPENDENT annotations (GO, TF targets, a held-out condition), never the temporal criterion used to select.
Run GO/GSEA per cluster to name programs, but the enrichment BACKGROUND (universe) must be the INPUT gene set that was clustered (the temporally variable genes), NOT the whole genome. Genome-as-background makes every cluster light up for the generic biology of "being a dynamic/expressed gene" (translation, stress, cell cycle) - that signal comes from the SELECTION step, not the cluster, and re-tests what was already done (mirrors the circularity trap). Testing cluster-vs-(rest-of-input) isolates what makes THIS shape distinct.
The examples cluster on replicate-AVERAGED profiles (standard and simple), but averaging DISCARDS uncertainty the DE step had: two genes with identical means but very different within-timepoint variance are treated as equally reliable. degPatterns makes the collapse explicit (mean within each time/col group) but still computes similarity on group means. The rigorous-but-rare alternative is a variance-aware/weighted distance; at minimum, state that averaging is a known limitation.
| Method | Clustering | Distance | Best for |
|---|---|---|---|
| Mfuzz | Soft (fuzzy c-means) | Euclidean on z-score | standard soft temporal profiling |
| TCseq | Soft (cm) or hard (km/pam/hc) | Euclidean on z-score | RNA-seq/ATAC time courses |
| DEGreport | Hierarchical, auto-k | Correlation | design-aware, quick auto-k |
| tslearn | Hard k-means | Euclidean / DTW / soft-DTW | phase-shifted profiles (constrained DTW) |
| Trap | Why it is wrong | Fix |
|---|---|---|
| Clustering ALL genes (incl. flat) | no null -> always returns clusters; z-score amplifies flat-gene noise into fake programs | prefilter to timeseries-DE hits or filter.std/top-variance FIRST |
| Skipping z-score | magnitude dominates shape; abundance clusters, not dynamics | standardise() / standardize=TRUE / TimeSeriesScalerMeanVariance() |
Hardcoding m=2 or trusting mestimate() blindly | m=2 over-fuzzy for many timepoints; mestimate degenerates at extreme D | inspect returned m + membership fraction; check it does not cluster randomized data |
| Treating k as having a "true" value | indices disagree; clustering has no true count | triangulate indices + biology + bootstrap stability |
| Unconstrained DTW | singularities invent structure from noise | set global_constraint='sakoe_chiba'; use DTW only for real phase shifts |
| "soft-DTW is just faster DTW" | still quadratic; its value is differentiability/barycenters | use soft-DTW for smooth averaging, not speed |
| Euclidean silhouette to pick k for DTW clusters | scores a different geometry than formed the clusters -> mis-ranks k | tslearn.clustering.silhouette_score(..., metric='dtw'), or cluster Euclidean throughout |
| Testing clusters for the temporal signal selected on | circular / double-dipping; p-values inflated | test only INDEPENDENT annotations |
| GO enrichment vs whole-genome background | re-detects "being dynamic" from the selection step | background = the clustered input gene set |
| Reporting centroids as if genes follow them exactly | centroid is an average; membership/spread varies | report membership (acore) fraction + within-cluster spread |
mestimate(); fuzzifier depends on the number of timepoints.)© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in temporal-genomics/temporal-clustering of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Temporal Genomics Temporal Clustering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Temporal Genomics Temporal Clustering this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.2k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Singlecell Qcxuzhougeng/wisp-science | 1k | — | ~1.6k | Automated safety check: Pass | AGPL-3.0 | |
| Trackplotygidtu/trackplot | 109 | — | ~1.9k | Automated safety check: Pass | BSD-3-Clause | |
| UniProt Database Accessdavila7/claude-code-templates | 33k | 14 repos | ~1.7k | Automated safety check: Pass | MIT |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
xuzhougeng/wisp-science
A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.
ygidtu/trackplot
Generate sashimi-style genome visualization plots (coverage, line, heatmap, IGV read-by-read, HiC, circRNA, motif) from BAM/bigWig/depth/HiC inputs.
davila7/claude-code-templates
Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.
QING1105/ezST
End-to-end 10x Visium spatial transcriptomics analysis workflow with staged execution and human review gates.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Works with
Categories
Clusters temporally variable genes by expression-profile SHAPE (not significance) using Mfuzz fuzzy c-means, TCseq, DEGreport degPatterns, and tslearn DTW/soft-DTW. Bio Temporal Genomics Temporal Clustering is an agent skill from GPTomics/bioSkills. Clusters temporally variable genes by expression-profile SHAPE (not significance) using Mfuzz fuzzy c-means, TCseq, DEGreport degPatterns, and tslearn DTW/soft-DTW.
Bio Temporal Genomics Temporal Clustering fits situations like: grouping pre-selected time-course genes into shared trajectory programs (co-expression modules); choosing between soft vs hard clustering; selecting a distance metric (Euclidean/correlation/DTW); interpreting clusters with per-cluster enrichment.
Run `npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a claude-code`. Or copy the skill folder (temporal-genomics/temporal-clustering in GPTomics/bioSkills) into .claude/skills/bio-temporal-genomics-temporal-clustering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a codex`. Or copy the skill folder (temporal-genomics/temporal-clustering in GPTomics/bioSkills) into .agents/skills/bio-temporal-genomics-temporal-clustering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-temporal-clustering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-temporal-genomics-temporal-clustering, .gemini/skills/bio-temporal-genomics-temporal-clustering, .github/skills/bio-temporal-genomics-temporal-clustering and .opencode/skills/bio-temporal-genomics-temporal-clustering in your project.
Going by SKILL.md and its folder, Bio Temporal Genomics Temporal Clustering needs R and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Temporal Genomics Temporal Clustering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Temporal Genomics Temporal Clustering: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Singlecell Qc (xuzhougeng/wisp-science, 1k stars) and Trackplot (ygidtu/trackplot, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.