Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Models continuous temporal trajectories from BULK or time-resolved omics where the x-axis is measured experimental time: penalized GAMs (mgcv) for smooth trends and changepoint detection (segmented…
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-trajectory-modeling --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/temporal-genomics/trajectory-modeling .claude/skills/bio-temporal-genomics-trajectory-modeling && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-temporal-genomics-trajectory-modeling" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/trajectory-modeling into .claude/skills/bio-temporal-genomics-trajectory-modeling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-trajectory-modeling", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/trajectory-modelingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-trajectory-modeling --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/temporal-genomics/trajectory-modeling .agents/skills/bio-temporal-genomics-trajectory-modeling && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-temporal-genomics-trajectory-modeling" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/trajectory-modeling into .agents/skills/bio-temporal-genomics-trajectory-modeling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-trajectory-modeling", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-trajectory-modeling --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/temporal-genomics/trajectory-modeling .cursor/skills/bio-temporal-genomics-trajectory-modeling && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-temporal-genomics-trajectory-modeling" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/trajectory-modeling into .cursor/skills/bio-temporal-genomics-trajectory-modeling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-trajectory-modeling", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path temporal-genomics/trajectory-modeling--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-trajectory-modeling --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/temporal-genomics/trajectory-modeling .gemini/skills/bio-temporal-genomics-trajectory-modeling && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-temporal-genomics-trajectory-modeling" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/trajectory-modeling into .gemini/skills/bio-temporal-genomics-trajectory-modeling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-trajectory-modeling", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-trajectory-modelingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/temporal-genomics/trajectory-modeling .github/skills/bio-temporal-genomics-trajectory-modeling && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-temporal-genomics-trajectory-modeling" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/trajectory-modeling into .github/skills/bio-temporal-genomics-trajectory-modeling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-trajectory-modeling", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-temporal-genomics-trajectory-modeling --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/temporal-genomics/trajectory-modeling .opencode/skills/bio-temporal-genomics-trajectory-modeling && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-temporal-genomics-trajectory-modeling" agent skill from https://github.com/GPTomics/bioSkills/tree/main/temporal-genomics/trajectory-modeling into .opencode/skills/bio-temporal-genomics-trajectory-modeling/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-temporal-genomics-trajectory-modeling", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-temporal-genomics-trajectory-modelingModels continuous temporal trajectories from BULK or time-resolved omics where the x-axis is measured experimental time: penalized GAMs (mgcv) for smooth trends and changepoint detection (segmented…
Bio Temporal Genomics Trajectory Modeling is an agent skill from GPTomics/bioSkills. Models continuous temporal trajectories from BULK or time-resolved omics where the x-axis is measured experimental time: penalized GAMs (mgcv) for smooth trends and changepoint detection (segmented, ruptures) for abrupt regime shifts. Use when deciding between a smooth GAM and a changepoint model; choosing the GAM distribution (nb() plus a library-size offset for raw counts vs Gaussian on vst/log-CPM); setting the basis-dimension ceiling k below the number of timepoints and letting REML pick wiggliness; handling…
Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/changepoint_detection.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python and R), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Temporal Genomics Trajectory Modeling loads about 5.5k tokens when it runs. Until then it costs about 215 tokens; SKILL.md has 1,764 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,764 words, ~5,451 tokens.
.claude/skills/bio-temporal-genomics-trajectory-modeling/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: mgcv 1.9+, tradeSeq 1.16+, segmented 2.0+, ruptures 1.1+, numpy 1.26+, pandas 2.2+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturespackageVersion('<pkg>') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Note: the modeled quantity must be on a scale the family assumes. A Gaussian GAM is valid only on variance-stabilized/log-transformed expression (vst, rlog, log-CPM); raw RNA-seq counts require family=nb() with a offset(log(library_size)). Successive timepoints are correlated, so a plain gam() (which assumes independent residuals) inflates smooth-term significance unless the AR structure is modeled or the residual ACF is checked.
"Fit smooth curves to my gene expression over real time, compare trajectories, and find abrupt shifts" -> model a continuous function of MEASURED time f(time), test whether it changes / differs between conditions, and locate discrete regime changes.
mgcv::gam()/gamm()/bam() for penalized-spline GAMs; segmented::segmented() for slope breaksruptures for level/distribution changepointsThis skill models trajectories where the x-axis is ACTUAL experimental time (hours, days, developmental stage) or a pseudobulk value aggregated over real time. Time is measured and shared across every sample at a timepoint, so replicates are exchangeable, timepoints are few and fixed, and residual autocorrelation across ordered timepoints is real. This is categorically different from single-cell pseudotime, which is a latent per-cell ordering estimated with error and belongs to single-cell/trajectory-inference (and to tradeSeq's native use case). Conflating the two is the deepest error in the area.
Three decisions dominate correctness before any p-value is read:
nb() and a library-size offset, or fit Gaussian on a variance-stabilized scale.gam() under-estimates SEs and over-calls temporal trends. Name the assumption and model it (corAR1/bam(rho=)) or at least inspect the residual ACF.| Question | Model | Use when | Do NOT use when |
|---|---|---|---|
| Does expression change / curve over time? | GAM s(time) (mgcv) | process is gradual (induction/decay kinetics, developmental ramps) | the process is a discrete switch -> a smooth smears the discontinuity |
| Do two conditions' trajectories diverge? | ordered-factor difference smooth (mgcv) | testing whether treated shape departs from control | groups have no shared reference / only a constant offset differs (use a parametric term) |
| When does the regime shift (slope break)? | segmented (broken-line) | continuous piecewise-LINEAR change in slope | the shift is a level jump, not a slope change (use ruptures l2) |
| When does the regime shift (level/distribution)? | ruptures Pelt/Binseg | step change in mean (l2) or distribution (rbf) | the curve is smooth -> any liberal penalty fabricates breaks |
| Is a curve even warranted? | AIC/edf of s(time) vs linear | deciding non-linear vs linear is enough | over-interpreting edf near 1 as a real curve |
Methodology evolves; before committing verify current defaults and recommendations against the latest mgcv/ruptures/segmented documentation.
Goal: Fit a smooth non-linear curve to expression over measured time and test whether it changes, on the correct distributional scale.
Approach: Use a penalized regression spline s(time); set the basis-dimension ceiling k generously but below the number of unique timepoints and let REML choose the realized wiggliness; use nb() + a library-size offset for raw counts, or Gaussian on a variance-stabilized scale.
library(mgcv)
# k is a CEILING (max basis dimension), NOT the number of knots the curve will use.
# The REML-chosen penalty picks the realized wiggliness; edf (below) reports it.
# k must be < number of unique timepoints; realistically k <= (#timepoints - 1).
# method='REML': better-behaved objective than GCV, resists under/over-smoothing (Wood 2011).
fit <- gam(expression ~ s(time, k = 6, bs = 'tp'), data = gene_df, method = 'REML')
summary(fit)
# s.table columns: edf, Ref.df, F (Gaussian/unknown scale), p-value.
# edf ~ 1 => penalty shrank the smooth to linear; edf near k-1 => nearly full flexibility.
# The smooth-term p-value is APPROXIMATE (conditional on estimated lambda; Wood 2013):
# treat it as categorical significant/not, do not compare tiny magnitudes.Goal: Model overdispersed RNA-seq counts on the count scale without violating the Gaussian assumption.
Approach: Fit family=nb() (mgcv estimates theta by REML) with offset(log(library_size)) so the smooth describes rate, not depth.
# Counts are mean-variance coupled and overdispersed: Gaussian is wrong on raw counts.
# nb() estimates theta jointly with the smoothing parameters under REML.
# offset(log(libsize)) absorbs sequencing depth so s(time) models expression rate.
fit_nb <- gam(counts ~ s(time, k = 6) + offset(log(library_size)),
data = gene_df, family = nb(), method = 'REML')
# With a known-scale family the test column becomes Chi.sq, not F.
summary(fit_nb)$s.tableGoal: Prevent inflated smooth-term significance caused by correlation between successive timepoints.
Approach: Model a lag-1 AR structure grouped by the replication unit with gamm(correlation=corAR1()), or fix rho in bam() for genome-wide fits after reading the lag-1 residual ACF.
# Plain gam() assumes independent residuals; positive AR(1) shrinks effective n,
# under-estimates SEs, and lets the smooth get too wiggly -> false temporal trends.
# corAR1 models e_t = rho * e_{t-1} + noise; group by subject/animal/plate.
fit_ar <- gamm(expression ~ s(time, k = 6),
correlation = corAR1(form = ~ time | subject),
data = gene_df, method = 'REML')
# fit_ar$gam holds the smooth; fit_ar$lme holds the correlation estimate.
# Genome-wide alternative: fit once without AR, read lag-1 residual ACF, set rho, refit.
# AR.start flags the first observation of each independent series.
fit_bam <- bam(expression ~ s(time, k = 6), data = gene_df,
rho = 0.4, AR.start = series_start, method = 'fREML')With independent biological replicates AT EACH timepoint the correlation is often weak or unidentifiable and plain gam() is defensible; a single series sampled repeatedly over many timepoints is where AR bites hardest. Always inspect the residual ACF before trusting the smooth p-value.
Goal: Directly test whether the treated trajectory's shape diverges from control, with its own p-value.
Approach: Make the grouping an ORDERED factor so s(time, by=grp) becomes a difference smooth (level minus reference); keep the reference global smooth AND the parametric main effect.
# Unordered by= gives each group's curve vs zero -- NOT a divergence test.
# Ordered factor: s(time, by=grp) is the difference smooth; its single p-value
# directly tests whether the trajectories diverge. Extends cleanly to >2 groups.
gene_df$condition <- as.ordered(gene_df$condition)
fit_diff <- gam(expression ~ condition + s(time, k = 6) + s(time, k = 6, by = condition),
data = gene_df, method = 'REML')
# The parametric 'condition' term is REQUIRED: centered smooths cannot carry the
# group's overall level, so without it a constant offset is misattributed to the smooth.
summary(fit_diff)A numeric 0/1 by=is_treated indicator is a valid shortcut for a single 2-level contrast (the second smooth is the treatment deviation), but the ordered-factor form is the general, canonical idiom.
Goal: Decide whether the basis is adequate and whether smooth terms are mutually identifiable.
Approach: Read gam.check()/k.check(); respond to a low k-index by doubling k and refitting, not by reflexively cranking k; use concurvity() only for multi-smooth models.
gam.check(fit) # 4 residual plots + k.check(): reports k', edf, k-index, p-value
# k-index < 1 with a small p means residual pattern the basis is too rigid to capture.
# CORRECT diagnostic = double k and refit: if edf rises substantially, k was too low;
# if edf barely moves, k was fine and the low k-index reflects autocorrelation or
# a distributional problem -- do not just raise k.
# Concurvity = the smooth analog of collinearity; only meaningful for multi-term models.
# Near 1 => partial attribution between smooths is unstable; > 0.8 is a worry, not a cutoff.
concurvity(fit_diff, full = TRUE)Goal: Visualize the fitted trajectory with an uncertainty band, within the sampled range only.
Approach: Predict on a fine grid with se.fit=TRUE; band = fit +/- 1.96*SE (pointwise, not simultaneous); never extrapolate.
grid <- data.frame(time = seq(min(gene_df$time), max(gene_df$time), length.out = 200))
pred <- predict(fit, newdata = grid, se.fit = TRUE)
grid$fitted <- pred$fit
# 1.96*SE is a POINTWISE 95% band; whole-curve (simultaneous) coverage is < 95%,
# so overlapping condition bands are NOT a formal test -- use the difference smooth p-value.
grid$lower <- pred$fit - 1.96 * pred$se.fit
grid$upper <- pred$fit + 1.96 * pred$se.fit
# Beyond [min(time), max(time)] the spline and its SE diverge -- do not predict outside range.Goal: Rank genes by temporal significance across the transcriptome.
Approach: Fit s(time) per gene, collect the smooth p-value, apply BH across genes (the per-gene p-values are approximate, so the FDR is approximate; permutation calibration is the gold standard for strong claims).
# expr_mat here must be variance-stabilized (vst / log-CPM); for raw counts use the family=nb() + offset
# form above instead of this default-Gaussian fit, or the per-gene SEs and p-values are invalid.
results <- data.frame()
for (gene in rownames(expr_mat)) {
df <- data.frame(expression = as.numeric(expr_mat[gene, ]), time = timepoints)
fit <- gam(expression ~ s(time, k = 6), data = df, method = 'REML')
s_tab <- summary(fit)$s.table
results <- rbind(results, data.frame(gene = gene, edf = s_tab[, 'edf'],
p_value = s_tab[, 'p-value']))
}
results$q_value <- p.adjust(results$p_value, method = 'BH') # q<0.05: standard FDR floor
temporal_genes <- results[results$q_value < 0.05, ]tradeSeq is BUILT for single-cell pseudotime lineages, not bulk real-time. fitGAM(counts, pseudotime, cellWeights, nknots) expects a gene x cell count matrix, a cell x lineage pseudotime matrix, and cell x lineage soft-assignment weights, and fits an NB GAM per gene per lineage. Its tests (associationTest, startVsEndTest, conditionTest, patternTest) are keyed to pseudotime lineages.
# Only reasonable when the design is genuinely a pseudobulk-over-a-lineage.
# For a standard bulk time-course with replicates at fixed timepoints, use mgcv directly:
# the same NB GAM, with full control over by= contrasts, offsets, and AR structure,
# and without single-cell scaffolding. nknots plays the same ceiling role as k (choose it with evaluateK()).
library(tradeSeq)
sce <- fitGAM(counts = count_mat, pseudotime = pt_mat, cellWeights = cw_mat, nknots = 6)
assoc_res <- associationTest(sce) # matrix of Wald stat + df + p-value, NOT an SCEGoal: Locate a continuous change in SLOPE and test whether a break exists at all.
Approach: Pre-test with davies.test before fitting a break; estimate the breakpoint with segmented() from a starting value; limit to one break unless the data are dense.
library(segmented)
lm_fit <- lm(expression ~ time, data = gene_df)
# davies.test H0: difference-in-slopes = 0 (no breakpoint). It searches candidate
# locations and corrects the minimum p for the search (slightly conservative).
# Decide WHETHER a break exists before interpreting one. pscore.test() is more powerful
# for a single break.
davies.test(lm_fit, seg.Z = ~time)
# segmented models a CONTINUOUS slope change (not a level jump); psi = starting value(s),
# NA auto-initializes. The estimator is a local search: sensitive to psi, fragile with
# multiple breaks (closely spaced breaks are non-identifiable).
seg_fit <- segmented(lm_fit, seg.Z = ~time, psi = NA)
summary(seg_fit)$psi # breakpoint estimate + SE (break CI is approximate/often too narrow)Goal: Detect discrete times where the mean (or whole distribution) shifts.
Approach: Factorize as (search method) x (cost model) x (penalty); the penalty choice IS the number-of-changepoints choice; match the cost model to the shift type and estimate the noise variance, not the total variance.
import numpy as np
import ruptures as rpt
signal = np.asarray(expression_values)
# model='l2' detects changes in MEAN (piecewise-constant level -- the natural choice for
# step-like expression regimes). 'rbf' detects changes in the whole distribution
# (mean AND variance) -- more general, hungrier for data, less interpretable.
# min_size=2: minimum segment length. Pelt is exact (O(n) via pruning); it takes a penalty
# and RETURNS the number+location of breaks -- there is no separate 'how many' knob.
n = len(signal)
# Noise variance, NOT total variance: np.var(signal) includes between-regime variation, so
# a BIC-style penalty log(n)*np.var(signal) is too large and UNDER-detects. Estimate noise
# from lag-1 differences instead. BIC is derived for the l2/Gaussian-mean cost -- pairing it
# with model='rbf' is theoretically mismatched; use l2 with BIC, or calibrate rbf empirically.
sigma2 = np.var(np.diff(signal)) / 2.0
penalty = np.log(n) * sigma2
bkps = rpt.Pelt(model='l2', min_size=2).fit(signal).predict(pen=penalty)
# predict returns break indices INCLUDING the terminal index n:
n_changepoints = len(bkps) - 1 # bkps[:-1] are the actual break locations
# Binseg: greedy, approximate, fast; takes a KNOWN number of breaks (sanity check vs Pelt).
bkps_binseg = rpt.Binseg(model='l2', min_size=2).fit(signal).predict(n_bkps=2)Guard against fabricated breaks: a liberal penalty always "finds" changepoints in a smooth ramp. Require a pre-test (a break exists) or compare a piecewise fit against a smooth-GAM fit by AIC -- if the smooth wins, the "changepoint" is a sampling-noise artifact. With few timepoints, be extremely skeptical of more than one break.
| Symptom | Cause | Fix |
|---|---|---|
| p-values wrong / fitted curve predicts negative expression | Gaussian GAM on raw overdispersed counts | family=nb() + offset(log(library_size)), or fit Gaussian on vst/log-CPM |
| Many genes "significantly change over time" implausibly | residual autocorrelation inflates smooth-term significance | `gamm(..., correlation=corAR1(form=~time |
| Treating k as "the number of bends I want" | k is the flexibility CEILING, not realized complexity | set k generously (< #timepoints), let REML pick lambda; read edf, not k |
| Cranking k whenever k-index < 1 | low k-index can mean autocorrelation/heteroscedasticity, not low basis | double k and refit -- if edf jumps, raise k; if not, look at correlation/distribution |
| Reading p=1e-30 as thirty orders of certainty | smooth p-values are approximate (ignore full lambda uncertainty) | treat as categorical significant/not; apply BH FDR across genes |
Unordered by= used "to test if curves differ" | it gives each group vs zero, not a divergence test | as.ordered(condition) -> the difference smooth's p-value IS the divergence test |
by= smooth without the parametric main effect | centered smooths cannot carry the group level | include condition + alongside s(time, by=condition) |
tradeSeq::fitGAM on a plain bulk time-course | tradeSeq is single-cell pseudotime machinery (needs cellWeights) | use mgcv directly for bulk real-time; reserve tradeSeq for pseudobulk lineages |
| ruptures under-detects real changepoints | pen=log(n)*np.var(signal) uses TOTAL variance -> penalty too large | estimate noise from np.var(np.diff(signal))/2; sweep the penalty |
model='rbf' with a BIC (log n * var) penalty | BIC penalty is derived for the l2/Gaussian-mean cost | use model='l2' with BIC, or calibrate the rbf penalty empirically |
| Changepoints "found" in a clearly smooth ramp | a liberal penalty fabricates breaks in gradual data | pre-test with davies.test; compare piecewise vs smooth-GAM AIC; the smooth often wins |
| Fitted curve behaves wildly past the last timepoint | extrapolating a penalized spline beyond the data | predict only within [min(time), max(time)] |
| "The condition bands overlap, so no difference" | 1.96*SE bands are pointwise, not simultaneous | use the difference-smooth p-value or simultaneous (posterior-simulation) intervals |
segmented with several psi gives unstable breaks | multiple breakpoints are weakly identifiable with few noisy points | limit to 1 break unless data are dense; supply good psi starts; check convergence |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in temporal-genomics/trajectory-modeling of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Temporal Genomics Trajectory Modeling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Temporal Genomics Trajectory Modeling this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.5k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Models continuous temporal trajectories from BULK or time-resolved omics where the x-axis is measured experimental time: penalized GAMs (mgcv) for smooth trends and changepoint detection (segmented…. Bio Temporal Genomics Trajectory Modeling is an agent skill from GPTomics/bioSkills. Models continuous temporal trajectories from BULK or time-resolved omics where the x-axis is measured experimental time: penalized GAMs (mgcv) for smooth trends and changepoint detection (segmented, ruptures) for abrupt regime shifts.
Bio Temporal Genomics Trajectory Modeling fits situations like: deciding between a smooth GAM and a changepoint model; choosing the GAM distribution (nb() plus a library-size offset for raw counts vs Gaussian on vst/log-CPM); setting the basis-dimension ceiling k below the number of timepoints and letting REML pick wiggliness; handling residual autocorrelation across timepoints with corAR1/bam(rho=).
Run `npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a claude-code`. Or copy the skill folder (temporal-genomics/trajectory-modeling in GPTomics/bioSkills) into .claude/skills/bio-temporal-genomics-trajectory-modeling in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a codex`. Or copy the skill folder (temporal-genomics/trajectory-modeling in GPTomics/bioSkills) into .agents/skills/bio-temporal-genomics-trajectory-modeling in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-temporal-genomics-trajectory-modeling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-temporal-genomics-trajectory-modeling, .gemini/skills/bio-temporal-genomics-trajectory-modeling, .github/skills/bio-temporal-genomics-trajectory-modeling and .opencode/skills/bio-temporal-genomics-trajectory-modeling in your project.
Going by SKILL.md and its folder, Bio Temporal Genomics Trajectory Modeling needs Python and R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Temporal Genomics Trajectory Modeling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Temporal Genomics Trajectory Modeling: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.