Agent skill

Safe Optimization

by GuyTeichman in GuyTeichman/RNAlysis

Make RNAlysis code faster while proving the output does not change.

MITAuto-check passedResearch & Science

Install Safe Optimization

skills CLI
$ npx skills add GuyTeichman/RNAlysis --skill safe-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GuyTeichman/RNAlysis safe-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GuyTeichman/RNAlysis.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/safe-optimization .claude/skills/safe-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
safe-optimization
GitHub stars
140
Token cost
~2.8k tokens
SKILL.md length
1,379 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Make RNAlysis code faster while proving the output does not change.

  • Works in 6 steps: Profile first. Do not guess. → Establish a baseline → Optimize → …
  • Asked to benchmark
  • SKILL.md covers When to use / skip, Step 1 -- Profile first. Do…, Step 2 -- Establish a baseline and Step 3 -- Optimize, plus 4 more sections
  • Calls git and python

What it does

Safe Optimization is an agent skill from GuyTeichman/RNAlysis. Make RNAlysis code faster while proving the output does not change. Use before touching any code for performance reasons -- a slow function, a profiling request, "speed this up", "this is too slow on large datasets", vectorizing a loop, parallelizing across CPU cores, adding/ changing a numba @jit, or swapping an algorithm/data structure for a faster one. Also use when asked to "benchmark", "profile", or write a perf PR. Skip for changes that are not primarily about speed (a correctness fix, a new feature) even…

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: Analyze your RNA sequencing data without writing a single line of code. The licence is MIT.

When your agent uses it

  • Asked to benchmark
  • Write a perf PR

Example prompts

  • “speed this up”
  • “this is too slow on large datasets”
  • “benchmark”
  • “/safe-optimization”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Profile first. Do not guess.
  2. Establish a baseline
  3. Optimize
  4. Prove bit-identical output
  5. Benchmark and record the speedup
  6. Known gotchas to check

What it can do on your machine

Read from SKILL.md and the folder at commit 0c70cc2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Safe Optimization loads about 2.8k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 1,379 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~142
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GuyTeichman/RNAlysis at commit 0c70cc2, republished under its MIT licence (© GuyTeichman). 1,379 words, ~2,763 tokens.

Download SKILL.mdSave it as .claude/skills/safe-optimization/SKILL.md (or your agent's skills folder).
name
safe-optimization
description
Make RNAlysis code faster while proving the output does not change. Use before touching any code for performance reasons -- a slow function, a profiling request, "speed this up", "this is too slow on large datasets", vectorizing a loop, parallelizing across CPU cores, adding/ changing a numba @jit, or swapping an algorithm/data structure for a faster one. Also use when asked to "benchmark", "profile", or write a perf PR. Skip for changes that are not primarily about speed (a correctness fix, a new feature) even if they happen to also be faster.

Safe performance optimization

Correctness and reproducibility are non-negotiable in RNAlysis (CLAUDE.md rule #5, AGENTS.md rule #5): "Results must reproduce across versions. A given analysis with given parameters must produce the same output." A performance change that is fast but wrong, or fast but silently different, is a regression dressed up as an improvement. This skill is the procedure for making code faster and proving it did not change what the code computes.

The engine for step 4 (proof) is a plain, provider-neutral script: packaging/bench_equal.py. This skill adds the when and the procedure around it.


When to use / skip

Use this skill for any change whose primary motivation is speed: a slow function, a profiling request, vectorizing a loop, parallelizing across CPU cores, adding/changing a numba @jit, swapping an algorithm or data structure, introducing Polars lazy evaluation to a hot path, etc.

Skip it for changes that are not primarily about speed -- a correctness fix, a new feature, a refactor for readability -- even if they incidentally run faster. Those still go through the normal tdd workflow; retrofitting this skill's ceremony onto them is not the point. (If a correctness fix happens to also touch a real hotspot, it is fine to fold in a benchmark, but the red-green test for the bug comes first.)


Step 1 -- Profile first. Do not guess.

Optimizing the wrong thing wastes effort and adds risk to the codebase for nothing. This codebase's own history has punished guessing more than once:

  • In pca/clustering, Box-Cox (rnalysis/utils/generic.py::box_cox), not PCA itself, was the dominant per-gene cost -- the obvious suspect (PCA) was not the real one.
  • In GO/KEGG enrichment, set-intersection work was not a hotspot at all; the real cost was the stats-test computation (recomputed per ontology term) and an elim-propagation deep-copy repeated on every run (rnalysis/utils/enrichment_runner.py) -- see HISTORY.rst's entry on memoizing the hypergeometric p-value and dropping the elim deep-copy.

Profile with whatever tool fits the question -- cProfile/pstats for a first cut ("which function"), line_profiler for "which line", py-spy/scalene for sampling a running process without instrumentation overhead. A couple of representative inputs at realistic scale (not a toy 3-row table) is what surfaces real hotspots; a tiny input mostly measures constant overhead. Only once profiling data names a specific function/line should you move to step 2.

Step 2 -- Establish a baseline

Before changing anything, capture the current code's output on one or more representative inputs. "Representative" means inputs that exercise the real shape of the data this function sees in production: realistic row/column counts, some NaN/null values if the real data has them, edge cases (empty input, a single row, all-identical values) if those are plausible.

In practice this is either:

  • a fixture already in tests/test_files/, loaded the way the matching tests/test_*.py module loads it, or
  • synthetic data built to match the real shape (np.random.default_rng(<fixed seed>) / pl.DataFrame(...)) when no fixture is representative enough.

You do not need to persist the baseline anywhere durable -- bench_equal.compare() (step 4) calls the current code itself as the baseline and compares it against your optimized version in the same run, so "capturing" it is really just picking the inputs and keeping a reference to the pre-optimization function (e.g. via git stash/a second import/a renamed copy) long enough to run the comparison.

Step 3 -- Optimize

Make the change the profiling data justified. Common levers in this codebase: vectorize a Python loop with NumPy/Polars, push work into a Polars lazy pipeline instead of eager, memoize a value recomputed across many iterations, parallelize across CPU cores (mind the frozen-vs-source gotcha below), or add a numba @jit to a numeric inner loop (mind the RNG gotcha below).

Step 4 -- Prove bit-identical output

This is the step that turns "I made it faster" into "I made it faster, safely." Use packaging/bench_equal.py's assert_equal/compare to compare the old and new implementations on the same inputs from step 2:

python
import importlib.util
import sys
from pathlib import Path

spec = importlib.util.spec_from_file_location('bench_equal', Path('packaging/bench_equal.py'))
bench_equal = importlib.util.module_from_spec(spec)
sys.modules[spec.name] = bench_equal
spec.loader.exec_module(bench_equal)

result = bench_equal.compare(old_box_cox, new_box_cox, args=(representative_data,))
print(f'{result.speedup:.1f}x faster, bit-identical')

or, for several representative inputs at once, write a small throwaway "bench spec" file and run it from the CLI (python packaging/bench_equal.py my_bench_spec.py --repeats 5) -- see bench_equal.py's module docstring for the spec-file format (BASELINE, CANDIDATE, INPUTS).

Exact equality is the default and should stay the default. assert_equal/compare compare NumPy arrays and Polars DataFrame/Series element-for-element (shape and dtype must match too), with NaN/null treated as equal to itself -- the same standard the test suite already holds analysis output to.

Only pass rtol/atol when the optimization itself legitimately reorders floating-point operations (a parallel reduction, a different BLAS/LAPACK code path, float32 vs float64 accumulation) such that bit-identical output is not achievable even though the computation is equivalent. HISTORY.rst already has a real precedent for this: the parallelized Box-Cox transform is documented as "Results are unchanged up to floating-point precision (verified against the test suite's reference outputs)" -- not "unchanged", because parallel reduction order isn't associative in floating point. If you need a tolerance, that is itself the reproducibility event from hard invariant #5: it must be intentional, justified, and called out in HISTORY.rst (state the tolerance and why exact equality isn't achievable), not a quiet flag you flip to make a red assertion go green.

If compare()/assert_equal raises AssertionError: the optimization is not safe yet. Either the optimization has a bug, or it is a genuine behavior change that needs its own justification and sign-off (per the plan-first rule for risky changes) -- not a benchmark to report.

Show full SKILL.md (501 more words)Show less

Step 5 -- Benchmark and record the speedup

Once equality holds, report the number. BenchmarkResult.speedup (from compare()) gives it to you directly, using best-of-N wall-clock time on each side (robust to one noisy call). Quote it in the PR description and, for a user-visible perf change, in HISTORY.rst -- concretely, e.g.: "the hypergeometric test and elim propagation benefit the most (up to ~20x and ~2-9x faster respectively on large ontologies)". Prefer a range across the representative inputs from step 2 over a single cherry-picked best case.

Step 6 -- Known gotchas to check

numba RNG must be seeded inside the jitted function. numba compiles its own internal RNG state that is separate from NumPy's global RNG -- seeding NumPy's global RNG from ordinary Python code before calling into @jit(nopython=True) code is a no-op for any np.random.* call made inside that jitted function. This codebase has a live example of the trap: rnalysis/utils/enrichment_runner.py's PermutationTest.run() calls np.random.seed(...) in plain Python, then calls the jitted _calc_permutation_pval (decorated @generic.numba.jit(nopython=True)), which itself calls np.random.choice(...) -- the seed set in run() does not reach the random draws inside _calc_permutation_pval. This was the root cause of a previously flaky "reproducibility" test. If you jit a function that needs reproducible randomness, seed it with np.random.seed(...) (or pass the seed in and call it) from inside the jitted function itself, and prove it with a test that calls the jitted function twice with the same seed and asserts equal output.

The frozen-vs-source multiprocessing split. The app runs both from a source checkout and as a frozen PyInstaller executable (RNAlysis.exe/.dmg), and they do not support the same parallel backends. rnalysis/__init__.py sets FROZEN_ENV = getattr(sys, 'frozen', False) and hasattr(sys, '_MEIPASS'); rnalysis/utils/param_typing.py gates on it directly:

python
PARALLEL_BACKENDS = ('multiprocessing', 'sequential') if FROZEN_ENV else (
    'multiprocessing', 'loky', 'threading', 'sequential')

loky/threading are only available from source -- a frozen build cannot use them. If your optimization parallelizes something, either accept a parallel_backend parameter typed Literal[PARALLEL_BACKENDS] (as the existing filtering/enrichment functions do) so the GUI exposes only the legal choices per environment, or, if the backend is chosen internally rather than user-facing, follow the pattern in rnalysis/utils/generic.py::box_cox_parallel_backend() (picks 'multiprocessing' when FROZEN_ENV, 'loky' otherwise) rather than hardcoding a backend that breaks one of the two shipping forms. Reason through both environments before touching parallelism -- you cannot test the frozen build's behavior by running from source.


Definition of done

A performance change is not done until, in the PR:

  1. Profiling evidence names the actual hotspot (not an assumption) -- a one-line summary or a cProfile/line_profiler excerpt is enough.
  2. The baseline and candidate were compared on representative inputs via bench_equal.compare()/assert_equal -- exact by default; any rtol/atol used is justified in the PR text.
  3. A measured speedup number is reported (best-of-N, from BenchmarkResult.speedup or the CLI's printed ratio).
  4. If output changed at all (even "only" up to floating-point precision), HISTORY.rst has an entry stating that plainly -- per hard invariant #5, this is never a silent side effect.
  5. The relevant tests/test_*.py module(s) still pass, per the normal tdd/finishing-a-change workflow in .claude/workflows.md.
  6. If the optimization touches parallelism, the frozen-vs-source split (Step 6) was reasoned through, not ignored.

© GuyTeichman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/safe-optimization of GuyTeichman/RNAlysis.

Open the folder on GitHubat commit 0c70cc2

Compare with similar skills

Safe Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Safe Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Safe Optimization this skillGuyTeichman/RNAlysis140—~2.8kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GuyTeichman/RNAlysis

  • External API Change

    GuyTeichman/RNAlysis

    Workflow for fixing or changing RNAlysis code that talks to an EXTERNAL WEB SERVICE — UniProt, Ensembl, PANTHER, PhylomeDB, OrthoInspector, KEGG, or GO.

    140 GitHub stars~1.8k tokensUpdated 9 days ago
    Auto-check passed
  • Gui Screenshots

    GuyTeichman/RNAlysis

    Capture and attach RNAlysis GUI screenshots to a PR whenever a change makes a VISIBLE difference to a GUI dialog.

    140 GitHub stars~2k tokensUpdated 9 days ago
    Auto-check passed

Questions about Safe Optimization

What does Safe Optimization do?

Make RNAlysis code faster while proving the output does not change. Safe Optimization is an agent skill from GuyTeichman/RNAlysis. Make RNAlysis code faster while proving the output does not change.

When should I use Safe Optimization?

Safe Optimization fits situations like: asked to benchmark; write a perf PR.

How do I install Safe Optimization in Claude Code?

Run `npx skills add GuyTeichman/RNAlysis --skill safe-optimization -a claude-code`. Or copy the skill folder (.claude/skills/safe-optimization in GuyTeichman/RNAlysis) into .claude/skills/safe-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Safe Optimization in Codex?

Run `npx skills add GuyTeichman/RNAlysis --skill safe-optimization -a codex`. Or copy the skill folder (.claude/skills/safe-optimization in GuyTeichman/RNAlysis) into .agents/skills/safe-optimization in your project. Codex loads it when a task matches its description.

Can I use Safe Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GuyTeichman/RNAlysis --skill safe-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/safe-optimization, .gemini/skills/safe-optimization, .github/skills/safe-optimization and .opencode/skills/safe-optimization in your project.

What does Safe Optimization need to run?

Going by SKILL.md and its folder, Safe Optimization needs the command-line tools its instructions call (git and python). Our summary lists: Python 3.

Does Safe Optimization access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Safe Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Safe Optimization use?

Safe Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Safe Optimization use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Safe Optimization?

Skills that share tags, products or a category with Safe Optimization: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Safe Optimization?

GuyTeichman (a GitHub user) maintains it in GuyTeichman/RNAlysis, which has 140 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 28, 2026.

Source: GuyTeichman/RNAlysis on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.