Agent skill

Bio Crispr Screens Library Design

by GPTomics in GPTomics/bioSkills

Designs pooled sgRNA libraries for CRISPR knockout, interference (CRISPRi), activation (CRISPRa), Cas12a multiplex, base-editor, and prime-editor screens.

MITAuto-check passedResearch & Science

Install Bio Crispr Screens Library Design

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-crispr-screens-library-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-crispr-screens-library-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/crispr-screens/library-design .claude/skills/bio-crispr-screens-library-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-crispr-screens-library-design
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6k tokens
SKILL.md length
2,567 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Designs pooled sgRNA libraries for CRISPR knockout, interference (CRISPRi), activation (CRISPRa), Cas12a multiplex, base-editor, and prime-editor screens.

  • Choosing a genome-wide library (GeCKOv2 vs Avana vs Brunello vs TKOv3 vs Inzolia)
  • SKILL.md covers Version Compatibility, sgRNA Library Design, Library Chemistry Decision Tree and On-Target Scoring: Algorithmic…, plus 14 more sections
  • Runs Python scripts from its folder
  • Designing a focused

What it does

Bio Crispr Screens Library Design is an agent skill from GPTomics/bioSkills. Designs pooled sgRNA libraries for CRISPR knockout, interference (CRISPRi), activation (CRISPRa), Cas12a multiplex, base-editor, and prime-editor screens. Covers on-target scoring (Rule Set 2, Azimuth, DeepSpCas9, CRISPRon), off-target scoring (CFD, MIT), TSS-relative positioning for CRISPRi/a (Horlbeck, Dolcetto, Calabrese), PAM-variant chemistries, control-guide composition, oligo cloning architecture, and library QC. Use when choosing a genome-wide library (GeCKOv2 vs Avana vs Brunello vs TKOv3 vs Inzolia)…

Its SKILL.md is about 6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/design_library.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Choosing a genome-wide library (GeCKOv2 vs Avana vs Brunello vs TKOv3 vs Inzolia)
  • Designing a focused
  • Paralog-focused custom library
  • Picking CRISPRi vs CRISPRa TSS windows

Example prompts

  • “Use the bio-crispr-screens-library-design skill to design pooled sgRNA libraries for CRISPR knockout, interference (CRISPRi), activation (CRISPRa)…”
  • “/bio-crispr-screens-library-design”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Crispr Screens Library Design loads about 6k tokens when it runs. Until then it costs about 186 tokens; SKILL.md has 2,567 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~186
When it runs · the whole SKILL.md, loaded when a task matches
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,567 words, ~6,025 tokens.

Download SKILL.mdSave it as .claude/skills/bio-crispr-screens-library-design/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-crispr-screens-library-design
description
Designs pooled sgRNA libraries for CRISPR knockout, interference (CRISPRi), activation (CRISPRa), Cas12a multiplex, base-editor, and prime-editor screens. Covers on-target scoring (Rule Set 2, Azimuth, DeepSpCas9, CRISPRon), off-target scoring (CFD, MIT), TSS-relative positioning for CRISPRi/a (Horlbeck, Dolcetto, Calabrese), PAM-variant chemistries, control-guide composition, oligo cloning architecture, and library QC. Use when choosing a genome-wide library (GeCKOv2 vs Avana vs Brunello vs TKOv3 vs Inzolia), designing a focused or paralog-focused custom library, picking CRISPRi vs CRISPRa TSS windows, deciding control-guide proportions, or diagnosing library skew and dropout in a freshly cloned pool.
tool_type
mixed
primary_tool
CRISPOR

Version Compatibility

Reference examples tested with: CRISPOR 5.01+, BioPython 1.83+, pandas 2.2+, numpy 1.26+, Azimuth 2.0+ (Doench 2016), CRISPRon 1.0+ (Xiang 2021), DeepSpCas9 1.0+ (Kim 2019).

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: crispor.py --help from the crisporWebsite clone
  • Python: Azimuth has no console script; call azimuth.model_comparison.predict(...)

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

sgRNA Library Design

"Design a CRISPR library for my screen" -> Pick a chemistry (Cas9 KO, CRISPRi, CRISPRa, Cas12a, base or prime editor), score candidate guides for on-target activity and off-target liability, position them relative to gene/TSS, add appropriate controls, lay out the oligo for synthesis, and validate the cloned pool.

  • Python: crispor.py (web + CLI) for batch genome-wide guide scoring with CFD+MIT off-target
  • Python: azimuth (Microsoft Research) for Rule Set 2 on-target predictions (Brunello-style)
  • Python: CRISPRon, DeepSpCas9 for modern deep-learning predictors
  • R: crisprDesign (Bioconductor) for integrated annotation-aware design

Library Chemistry Decision Tree

GoalChemistryCanonical libraryGuides/geneTSS / target window
Loss-of-function essentiality, fitnessSpCas9 KOBrunello, TKOv3, Avana4 (Brunello), 4 (TKOv3), 6 (Avana)Constitutive exons, prefer aa 5-65% from N-terminus
Knockdown of non-cuttable genes, dosage-sensitivedCas9-KRAB (CRISPRi)Dolcetto, Horlbeck v26 (Dolcetto), 5 (Horlbeck)Optimum +25 to +75 downstream of FANTOM5 TSS, searched out to -50/+300 (Dolcetto, Sanson 2018); -25 to +500 (Horlbeck v2)
Gain-of-function, gene activationdCas9-VP64 / SAM / SunTag (CRISPRa)Calabrese, Horlbeck-CRISPRa6 (Calabrese), 5 (Horlbeck)-150 to -75 from TSS (Calabrese); -550 to -25 (Horlbeck v2)
Paralog buffering, GI screensenAsCas12a multiplexInzolia, in4mer4-guide arraysConstitutive exons
Variant function, SNV scanningCBE / ABECustom tiling libraryTile editing windowsEditing window pos 4-8 from PAM-distal end
Precise edit, indel-freePrime editorCustom PRIDICT-designedTile pegRNAsAnywhere with NGG PAM within 30 nt of edit

Fails when:

  • CRISPRi/a targeting wrong TSS: any TSS without FANTOM5 CAGE evidence is suspect; guides positioned against the wrong TSS lose most of their knockdown.
  • Cas9 KO of essential paralogs: single-KO buffering hides paralog-redundant essentials (42% of constitutively expressed genes never score, Dede 2020); switch to Cas12a multiplex.
  • Base editor over an exon-intron boundary: editing-window bystanders create splice variants instead of the intended SNV.

On-Target Scoring: Algorithmic Taxonomy

PredictorYearTraining setStrengthsFails when
Doench Rule Set 12014Flow-sorted GFP+ knockoutsSimple, interpretableLimited training data; sub-optimal at >NGG context
Doench Rule Set 2 / Azimuth 2.020161,841 flow-cytometry guides (Doench 2014) plus new guides tiling additional genesGold-standard for SpCas9; basis of BrunelloTrained on dropouts; under-predicts efficacy for nuclear-localized targets
DeepSpCas9201912,832 synthetic targets integrated in HEK293TSpearman ~0.77 vs measured indel frequency on held-out dataBlack-box; sensitive to chromatin/context features it wasn't trained on
CRISPRon2021High-throughput indel sequencingBest for therapeutic-grade target nominationSlow per-guide; over-fits to its specific cell line
DeepHF2019~171k guides in HEK293T (WT 55,604; eSpCas9(1.1) 58,167; HF1 56,888)Separate model per enzyme variant, including WTPick the model matching the enzyme actually used

Reconciliation: When predictors disagree, prefer the model whose training cell line matches the screen line (DeepSpCas9 was trained on synthetic targets integrated in HEK293T). For Brunello selection, Azimuth/Rule Set 2 is sufficient because the library was built with it -- introducing a different scorer creates apples-to-oranges ranking with the original library.

Off-Target Scoring

ScoreYearMathCutoff convention
MIT (Hsu)2013Position-weighted mismatch penaltySpecificity score 0-100, higher is better; CRISPOR treats >=50 as a good guide
CFD (Doench)2016Position+nucleotide-specific penalty fit on BrunelloPer-site CFD >0.2 counts a candidate off-target (Doench 2016); CRISPOR's aggregate CFD specificity score is 0-100, higher is better
Elevation2018ML on CFD + mismatch positionsTighter than CFD

CFD remains the default for genome-wide library design. Critical pitfall: CFD penalizes only mismatches, not bulges; for ≤1 mismatch + 1-bp bulge off-targets, validate empirically with GUIDE-seq or CIRCLE-seq. CRISPOR reports both the MIT (Hsu) and CFD guide specificity scores in a single output.

Score and Rank sgRNAs for a Target Gene

Goal: Generate ranked sgRNA candidates for a single gene, jointly scored on on-target activity (Rule Set 2 / Azimuth) and off-target liability (CFD).

Approach: Identify all PAM-adjacent 20-nt protospacers in the target gene's coding sequence, retain only those in the first 5-65% of the protein (constitutive-exon convention from Brunello), filter on GC 30-70% and absence of poly-T (≥4 Ts terminates U6), call Azimuth for on-target and CRISPOR for off-target, and select the top N satisfying both criteria.

python
import re
import pandas as pd
import numpy as np
from Bio.Seq import Seq

def find_sgrna_candidates(cds_sequence, pam='NGG', guide_length=20):
    '''Return all protospacer candidates with PAM coordinates on + strand.
    Caller must filter by exon position and Azimuth/CFD score.'''
    pam_pattern = re.compile(f'(?=([ACGT]{{{guide_length}}}{pam.replace("N", "[ACGT]")}))')
    candidates = []
    for strand, seq in [('+', cds_sequence), ('-', str(Seq(cds_sequence).reverse_complement()))]:
        for m in pam_pattern.finditer(seq):
            spacer = m.group(1)[:guide_length]
            if 'TTTT' in spacer or spacer.count('G') + spacer.count('C') not in range(6, 15):
                continue
            candidates.append({'spacer': spacer, 'strand': strand,
                               'pos_in_cds': m.start() if strand == '+' else len(seq) - m.start() - 23,
                               'gc_frac': (spacer.count('G') + spacer.count('C')) / guide_length})
    return pd.DataFrame(candidates)

def annotate_exon_position(candidates_df, cds_length):
    '''Filter to protospacers within first 5-65% of CDS (Brunello convention).
    Reason: N-terminal indels truncate protein; very-N-terminal hits alt initiation;
    C-terminal hits miss functional domains (Doench 2016 Nat Biotech).'''
    lo, hi = 0.05 * cds_length, 0.65 * cds_length
    return candidates_df[(candidates_df['pos_in_cds'] >= lo) & (candidates_df['pos_in_cds'] <= hi)].copy()

CRISPRi / CRISPRa TSS Targeting

Goal: Position guides relative to the empirical TSS for maximum knockdown (CRISPRi) or activation (CRISPRa).

Approach: Resolve TSS from FANTOM5 CAGE peaks (highest-ranked peak per gene; fall back to Ensembl/RefSeq if absent), define the modality-specific window, score candidate spacers in that window with Rule Set 2 plus the Horlbeck/Sanson CRISPRi/a-tailored rules, and select 5-6 guides per gene biased toward the window center.

python
def crispri_window(tss_coord, strand='+'):
    '''Dolcetto convention: search -50 to +300 around the FANTOM5 highest-rank CAGE peak.
    Reason: Sanson 2018 found +25 to +75 nt downstream of the TSS optimal for CRISPRi,
    so rank candidates toward that band; the search is relaxed outward to fill the
    per-gene guide quota when poorly-annotated TSSs leave too few candidates.'''
    if strand == '+':
        return (tss_coord - 50, tss_coord + 300)
    return (tss_coord - 300, tss_coord + 50)

def crispra_window(tss_coord, strand='+'):
    '''Calabrese convention: -150 to -75 upstream of TSS.
    Reason: dCas9-VP64 (and SAM, SunTag) activate maximally when bound
    just upstream of Pol II loading. Horlbeck v2 CRISPRa uses -550 to -25
    (broader, lower per-guide signal). For SAM, prefer Calabrese tightness;
    for SunTag, Horlbeck width is acceptable.'''
    if strand == '+':
        return (tss_coord - 150, tss_coord - 75)
    return (tss_coord + 75, tss_coord + 150)

Critical nuance: Cell-type-specific TSSs differ from the FANTOM5 consensus in ~15% of genes. For tissue-specific screens (e.g., neuron, hepatocyte), re-derive TSSs from a matched CAGE / GRO-seq / PRO-seq dataset before locking guide positions, or knockdown efficiency drops several-fold. The single most common cause of "weak" CRISPRi hits is mis-positioned guides against an alternative TSS.

Genome-Wide Library Selection

LibraryYearModalitySize (genes x guides)sgRNA rulesNotable
GeCKOv22014Cas9 KO~19k x 6 (~123k)Exon position + off-target specificity (predates Rule Set 1)Older; legacy datasets still use it
Avana2016Cas9 KO110,257 as published; DepMap screens a ~4-guide subset (Meyers 2017: 70,086 after filtering, 17,670 genes)Rule Set 1Still the Broad's primary Cas9 library; CERES->Chronos changed in 2021, not the library
Brunello2016Cas9 KO~19k x 4 (~77k)Rule Set 2 + CFDModern standard for new screens
TKOv32017Cas9 KO~18k x 4 (~71k)Hart on/off-targetBagel/BAGEL2-optimized
Humagne2020enAsCas12a~19.8k x 1 dual-guide construct (~20k per set)enAsCas12a rulesCompact Cas12a sets C and D
Horlbeck CRISPRi v22016dCas9-KRAB~18k x 5 (~104k)Horlbeck CRISPRi rulesFirst-gen, still widely used
Dolcetto2018dCas9-KRAB~19k x 3 per set (114,061 across Sets A+B)Horlbeck + Rule Set 2Modern CRISPRi standard
Horlbeck CRISPRa2016dCas9-VP64~18k x 5 (~104k)Horlbeck CRISPRa rulesOriginal CRISPRa
Calabrese2018dCas9-VP64~18.9k x 3 per set (113,238 across Sets A+B)Tight TSS windowModern CRISPRa standard
Inzolia2024enAsCas12a~49k arrays: 19,687 genes (2 arrays each) plus ~4,435 paralog pairsenAsCas12a rulesParalog-pair multiplex; ~30% smaller than a typical Cas9 library
in4mer2024Cas12a (4-guide)CustomenAsCas12a multiplexTriple/quadruple KO per cassette

dAUC trajectory (essentiality benchmark): GeCKOv2 < Avana < Brunello/TKOv3 (Doench 2016 + Hart 2017). Moving from 4 to 6 sgRNAs/gene gives diminishing returns; the larger gain is moving from Rule Set 1 to Rule Set 2.

Cost-coverage tradeoff: A 77k-guide Brunello at 500x cells/sgRNA needs 38.5M cells in pool, scalable. A 117k-guide Calabrese at 500x needs 59M cells -- often the deciding factor against CRISPRa for difficult-to-grow lines.

PAM Variants and Alternative Cas Enzymes

EnzymePAMSpacer lengthBest for
SpCas9 (WT)NGG20 ntStandard pooled screens; broadest library support
eSpCas9, SpCas9-HF1NGG20 ntLower off-target rate; use for therapeutic-grade nomination
SpCas9-NGNG20 ntExpanded targeting (~4x coverage); accept lower activity per guide
SpRYNRN / NYN20 ntNear-PAMless; coverage at every position; ~50% lower per-guide activity
SaCas9NNGRRT21 ntAAV-packageable (small ORF); rarely used in pooled screens
AsCas12a, LbCas12aTTTV23 ntAT-rich regions; staggered cut; lower expression noise
enAsCas12a (DeWeirdt 2021)Expanded TTTV + several non-canonical23 ntCombinatorial / paralog screens

Decision rule: If the screen requires every possible TSS position (saturation tiling, dense regulatory dissection), use SpRY despite lower activity; otherwise, NGG is best because the on-target predictors were trained on it.

Control Guides

A genome-wide library should include:

Control typeCountPurpose
Non-targeting (scrambled, no genomic match)500-1,000 (~1% of library)Primary null distribution for CRISPRi/a; safe baseline for normalization
Safe-harbor (AAVS1, ROSA26-equivalent)50-100Cas9-only: absorbs cut-toxicity baseline (matters for amplicon-correction)
Olfactory receptors (presumed non-expressed)50-100Second null set for orthogonal normalization
Reference essentials (CEGv2 subset: e.g. RPS3, RPL11, EIF3A, POLR2A)50-100Internal positive control; QC dropout signal
Reference non-essentials (NEGv1 subset)50-100Internal negative control; BAGEL2 calibration

Critical pitfall: Using only AAVS1 as the negative control in a Cas9 screen creates a normalization baseline biased toward "any cut is bad." Always add NTCs or non-essentials so that downstream median normalization and PR-AUC against CEGv2 work without baseline-shift artifacts.

Library Composition for Specialized Screens

Paralog buffering (Cas12a multiplex): Build 4-guide arrays where positions 1-2 target gene A and positions 3-4 target paralog gene B. Inzolia covers ~4,435 paralog pairs within ~49k arrays. Singleton controls (gene A alone, gene B alone) must be included to score genetic interaction = double_KO_LFC - sum(single_KO_LFC).

Base editor screens (tiling-library design): Tile NGG-adjacent spacers across exons; ensure editing window (positions 4-8 from PAM-distal end) lands inside coding exons; flag bystander Cs/As in the window for downstream interpretation. Restrict to 50-90% editing efficiency a priori (filter out predicted low-efficacy guides) -- see [[base-editing-analysis]].

Tiling / regulatory dissection: Dense (every 5-10 bp) CRISPRi or CRISPRa guides across the candidate region; CRISPRi has broader signal width (good for enhancer discovery) but Cas9-indel tiling has sharper resolution (good for pinpointing critical bases). Pair with CRISPR-SURF deconvolution.

Show full SKILL.md (1,046 more words)Show less

Oligo Design for Pooled Synthesis

Goal: Generate the final oligo sequence ready for chip-based synthesis. Vendor limits differ: Twist oligo pools cap at ~300 nt per oligo with no fixed pool size, GenScript's 92K format spans 20-170 nt, and Agilent OLS 244K spans 30-230 nt.

Approach: Add subpool PCR primers (so multiple sublibraries can share a synthesis array), the BsmBI/Esp3I overhang for golden-gate cloning into LentiGuide-Puro (Addgene 52963) or LentiCRISPRv2, and append the tracrRNA scaffold if the array length permits.

python
def build_oligo(spacer, vector='lentiGuide-Puro', subpool_idx=None):
    '''Construct final oligo for pooled synthesis.

    LentiGuide-Puro / LentiCRISPRv2 use BsmBI (Esp3I) with these overhangs:
        forward: 5'-CACCG[spacer]-3'
        reverse: 5'-AAAC[revcomp(spacer)]C-3'
    For chip synthesis, the spacer is flanked by subpool-specific PCR primers.'''
    subpool_fwd = {
        1: 'GGAAAGGACGAAACACCG',   # subpool 1 forward primer + BsmBI overhang
        2: 'GAGGCACTGGGCAGGTACCG',
    }.get(subpool_idx, 'GGAAAGGACGAAACACCG')
    # First 33 nt of the Chen 2013 sgRNA(F+E) optimized scaffold. NOTE: lentiGuide-Puro (#52963)
    # and lentiCRISPRv2 (#52961) carry the ORIGINAL scaffold; F+E belongs to lentiCRISPRv2-Opti (#163126).
    scaffold_short = 'GTTTAAGAGCTATGCTGGAAACAGCATAGCAAG'
    oligo = subpool_fwd + spacer + scaffold_short
    if len(oligo) > 200:
        raise ValueError(f'Oligo length {len(oligo)} exceeds the 200 nt design budget; check the vendor limit')
    return oligo

Subpool design: A large synthesis pool can be partitioned into multiple sublibraries via subpool primers; each sub-PCR amplifies its subpool, allowing one synthesis batch to serve several screens. Typical subpool size: 10k-20k oligos.

Library QC After Cloning

MetricTargetFailure mode if missed
sgRNA detection (>25 reads/guide in plasmid pool)≥99%Founder effect: missing guides cannot be screened; dropout impossible to distinguish from missing
Gini coefficient of plasmid pool<0.1Synthesis defects or PCR bias; pool unfit for screening at standard 500x coverage
Skew ratio (top 10% / bottom 10%)<2 (good), <5 (acceptable)Skew >5 means underrepresented guides cannot generate statistical signal even at 1000x
% zero-count sgRNAs in plasmid pool<0.5%Plasmid bottleneck during cloning; re-amplify or re-clone
Replicate Pearson on plasmid pool (between sequencing technical replicates)>0.99Sequencing artifact, not biology

Plasmid pool sequencing convention: 200-500 reads per sgRNA before any biology (i.e. 15-40M reads for a 77k Brunello). This is the baseline against which all downstream depletion is computed; sequencing the plasmid is non-negotiable.

Failure Modes

Wrong TSS in CRISPRi/a library

Trigger: Using Ensembl/RefSeq TSS instead of empirical CAGE peak for genes with broad or non-canonical promoters. Mechanism: dCas9-KRAB knockdown is maximal within ±100 bp of the actual Pol II loading site; canonical annotation can be off by 1-10 kb. Symptom: "Easy" essentials (RPS, RPL, EIF) show normal dropout but newer genes don't; library validates poorly against CEGv2. Fix: Re-derive TSS from FANTOM5 CAGE highest-rank peak; for tissue-specific lines, use matched CAGE or GRO-seq.

Library skew from PCR bias during amplification

Trigger: Amplifying the cloned plasmid pool with too many PCR cycles (>20) or with high-GC-bias polymerase. Mechanism: GC-extreme guides amplify nonlinearly; high-GC guides dominate, low-GC guides drop out. Symptom: Gini >0.2 on plasmid pool; sgRNAs with GC <30% systematically depleted. Fix: Cap PCR at 15 cycles; use Q5 or NEBNext Ultra II (low-bias); sequence at 500x post-amp to confirm Gini.

Oligo-synthesis dropouts in low-complexity guides

Trigger: Chip-synthesis errors at homopolymer runs or guides starting with GGGG. Mechanism: Synthesis chemistry has higher error rate at low-complexity regions; missing oligos cannot be cloned. Symptom: Specific guides absent from plasmid pool despite no design-rule violation. Fix: Re-design replacement guides; for production runs, request 2-3x synthesis depth so dropouts are buffered.

Polyclonality from high MOI

Trigger: Infection at MOI >0.5 to "save cells." Mechanism: Poisson math: at MOI 0.3, 26% of all cells are infected and 4% carry >=2 sgRNAs (14% of the infected fraction); at MOI 0.5, 39% are infected and 9% carry >=2. Symptom: Hits include neutral genes that co-infect with true essentials. Fix: MOI 0.3 strict; titer Cas9-positive cells specifically; re-check by qPCR of integration.

Wrong control proportion

Trigger: <100 non-targeting controls in a 70k library. Mechanism: Null distribution for normalization and FDR rests on the NTC variance; too few NTCs yields unstable median and inflated FDR. Symptom: Erratic gene-level p-values; MAGeCK FDR fluctuates wildly between runs. Fix: ~1% of library (500-1,000) NTCs; supplement with non-essential-gene controls.

Quantitative Thresholds

ThresholdValueSource / Rationale
GC content30-70%Doench 2016 Nat Biotech: guides outside this range have low activity
Poly-T avoidance≤3 consecutive TU6 Pol III terminator; ≥4 Ts terminates sgRNA transcription
Guides per gene (Cas9)4 (Brunello/TKOv3 standard); up to 6 (Avana, older)Doench 2016 reports diminishing gene recovery below 4 sgRNAs/gene; returns flatten above 6
CRISPRi window-50 to +300 search; +25 to +75 optimumSanson 2018 (Dolcetto); Horlbeck v2 uses -25 to +500
CRISPRa window-150 to -75 from TSSSanson 2018 (Calabrese); narrower than Horlbeck v2 (-550 to -25)
NTCs in library~1% (500-1,000 in a 70k library)DepMap library design notes; rule-of-thumb for stable null
MOI0.3Poisson: P(>=2 sgRNAs/cell) = 4% at MOI 0.3
Coverage at infection500 cells/sgRNADepMap convention; 200x minimum, 1000x for noisy / in-vivo
CRISPOR MIT specificity score>=50 (higher = more specific)CRISPOR convention (Haeussler 2016)
Library skew (top 10% / bottom 10%)<2 ideal, <5 acceptableJoung 2017 Nat Protoc

Common Errors

Error / symptomCauseSolution
sgRNA fails to expressPoly-T in spacer terminates U6Filter TTTT in design; this is the #1 silent failure
CRISPRi guide gives no knockdownWrong TSS usedRe-derive TSS from FANTOM5 / matched CAGE
Library Gini >0.3 in plasmid poolPCR over-amplification or synthesis defectCap at 15 cycles; re-sequence plasmid; consider re-synthesis
Hits include amplified loci (e.g. ERBB2 in HER2+)Copy-number amplicon false-essentialitySee [[copy-number-correction]]
Paralog gene absent from hit list despite expressionCas9 single-KO bufferingSwitch to Cas12a multiplex; see [[combinatorial-screens]]
Cas12a oligo doesn't cutForgot Cas12a's TTTV PAM is 5' of spacer, not 3'Re-orient: PAM-then-spacer for Cas12a, opposite of Cas9

References

  • Doench JG et al. 2014. Nat Biotechnol 32:1262. Rule Set 1.
  • Doench JG et al. 2016. Nat Biotechnol 34:184. Rule Set 2, CFD, Brunello/Avana libraries.
  • Sanjana NE et al. 2014. Nat Methods 11:783. GeCKOv2.
  • Hart T et al. 2017. G3 7:2719. TKOv3 library; CEGv2/NEGv1 reference essentiality gene sets.
  • Sanson KR et al. 2018. Nat Commun 9:5416. Dolcetto + Calabrese libraries; CRISPRi/a TSS rules.
  • Horlbeck MA et al. 2016. eLife 5:e19760. CRISPRi/a design rules; Horlbeck v2 library.
  • Kim HK et al. 2019. Sci Adv 5:eaax9249. DeepSpCas9.
  • Xiang X et al. 2021. Nat Commun 12:3238. CRISPRon.
  • Tycko J et al. 2019. Nat Commun 10:4063. Off-target toxicity mitigation in CRISPR screens.
  • DeWeirdt PC et al. 2021. Nat Biotechnol 39:94. enAsCas12a optimization.
  • Esmaeili Anvar N et al. 2024. Nat Commun 15:3577. Inzolia / in4mer paralog library.
  • Dede M et al. 2020. Genome Biol 21:262. Paralog buffering invisible to Cas9 single-KO.
  • Joung J et al. 2017. Nat Protoc 12:828. Genome-wide library screen protocol.
  • Shalem O et al. 2014. Science 343:84. Original GeCKO genome-scale knockout library design.
  • crispr-screens/screen-qc - Validate library skew, Gini, replicate correlation
  • crispr-screens/mageck-analysis - Analyze screens run with the designed library
  • crispr-screens/combinatorial-screens - Cas12a multiplex / paralog-pair library design
  • crispr-screens/base-editing-analysis - base-editor library design
  • crispr-screens/prime-editing-screens - PRIDICT2-optimized pegRNA libraries
  • crispr-screens/copy-number-correction - Filter amplicon-driven artifacts in cancer-cell-line screens

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in crispr-screens/library-design of GPTomics/bioSkills.

  • SKILL.md
  • examples/design_library.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Crispr Screens Library Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Crispr Screens Library Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Crispr Screens Library Design this skillGPTomics/bioSkills1.2k2 repos~6kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Singlecell Qcxuzhougeng/wisp-science1k—~1.6kAutomated safety check: PassAGPL-3.0
Trackplotygidtu/trackplot109—~1.9kAutomated safety check: PassBSD-3-Clause
UniProt Database Accessdavila7/claude-code-templates33k14 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Trackplot

    ygidtu/trackplot

    Generate sashimi-style genome visualization plots (coverage, line, heatmap, IGV read-by-read, HiC, circRNA, motif) from BAM/bigWig/depth/HiC inputs.

    109 GitHub stars~1.9k tokensUpdated 15 days ago
    Research & ScienceAuto-check passed
  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    33k GitHub starsUsed in 14 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • End-to-end 10x Visium spatial transcriptomics analysis workflow with staged execution and human review gates.

    101 GitHub stars~1.4k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Crispr Screens Library Design

What does Bio Crispr Screens Library Design do?

Designs pooled sgRNA libraries for CRISPR knockout, interference (CRISPRi), activation (CRISPRa), Cas12a multiplex, base-editor, and prime-editor screens. Bio Crispr Screens Library Design is an agent skill from GPTomics/bioSkills. Designs pooled sgRNA libraries for CRISPR knockout, interference (CRISPRi), activation (CRISPRa), Cas12a multiplex, base-editor, and prime-editor screens.

When should I use Bio Crispr Screens Library Design?

Bio Crispr Screens Library Design fits situations like: choosing a genome-wide library (GeCKOv2 vs Avana vs Brunello vs TKOv3 vs Inzolia); designing a focused; paralog-focused custom library; picking CRISPRi vs CRISPRa TSS windows.

How do I install Bio Crispr Screens Library Design in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-crispr-screens-library-design -a claude-code`. Or copy the skill folder (crispr-screens/library-design in GPTomics/bioSkills) into .claude/skills/bio-crispr-screens-library-design in your project. Claude Code loads it when a task matches its description.

How do I install Bio Crispr Screens Library Design in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-crispr-screens-library-design -a codex`. Or copy the skill folder (crispr-screens/library-design in GPTomics/bioSkills) into .agents/skills/bio-crispr-screens-library-design in your project. Codex loads it when a task matches its description.

Can I use Bio Crispr Screens Library Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-crispr-screens-library-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-crispr-screens-library-design, .gemini/skills/bio-crispr-screens-library-design, .github/skills/bio-crispr-screens-library-design and .opencode/skills/bio-crispr-screens-library-design in your project.

What does Bio Crispr Screens Library Design need to run?

Going by SKILL.md and its folder, Bio Crispr Screens Library Design needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Bio Crispr Screens Library Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Crispr Screens Library Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Crispr Screens Library Design use?

Bio Crispr Screens Library Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Crispr Screens Library Design use?

About 6k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Crispr Screens Library Design?

Skills that share tags, products or a category with Bio Crispr Screens Library Design: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Singlecell Qc (xuzhougeng/wisp-science, 1k stars) and Trackplot (ygidtu/trackplot, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Crispr Screens Library Design?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.