Alphafold Database Fetch And Analyze
google-deepmind/science-skills
Retrieve and analyze AlphaFold predicted structures for a protein.
Judges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB.
$ npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-structural-biology-structure-validation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/structural-biology/structure-validation .claude/skills/bio-structural-biology-structure-validation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-structural-biology-structure-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/structural-biology/structure-validation into .claude/skills/bio-structural-biology-structure-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-structural-biology-structure-validation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/structural-biology/structure-validationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-structural-biology-structure-validation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/structural-biology/structure-validation .agents/skills/bio-structural-biology-structure-validation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-structural-biology-structure-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/structural-biology/structure-validation into .agents/skills/bio-structural-biology-structure-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-structural-biology-structure-validation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-structural-biology-structure-validation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/structural-biology/structure-validation .cursor/skills/bio-structural-biology-structure-validation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-structural-biology-structure-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/structural-biology/structure-validation into .cursor/skills/bio-structural-biology-structure-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-structural-biology-structure-validation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path structural-biology/structure-validation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-structural-biology-structure-validation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/structural-biology/structure-validation .gemini/skills/bio-structural-biology-structure-validation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-structural-biology-structure-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/structural-biology/structure-validation into .gemini/skills/bio-structural-biology-structure-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-structural-biology-structure-validation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-structural-biology-structure-validationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/structural-biology/structure-validation .github/skills/bio-structural-biology-structure-validation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-structural-biology-structure-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/structural-biology/structure-validation into .github/skills/bio-structural-biology-structure-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-structural-biology-structure-validation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-structural-biology-structure-validation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/structural-biology/structure-validation .opencode/skills/bio-structural-biology-structure-validation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-structural-biology-structure-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/structural-biology/structure-validation into .opencode/skills/bio-structural-biology-structure-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-structural-biology-structure-validation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-structural-biology-structure-validationJudges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB.
Bio Structural Biology Structure Validation is an agent skill from GPTomics/bioSkills. Judges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB. Use when deciding if a structure or a specific region is trustworthy before docking/mechanism/measurement; reading resolution, R-work vs R-free and the R-free-minus-R-work overfitting gap; sanity-checking per-residue and mean B-factors; flagging clashscore, Ramachandran and rotamer outliers and cis non-proline peptides…
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/validate_structure.py` and `usage-guide.md`).
It sits in Research & Science, covering Protein structure and design. It works with AlphaFold. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
files.rcsb.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Structural Biology Structure Validation loads about 5.9k tokens when it runs. Until then it costs about 233 tokens; SKILL.md has 2,245 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,245 words, ~5,867 tokens.
.claude/skills/bio-structural-biology-structure-validation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: biopython 1.85+, numpy 1.26+
MolProbity, phenix (phenix.molprobity, phenix.process_predicted_model), and DSSP (mkdssp) are external CLI tools invoked via subprocess, not pip packages; install them separately and confirm they are on PATH before use.
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signatures<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Is this structure good enough to build on?" -> Read the data-quality metadata, run geometry validation, and decide per-region rather than per-file.
Bio.PDB.MMCIF2Dict (resolution/R-free), Bio.PDB.calc_dihedral (phi/psi/omega), subprocess -> phenix.molprobity (clashscore, rotamer, Ramachandran).A coordinate file is an INTERPRETED MODEL fit to experimental data, not the data and not ground truth, and its reliability varies PER-ATOM within one file. The single global resolution is a DATA CEILING on what the map can resolve, never a per-region quality certificate: a 1.5 Angstrom structure can still carry a guesswork surface loop, and a 3.2 Angstrom structure can have a rigid, locally excellent active site. The number that answers "can I trust THIS region" is the LOCAL signal - the per-residue B-factor and the real-space fit (RSRZ) for the residues actually used - not the headline resolution. Validate before any geometric interpretation: clashscore, Ramachandran and rotamer outliers, cis-peptides, and bond/angle deviations tell whether the model is even self-consistent before a distance or an angle computed on it means anything.
R-free, not R-work, is the cross-validation statistic. R-work is computed on the reflections used in refinement and can be driven down by adding parameters (waters, alt-confs, loose restraints) that fit noise; R-free is the same metric on a held-out test set that never entered refinement (Brunger 1992 Nature 355:472-475). The R-free-MINUS-R-work GAP is the overfitting flag - a gap much larger than expected for the resolution signals a model fitting its own noise, and a suspiciously small gap signals test-set leakage (Kleywegt & Brunger 1996 Structure 4:897-904). B-factors conflate genuine thermal motion, static disorder, and MODEL ERROR into one number, so they are a within-structure relative signal ("which atoms are least certain here"), NOT a portable cross-structure dynamics readout; normalize before comparing across structures and exclude high-B atoms (say >60-80 Angstrom^2 at moderate resolution) from precise geometric claims.
For a PREDICTED model the confidence scores bound SELF-CONSISTENCY, NOT correctness. pLDDT is per-residue and PAE is inter-residue confidence in the model's OWN frame (Jumper et al. 2021 Nature 596:583-589); a confident model can be confidently wrong when AlphaFold modeled a monomer/apo state where the truth is a complex/holo/alternative state. Trim low-pLDDT residues and split by PAE-defined domains before docking or molecular replacement (Oeffner et al. 2022 Acta Crystallogr D 78:1303-1314). Cryo-EM has GLOBAL vs LOCAL resolution - flexible peripheries in a big map are often docked hypotheses at 6-8 Angstrom local resolution inside a "2.5 Angstrom" map - and TWO different FSC thresholds: 0.143 on the gold-standard half-map FSC for the resolution claim, 0.5 for map-vs-model agreement (Rosenthal & Henderson 2003 J Mol Biol 333:721-745). An NMR deposit is an ENSEMBLE of models; validate per-model and report the per-residue spread - never average coordinates, which produces a physically impossible structure with distorted bonds and clashes (Montelione et al. 2013 Structure 21:1563-1570).
| Resolution (A) | Reliably resolvable | Do NOT over-read |
|---|---|---|
| < 1.2 (atomic) | Individual atoms, H atoms, anisotropic ADPs, alt-confs | - |
| 1.2-1.8 | Side-chain rotamers, ordered waters, alt-confs | Surface/loop atoms with high B still uncertain |
| 1.8-2.5 (typical) | Backbone plus most side chains; fold solid | Long polar rotamers, water networks, weak ligand density are model-dependent |
| 2.5-3.5 | Backbone trace, secondary structure, domain arrangement | Side-chain conformations, exact ligand pose, H-bond geometry are interpretive |
| > 3.5 | Overall fold, gross assembly | Individual side chains largely placed by geometry; treat atomic detail as hypothesis |
| Resolution | Typical R-work | Typical R-free | Concern if gap (R-free - R-work) exceeds ~ |
|---|---|---|---|
| ~1.5 A | 0.13-0.17 | 0.15-0.19 | 0.03-0.04 |
| ~2.0 A | 0.16-0.20 | 0.19-0.24 | 0.04-0.05 |
| ~2.5-3.0 A | 0.18-0.24 | 0.23-0.29 | 0.05-0.06 |
Absolute R rises with resolution (weaker high-angle reflections are noisier); a gap wider than the row's threshold flags overfitting, a near-zero gap flags work/test contamination. These are heuristics - verify against contemporaneous PDB statistics for the resolution.
| Metric | Measures | Target (goal) | Concern | Source |
|---|---|---|---|---|
| Clashscore | Serious all-atom overlaps (>0.4 A) per 1000 atoms | low, high same-resolution percentile | high vs same-resolution peers | Chen 2010 |
| Ramachandran favored | % residues in favored phi/psi basins | > 98% | < 95% | Williams 2018 |
| Ramachandran outliers | % residues in disallowed phi/psi | < 0.2% (goal 0) | > 0.5% | Williams 2018 |
| Poor rotamers | % side chains in disallowed rotamers | < 0.3% | > 1.5% | Williams 2018 |
| RSRZ outliers (X-ray) | Residues poorly fitting local density (RSRZ > 2) | few, dispersed | many, or clustered in the region of interest | Gore 2017 |
| MolProbity score | Composite mapped to a resolution-equivalent | <= deposited resolution | >> deposited resolution | Chen 2010 |
The poor-rotamer target tightened from ~1.0% to 0.3% with the updated reference distributions (Williams 2018). Read the wwPDB report's PERCENTILE sliders alongside the raw numbers: they rank an entry against the whole archive AND against same-resolution entries (Gore et al. 2017 Structure 25:1916-1927).
| The model is | Trust question | What to read | Authoritative route |
|---|---|---|---|
| X-ray | Is the region well-determined? | resolution + local B + RSRZ; R-free and its gap | wwPDB report + phenix.molprobity |
| Cryo-EM | Is the region rigid or a docked guess? | LOCAL resolution map, not global; map-model FSC (0.5) | EMDB local-resolution + map-model FSC |
| NMR | How wide is the conformational spread? | per-residue RMSD across all models | validate each model; report spread |
| Predicted (AlphaFold/ESMFold) | Is it self-consistent AND in the right biological context? | pLDDT bands + PAE blocks; is it apo/monomer where truth is holo/complex? | phenix.process_predicted_model (trim + PAE split) |
Goal: Pull resolution, R-work, R-free, and method from a deposited entry and flag the overfitting gap - the numbers Bio.PDB's thin structure.header omits.
Approach: Read the raw mmCIF categories with MMCIF2Dict (which reaches anything in the file), cast the strings, and compare the R-free-minus-R-work gap against a resolution-scaled expectation.
from Bio.PDB.MMCIF2Dict import MMCIF2Dict
def read_refinement_metadata(cif_path):
meta = MMCIF2Dict(cif_path)
def first(key):
val = meta.get(key, ['NA'])[0]
try:
return float(val)
except ValueError:
return val
method = meta.get('_exptl.method', ['NA'])[0]
resolution = first('_refine.ls_d_res_high')
r_work = first('_refine.ls_R_factor_R_work')
r_free = first('_refine.ls_R_factor_R_free')
gap = r_free - r_work if isinstance(r_free, float) and isinstance(r_work, float) else None
# A gap wider than ~0.05 flags overfitting at typical (~2 A) resolution (Kleywegt & Brunger 1996).
overfit_flag = gap is not None and gap > 0.05
return {'method': method, 'resolution': resolution, 'r_work': r_work, 'r_free': r_free, 'gap': gap, 'overfit_flag': overfit_flag}Goal: Locate the least-certain atoms of THIS model so precise-distance claims avoid them - a relative, within-structure read, never a cross-structure comparison.
Approach: Collect per-residue mean B for a chain, then flag residues whose B sits far above the structure's own median as low-confidence.
import numpy as np
from Bio.PDB import PDBParser
def bfactor_outliers(structure, chain_id, z_cut=2.0):
residue_b = {}
for residue in structure[0][chain_id]:
if residue.id[0] != ' ': # skip hetero/water; validate the polymer only
continue
residue_b[residue.id] = np.mean([a.get_bfactor() for a in residue])
vals = np.array(list(residue_b.values()))
median, mad = np.median(vals), np.median(np.abs(vals - np.median(vals))) + 1e-9
# Robust z on B: |B - median| / (1.4826*MAD); high B marks disorder/model error, not portable dynamics.
return {rid: b for rid, b in residue_b.items() if (b - median) / (1.4826 * mad) > z_cut}Goal: Catch backbone geometry that is almost always a modeling error - residues in disallowed phi/psi and cis peptide bonds that are not proline.
Approach: Get per-residue phi/psi from PPBuilder, coarsely classify against the canonical basins, and compute omega (Ca-C-N-Ca) directly to find cis bonds; a cis assignment at a non-proline is a validation flag until proven by density.
import numpy as np
from Bio.PDB import PDBParser, PPBuilder, calc_dihedral
# Coarse general-allowed basins (deg): alpha, beta/PPII, left-handed. Authoritative
# favored/allowed/outlier percentages need MolProbity's rama8000 contours (Williams 2018);
# this only screens gross outliers to decide whether to run MolProbity.
_BASINS = [(-160, -40, -80, 30), (-180, -40, 90, 180), (30, 90, -30, 90)]
def is_rama_allowed(phi, psi):
d = np.degrees([phi, psi])
return any(lo_p <= d[0] <= hi_p and lo_s <= d[1] <= hi_s for lo_p, hi_p, lo_s, hi_s in _BASINS)
def geometry_flags(structure):
outliers, cis_nonpro = [], []
ppb = PPBuilder()
for pp in ppb.build_peptides(structure[0]):
for residue, (phi, psi) in zip(pp, pp.get_phi_psi_list()):
if phi is not None and psi is not None and not is_rama_allowed(phi, psi):
outliers.append((residue.get_parent().id, residue.id[1], residue.resname))
residues = list(pp)
for prev, curr in zip(residues, residues[1:]):
omega = calc_dihedral(prev['CA'].get_vector(), prev['C'].get_vector(), curr['N'].get_vector(), curr['CA'].get_vector())
# omega ~180 = trans, ~0 = cis; |omega| < 30 deg is cis. cis at non-Pro is rare and usually an error.
if abs(np.degrees(omega)) < 30 and curr.resname != 'PRO':
cis_nonpro.append((curr.get_parent().id, curr.id[1], curr.resname))
return {'rama_outliers': outliers, 'cis_nonproline': cis_nonpro}Goal: Get archive-calibrated clashscore, rotamer, and Ramachandran outlier percentages rather than the coarse Python screen above.
Approach: Shell out to phenix.molprobity (or the MolProbity web service / molprobity.molprobity), which carries the reference contour and rotamer distributions the Python screen cannot reproduce.
import subprocess
def run_molprobity(model_path):
# phenix.molprobity writes molprobity.out with clashscore, rotamer_outliers,
# ramachandran_outliers, ramachandran_favored, molprobity_score. Parse that file.
result = subprocess.run(['phenix.molprobity', model_path], capture_output=True, text=True)
return result.stdoutThe coarse Python screen decides WHETHER to run MolProbity; MolProbity (Chen et al. 2010 Acta Crystallogr D 66:12-21; Williams et al. 2018 Protein Sci 27:293-315) gives the numbers to report. For a deposited entry, prefer the pre-computed wwPDB validation report (https://files.rcsb.org/pub/pdb/validation_reports/<xy>/<id>/<id>_validation.xml.gz) - it already carries the percentile sliders and per-residue RSRZ.
Goal: Turn an AlphaFold/ESMFold model into a trustworthy input by trimming low-confidence residues and reading inter-domain confidence - because pLDDT/PAE bound self-consistency, not correctness.
Approach: Read pLDDT from the B-factor column into bands, read the PAE matrix for domain segmentation, then hand the raw file to phenix.process_predicted_model (which converts pLDDT to a pseudo-B, trims, and splits by PAE-defined domains).
import json
import numpy as np
import subprocess
from Bio.PDB import MMCIFParser
_PLDDT_BANDS = [(90, 'very_high'), (70, 'confident'), (50, 'low'), (0, 'very_low')]
def plddt_bands(cif_path):
# pLDDT rides in the B-factor column but is confidence (high = good), OPPOSITE polarity to a real B-factor.
structure = MMCIFParser(QUIET=True).get_structure('pred', cif_path)
counts = {label: 0 for _, label in _PLDDT_BANDS}
for residue in structure[0].get_residues():
if 'CA' not in residue:
continue
score = residue['CA'].get_bfactor()
counts[next(label for cut, label in _PLDDT_BANDS if score >= cut)] += 1
return counts # a long very_low run usually flags an intrinsically disordered region, not an error
def pae_interdomain_confident(pae_json, block_a, block_b, cutoff=5.0):
# Off-diagonal PAE (A) between two domain blocks; low = relative orientation trusted, high = independent guess.
pae = np.array(json.load(open(pae_json))[0]['predicted_aligned_error'])
return pae[np.ix_(block_a, block_b)].mean() < cutoff
def process_for_mr(model_path):
# Trims below ~0.7 fractional pLDDT, converts pLDDT->pseudo-B, splits into PAE-defined domains (Oeffner 2022).
return subprocess.run(['phenix.process_predicted_model', model_path], capture_output=True, text=True).stdoutGoal: Turn a multi-model NMR deposit into a per-residue uncertainty map instead of over-claiming precision from model 1.
Approach: Superpose all models on a reference and report per-residue Ca RMSD across the ensemble; wide spread marks flexible or under-restrained regions.
import numpy as np
from Bio.PDB import PDBParser
def ensemble_ca_spread(structure, chain_id):
coords = []
for model in structure:
coords.append(np.array([res['CA'].get_coord() for res in model[chain_id] if 'CA' in res]))
stack = np.stack(coords) # (n_models, n_residues, 3); assumes consistent residue set across models
# Per-residue spread = mean distance of each model's Ca from the ensemble mean position.
return np.linalg.norm(stack - stack.mean(axis=0), axis=2).mean(axis=0)Secondary-structure validation (DSSP, Bio.PDB.DSSP(model, path, dssp='mkdssp')) needs backbone geometry to place the amide H and computes its own H-bond energy; it processes only the first model and its output drifts across the dssp->mkdssp v2->v4 rewrites, so name the version (Kabsch & Sander 1983 Biopolymers 22:2577-2637). See geometric-analysis for dihedral and DSSP mechanics.
| Symptom | Cause | Fix |
|---|---|---|
| "It is a 1.5 A structure so every atom is accurate" | Global resolution read as per-region quality | Read the local B-factor and RSRZ for the specific residues used; resolution is a data ceiling |
| Low R-work reported as proof of a good model | R-work is fit on refinement reflections and rewards overfitting | Report R-free (held-out) and the R-free-minus-R-work gap; a wide gap flags overfitting |
| B-factors of two structures compared at face value | B conflates thermal motion, disorder, and model error and is not portable | Compare within one structure only; normalize (z/percentile) before any cross-structure claim |
| Coloring an AlphaFold model "by B-factor" to infer flexibility | pLDDT rides in the B-factor column with OPPOSITE polarity (high = confident) | Read the column as pLDDT bands; high value means high confidence, not high motion |
| Deleting all low-pLDDT residues as junk | Low pLDDT often marks a real intrinsically disordered region | Distinguish disorder (keep, annotate) from misfold; trim only for MR/geometry pipelines |
| Docking straight into an AlphaFold pocket | Pocket is apo, side-chain rotamers are the least reliable atoms, may be wrong state | Prefer an experimental holo structure; if using the model, ensemble/flexible-side-chain dock and treat hits as hypotheses |
| Confident predicted model trusted for a complex | pLDDT/PAE bound self-consistency, not biological correctness | Ask what context AF could not see (partner, ligand, PTM); a monomer/apo model can be confidently wrong |
| Cryo-EM peripheral domain trusted at the headline resolution | Global resolution hides a low-local-resolution flexible arm | Consult the EMDB local-resolution map; treat low-local-res regions as docked hypotheses |
| Quoting FSC 0.5 as the cryo-EM resolution | 0.143 is the half-map criterion; 0.5 is the map-vs-model curve | Use gold-standard half-map FSC at 0.143 for resolution; 0.5 for checking the model against the map |
| Averaging NMR model coordinates into one structure | The mean of two valid conformers is physically impossible (distorted bonds, clashes) | Compute on each model and report the spread, or pick a representative/medoid model |
| cis peptide flagged everywhere, or missed entirely | omega near 0 (cis) vs 180 (trans) not checked; cis-Pro is common but cis non-Pro is rare | Compute omega; treat cis non-proline as a validation flag pending density, cis-Pro as plausible |
| Python Ramachandran percentages reported as authoritative | Coarse basin boxes are not MolProbity's rama8000 contours | Use the screen to decide whether to run phenix.molprobity; report MolProbity's numbers |
resolution/R-free come back None from structure.header | Bio.PDB's header dict is thin | Read _refine.ls_d_res_high, _refine.ls_R_factor_R_free, _exptl.method via MMCIF2Dict |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in structural-biology/structure-validation of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Structural Biology Structure Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Structural Biology Structure Validation this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.9k | Automated safety check: Pass | MIT | |
| Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills | 3.2k | 2 repos | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Alphafoldadaptyvbio/protein-design-skills | 164 | 3 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Biopipelineslocbp-uzh/biopipelines | 109 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Chaiadaptyvbio/protein-design-skills | 164 | 3 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Bio DB ToolsDrugClaw/DrugClaw | 125 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 |
google-deepmind/science-skills
Retrieve and analyze AlphaFold predicted structures for a protein.
adaptyvbio/protein-design-skills
Validate protein designs using AlphaFold2 structure prediction.
locbp-uzh/biopipelines
Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…
adaptyvbio/protein-design-skills
Structure prediction using Chai-1, a foundation model for molecular structure.
DrugClaw/DrugClaw
Query public biology databases and APIs including UniProt, RCSB PDB, AlphaFold DB, ClinVar, dbSNP, gnomAD, Ensembl, GEO, InterPro, KEGG, OpenTargets, Reactome, and STRING.
davila7/claude-code-templates
CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Works with
Categories
Judges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB. Bio Structural Biology Structure Validation is an agent skill from GPTomics/bioSkills.PDB.
Bio Structural Biology Structure Validation fits situations like: deciding if a structure; A specific region is trustworthy before docking/mechanism/measurement; reading resolution; R-work vs R-free and the R-free-minus-R-work overfitting gap.
Run `npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a claude-code`. Or copy the skill folder (structural-biology/structure-validation in GPTomics/bioSkills) into .claude/skills/bio-structural-biology-structure-validation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a codex`. Or copy the skill folder (structural-biology/structure-validation in GPTomics/bioSkills) into .agents/skills/bio-structural-biology-structure-validation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-structural-biology-structure-validation, .gemini/skills/bio-structural-biology-structure-validation, .github/skills/bio-structural-biology-structure-validation and .opencode/skills/bio-structural-biology-structure-validation in your project.
Going by SKILL.md and its folder, Bio Structural Biology Structure Validation needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: files.rcsb.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Structural Biology Structure Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Structural Biology Structure Validation: Alphafold Database Fetch And Analyze (google-deepmind/science-skills, 3.2k stars), Alphafold (adaptyvbio/protein-design-skills, 164 stars), Biopipelines (locbp-uzh/biopipelines, 109 stars) and Chai (adaptyvbio/protein-design-skills, 164 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.