Agent skill

Bio Structural Biology Structure Validation

by GPTomics in GPTomics/bioSkills

Judges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB.

MITAuto-check passedResearch & Science

Install Bio Structural Biology Structure Validation

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-structural-biology-structure-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/structural-biology/structure-validation .claude/skills/bio-structural-biology-structure-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-structural-biology-structure-validation
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.9k tokens
SKILL.md length
2,245 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Judges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB.

  • Deciding if a structure
  • SKILL.md covers Version Compatibility, Governing Principle, Decision: what the resolution… and Decision: R-free and the…, plus 11 more sections
  • Runs Python scripts from its folder; calls pip; reaches files.rcsb.org
  • A specific region is trustworthy before docking/mechanism/measurement

What it does

Bio Structural Biology Structure Validation is an agent skill from GPTomics/bioSkills. Judges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB. Use when deciding if a structure or a specific region is trustworthy before docking/mechanism/measurement; reading resolution, R-work vs R-free and the R-free-minus-R-work overfitting gap; sanity-checking per-residue and mean B-factors; flagging clashscore, Ramachandran and rotamer outliers and cis non-proline peptides…

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/validate_structure.py` and `usage-guide.md`).

It sits in Research & Science, covering Protein structure and design. It works with AlphaFold. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Deciding if a structure
  • A specific region is trustworthy before docking/mechanism/measurement
  • Reading resolution
  • R-work vs R-free and the R-free-minus-R-work overfitting gap

Example prompts

  • “Use the bio-structural-biology-structure-validation skill to judge whether a macromolecular model (or a region of it) is reliable enough to build…”
  • “/bio-structural-biology-structure-validation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • files.rcsb.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Structural Biology Structure Validation loads about 5.9k tokens when it runs. Until then it costs about 233 tokens; SKILL.md has 2,245 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~233
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,245 words, ~5,867 tokens.

Download SKILL.mdSave it as .claude/skills/bio-structural-biology-structure-validation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-structural-biology-structure-validation
description
Judges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB. Use when deciding if a structure or a specific region is trustworthy before docking/mechanism/measurement; reading resolution, R-work vs R-free and the R-free-minus-R-work overfitting gap; sanity-checking per-residue and mean B-factors; flagging clashscore, Ramachandran and rotamer outliers and cis non-proline peptides; validating a PREDICTED (AlphaFold/ESMFold) model via pLDDT bands and PAE before docking or molecular replacement; and interpreting cryo-EM global-vs-local resolution (FSC 0.143 half-map vs 0.5 map-model) or an NMR ensemble spread. Keywords validation, resolution, R-free, B-factor, MolProbity, clashscore, Ramachandran, rotamer, pLDDT, PAE, wwPDB, cryo-EM local resolution.
tool_type
python
primary_tool
Bio.PDB

Version Compatibility

Reference examples tested with: biopython 1.85+, numpy 1.26+

MolProbity, phenix (phenix.molprobity, phenix.process_predicted_model), and DSSP (mkdssp) are external CLI tools invoked via subprocess, not pip packages; install them separately and confirm they are on PATH before use.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Structure Validation

"Is this structure good enough to build on?" -> Read the data-quality metadata, run geometry validation, and decide per-region rather than per-file.

  • Python: Bio.PDB.MMCIF2Dict (resolution/R-free), Bio.PDB.calc_dihedral (phi/psi/omega), subprocess -> phenix.molprobity (clashscore, rotamer, Ramachandran).

Governing Principle

A coordinate file is an INTERPRETED MODEL fit to experimental data, not the data and not ground truth, and its reliability varies PER-ATOM within one file. The single global resolution is a DATA CEILING on what the map can resolve, never a per-region quality certificate: a 1.5 Angstrom structure can still carry a guesswork surface loop, and a 3.2 Angstrom structure can have a rigid, locally excellent active site. The number that answers "can I trust THIS region" is the LOCAL signal - the per-residue B-factor and the real-space fit (RSRZ) for the residues actually used - not the headline resolution. Validate before any geometric interpretation: clashscore, Ramachandran and rotamer outliers, cis-peptides, and bond/angle deviations tell whether the model is even self-consistent before a distance or an angle computed on it means anything.

R-free, not R-work, is the cross-validation statistic. R-work is computed on the reflections used in refinement and can be driven down by adding parameters (waters, alt-confs, loose restraints) that fit noise; R-free is the same metric on a held-out test set that never entered refinement (Brunger 1992 Nature 355:472-475). The R-free-MINUS-R-work GAP is the overfitting flag - a gap much larger than expected for the resolution signals a model fitting its own noise, and a suspiciously small gap signals test-set leakage (Kleywegt & Brunger 1996 Structure 4:897-904). B-factors conflate genuine thermal motion, static disorder, and MODEL ERROR into one number, so they are a within-structure relative signal ("which atoms are least certain here"), NOT a portable cross-structure dynamics readout; normalize before comparing across structures and exclude high-B atoms (say >60-80 Angstrom^2 at moderate resolution) from precise geometric claims.

For a PREDICTED model the confidence scores bound SELF-CONSISTENCY, NOT correctness. pLDDT is per-residue and PAE is inter-residue confidence in the model's OWN frame (Jumper et al. 2021 Nature 596:583-589); a confident model can be confidently wrong when AlphaFold modeled a monomer/apo state where the truth is a complex/holo/alternative state. Trim low-pLDDT residues and split by PAE-defined domains before docking or molecular replacement (Oeffner et al. 2022 Acta Crystallogr D 78:1303-1314). Cryo-EM has GLOBAL vs LOCAL resolution - flexible peripheries in a big map are often docked hypotheses at 6-8 Angstrom local resolution inside a "2.5 Angstrom" map - and TWO different FSC thresholds: 0.143 on the gold-standard half-map FSC for the resolution claim, 0.5 for map-vs-model agreement (Rosenthal & Henderson 2003 J Mol Biol 333:721-745). An NMR deposit is an ENSEMBLE of models; validate per-model and report the per-residue spread - never average coordinates, which produces a physically impossible structure with distorted bonds and clashes (Montelione et al. 2013 Structure 21:1563-1570).

Decision: what the resolution buys (X-ray / cryo-EM global)

Resolution (A)Reliably resolvableDo NOT over-read
< 1.2 (atomic)Individual atoms, H atoms, anisotropic ADPs, alt-confs-
1.2-1.8Side-chain rotamers, ordered waters, alt-confsSurface/loop atoms with high B still uncertain
1.8-2.5 (typical)Backbone plus most side chains; fold solidLong polar rotamers, water networks, weak ligand density are model-dependent
2.5-3.5Backbone trace, secondary structure, domain arrangementSide-chain conformations, exact ligand pose, H-bond geometry are interpretive
> 3.5Overall fold, gross assemblyIndividual side chains largely placed by geometry; treat atomic detail as hypothesis

Decision: R-free and the overfitting gap

ResolutionTypical R-workTypical R-freeConcern if gap (R-free - R-work) exceeds ~
~1.5 A0.13-0.170.15-0.190.03-0.04
~2.0 A0.16-0.200.19-0.240.04-0.05
~2.5-3.0 A0.18-0.240.23-0.290.05-0.06

Absolute R rises with resolution (weaker high-angle reflections are noisier); a gap wider than the row's threshold flags overfitting, a near-zero gap flags work/test contamination. These are heuristics - verify against contemporaneous PDB statistics for the resolution.

Decision: geometry validation targets (MolProbity)

MetricMeasuresTarget (goal)ConcernSource
ClashscoreSerious all-atom overlaps (>0.4 A) per 1000 atomslow, high same-resolution percentilehigh vs same-resolution peersChen 2010
Ramachandran favored% residues in favored phi/psi basins> 98%< 95%Williams 2018
Ramachandran outliers% residues in disallowed phi/psi< 0.2% (goal 0)> 0.5%Williams 2018
Poor rotamers% side chains in disallowed rotamers< 0.3%> 1.5%Williams 2018
RSRZ outliers (X-ray)Residues poorly fitting local density (RSRZ > 2)few, dispersedmany, or clustered in the region of interestGore 2017
MolProbity scoreComposite mapped to a resolution-equivalent<= deposited resolution>> deposited resolutionChen 2010

The poor-rotamer target tightened from ~1.0% to 0.3% with the updated reference distributions (Williams 2018). Read the wwPDB report's PERCENTILE sliders alongside the raw numbers: they rank an entry against the whole archive AND against same-resolution entries (Gore et al. 2017 Structure 25:1916-1927).

Decision: how to validate this model (experimental vs predicted fork)

The model isTrust questionWhat to readAuthoritative route
X-rayIs the region well-determined?resolution + local B + RSRZ; R-free and its gapwwPDB report + phenix.molprobity
Cryo-EMIs the region rigid or a docked guess?LOCAL resolution map, not global; map-model FSC (0.5)EMDB local-resolution + map-model FSC
NMRHow wide is the conformational spread?per-residue RMSD across all modelsvalidate each model; report spread
Predicted (AlphaFold/ESMFold)Is it self-consistent AND in the right biological context?pLDDT bands + PAE blocks; is it apo/monomer where truth is holo/complex?phenix.process_predicted_model (trim + PAE split)

Read Validation Metadata From the mmCIF Header

Goal: Pull resolution, R-work, R-free, and method from a deposited entry and flag the overfitting gap - the numbers Bio.PDB's thin structure.header omits.

Approach: Read the raw mmCIF categories with MMCIF2Dict (which reaches anything in the file), cast the strings, and compare the R-free-minus-R-work gap against a resolution-scaled expectation.

python
from Bio.PDB.MMCIF2Dict import MMCIF2Dict

def read_refinement_metadata(cif_path):
    meta = MMCIF2Dict(cif_path)
    def first(key):
        val = meta.get(key, ['NA'])[0]
        try:
            return float(val)
        except ValueError:
            return val
    method = meta.get('_exptl.method', ['NA'])[0]
    resolution = first('_refine.ls_d_res_high')
    r_work = first('_refine.ls_R_factor_R_work')
    r_free = first('_refine.ls_R_factor_R_free')
    gap = r_free - r_work if isinstance(r_free, float) and isinstance(r_work, float) else None
    # A gap wider than ~0.05 flags overfitting at typical (~2 A) resolution (Kleywegt & Brunger 1996).
    overfit_flag = gap is not None and gap > 0.05
    return {'method': method, 'resolution': resolution, 'r_work': r_work, 'r_free': r_free, 'gap': gap, 'overfit_flag': overfit_flag}

Sanity-Check B-factors Within One Structure

Goal: Locate the least-certain atoms of THIS model so precise-distance claims avoid them - a relative, within-structure read, never a cross-structure comparison.

Approach: Collect per-residue mean B for a chain, then flag residues whose B sits far above the structure's own median as low-confidence.

python
import numpy as np
from Bio.PDB import PDBParser

def bfactor_outliers(structure, chain_id, z_cut=2.0):
    residue_b = {}
    for residue in structure[0][chain_id]:
        if residue.id[0] != ' ':  # skip hetero/water; validate the polymer only
            continue
        residue_b[residue.id] = np.mean([a.get_bfactor() for a in residue])
    vals = np.array(list(residue_b.values()))
    median, mad = np.median(vals), np.median(np.abs(vals - np.median(vals))) + 1e-9
    # Robust z on B: |B - median| / (1.4826*MAD); high B marks disorder/model error, not portable dynamics.
    return {rid: b for rid, b in residue_b.items() if (b - median) / (1.4826 * mad) > z_cut}

Flag Ramachandran Outliers and cis Non-Proline Peptides

Goal: Catch backbone geometry that is almost always a modeling error - residues in disallowed phi/psi and cis peptide bonds that are not proline.

Approach: Get per-residue phi/psi from PPBuilder, coarsely classify against the canonical basins, and compute omega (Ca-C-N-Ca) directly to find cis bonds; a cis assignment at a non-proline is a validation flag until proven by density.

python
import numpy as np
from Bio.PDB import PDBParser, PPBuilder, calc_dihedral

# Coarse general-allowed basins (deg): alpha, beta/PPII, left-handed. Authoritative
# favored/allowed/outlier percentages need MolProbity's rama8000 contours (Williams 2018);
# this only screens gross outliers to decide whether to run MolProbity.
_BASINS = [(-160, -40, -80, 30), (-180, -40, 90, 180), (30, 90, -30, 90)]

def is_rama_allowed(phi, psi):
    d = np.degrees([phi, psi])
    return any(lo_p <= d[0] <= hi_p and lo_s <= d[1] <= hi_s for lo_p, hi_p, lo_s, hi_s in _BASINS)

def geometry_flags(structure):
    outliers, cis_nonpro = [], []
    ppb = PPBuilder()
    for pp in ppb.build_peptides(structure[0]):
        for residue, (phi, psi) in zip(pp, pp.get_phi_psi_list()):
            if phi is not None and psi is not None and not is_rama_allowed(phi, psi):
                outliers.append((residue.get_parent().id, residue.id[1], residue.resname))
        residues = list(pp)
        for prev, curr in zip(residues, residues[1:]):
            omega = calc_dihedral(prev['CA'].get_vector(), prev['C'].get_vector(), curr['N'].get_vector(), curr['CA'].get_vector())
            # omega ~180 = trans, ~0 = cis; |omega| < 30 deg is cis. cis at non-Pro is rare and usually an error.
            if abs(np.degrees(omega)) < 30 and curr.resname != 'PRO':
                cis_nonpro.append((curr.get_parent().id, curr.id[1], curr.resname))
    return {'rama_outliers': outliers, 'cis_nonproline': cis_nonpro}

Run MolProbity for Authoritative Geometry (the real answer)

Goal: Get archive-calibrated clashscore, rotamer, and Ramachandran outlier percentages rather than the coarse Python screen above.

Approach: Shell out to phenix.molprobity (or the MolProbity web service / molprobity.molprobity), which carries the reference contour and rotamer distributions the Python screen cannot reproduce.

python
import subprocess

def run_molprobity(model_path):
    # phenix.molprobity writes molprobity.out with clashscore, rotamer_outliers,
    # ramachandran_outliers, ramachandran_favored, molprobity_score. Parse that file.
    result = subprocess.run(['phenix.molprobity', model_path], capture_output=True, text=True)
    return result.stdout

The coarse Python screen decides WHETHER to run MolProbity; MolProbity (Chen et al. 2010 Acta Crystallogr D 66:12-21; Williams et al. 2018 Protein Sci 27:293-315) gives the numbers to report. For a deposited entry, prefer the pre-computed wwPDB validation report (https://files.rcsb.org/pub/pdb/validation_reports/<xy>/<id>/<id>_validation.xml.gz) - it already carries the percentile sliders and per-residue RSRZ.

Validate a Predicted Model Before Docking or MR

Goal: Turn an AlphaFold/ESMFold model into a trustworthy input by trimming low-confidence residues and reading inter-domain confidence - because pLDDT/PAE bound self-consistency, not correctness.

Approach: Read pLDDT from the B-factor column into bands, read the PAE matrix for domain segmentation, then hand the raw file to phenix.process_predicted_model (which converts pLDDT to a pseudo-B, trims, and splits by PAE-defined domains).

python
import json
import numpy as np
import subprocess
from Bio.PDB import MMCIFParser

_PLDDT_BANDS = [(90, 'very_high'), (70, 'confident'), (50, 'low'), (0, 'very_low')]

def plddt_bands(cif_path):
    # pLDDT rides in the B-factor column but is confidence (high = good), OPPOSITE polarity to a real B-factor.
    structure = MMCIFParser(QUIET=True).get_structure('pred', cif_path)
    counts = {label: 0 for _, label in _PLDDT_BANDS}
    for residue in structure[0].get_residues():
        if 'CA' not in residue:
            continue
        score = residue['CA'].get_bfactor()
        counts[next(label for cut, label in _PLDDT_BANDS if score >= cut)] += 1
    return counts  # a long very_low run usually flags an intrinsically disordered region, not an error

def pae_interdomain_confident(pae_json, block_a, block_b, cutoff=5.0):
    # Off-diagonal PAE (A) between two domain blocks; low = relative orientation trusted, high = independent guess.
    pae = np.array(json.load(open(pae_json))[0]['predicted_aligned_error'])
    return pae[np.ix_(block_a, block_b)].mean() < cutoff

def process_for_mr(model_path):
    # Trims below ~0.7 fractional pLDDT, converts pLDDT->pseudo-B, splits into PAE-defined domains (Oeffner 2022).
    return subprocess.run(['phenix.process_predicted_model', model_path], capture_output=True, text=True).stdout

Report an NMR Ensemble Spread (never average coordinates)

Goal: Turn a multi-model NMR deposit into a per-residue uncertainty map instead of over-claiming precision from model 1.

Approach: Superpose all models on a reference and report per-residue Ca RMSD across the ensemble; wide spread marks flexible or under-restrained regions.

python
import numpy as np
from Bio.PDB import PDBParser

def ensemble_ca_spread(structure, chain_id):
    coords = []
    for model in structure:
        coords.append(np.array([res['CA'].get_coord() for res in model[chain_id] if 'CA' in res]))
    stack = np.stack(coords)  # (n_models, n_residues, 3); assumes consistent residue set across models
    # Per-residue spread = mean distance of each model's Ca from the ensemble mean position.
    return np.linalg.norm(stack - stack.mean(axis=0), axis=2).mean(axis=0)

Secondary-structure validation (DSSP, Bio.PDB.DSSP(model, path, dssp='mkdssp')) needs backbone geometry to place the amide H and computes its own H-bond energy; it processes only the first model and its output drifts across the dssp->mkdssp v2->v4 rewrites, so name the version (Kabsch & Sander 1983 Biopolymers 22:2577-2637). See geometric-analysis for dihedral and DSSP mechanics.

Show full SKILL.md (865 more words)Show less

Common Errors

SymptomCauseFix
"It is a 1.5 A structure so every atom is accurate"Global resolution read as per-region qualityRead the local B-factor and RSRZ for the specific residues used; resolution is a data ceiling
Low R-work reported as proof of a good modelR-work is fit on refinement reflections and rewards overfittingReport R-free (held-out) and the R-free-minus-R-work gap; a wide gap flags overfitting
B-factors of two structures compared at face valueB conflates thermal motion, disorder, and model error and is not portableCompare within one structure only; normalize (z/percentile) before any cross-structure claim
Coloring an AlphaFold model "by B-factor" to infer flexibilitypLDDT rides in the B-factor column with OPPOSITE polarity (high = confident)Read the column as pLDDT bands; high value means high confidence, not high motion
Deleting all low-pLDDT residues as junkLow pLDDT often marks a real intrinsically disordered regionDistinguish disorder (keep, annotate) from misfold; trim only for MR/geometry pipelines
Docking straight into an AlphaFold pocketPocket is apo, side-chain rotamers are the least reliable atoms, may be wrong statePrefer an experimental holo structure; if using the model, ensemble/flexible-side-chain dock and treat hits as hypotheses
Confident predicted model trusted for a complexpLDDT/PAE bound self-consistency, not biological correctnessAsk what context AF could not see (partner, ligand, PTM); a monomer/apo model can be confidently wrong
Cryo-EM peripheral domain trusted at the headline resolutionGlobal resolution hides a low-local-resolution flexible armConsult the EMDB local-resolution map; treat low-local-res regions as docked hypotheses
Quoting FSC 0.5 as the cryo-EM resolution0.143 is the half-map criterion; 0.5 is the map-vs-model curveUse gold-standard half-map FSC at 0.143 for resolution; 0.5 for checking the model against the map
Averaging NMR model coordinates into one structureThe mean of two valid conformers is physically impossible (distorted bonds, clashes)Compute on each model and report the spread, or pick a representative/medoid model
cis peptide flagged everywhere, or missed entirelyomega near 0 (cis) vs 180 (trans) not checked; cis-Pro is common but cis non-Pro is rareCompute omega; treat cis non-proline as a validation flag pending density, cis-Pro as plausible
Python Ramachandran percentages reported as authoritativeCoarse basin boxes are not MolProbity's rama8000 contoursUse the screen to decide whether to run phenix.molprobity; report MolProbity's numbers
resolution/R-free come back None from structure.headerBio.PDB's header dict is thinRead _refine.ls_d_res_high, _refine.ls_R_factor_R_free, _exptl.method via MMCIF2Dict
  • structure-io - Read resolution/R-free via MMCIF2Dict and fetch the biological assembly the validation applies to
  • structure-navigation - Resolve altlocs, insertion codes, and multi-model NMR files before validating per-model
  • geometric-analysis - Compute the dihedrals and DSSP secondary structure this skill validates; measure only after validation passes
  • structure-modification - Trim low-pLDDT residues or edit B-factors once a predicted model is validated
  • structure-preparation - Add hydrogens, protonation, and missing atoms after validation and before docking/MD
  • alphafold-predictions - Download the AlphaFold model plus its PAE JSON that this skill reads for confidence
  • modern-structure-prediction - Reconcile a re-run prediction with pLDDT/PAE/pTM when the AFDB entry is untrustworthy
  • interface-analysis - Validate the assembly before interpreting an interface that only exists in it
  • database-access/uniprot-access - Map validated residues back to a UniProt reference sequence

References

  • Ramachandran GN, Ramakrishnan C, Sasisekharan V (1963). Stereochemistry of polypeptide chain configurations. J Mol Biol 7:95-99. DOI 10.1016/S0022-2836(63)80023-6.
  • Brunger AT (1992). Free R value: a novel statistical quantity for assessing the accuracy of crystal structures. Nature 355(6359):472-475. DOI 10.1038/355472a0.
  • Kabsch W, Sander C (1983). Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features. Biopolymers 22(12):2577-2637. DOI 10.1002/bip.360221211.
  • Kleywegt GJ, Brunger AT (1996). Checking your imagination: applications of the free R value. Structure 4(8):897-904. DOI 10.1016/S0969-2126(96)00097-4.
  • Rosenthal PB, Henderson R (2003). Optimal determination of particle orientation, absolute hand, and contrast loss in single-particle electron cryomicroscopy. J Mol Biol 333(4):721-745. DOI 10.1016/j.jmb.2003.07.013.
  • Chen VB, Arendall WB III, Headd JJ, Keedy DA, Immormino RM, Kapral GJ, Murray LW, Richardson JS, Richardson DC (2010). MolProbity: all-atom structure validation for macromolecular crystallography. Acta Crystallogr D 66(1):12-21. DOI 10.1107/S0907444909042073.
  • Read RJ, Adams PD, Arendall WB III, Brunger AT, Emsley P, Joosten RP, Kleywegt GJ, Krissinel EB, Luetteke T, Otwinowski Z, Perrakis A, Richardson JS, Sheffler WH, Smith JL, Tickle IJ, Vriend G, Zwart PH (2011). A new generation of crystallographic validation tools for the Protein Data Bank. Structure 19(10):1395-1412. DOI 10.1016/j.str.2011.08.006.
  • Gore S, Sanz Garcia E, Hendrickx PMS, Gutmanas A, Westbrook JD, Yang H, Feng Z, Baskaran K, Berrisford JM, et al. (2017). Validation of structures in the Protein Data Bank. Structure 25(12):1916-1927. DOI 10.1016/j.str.2017.10.009.
  • Williams CJ, Headd JJ, Moriarty NW, Prisant MG, Videau LL, Deis LN, Verma V, Keedy DA, Hintze BJ, Chen VB, Jain S, Lewis SM, Arendall WB III, Snoeyink J, Adams PD, Lovell SC, Richardson JS, Richardson DC (2018). MolProbity: more and better reference data for improved all-atom structure validation. Protein Sci 27(1):293-315. DOI 10.1002/pro.3330.
  • Jumper J, Evans R, Pritzel A, et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596(7873):583-589. DOI 10.1038/s41586-021-03819-2.
  • Oeffner RD, Croll TI, Millan C, Poon BK, Schlicksup CJ, Read RJ, Terwilliger TC (2022). Putting AlphaFold models to work with phenix.process_predicted_model and ISOLDE. Acta Crystallogr D 78(11):1303-1314. DOI 10.1107/S2059798322010026.
  • Montelione GT, Nilges M, Bax A, et al. (2013). Recommendations of the wwPDB NMR Validation Task Force. Structure 21(9):1563-1570. DOI 10.1016/j.str.2013.07.021.

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in structural-biology/structure-validation of GPTomics/bioSkills.

  • SKILL.md
  • examples/validate_structure.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Structural Biology Structure Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Structural Biology Structure Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Structural Biology Structure Validation this skillGPTomics/bioSkills1.2k1 repos~5.9kAutomated safety check: PassMIT
Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills3.2k2 repos~1.2kAutomated safety check: PassApache-2.0
Alphafoldadaptyvbio/protein-design-skills1643 repos~1.2kAutomated safety check: PassMIT
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT
Chaiadaptyvbio/protein-design-skills1643 repos~1.5kAutomated safety check: PassMIT
Bio DB ToolsDrugClaw/DrugClaw125—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Alphafold

    adaptyvbio/protein-design-skills

    Validate protein designs using AlphaFold2 structure prediction.

    164 GitHub starsUsed in 3 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Chai

    adaptyvbio/protein-design-skills

    Structure prediction using Chai-1, a foundation model for molecular structure.

    164 GitHub starsUsed in 3 repos~1.5k tokens
    Research & ScienceAuto-check passed
  • Bio DB Tools

    DrugClaw/DrugClaw

    Query public biology databases and APIs including UniProt, RCSB PDB, AlphaFold DB, ClinVar, dbSNP, gnomAD, Ensembl, GEO, InterPro, KEGG, OpenTargets, Reactome, and STRING.

    125 GitHub stars~1.4k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~6.3k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Structural Biology Structure Validation

What does Bio Structural Biology Structure Validation do?

Judges whether a macromolecular model (or a region of it) is reliable enough to build on, using resolution, R-free, B-factors, MolProbity geometry, and predicted-model confidence with Bio.PDB. Bio Structural Biology Structure Validation is an agent skill from GPTomics/bioSkills.PDB.

When should I use Bio Structural Biology Structure Validation?

Bio Structural Biology Structure Validation fits situations like: deciding if a structure; A specific region is trustworthy before docking/mechanism/measurement; reading resolution; R-work vs R-free and the R-free-minus-R-work overfitting gap.

How do I install Bio Structural Biology Structure Validation in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a claude-code`. Or copy the skill folder (structural-biology/structure-validation in GPTomics/bioSkills) into .claude/skills/bio-structural-biology-structure-validation in your project. Claude Code loads it when a task matches its description.

How do I install Bio Structural Biology Structure Validation in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a codex`. Or copy the skill folder (structural-biology/structure-validation in GPTomics/bioSkills) into .agents/skills/bio-structural-biology-structure-validation in your project. Codex loads it when a task matches its description.

Can I use Bio Structural Biology Structure Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-structural-biology-structure-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-structural-biology-structure-validation, .gemini/skills/bio-structural-biology-structure-validation, .github/skills/bio-structural-biology-structure-validation and .opencode/skills/bio-structural-biology-structure-validation in your project.

What does Bio Structural Biology Structure Validation need to run?

Going by SKILL.md and its folder, Bio Structural Biology Structure Validation needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Structural Biology Structure Validation access the network?

SKILL.md names 1 domain. In commands or code: files.rcsb.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Structural Biology Structure Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Structural Biology Structure Validation use?

Bio Structural Biology Structure Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Structural Biology Structure Validation use?

About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Structural Biology Structure Validation?

Skills that share tags, products or a category with Bio Structural Biology Structure Validation: Alphafold Database Fetch And Analyze (google-deepmind/science-skills, 3.2k stars), Alphafold (adaptyvbio/protein-design-skills, 164 stars), Biopipelines (locbp-uzh/biopipelines, 109 stars) and Chai (adaptyvbio/protein-design-skills, 164 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Structural Biology Structure Validation?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.