Setup Workshop
brevdev/workshop-build-an-agent
This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a.
Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM.
$ npx skills add NVIDIA/skills --skill msa-search-nim -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills msa-search-nim --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bionemo-msa-search-nim .claude/skills/msa-search-nim && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "msa-search-nim" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/bionemo-msa-search-nim into .claude/skills/msa-search-nim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msa-search-nim", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/bionemo-msa-search-nimType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill msa-search-nim -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills msa-search-nim --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/bionemo-msa-search-nim .agents/skills/msa-search-nim && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "msa-search-nim" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/bionemo-msa-search-nim into .agents/skills/msa-search-nim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msa-search-nim", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill msa-search-nim -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills msa-search-nim --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/bionemo-msa-search-nim .cursor/skills/msa-search-nim && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "msa-search-nim" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/bionemo-msa-search-nim into .cursor/skills/msa-search-nim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msa-search-nim", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/bionemo-msa-search-nim--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill msa-search-nim -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills msa-search-nim --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/bionemo-msa-search-nim .gemini/skills/msa-search-nim && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "msa-search-nim" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/bionemo-msa-search-nim into .gemini/skills/msa-search-nim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msa-search-nim", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills msa-search-nimInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill msa-search-nim -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/bionemo-msa-search-nim .github/skills/msa-search-nim && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "msa-search-nim" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/bionemo-msa-search-nim into .github/skills/msa-search-nim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msa-search-nim", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill msa-search-nim -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills msa-search-nim --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/bionemo-msa-search-nim .opencode/skills/msa-search-nim && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "msa-search-nim" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/bionemo-msa-search-nim into .opencode/skills/msa-search-nim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "msa-search-nim", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
msa-search-nimGenerate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM.
Msa Search Nim is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. Use for homolog search, UniRef30/ColabFold env searches, A3M or FASTA alignments, paired MSA search for complexes, PDB70 structural templates, hosted NVIDIA API calls, or local Docker deployment. For local deployment, download the databases in parallel with aria2c and launch via NIMMODELNAME (the recommended default fast path, ~14 min vs over 80 min for the built-in downloader); a plain docker run uses the slow…
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts and reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yml` and `evals/config.yml`). Compatibility notes: requests=2.28
It sits in DevOps & Cloud, covering Containers and Deployment. It works with NVIDIA AI Platform and Docker. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteAskUserQuestionFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
dockercurlpython3jqpythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
health.api.nvidia.comapi.ngc.nvidia.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NGC_API_KEYNVIDIA_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
requests>=2.28
From compatibility in the SKILL.md frontmatter.
Msa Search Nim loads about 4.6k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 139 tokens; SKILL.md has 1,397 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
[ -f .env ] && . ./.envallowed-tools: Bash, Read, Write, AskUserQuestionAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,397 words, ~4,555 tokens.
.claude/skills/msa-search-nim/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.Generate protein MSAs with GPU-accelerated MMSeqs2. Use this guide for first-pass hosted/local usage; load supplemental files only when needed:
references/api.md: exact endpoints, schemas, Docker flags, response fields.references/science.md: MSA purpose, pairing/templates, limits, handoffs.references/parameters.md: database, pairing, depth, and template tuning.references/validation.md: alignment, template, and artifact checks.references/examples.md: compact hosted/local request patterns.Ask only when context is unclear:
Hosted NVIDIA API or local Docker NIM?
https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predicthttps://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predicthttp://localhost:8000/biology/colabfold/msa-search/predicthttp://localhost:8000/biology/colabfold/msa-search/paired/predicthttp://localhost:8000/biology/colabfold/msa-search/structure-templates/predictLocal inference paths do not include /v1/. Hosted requests use Authorization: Bearer $NGC_API_KEY. Supported local Docker
startup uses NGC_API_KEY (or NVIDIA_API_KEY via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with -e NGC_API_KEY. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
The hosted template path returned HTTP 404 in validation, so use local Docker
for template search unless the hosted docs/service changes.
Default local deployment = parallel download + NIM_MODEL_NAME. The first recipe below
is the one to use for real workflows. It downloads the database(s) with a range-parallel
downloader (aria2c) and starts the NIM against those files — ~14 min for UniRef30 vs >80 min
for the NIM's built-in downloader (measured, H100). Do not reach for the plain docker run
(the "Fallback" subsection) unless you only want a databases:pdb70 smoke test or you
deliberately want the NIM to manage its own blob cache.
Local setup requires a GPU. Size the NVMe volume to the profile you pick (UniRef30 ~490 GB;
full set ~1.4 TB). For setup answers, include env preflight, docker login, the parallel
download, NIM_MODEL_NAME launch, readiness, and then no-auth local inference. Do not invent a
cache default or drop the NVIDIA_API_KEY fallback.
# --- env preflight (do not drop the NVIDIA_API_KEY fallback) ---
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${DB_DIR:=/data/fast-db}" # where the parallel download lands
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
# --- 1) pick the DB version(s) you need (paired/complex work = uniref30 only) ---
DB_VERSION=uniref30_2302-m18v1
command -v aria2c >/dev/null || { echo "aria2c required; install it (e.g. apt-get install -y aria2) and re-run"; exit 1; }
mkdir -p "$DB_DIR"
# --- 2) parallel download from NGC (see "Parallel Download" section for the all-DB loop) ---
curl -fsS -H "Authorization: Bearer $NGC_API_KEY" \
"https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/${DB_VERSION}/files" \
-o /tmp/files.json
DB_DIR="$DB_DIR" python3 - <<'PY'
import json, os
d = json.load(open("/tmp/files.json")); dbdir = os.environ["DB_DIR"]; lines = []
for url, path in zip(d["urls"], d["filepath"]):
lines += [url.strip(), f" dir={dbdir}", f" out={path}"]
open("/tmp/aria.in", "w").write("\n".join(lines) + "\n")
PY
aria2c -i /tmp/aria.in --max-concurrent-downloads=4 --max-connection-per-server=16 \
--split=16 --min-split-size=1M --continue=true --file-allocation=none
# --- 3) launch the NIM against the downloaded files (skips the slow built-in download) ---
docker run -d --name msa-search --runtime=nvidia --gpus all \
-e NGC_API_KEY \
-e NIM_MODEL_NAME=/databases \
-v "${DB_DIR}:/databases" \
-p 8000:8000 \
nvcr.io/nim/colabfold/msa-search:2Readiness:
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; doneIf the DB is already present in $DB_DIR, skip steps 1-2 — the launch alone is a ~20 s warm
start. See "Parallel Download For Any Database Set" for the multi-database (databases:all)
loop and the full rationale.
Use this only for a quick databases:pdb70 smoke test, or when you specifically want the NIM
to manage its own blob cache. It uses the built-in downloader, which is slow on large profiles
(UniRef30 stalled past 80 min in testing). Pin the smallest profile with NIM_MODEL_PROFILE
(see "Faster Startup") so it does not fetch the full 1.4 TB.
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
mkdir -p "${LOCAL_NIM_CACHE}"; chmod 755 "${LOCAL_NIM_CACHE}"
docker run --rm --name msa-search \
--runtime=nvidia --gpus all \
-e NGC_API_KEY \
-e NIM_MODEL_PROFILE=<hash-from-list-model-profiles> \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/colabfold/msa-search:2The full database download is ~1.4 TB and can take well over an hour on first launch. If you only need some databases, select a task-specific profile so the NIM downloads just those. This is the single biggest lever on local startup time.
List the profiles your image actually ships (hashes change between releases — never hardcode them):
docker run --rm --entrypoint list-model-profiles nvcr.io/nim/colabfold/msa-search:2Then pass the chosen hash with NIM_MODEL_PROFILE:
docker run --rm --name msa-search \
--runtime=nvidia --gpus all \
-e NGC_API_KEY \
-e NIM_MODEL_PROFILE=<hash-from-list-model-profiles> \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/colabfold/msa-search:2Profiles available in this image (confirm hashes with list-model-profiles):
| Profile tags | Databases | Best for | Storage |
|---|---|---|---|
databases:pdb70 | PDB70 | Quick testing / smoke check | ~100 MB |
databases:uniref30 | UniRef30 | Paired MSA search for complexes — UniRef30 is the only DB used for species-based pairing | ~500 GB |
databases:uniref30,pdb70,pdb | UniRef30 + PDB70 + PDB structures | Structural template search | ~700 GB |
databases:all (default) | UniRef30 + ColabFold envdb + PDB70 + PDB100 + PDB structures | Full sensitivity, all databases | ~1.2 TB |
Verify the loaded profile after readiness:
curl -s localhost:8000/v1/metadata | jqNotes:
databases parameter only selects among databases already
downloaded; it does NOT change what is fetched at startup. Startup footprint is set by
NIM_MODEL_PROFILE alone.colabfold_envdb_202108 has no taxonomy and cannot
be used for pairing, so databases:uniref30 is the correct, smallest profile for
complex/paired workflows — it skips the envdb, the largest part of the full set.databases:all;
there is no envdb-inclusive profile smaller than the full set.To use a single manually downloaded database (or your own MMSeqs2 DB), download it from NGC
and point the NIM at the mount with NIM_MODEL_NAME instead of a profile:
ngc registry model download-version nim/colabfold/msa-search:uniref30_2302-m18v1
# then mount the directory and set -e NIM_MODEL_NAME=/databasesNIM_MODEL_NAME replaces the profile databases entirely — the NIM uses only what is
under that directory (discovered by scanning for **/*.idx). Mount multiple databases under
one parent to combine them. NGC-downloaded databases are pre-indexed for GPU Server; custom
databases must be indexed with mmseqs createindex first. Individually downloadable NGC model
versions: uniref30_2302-m18v1, colabfold_envdb_202108-m18v1, pdb70_220313-m18v1,
pdb100_230517-m18v1, pdb_20251028_zip-m18v1.
This is the recommended way to download the databases at all — for any profile,
including the full databases:all set. Task-specific profiles cut what you
download; this parallel downloader cuts how long that download takes. Use it whether
you need one database or all of them. The gain is largest for the databases:uniref30 profile, which is ~490 GB
dominated by two very large files (a ~241 GB GPU index and a ~134 GB sequence DB).
The NIM's built-in downloader parallelizes across files (max_parallel_files=10) but pulls
each file over roughly one connection. The NGC CDN throttles a single connection to ~20–25
MB/s, so while the downloader is fetching one of the two giant files, most of its parallel
slots sit idle and throughput collapses to that single-flow rate. Measured on an H100 node,
the built-in path did not reach /health/ready in over 80 minutes.
A range-parallel downloader splits each file into many byte-range segments (the NGC CDN
advertises accept-ranges: bytes), so a single 241 GB file is pulled over 16 connections at
once — ~15× the single-flow rate. Same node, aria2c fetched the full ~490 GB in ~13.5
minutes.
Workflow (download once with aria2, then start the NIM against the files via NIM_MODEL_NAME):
# 1) Get presigned file URLs for the individual database model version from NGC.
# (Requires NGC_API_KEY. The response arrays `urls` and `filepath` are positionally paired.)
curl -s -H "Authorization: Bearer $NGC_API_KEY" \
'https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/uniref30_2302-m18v1/files' \
-o files.json
# 2) Build an aria2 input file (URL + target filename per entry) and download in parallel.
python3 - <<'PY'
import json
d = json.load(open("files.json"))
lines = []
for url, path in zip(d["urls"], d["filepath"]):
lines += [url.strip(), " dir=/data/fast-db", f" out={path}"]
open("aria.in", "w").write("\n".join(lines) + "\n")
PY
aria2c -i aria.in \
--max-concurrent-downloads=4 --max-connection-per-server=16 --split=16 \
--min-split-size=1M --continue=true --file-allocation=none
# 3) Start the NIM against the downloaded directory. NIM_MODEL_NAME makes the NIM discover
# databases by scanning for **/*.idx, bypassing the profile/blob cache entirely.
docker run -d --name msa-search --runtime=nvidia --gpus all \
-e NGC_API_KEY \
-e NIM_MODEL_NAME=/databases \
-v /data/fast-db:/databases \
-p 8000:8000 \
nvcr.io/nim/colabfold/msa-search:2For all databases (equivalent to databases:all), repeat step 1 for each individual DB
version and download them into sibling directories under one parent, then point
NIM_MODEL_NAME at that parent — the NIM discovers every DB by scanning **/*.idx:
# fetch each DB's file list into /data/all-db/<db>/ ... then one aria2c per list, e.g.:
for V in uniref30_2302-m18v1 colabfold_envdb_202108-m18v1 pdb70_220313-m18v1 \
pdb100_230517-m18v1 pdb_20251028_zip-m18v1; do
curl -s -H "Authorization: Bearer $NGC_API_KEY" \
"https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/$V/files" \
-o "files_$V.json"
# build an aria2 input from files_$V.json (dir=/data/all-db) and run aria2c on it
done
# then launch once against the parent:
# docker run -d ... -e NIM_MODEL_NAME=/databases -v /data/all-db:/databases ...The per-connection CDN throttle is the same for every database, so parallel download helps the full set proportionally — the more you download, the more absolute time it saves.
Notes:
files.json.uniref30_2302/…); the
filepath values already encode it. The NIM needs the .idx file plus its companion files
and the small .UNIREF30_READY / *.tar.gz.unpacked markers.--split / --max-connection-per-server
helps only up to the node's aggregate egress ceiling.fast-db directory) and mount it on future nodes for a ~20 s warm start with no re-download.Use exact case-sensitive database names and response keys.
For a hosted standard search, run the bundled client from this skill's directory. It submits the real request, validates both database results, and saves the raw JSON and A3M files. Choose a new output directory for each run:
python scripts/hosted_search.py \
--sequence SGSMKTAISLPDETFDRVSRRASELGMSRSEFFTKAAQR \
--output-dir msa-outputThe client reads NGC_API_KEY or NVIDIA_API_KEY from the environment. It permits
at most two requests, each with a 10-second connection timeout and a 300-second
read timeout, with five seconds between attempts. If it exits nonzero, report the
service failure and stop. Do not restart it repeatedly, extend timeouts beyond the
task budget, or replace the missing response with synthetic alignments.
The underlying request format, also usable with a running local NIM, is:
import os
import requests
HOSTED = True
url = (
"https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict"
if HOSTED else "http://localhost:8000/biology/colabfold/msa-search/predict"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.getenv('NGC_API_KEY')}"
payload = {
"sequence": "SGSMKTAISLPDETFDRVSRRASELGMSRSEFFTKAAQR",
"databases": ["Uniref30_2302", "colabfold_envdb_202108"],
"e_value": 0.0001,
"output_alignment_formats": ["a3m"],
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()Use paired search for protein complexes; payload field is sequences plural,
and output is alignments_by_chain.
url = (
"https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict"
if HOSTED else "http://localhost:8000/biology/colabfold/msa-search/paired/predict"
)
payload = {
"sequences": [chain_a_sequence, chain_b_sequence],
"e_value": 0.0001,
"output_alignment_formats": ["a3m"],
}Use local Docker for structural templates. Set max_msa_sequences=500 unless
NIM_GLOBAL_MAX_MSA_DEPTH was changed.
url = "http://localhost:8000/biology/colabfold/msa-search/structure-templates/predict"
headers = {"Content-Type": "application/json"}
payload = {
"sequence": "VLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQVKGHGKKVADALTNAVA",
"structural_template_databases": ["pdb70_220313"],
"max_structures": 20,
"max_msa_sequences": 500,
}# Standard MSA: result["alignments"][database][format]["alignment"]
for db_name, formats in result.get("alignments", {}).items():
for fmt_name, data in formats.items():
with open(f"msa_{db_name}.{fmt_name}", "w", encoding="utf-8") as handle:
handle.write(data["alignment"])
# Paired MSA: one alignment set per chain
for chain_id, chain_data in result.get("alignments_by_chain", {}).items():
for db_name, formats in chain_data.items():
for fmt_name, data in formats.items():
with open(f"msa_chain_{chain_id}_{db_name}.{fmt_name}", "w", encoding="utf-8") as handle:
handle.write(data["alignment"])
# Template search: save mmCIF structures and M8 hit tables
for name, cif in result.get("structures", {}).items():
open(f"template_{name}.cif", "w", encoding="utf-8").write(cif)
for name, hit_table in result.get("search_hits", {}).items():
open(f"template_hits_{name}.m8", "w", encoding="utf-8").write(hit_table)A3M output can feed OpenFold3, AlphaFold2, or RoseTTAFold. For alignment depth,
template, and sequence sanity checks, read references/validation.md.
X works since v2.3.0.max_msa_sequences: 1-500; local GPU server default must match
NIM_GLOBAL_MAX_MSA_DEPTH./v1/ prefix.LOCAL_NIM_CACHE.health.api.nvidia.com. The published standard
MSA example uses synchronous POST; a pending response needs a documented
service-specific completion mechanism before it can count as a result.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 14 other files (scripts, references) in skills/bionemo-msa-search-nim of NVIDIA/skills.
Open the folder on GitHubat commit dfdd080
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/skills, which our catalogue first saw on October 7, 2026.
Msa Search Nim next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Msa Search Nim this skillNVIDIA/skills | 3.5k | 1 repos | ~4.6k | Automated safety check: Notes | Apache-2.0 | |
| Setup Workshopbrevdev/workshop-build-an-agent | 146 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | |
| GreptimeDB Dev Docker ImageGreptimeTeam/greptimedb | 6.7k | — | ~4k | Automated safety check: Notes | Apache-2.0 | |
| Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit | 260 | 6 repos | ~1.1k | Automated safety check: Notes | Custom licence | |
| LangBot Deployment Guidelangbot-app/LangBot | 18k | — | ~1.2k | Automated safety check: Notes | Apache-2.0 | |
| Reflexo ReleaseMyriad-Dreamin/typst.ts | 1.2k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 |
brevdev/workshop-build-an-agent
This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a.
GreptimeTeam/greptimedb
Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
langbot-app/LangBot
Deploys and configures a LangBot instance with Docker Compose or Kubernetes, covering config.yaml, the Box sandbox runtime, the plugin runtime and the global API key.
Myriad-Dreamin/typst.ts
Guide Reflexo/typst.ts release preparation and operator handoffs.
Mr-funny/hbg-classical-poem-silk-video
Turn Chinese classical poems and ci into coherent vertical Chinese-art videos with poem-driven scene grouping, GPT ImageGen stills, Docker-only Gemini I2V, retained model-generated ambience, Gemini…
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. Msa Search Nim is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM.
Msa Search Nim fits situations like: uniRef30/ColabFold env searches; FASTA alignments; paired MSA search for complexes; PDB70 structural templates.
Run `npx skills add NVIDIA/skills --skill msa-search-nim -a claude-code`. Or copy the skill folder (skills/bionemo-msa-search-nim in NVIDIA/skills) into .claude/skills/msa-search-nim in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill msa-search-nim -a codex`. Or copy the skill folder (skills/bionemo-msa-search-nim in NVIDIA/skills) into .agents/skills/msa-search-nim in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill msa-search-nim -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/msa-search-nim, .gemini/skills/msa-search-nim, .github/skills/msa-search-nim and .opencode/skills/msa-search-nim in your project.
Going by SKILL.md and its folder, Msa Search Nim needs Python for the scripts in its folder, the command-line tools its instructions call (docker, curl, python3, jq and python) and credentials named NGC_API_KEY and NVIDIA_API_KEY. Our summary lists: Python 3; Docker; A credential in NGC_API_KEY; A credential in NVIDIA_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, AskUserQuestion. Compatibility (from SKILL.md): requests>=2.28.
SKILL.md names 2 domains. In commands or code: health.api.nvidia.com and api.ngc.nvidia.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Msa Search Nim is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Msa Search Nim: Setup Workshop (brevdev/workshop-build-an-agent, 146 stars), GreptimeDB Dev Docker Image (GreptimeTeam/greptimedb, 6.7k stars), Senior DevOps Toolkit (maslennikov-ig/claude-code-orchestrator-kit, 260 stars) and LangBot Deployment Guide (langbot-app/LangBot, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.