GitHub Review Iteration
prisma/orm
Runs a loop on a GitHub pull request: fetch review state, triage comments into actions, implement them and resolve threads, repeating until nothing actionable is left.
Implements a new bioinformatics tool wrapper in proto-tools using a parallelized agent pipeline.
$ npx skills add evo-design/proto-tools --skill implement-tool -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install evo-design/proto-tools implement-tool --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/evo-design/proto-tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/implement-tool .claude/skills/implement-tool && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "implement-tool" agent skill from https://github.com/evo-design/proto-tools/tree/main/.claude/skills/implement-tool into .claude/skills/implement-tool/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement-tool", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/evo-design/proto-tools/tree/main/.claude/skills/implement-toolType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add evo-design/proto-tools --skill implement-tool -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install evo-design/proto-tools implement-tool --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/evo-design/proto-tools.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/implement-tool .agents/skills/implement-tool && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "implement-tool" agent skill from https://github.com/evo-design/proto-tools/tree/main/.claude/skills/implement-tool into .agents/skills/implement-tool/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement-tool", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add evo-design/proto-tools --skill implement-tool -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install evo-design/proto-tools implement-tool --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/evo-design/proto-tools.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/implement-tool .cursor/skills/implement-tool && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "implement-tool" agent skill from https://github.com/evo-design/proto-tools/tree/main/.claude/skills/implement-tool into .cursor/skills/implement-tool/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement-tool", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/evo-design/proto-tools.git --path .claude/skills/implement-tool--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add evo-design/proto-tools --skill implement-tool -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install evo-design/proto-tools implement-tool --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/evo-design/proto-tools.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/implement-tool .gemini/skills/implement-tool && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "implement-tool" agent skill from https://github.com/evo-design/proto-tools/tree/main/.claude/skills/implement-tool into .gemini/skills/implement-tool/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement-tool", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install evo-design/proto-tools implement-toolInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add evo-design/proto-tools --skill implement-tool -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/evo-design/proto-tools.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/implement-tool .github/skills/implement-tool && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "implement-tool" agent skill from https://github.com/evo-design/proto-tools/tree/main/.claude/skills/implement-tool into .github/skills/implement-tool/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement-tool", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add evo-design/proto-tools --skill implement-tool -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install evo-design/proto-tools implement-tool --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/evo-design/proto-tools.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/implement-tool .opencode/skills/implement-tool && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "implement-tool" agent skill from https://github.com/evo-design/proto-tools/tree/main/.claude/skills/implement-tool into .opencode/skills/implement-tool/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "implement-tool", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
implement-toolImplements a new bioinformatics tool wrapper in proto-tools using a parallelized agent pipeline.
Implement Tool is an agent skill from evo-design/proto-tools. Implements a new bioinformatics tool wrapper in proto-tools using a parallelized agent pipeline. Orchestrates phases: Research, Contract (core tool file), Fan-out (5 parallel subagents), Verify, then the decisive gates — Config Field Audit, Temp Integration & Stress, Docs Verification, Self-Audit & Full PR Review — and Ship. Use when creating tools from GitHub issues, wrapping models, or implementing new tool wrappers end-to-end.
Its SKILL.md is about 15k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `PATTERNS.md` and `TEMPLATES.md`).
It sits in Research & Science, covering Bioinformatics, Pull requests and Subagents. It works with GitHub. The repository describes itself as: A universal infrastructure layer for generative biology. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit dbc0491. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBashGlobGrepTaskWebFetchWebSearchAskUserQuestionFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitpython3uvghwgetpytestpipjupyterFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comhuggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Implement Tool loads about 15k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 5,353 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Write, Edit, Bash, Glob, Grep, Task, WebFetch, WebSearch, AskUserQuestionAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from evo-design/proto-tools at commit dbc0491, republished under its MIT licence (© evo-design). 5,353 words, ~15,269 tokens.
.claude/skills/implement-tool/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.You are implementing a new bioinformatics tool in the proto-tools codebase. This skill orchestrates the full lifecycle using parallel subagents for speed.
Pipeline overview:
Phase 1: Research → Phase 2: Contract → Phase 3: Fan-out (5 parallel agents) → Phase 4: Verify → Phase 4.5: Config Field Audit → Phase 4.6: Temp Integration & Stress → Phase 4.7: Docs Verification → Phase 4.8: Self-Audit & Full PR Review → Phase 5: ShipThe Phase 4.x gates are decisive and non-negotiable — they catch the failure modes that the fast fan-out reliably misses (invented config parameters, untested real workloads, hallucinated citations/licenses, missing license.yaml/links.yaml). Do not skip them to ship faster. Ground every decision in upstream source or a credible reference — never assume.
Key principle: The core tool file (Input/Config/Output + @tool() + run_*()) is the contract everything else depends on. Write it first (sequential), then fan out to subagents (parallel).
CRITICAL: Standalone scripts run in isolated environments and MUST NOT import from proto_tools:
from proto_tools.utils import ... — will fail at runtimefrom proto_tools.entities import ... — will fail at runtimefrom standalone_helpers import get_subprocess_device_env — published on PYTHONPATH (OK)import json, import subprocess, etc. (OK)requirements.txt: import torch, import numpy, etc. (OK)proto_tools in standalone environments (creates circular dependency, breaks isolation)The user provides EITHER:
#<issue-number> or https://github.com/evo-design/proto-tools/issues/<issue-number>)If a GitHub issue is provided:
gh issue view <number> --json title,body,labelsIf direct tool info is provided:
Output of Phase 0: A clear specification:
esmif)inverse_folding)sample, score)Goal: Understand the tool's API, inputs, outputs, and dependencies.
Steps:
Fetch source repo README — Use WebFetch or gh to read the README of the research repo. Extract:
LICENSE/COPYING file (not just the README badge) and the weights license if it differs. Record the SPDX id (or note it's non-standard) and the file URL. Feeds license.yaml and the README License callout.links.yaml. Verify each resolves.from_pretrained → no weight code needed (HF_HOME set automatically). Example: tools/masked_models/esm2/PROTO_MODEL_CACHE pattern. Example: tools/inverse_folding/fampnn/standalone/setup.shPROTO_MODEL_CACHE pattern. Example: tools/inverse_folding/ligandmpnn/standalone/setup.shproto_check_gated_hf_repo for those). For not-fetchable assets, the user must place files on disk themselves; setup.sh doesn't fetch. → precheck at the top of setup.sh: when the asset is missing, print [proto-tools] ASSET_NOT_AVAILABLE: <toolkit>:<asset_kind> and exit 64. Tests skip cleanly on hosts without weights; misconfigured PROTO_<TOOLKIT>_WEIGHTS_DIR still fails (exit 1). Example: tools/structure_prediction/x3dna/standalone/setup.sh. See PATTERNS.md → "standalone/setup.sh for tools whose assets are NOT automatically downloaded".setup.sh and record the restrictions in license.yaml (commercial_use, redistribution, weights.text) instead of weights.access. tools/structure_prediction/alphafold3/standalone/setup.sh does this: public download, non-commercial-only terms.Read key source files — Find and read the research repo's inference/prediction scripts to understand:
Read existing tools in the same category — Find the reference implementation:
ls proto_tools/tools/{category}/Read the reference tool's:
@tool() decorated function)standalone/inference.py (or run.py)standalone/setup.shstandalone/requirements.txtstandalone/python_version.txt (required — the Python version pin for the env)__init__.py (at tool and category level)README.mdcite.bibtests/{category}_tests/Check for shared data models — If the category has a shared_data_models.py, read it. The new tool should extend these base classes rather than creating new ones. Input and Output classes may be reused directly across sibling tools. The Config must be per-tool: each tool needs its own config class so it carries its own tool_key (stamped at registration and checked by test_config_carries_its_tool_key). If the shape is identical to a sibling's, subclass it with a one-line docstring rather than registering the shared class twice.
Output of Phase 1: Mental model of:
Goal: Write the core tool file that defines the API surface. Everything else depends on this.
This phase is sequential — no subagents. The orchestrator writes this directly.
Steps:
Placeholder glossary (used throughout this skill — all distinct concepts, not interchangeable):
{toolkit} — snake_case directory name shared by all tools in the family (e.g., evo2, pyrosetta). Strict — drives the directory path and the dispatch/worker identifier.{tool_key} — kebab-case registration key, {toolkit}-{suffix} (e.g., evo2-sample, pyrosetta-energy). Strict — the value passed to @tool(key=...).{tool_key_snake} — snake_case form of {tool_key} (e.g., evo2_sample). Strict — the core tool file name, the run_* function, and the test file.{ToolName} — PascalCase class-name prefix for this tool's Input / Config / Output (e.g., Evo2Sample, ESMFoldPrediction). Developer's choice — typically the PascalCase of {tool_key}, but pick whatever reads cleanly (e.g., ESMFold over Esmfold) as long as it's specific to this tool.{tool_display_name} — human-readable label (e.g., "Evo 2"). Used in docstrings, commit messages, and PR titles.Create the tool directory structure:
mkdir -p proto_tools/tools/{category}/{toolkit}/standalone
mkdir -p proto_tools/tools/{category}/{toolkit}/examplesTarget file tree:
tools/{category}/
+-- shared_data_models.py # Shared Input/Config/Output base classes (if category has 2+ tools)
+-- {toolkit}/
| +-- __init__.py
| +-- {tool_key_snake}.py # Core tool file (Input/Config/Output + @tool + run_*); one per registered tool
| +-- cite.bib # BibTeX citation (required)
| +-- license.yaml # License metadata (required — enforced by test_license_consistency.py)
| +-- links.yaml # Canonical upstream links (required — github/website/paper/huggingface/...)
| +-- README.md
| +-- examples/
| | +-- example.ipynb # Working example notebook (required)
| +-- helpers.py # [optional] Self-contained plain-type file-format conversion helpers
| +-- standalone/
| | +-- setup.sh
| | +-- run.py OR inference.py # run.py for CPU tools, inference.py for AI models
| | +-- requirements.txt
| | +-- binary_config.py # [optional]
+-- __init__.pyWrite the core tool file proto_tools/tools/{category}/{toolkit}/{tool_key_snake}.py with:
BaseToolInput (or shared base) with Field() — extra="forbid"BaseConfig (or shared base) with ConfigField() — extra="forbid". Use reload_on_change=True on fields that require worker restart (model checkpoint, etc.). Use include_in_key=False on fields that don't affect computation results (device, verbose, timeout are already excluded on BaseConfig; tool-level overrides of device must also set include_in_key=False). include_in_key defaults to True. Use xor_group="<slug>" to mark mutually exclusive sibling fields (see "Mutual-exclusion fields (XOR groups)" below).BaseToolOutput (or shared base) with Field() — extra="ignore" (set by BaseToolOutput; the base also emits a warning for unexpected fields so computed-field JSON round-trips can validate cleanly)@tool() decorator with the 7 required kwargs (key, label, category, input_class, config_class, output_class, description) plus conventionally set uses_gpu. Optional: example_input, device_count, cacheable, stochastic, iterable_input_fields (a list — ["sequences"] for an ordinary iterable, or an index-parallel group like ["complexes", "msas"]), iterable_output_field, metrics_class, gpu_only, post_process_iterablerun_*() function that calls ToolInstance.dispatch()ToolInput = InverseFoldingInput)Metrics subclass (from proto_tools/utils/tool_io.py). Prefer a shared per-category class (MaskedModelScoringMetrics, InverseFoldingScoringMetrics, CausalModelScoringMetrics, etc.) when the metric set matches siblings. Otherwise declare a per-tool <Tool>Metrics(Metrics) subclass in the tool file with a metric_spec: ClassVar[dict[str, MetricSpec]] documenting type/range and a better_values_are direction ("higher", "lower", "in-range", or "context-dependent") for each metric — test_metric_spec_consistency.py requires a valid better_values_are on every metric. Verification lives in each tool's e2e test via assert_metrics_in_spec(result) (helper at tests/tool_infra_tests/_metric_helpers.py).Critical conventions:
{tool_key}): {toolkit}-{suffix} kebab-case (e.g., "esmif-sample")run_{tool_key_snake} (e.g., run_esmif_sample){ToolName}Input/Config/Output, PascalCase (e.g., ESMIFSampleInput)batch_size defaults to 1 for GPU tools@tool decorator handles errorslogging.getLogger(__name__), never print()output_format_options, output_format_default, _export_output()iterable_input_fields and iterable_output_field are both set, len(output.{iterable_output_field}) must equal the (primary) input list length, by position. The framework's per-item cache stitching, dedup, and diversification testing all rely on this. If the tool produces N samples per input (e.g. a num_sequences_per_structure or num_designs config parameter), bundle the N samples inside a single per-input wrapper class — do NOT flatten N×inputs into the output list. Canonical patterns to follow: ProteinMPNNSequences / LigandMPNNSequences (per-structure sequence bundles, num_sequences_per_structure config) and RFdiffusion3Designs (per-spec design bundles, n_batches * diffusion_batch_size per spec). See PATTERNS.md → "Iterable cardinality" for the full rule and code shape.iterable_input_fields. A per-item value that is usually uniform can be a broadcastable field — InputField(..., broadcastable=True), declared as list[...] | None, added to iterable_input_fields, normalized via broadcast_parallel_fields() (one value broadcasts to all, N is strict 1:1). Broadcast is opt-in per field and only when a single value applied to every item is genuinely correct (safe for binder_chain; wrong for msas). See PATTERNS.md → "Input = per-item, Config = shared".cpus_per_instance opt-in: BaseConfig.cpus_per_instance defaults to None — every CPU tool stays off ToolPool's CPU scheduler and runs as a single direct call. Most CPU tools should leave this alone. Only opt in (override to a positive int) when per-call work is heavy enough to amortize spinning up N persistent worker subprocesses — each holds its own venv in RAM and pays a startup tax, so cheap tools (short per-item compute, internal threading, network IO) lose more than they gain. The canonical opt-in is PyRosetta (heavy init, multi-second per pose, embarrassingly parallel poses → cpus_per_instance = 1). GPU tools (gpus_per_instance > 0) ignore cpus_per_instance entirely.Inherited field audit: When reusing a shared base config (e.g., InverseFoldingConfig) or base input, enumerate every inherited field and verify the target model can implement it. For each unsupported field, either implement support (e.g., logit masking for excluded_amino_acids) or override the field with a validator that raises ValueError("'{field_name}' is not supported by {tool_display_name}") when a non-default value is provided. Do not silently inherit fields that the model ignores.
No config fields for internal/ephemeral plumbing: Don't expose a config field that doesn't change the computation result — e.g., an internal job name, a scratch/output directory, or intermediate-file labels. The wrapper returns results via the Output object, and the export API (_export_output() / output.export(...)) owns user-facing output files and naming — so such fields add API surface without behavior, and (unless marked include_in_key=False) wrongly split the cache so identical calls miss each other. Hardcode the internal value instead. Contributors reflexively add these; drop them. Example: opendde removed its name (job label → ephemeral temp-file names) and root_dir (duplicated the managed weights-cache resolution) fields for exactly this reason.
Remote fail-fast for local-resource configs: If a config setting needs a local resource that can't be staged to the hosted service — a local database/index, an on-disk weights/artifact directory, or a local file path — override BaseConfig.remote_unsupported_reason(self, device: str) -> str | None so device='proto' fails fast at dispatch with a clear message instead of a late runtime error (a missing-file crash, or silently ignoring the override). Return a user-facing reason when the offending setting is active, None otherwise (the default means remote-compatible). Guard on the setting actually being non-default so ordinary remote runs are untouched:
def remote_unsupported_reason(self, device: str) -> str | None:
"""A local weights directory (``local_path``) isn't present on a hosted worker."""
if self.local_path: # optional override; unset by default → remote-OK
return "local_path points to a local weights directory not available on device='proto'. Unset it, or run locally with device='cpu'."
return NoneA tool whose only mode needs a local file (e.g. a required local DB/HMM input, like blast-search local mode or pyhmmer-hmmscan) returns the reason unconditionally. Scope this to intrinsically local configs — do not encode deployment-specific hosting knowledge here (e.g. which model checkpoints are pre-warmed on the hosted service); that stays in the private hosting layer. The router invokes this on any remote device path, after the local_cpu no-op and the license gate. Branch on device when a restriction reflects what Proto chooses to host rather than an intrinsically local resource: a caller running on their own compute is not bound by our hosting limits.
For "pick one" siblings: make each Optional with default=None, tag with the same xor_group="<slug>", and add a @model_validator(mode="after") to enforce at runtime.
mmseqs_db: str | None = InputField(default=None, xor_group="target",
description="Target DB (path/slug/AssetRef).")
target_sequences: list[str] | None = InputField(default=None, xor_group="target",
description="Inline target protein sequences.")
@model_validator(mode="after")
def exactly_one_target(self) -> "Mmseqs2SearchProteinsInput":
if (self.mmseqs_db is None) == (self.target_sequences is None):
raise ValueError("provide exactly one of `mmseqs_db` or `target_sequences`")
return selfmodel_checkpoint fieldWhen a tool exposes both a fixed set of known/bundled checkpoints and the ability to load an arbitrary checkpoint file, collapse them into a single model_checkpoint: str field — do not ship separate model_name (an enum) and checkpoint_path fields. Contributors reflexively add both; standardize on one. The single field takes either a bundled model name or a path to a checkpoint file:
setup.sh into the managed weights cache (proto_resolve_weights_dir → PROTO_MODEL_CACHE), so a user never fetches anything to use them. A small map is the single source of truth for "which names exist," each resolving to a file under the weights dir (or None = the framework's own default weights)."opendde_v2").model_checkpoint: str = ConfigField(
title="Model / Checkpoint", default="<default_name>",
description="Bundled model name ('<name_a>' or '<name_b>') or a path to a custom .pt checkpoint.",
)remote_unsupported_reason rejects only a non-bundled value (a local path isn't on a hosted worker); bundled names are remote-OK because they're provisioned.setup.sh's download list, and add a test asserting setup.sh fetches every bundled checkpoint.tools/structure_prediction/opendde (_BUNDLED_MODELS in the tool, _BUNDLED_CHECKPOINTS in standalone/inference.py).These ensure consistent formatting across all generated tools:
"""Tool description.""". No multi-line module docstrings.Path, Optional, numpy, etc.# Input:
ToolSampleInput = InverseFoldingInput
# Output:
ToolSampleOutput = InverseFoldingOutputToolInstance at module level, not lazily inside the run function:from proto_tools.utils.tool_instance import ToolInstancelogger.debug("Using local venv for {toolkit} {operation}")verbose, timeout, and device are inherited from BaseConfig. Don't redeclare them in tool-specific Config classes unless overriding the default value.__init__.py files — Sort __all__ alphabetically.For the complete tool file template, refer to .claude/skills/implement-tool/TEMPLATES.md.
For implementation patterns (CPU/GPU/compile-from-source), refer to .claude/skills/implement-tool/PATTERNS.md.
Goal: Generate all supporting files in parallel. Each subagent gets the contract (tool file) + one reference file + a focused prompt.
IMPORTANT: Launch all 5 Task() calls in a single message. If sent in separate messages, the first blocks and the rest wait.
Before launching subagents:
Then launch all 5 subagents simultaneously:
What it produces: standalone/inference.py (or run.py), standalone/setup.sh, standalone/requirements.txt, standalone/python_version.txt, optionally standalone/env_vars.txt
Prompt template:
You are implementing the standalone execution environment for a bioinformatics tool.
## Your Task
Create the standalone files for the {tool_display_name} tool in:
proto_tools/tools/{category}/{toolkit}/standalone/
## Contract (the tool file this standalone serves)
<paste full tool file content here>
## Reference Implementation
Here is the standalone from {reference_tool} to use as a structural template:
<paste reference inference.py content>
<paste reference setup.sh content>
<paste reference requirements.txt content>
## Research: Source Repo API
<paste key findings about how the research code works — model loading, inference API, etc.>
## Files to Create
### 1. inference.py (for GPU/AI tools) or run.py (for CPU tools)
Structure:
- Module docstring
- Imports (json, sys, os at top; heavy imports inside functions)
- **Module-level logger** (REQUIRED — see "Logger Convention" below):
```python
from standalone_helpers import get_logger
logger = get_logger(__name__)__init__: set self._loaded = Falseload(device): import heavy deps, load model weightsserialize_output() from standalone_helpers for tensor → list conversiondispatch(input_dict) function as entry point:input_dict["operation"] to model methodsif __name__ == "__main__": block reading/writing JSON filestests/style_consistency_tests/test_standalone_logger_consistency.py)Every .py file under proto_tools/tools/*/*/standalone/ — inference.py, run.py, binary_config.py, model.py, helper modules, etc. — must declare a module-level logger:
from standalone_helpers import get_logger
logger = get_logger(__name__)Why: standalones run inside isolated micromamba subprocesses where proto_tools is NOT importable. get_logger is shipped via standalone_helpers/proto_logging.py (published on PYTHONPATH). Records emitted through it are bridged back to the parent process for re-emission under proto_tools.worker.{toolkit}.*. Plain logging.getLogger(__name__) produces a logger outside the bridge namespace and its records are silently dropped.
__init__.py is the only exemption (no log calls there). The consistency test enforces this on every other .py file in the standalone tree.
For status updates that should drive the parent's spinner subtitle (and never clutter console output), use the dedicated method:
logger.update_status("Loading checkpoint") # spinner subtitle update; not shown in console
logger.info("...") # normal log lineIf the tool has file-format conversion helpers (e.g., writing MSA to Parquet, converting complexes to YAML/FASTA), implement them as plain-type functions in helpers.py at the tool directory level. These functions must be self-contained (no proto_tools imports) and take only plain types (str, list, dict). The tool layer imports and wraps them with typed signatures. Keeping them dependency-free lets the hosted execution backend reuse the same files, preserving a single source of truth. See esmfold, chai1, boltz2 for examples.
CRITICAL RULES:
.py file in standalone/ (except __init__.py) MUST declare logger = get_logger(__name__) from standalone_helpers — see "Logger Convention" above. Plain logging.getLogger(__name__) is forbidden and enforced by a consistency test.config.seed (raw Optional[int]) through the dispatch dict. In inference.py, call set_torch_seed(seed) unconditionally — the helper's None-check gates the expensive cuDNN flags. For any downstream sampler that needs a concrete int, do sampling_seed = seed if seed is not None else get_random_int() using the helper from standalone_helpers.stochastic=True: set when outputs depend on config.seed. Cacheable unseeded calls skip cache; iterable dispatches skip dedup so duplicate items diverge via per-item RNG advancement. The framework does not unroll multi-item dispatches — per-item seed handling is the tool's responsibility. Do not set for tools that accept but ignore the seed.chain_id="A" should come from the input, not be a constant.Sync note: These patterns are duplicated in PATTERNS.md (templates + reference section) because subagents can't see PATTERNS.md. If the protocol changes, update both files.
All standalone scripts MUST implement two module-level protocol functions for DeviceManager integration. Add these AFTER the dispatch() function (or after operation functions for CPU tools), BEFORE if __name__.
to_device(device: str) -> dictEnables DeviceManager to move models between GPUs and CPU.
Persistent PyTorch tools (global _model kept loaded):
def to_device(device: str) -> dict:
global _model
if _model is not None and _model._loaded:
_model.to_device(device)
return {"success": True, "device": device}
else:
return {"success": True, "device": device, "note": "model not loaded yet"}Non-persistent / CLI tools (no persistent model state):
def to_device(device: str) -> dict:
return {"success": True, "device": device, "note": "CLI tool, auto-unloads"}get_memory_stats() -> dictEnables DeviceManager to query GPU memory usage. Import the helper from standalone_helpers (published on PYTHONPATH).
PyTorch tools:
def get_memory_stats() -> dict:
from standalone_helpers import get_pytorch_memory_stats
global _model
device = _model.device if _model and hasattr(_model, "device") else 0
return get_pytorch_memory_stats(device)JAX tools:
def get_memory_stats() -> dict:
from standalone_helpers import get_jax_memory_stats
return get_jax_memory_stats(device_index=0)CPU-only tools:
def get_memory_stats() -> dict:
return {"available": False, "framework": "cpu", "note": "CPU tool"}GPU CLI tools (e.g., Boltz2, RFDiffusion3 — spawn subprocesses that use GPU):
def get_memory_stats() -> dict:
from standalone_helpers import get_pytorch_memory_stats
return get_pytorch_memory_stats(device=0)Key rules:
to_device() returns {"success": bool, "device": str}get_memory_stats() returns {"available": bool, "framework": str, ...}standalone_helpers (published on PYTHONPATH) — do NOT import from proto_toolsAll setup.sh scripts source standalone_helpers.sh, resolved off PATH inside the build subprocess, for shared infrastructure functions. Reference: tools/inverse_folding/fampnn/standalone/setup.sh.
uv is already installed. ToolInstance creates every tool env with pip and uv pinned to UV_VERSION (proto_tools/utils/tool_instance.py) before setup.sh runs. Call uv pip install directly; never pip install uv or guard on command -v uv. Only if the tool's build breaks on that uv, pin another with an optional standalone/uv_version.txt (default: <version>; see "uv Version Override" in notes/tool-environments.md).
#!/bin/bash
set -euo pipefail
source standalone_helpers.sh
# PyTorch tools (add extras like torchvision if needed):
proto_install_pytorch
# proto_install_pytorch "" torchvision # torch + version-matched torchvision
# JAX tools (with optional tool-specific override prefix):
# proto_install_jax MYTOOL
# If CUDA toolkit needed (for JIT compilation):
# proto_install_cuda_toolkit
uv pip install -r requirements.txt
# Non-HF tools that download weights:
proto_resolve_weights_dir {toolkit}
if [ ! -f "$WEIGHTS_DIR/model.pt" ]; then
wget -q -O "$WEIGHTS_DIR/model.pt" "https://example.com/model.pt"
fi
# Gated HF models (validates access before pip install):
# proto_check_gated_hf_repo "org/model-name" "https://huggingface.co/org/model-name"Available functions (see utils/standalone_helpers_source/standalone_helpers.sh):
proto_install_pytorch [spec] [extras...] — install PyTorch via RECOMMENDED_TORCH_SPEC, with optional version-matched extras (e.g., torchvision)proto_install_jax [TOOL_PREFIX] — install JAX with tool-specific overridesproto_install_cuda_toolkit [constraint] [extras...] — micromamba CUDA toolkitproto_resolve_weights_dir <toolkit> — set $WEIGHTS_DIR via PROTO_MODEL_CACHEproto_check_gated_hf_repo <repo_id> <license_url> [probe_file] — validate HF accessFor the standalone inference.py, non-HF tools must call resolve_weights_dir() to find weights at runtime:
from standalone_helpers import resolve_weights_dir
weights_dir = resolve_weights_dir("{toolkit}")
model_path = os.path.join(weights_dir, "model.pt")List all Python dependencies with version pins. Do NOT include torch/jax — those are installed by setup.sh using hardware-aware detection.
Required. Every tool must pin its Python version — ToolInstance raises FileNotFoundError at setup time for a tool without this file, and the style consistency tests fail if any tool is missing it. Always use the keyed format with a required default key. Comments (# to end of line) and blank lines are allowed.
default: 3.12Per-platform overrides (rarely needed): use only when an upstream dependency is unavailable for the default Python on a specific platform. Three-tier lookup, most specific wins:
{system}-{machine} (e.g. linux-aarch64){system} (e.g. linux)default (required catch-all)default: 3.11
linux-aarch64: 3.10 # only when forced by dep availability — see PyRosettaPick default to match the reference tool unless the upstream library has explicit Python-version constraints. The full spec lives in notes/tool-environments.md. Validation: every python_version.txt is checked by tests/style_consistency_tests/test_python_version_consistency.py, which enforces the format and the required default key on every shipped file.
[passthrough] HF_TOKEN [set]
**10C: Run all tests for the new tool**
Run all tests (functional + infra) filtered to the new tool:
```bash
pytest --all --ext -k "{tool_key}" -vThis runs the tool's functional tests AND the parametrized infra tests (example_input, device consistency, registry integration). A detailed log file is generated in logs/ (project root).
What it produces: README.md, cite.bib, license.yaml, links.yaml
license.yaml and links.yaml are required for every toolkit — tests/style_consistency_tests/test_license_consistency.py enforces license.yaml's existence and schema, and the registry/docs surface both files. A toolkit missing either fails CI. Every value in all four files must be grounded in a credible source (the upstream repo, its LICENSE/COPYING, the paper, the SPDX registry) — never guess a license, DOI, or URL. The deeper triple-check happens in Phase 4.7; this subagent's job is to get them right the first time.
Prompt template:
You are writing the documentation and metadata for a bioinformatics tool.
## Your Task
Create README.md, cite.bib, license.yaml, and links.yaml for the {tool_display_name}
tool in:
proto_tools/tools/{category}/{toolkit}/
## Contract (the tool's API)
<paste full tool file content here>
## Reference files
Here are the same files from {reference_tool} to use as structural templates:
<paste reference README.md, cite.bib, license.yaml, links.yaml content>
## Research Context
<paste tool description, paper info, biological context, AND the verified upstream
license + canonical links from Phase 1>
## README.md Structure
Follow the structured template in TEMPLATES.md exactly. Four canonical H2 sections in order:
1. `# {Toolkit Display Name}` (plus the badge row above the title and a `> [!NOTE] License: ...` callout)
2. `## Overview` — 2–3 sentences: who built it, what it does, why useful (link the upstream repo on first mention)
3. `## Background` — 1–3 paragraphs of biological/algorithmic context with inline paper citations; optional `### Learning Resources` subsection for user-facing explainers (blogs, talks, courses)
4. `## Tools` — one `### {Tool Display Name} (\`{tool-key}\`)` block per registered tool, each with `#### Applications` and `#### Usage Tips` (critical-parameter tips in bold)
5. `## Toolkit Notes` — Toolkit-wide guide badges row + bulleted notes that apply to every tool in the toolkit
DO NOT add other H2 sections — schemas, configs, and output specs are auto-generated from Pydantic field descriptions. Read `gene_annotation/pyhmmer/README.md` for the canonical example before drafting. The `> [!NOTE] License: ...` callout MUST agree with `license.yaml`.
## cite.bib
Look up the paper's BibTeX citation using the DOI (verify the DOI resolves — do not fabricate). Format:
@article{firstauthor_year_toolname,
title={...},
author={...},
journal={...},
year={...},
doi={...}
}
## license.yaml
Capture the upstream license verified from the repo's LICENSE/COPYING (and any
weights/data license, which can differ from the code license). Schema and the
SPDX allowlist are in TEMPLATES.md → "license.yaml Template". SPDX-allowlisted
licenses must NOT inline text (it lives in proto_tools/tools/_licenses/{spdx}.txt);
non-SPDX terms use a "Custom (...)" spdx string with inline `text:`.
## links.yaml
Canonical upstream links, verified to resolve. Common keys: `github`, `website`,
`paper` / `preprint` (DOI or arXiv URL), `huggingface`, `organizations` (list).
See TEMPLATES.md → "links.yaml Template".What it produces: examples/example.ipynb
Prompt template:
You are creating an example Jupyter notebook for a bioinformatics tool.
## Your Task
Create examples/example.ipynb for the {tool_display_name} tool in:
proto_tools/tools/{category}/{toolkit}/examples/
## Contract (the tool's API)
<paste full tool file content here>
## Instructions
Create a Jupyter notebook (.ipynb JSON) with these cells:
1. **Markdown title cell** — Tool name, brief description, paper link
2. **Code: Imports** — Exact imports from proto_tools.tools.{category}.{toolkit}
3. **Code: Input API Reference** — Call `display_api_reference("{tool_key}", "input", "run_{tool_key_snake}")` to auto-render the table. Never hand-write the table.
4. **Code: Config API Reference** — Same pattern with `"config"`.
5. **Code: Output API Reference** — Same pattern with `"output"`.
6. **Code: Basic Usage** — Minimal working example with realistic biological data
7. **Code: Advanced Usage** — Example with non-default config parameters
8. **Code: Export** — Demonstrate result.export()
Notebook metadata:
- kernelspec must be exactly `{"name": "python3", "display_name": "proto-tools"}`. Custom names like `proto-language` raise `NoSuchKernel` outside the original author's machine.
- `language_info.name` is `"python"`.
- Use realistic biological data (real protein sequences, not placeholder text)
- Include comments explaining each step
- Show output inspection (printing key fields)
Write the notebook as valid .ipynb JSON (the raw JSON format with cells array, metadata, etc.).
Note: the notebook will be executed during Phase 4 via `scripts/run_example_notebooks.py`, which strips widget/Plotly/Bokeh mime types that don't render in static viewers.What it produces: tests/{category}_tests/test_{tool_key_snake}.py
Prompt template:
You are writing tests for a bioinformatics tool.
## Your Task
Create tests/{category}_tests/test_{tool_key_snake}.py
## Contract (the tool's API)
<paste full tool file content here>
## Reference Tests
Here are the tests from {reference_tool} as a structural template:
<paste reference test file content>
## Test Structure
```python
"""Tests for {tool_display_name} tool."""
from pathlib import Path
import pytest
from proto_tools.entities.structures.structure import Structure # if needed
from proto_tools.tools import run_{tool_key_snake}, {ToolName}Input, {ToolName}Config
TEST_PDB_FILE = Path(__file__).parent.parent / "dummy_data" / "renin_af3.pdb" # if needed
# ── Validation ────────────────────────────────────────────────────────────────
def test_{tool_key_snake}_input_normalizes_single_item():
"""Test that single inputs are normalized to lists (custom validator)."""
# ── Integration ───────────────────────────────────────────────────────────────
@pytest.mark.uses_gpu # or @pytest.mark.integration for CPU tools
def test_{tool_key_snake}_basic_execution():
"""Test basic tool execution with default config."""
# Create input with realistic data
# Run tool
# Assert result.success
# Assert result.tool_id == "{tool_key}"
# Assert specific output fields are populated
@pytest.mark.uses_gpu
def test_{tool_key_snake}_export(tmp_path):
"""Test output export to supported formats."""CRITICAL RULES:
class Test*, use descriptive function namesextra="forbid" rejects unknown fields, or config stores default values are worthless — Pydantic already guarantees these. Only test custom validators, computed properties, normalization logic, and real tool behavior.test_*_rejects_missing_*, test_*_rejects_extra_fields, test_*_config_defaults, test_*_rejects_invalid_* (for simple ge=/le= constraints). These just re-test Pydantic.proto_tools.tools (top-level), not from deep paths@pytest.mark.uses_gpu for any test that runs the actual tool on GPU@pytest.mark.integration for CPU dispatch testsresult.success and result.tool_id
---
### Subagent 5: Export Chain (__init__.py files)
**What it produces:** Updated `__init__.py` at 3 levels (tool, category, master tools)
**Prompt template:**
You are updating the Python export chain for a new bioinformatics tool.
Update init.py files at 3 levels to export the new tool's classes and functions.
<paste full tool file content here>
<paste current category __init__.py content>
<paste current tools/__init__.py content>
File: proto_tools/tools/{category}/{toolkit}/init.py
from proto_tools.tools.{category}.{toolkit}.{tool_key_snake} import (
{ToolName}Config,
{ToolName}Input,
{ToolName}Output,
run_{tool_key_snake},
)
__all__ = [
"{ToolName}Config",
"{ToolName}Input",
"{ToolName}Output",
"run_{tool_key_snake}",
]File: proto_tools/tools/{category}/init.py Add import block and all entries for the new tool. Preserve all existing imports.
File: proto_tools/tools/init.py Add import block and all entries. Preserve all existing imports.
CRITICAL RULES:
__all__ alphabeticallyfrom proto_tools.tools import * — no changes needed
---
## Phase 4: Verify
**Goal:** Confirm all subagent outputs integrate correctly.
**Steps:**
1. **Check import chain** — Verify the tool imports correctly:
```bash
python3 -c "from proto_tools.tools import run_{tool_key_snake}, {ToolName}Input, {ToolName}Config; print('Import OK')"Run tests — Execute the test file:
python3 -m pytest tests/{category}_tests/test_{tool_key_snake}.py -v(Plain pytest runs everything the host can handle. Add --cpu-only to force-skip GPU tests, or --gpu-only to filter to GPU-marked tests.)
Validate tool registration — Check the tool appears in the registry:
python3 -c "from proto_tools.tools.tool_registry import ToolRegistry; specs = [s for s in ToolRegistry.list_all() if '{tool_key}' in s.key]; print([(s.key, s.label) for s in specs])"Execute the example notebook — Always route notebook execution through the sanitizer script, never jupyter nbconvert --execute directly. The script strips widget / Plotly / Bokeh mime types that a kernel emits but static viewers (VS Code, JupyterLab, GitHub, nbviewer) can't render, leaving only the text/plain / text/html fallbacks that survive the round-trip:
python3 scripts/run_example_notebooks.py --only {toolkit}Use --sanitize-only to clean a pre-executed notebook without re-running (fast; for notebooks that can't re-run on the current host).
Fix any failures — If imports fail, check the export chain. If tests fail, fix the tool file or test file directly.
Run the validation checklist:
logging.getLogger(__name__), never print()BaseToolInput, uses Field() (not ConfigField)BaseConfig, uses ConfigField() (not bare Field)BaseToolOutput, implements output_format_options, output_format_default, _export_output()def run_*(inputs: *Input, config: *Config) -> *Outputmetadata={} dict of key parameters@tool() has the 7 required kwargs (key, label, category, input_class, config_class, output_class, description); optional kwargs commonly set: uses_gpu (default False), example_input (factory for parametrized tests)@tool(): uses_gpu=True matches Config device="cuda" override@tool(): Optional device_count specifies expected device allocation ("1", "1-2", ">=1", etc.)@tool(): stochastic=True is set when outputs depend on config.seed; the tool advances its own RNG per item (framework does not unroll)@tool(): metrics_class=<Tool>Metrics is set when the tool emits scalar metric-like values__init__.py exports at all 4 levelsto_device(device: str) -> dict at module levelget_memory_stats() -> dict at module levelsetup.sh implements PROTO_MODEL_CACHE pattern (see fampnn/standalone/setup.sh)inference.py calls resolve_weights_dir() from standalone_helpers (see fampnn/standalone/inference.py)Goal: Guarantee every Input/Config field is real, correctly typed, correctly defaulted, and earns its place. The fast fan-out tends to invent parameters the upstream tool doesn't have, copy defaults/ranges blindly, and leave types loose (str where a Literal belongs).
When: After Phase 4 passes. Before Phase 4.6 — pruning/retyping a field changes the surface the temp tests exercise.
How: Walk every field, one at a time — do NOT batch-approve. Check each against the upstream parameter surface captured in Phase 1 (re-read the upstream source if Phase 1 didn't fully capture it). Make no assumptions; ground each decision in the source.
Per-field checklist:
ge/le/gt/lt bounds must reflect the real valid range from the source, not a guess.Literal[...] over str, and list[Literal[...]] over list[str], whenever upstream accepts a fixed enumerated set. Use bounded int/float over unbounded numerics. Use X | None only when None is genuinely meaningful.reload_on_change=True only for fields that require a worker restart (model checkpoint / anything that reloads the model).include_in_key=False for fields that don't affect results (device, verbose, timeout are already excluded on BaseConfig; a tool-level device override must re-set it).xor_group="<slug>" for mutually exclusive siblings, paired with the @model_validator.ValueError, never silently ignored.remote_unsupported_reason(device) so device='proto' fails fast (see "Cloud fail-fast for local-resource configs" in Phase 2). Skip for AssetRef-stageable uploads and Structure-typed inputs (the gateway stages those).Output: edit the contract directly. If you remove/rename/retype a field, propagate to standalone dispatch(), README, notebook, tests, and the export chain — leave no dangling references. Re-run Phase 4 import/test checks. Do this deliberately yourself, or delegate to a focused subagent given the Phase 1 upstream param surface + the contract.
Goal: Prove the tool behaves correctly under a real user's workload — realistic data, batch/large inputs, edge sizes — beyond the committed unit tests. These are throwaway: write, run, read, delete. The committed suite stays as Subagent 4 wrote it.
When: After the config surface is final (4.5), on a host that can actually run the tool (GPU if uses_gpu=True).
How:
scratch/drive_{toolkit}.py) that imports run_{tool_key_snake} and calls it exactly as a user would: realistic biological inputs, default config, then a couple of non-default configs exercising the parameters you just audited. Print key output fields and result.success. Run it; confirm the output is scientifically sensible, not just non-erroring.xor_group paths.@tool decorator owns errors; no try/except in tool logic) and keep comments/docstrings brief.Goal: Triple-check every external claim in cite.bib, links.yaml, license.yaml, and README.md against credible, verifiable sources. Hallucinated DOIs, wrong licenses, dead links, and invented capabilities are common fan-out failures and are user-facing.
How — verify, never trust:
LICENSE/COPYING and confirm the SPDX id, commercial_use, redistribution, and attribution_required. Confirm separate weights/data licenses where they differ. Confirm _licenses/{spdx}.txt exists for any SPDX id, and that the README > [!NOTE] License: callout agrees.Use WebFetch / WebSearch to verify and ground each fact in a source. Fix discrepancies in place, then re-run test_license_consistency.py / test_readme_consistency.py if touched.
Goal: Catch convention violations, dead fields, hardcoded assumptions, and consistency gaps — then do a final file-by-file review of the entire PR diff against the quality established in the rest of the repo. Last gate before Ship.
When: After Phases 4.5–4.7. Before Phase 5 (Ship).
Launch a single audit agent via Task():
You are auditing a newly implemented bioinformatics tool for quality issues.
## Your Task
Read all generated files for the {tool_display_name} tool and check for issues.
## Files to Read
- proto_tools/tools/{category}/{toolkit}/{tool_key_snake}.py (core tool file)
- proto_tools/tools/{category}/{toolkit}/standalone/inference.py (or run.py)
- proto_tools/tools/{category}/{toolkit}/standalone/setup.sh
- proto_tools/tools/{category}/{toolkit}/standalone/requirements.txt
- proto_tools/tools/{category}/{toolkit}/standalone/python_version.txt
- proto_tools/tools/{category}/{toolkit}/README.md
- proto_tools/tools/{category}/{toolkit}/examples/example.ipynb
- tests/{category}_tests/test_{tool_key_snake}.py
## Reference Tool (for comparison)
Also read the same files from {reference_tool} to compare patterns.
## Checks
1. **Module-level imports** — Heavy dependencies (torch, model libraries) must only be imported inside functions in standalone scripts, never at module level
2. **Inherited fields** — Every field inherited from base classes (shared_data_models.py) must be either implemented in the standalone code or raise ValueError when non-default values are provided
3. **Hardcoded values** — Flag any values in standalone code that are hardcoded but should vary with input (chain IDs, model paths, assumed dimensions, fixed sequence lengths)
4. **Contract-standalone consistency** — Every Config field with `reload_on_change=True` must be read and used in the standalone code. Every operation dispatched in the tool file must be handled in standalone dispatch()
5. **Notebook accuracy** — Import paths, class names, and field names in the example notebook must exactly match the contract
6. **Seed propagation** — If the tool uses PyTorch/JAX RNG, Config inherits `seed: int | None` from `BaseConfig`. Pass `config.seed` (raw `Optional[int]`) through the dispatch dict as `"seed": config.seed`. In `standalone/inference.py`, call `set_torch_seed(seed)` unconditionally — the helper no-ops when `seed is None`, so the expensive cuDNN determinism flags are only set on the explicit-seed path. If the tool's internal sampling code needs a concrete int (e.g., a model sampler that doesn't accept `None`), use `seed if seed is not None else get_random_int()` to fall back to a fresh random int from `standalone_helpers.get_random_int`
7. **Test coverage** — Every operation (sample, score, etc.) should have at least one test
## Output Format
Print a JSON array of issues:
[
{"severity": "fix", "file": "path", "line": "~N", "issue": "description"},
{"severity": "cosmetic", "file": "path", "line": "~N", "issue": "description"}
]
"fix" = must fix before shipping. "cosmetic" = nice to fix, won't cause bugs.
If no issues found, print: []Review the actual diff, not just the files in isolation, against the quality bar set by the rest of the repo.
git add -A && git diff --cached --stat
git diff --cached -- <each changed file> # review file-by-file{reference_tool}, and instructions to (a) flag any drift from established repo conventions, (b) re-check everything flagged earlier in this session, and (c) confirm the change reads like world-class code already in the repo. Prefer the repo's pr-review-toolkit agents where the diff warrants — code-reviewer (conventions/correctness), silent-failure-hunter (error handling), type-design-analyzer (the audited types), comment-analyzer (docstring/comment accuracy) — falling back to generic Task() review agents otherwise. Each returns the same JSON issue array as Part A.After the audit and review agents return:
"fix" severity issues directly"cosmetic" issues if quick (< 1 minute each)"fix" issues were found, save them to persistent memory so we can track how the pipeline fails over time. Write (or append) to the auto-memory directory at implement-tool-audits.md:## {toolkit} audit — {date}
- {N} fix issues, {M} cosmetic
- Fixes: {brief list of fix issues}
- These indicate systemic gaps in the pipeline.Goal: Commit, push, create PR.
Steps:
Stage all new files:
git add proto_tools/tools/{category}/{toolkit}/
git add tests/{category}_tests/test_{tool_key_snake}.py
# Also stage modified __init__.py files
git add proto_tools/tools/{category}/__init__.py
git add proto_tools/tools/__init__.pyCommit:
git commit -m "feat: implement {tool_display_name} tool wrapper
Implements {tool_display_name} ({operations}) in the {category} category.
Closes #{issue_number}"Push and create PR:
git push -u origin HEAD
gh pr create --title "feat: implement {tool_display_name} tool" --body "$(cat <<'EOF'
## Summary
- Implements {tool_display_name} tool wrapper ({operations})
- Category: {category}
- Follows universal tool pattern with standalone execution via ToolInstance
Closes #{issue_number}
## What's Included
- Core tool file with Input/Config/Output + @tool decorator
- Standalone environment (inference.py, setup.sh, requirements.txt)
- README with biological context and usage examples
- cite.bib with paper citation
- Example Jupyter notebook
- Test suite
- Full export chain (__init__.py at all levels)
## Test Plan
- [ ] `python3 -c "from proto_tools.tools import run_{tool_key_snake}"` imports successfully
- [ ] `pytest tests/{category}_tests/test_{tool_key_snake}.py` passes
- [ ] Tool appears in ToolRegistry.list_all()
- [ ] GPU execution verified (if applicable)
EOF
)"Report the PR URL to the user.
© evo-design, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in .claude/skills/implement-tool of evo-design/proto-tools.
Open the folder on GitHubat commit dbc0491
Implement Tool next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Implement Tool this skillevo-design/proto-tools | 134 | — | ~15k | Automated safety check: Notes | MIT | |
| GitHub Review Iterationprisma/orm | 48k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Cherry Studio PR ReviewCherryHQ/cherry-studio | 52k | — | ~3.9k | Automated safety check: Pass | AGPL-3.0 | |
| PR Cyclejaemk/cached | 2.1k | — | ~4.8k | Automated safety check: Notes | MIT | |
| PR Reviewjaemk/self_update | 961 | — | ~1.5k | Automated safety check: Notes | MIT | |
| PR Reviewjaemk/cached | 2.1k | — | ~2.5k | Automated safety check: Notes | MIT |
prisma/orm
Runs a loop on a GitHub pull request: fetch review state, triage comments into actions, implement them and resolve threads, repeating until nothing actionable is left.
CherryHQ/cherry-studio
Reviews Cherry Studio branches, pull requests, commits, files and docs against the project's own architecture, naming, API-boundary and UI rules, report-only by default.
jaemk/cached
PR review-and-update cycle — the orchestrator that takes a PR from review to resolved.
jaemk/self_update
Targeted, read-only review of a PR or checked-out branch. An agent skill from jaemk/self_update.
jaemk/cached
Targeted, read-only review of a PR or checked-out branch. An agent skill from jaemk/cached.
openinterpreter/openinterpreter
Run a final code review on a pull request
evo-design/proto-tools
Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…
Works with
Categories
Implements a new bioinformatics tool wrapper in proto-tools using a parallelized agent pipeline. Implement Tool is an agent skill from evo-design/proto-tools. Implements a new bioinformatics tool wrapper in proto-tools using a parallelized agent pipeline.
Implement Tool fits situations like: creating tools from GitHub issues; wrapping models; implementing new tool wrappers end-to-end.
Run `npx skills add evo-design/proto-tools --skill implement-tool -a claude-code`. Or copy the skill folder (.claude/skills/implement-tool in evo-design/proto-tools) into .claude/skills/implement-tool in your project. Claude Code loads it when a task matches its description.
Run `npx skills add evo-design/proto-tools --skill implement-tool -a codex`. Or copy the skill folder (.claude/skills/implement-tool in evo-design/proto-tools) into .agents/skills/implement-tool in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add evo-design/proto-tools --skill implement-tool -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/implement-tool, .gemini/skills/implement-tool, .github/skills/implement-tool and .opencode/skills/implement-tool in your project.
Going by SKILL.md and its folder, Implement Tool needs the command-line tools its instructions call (git, python3, uv, gh, wget and pytest) and credentials named HF_TOKEN. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep, Task, WebFetch, WebSearch, AskUserQuestion.
SKILL.md names 2 domains. In commands or code: github.com and huggingface.co; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Implement Tool is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 15k tokens (SKILL.md is roughly 61k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Implement Tool: GitHub Review Iteration (prisma/orm, 48k stars), Cherry Studio PR Review (CherryHQ/cherry-studio, 52k stars), PR Cycle (jaemk/cached, 2.1k stars) and PR Review (jaemk/self_update, 961 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
evo-design (a GitHub organization) maintains it in evo-design/proto-tools, which has 134 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 6, 2026.
Source: evo-design/proto-tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.