Hugging Face Transformers Usage
davila7/claude-code-templates
Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.
HuggingFace Transformers with biomedical LMs (BioBERT, PubMedBERT, BioGPT, BioMedLM) for scientific NLP: NER (genes, diseases, chemicals), relation extraction, QA, text classification, abstract…
$ npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills transformers-bio-nlp --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scientific-computing/transformers-bio-nlp .claude/skills/transformers-bio-nlp && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "transformers-bio-nlp" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/transformers-bio-nlp into .claude/skills/transformers-bio-nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transformers-bio-nlp", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/transformers-bio-nlpType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills transformers-bio-nlp --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/scientific-computing/transformers-bio-nlp .agents/skills/transformers-bio-nlp && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "transformers-bio-nlp" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/transformers-bio-nlp into .agents/skills/transformers-bio-nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transformers-bio-nlp", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills transformers-bio-nlp --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/scientific-computing/transformers-bio-nlp .cursor/skills/transformers-bio-nlp && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "transformers-bio-nlp" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/transformers-bio-nlp into .cursor/skills/transformers-bio-nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transformers-bio-nlp", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/scientific-computing/transformers-bio-nlp--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills transformers-bio-nlp --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/scientific-computing/transformers-bio-nlp .gemini/skills/transformers-bio-nlp && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "transformers-bio-nlp" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/transformers-bio-nlp into .gemini/skills/transformers-bio-nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transformers-bio-nlp", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills transformers-bio-nlpInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/scientific-computing/transformers-bio-nlp .github/skills/transformers-bio-nlp && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "transformers-bio-nlp" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/transformers-bio-nlp into .github/skills/transformers-bio-nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transformers-bio-nlp", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills transformers-bio-nlp --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/scientific-computing/transformers-bio-nlp .opencode/skills/transformers-bio-nlp && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "transformers-bio-nlp" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/scientific-computing/transformers-bio-nlp into .opencode/skills/transformers-bio-nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transformers-bio-nlp", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
transformers-bio-nlpHuggingFace Transformers with biomedical LMs (BioBERT, PubMedBERT, BioGPT, BioMedLM) for scientific NLP: NER (genes, diseases, chemicals), relation extraction, QA, text classification, abstract…
Transformers Bio NLP is an agent skill from jaechang-hits/SciAgent-Skills. HuggingFace Transformers with biomedical LMs (BioBERT, PubMedBERT, BioGPT, BioMedLM) for scientific NLP: NER (genes, diseases, chemicals), relation extraction, QA, text classification, abstract summarization. Covers loading, biomedical tokenization, inference pipelines, fine-tuning. Alternatives: spaCy encorescilg (rule-based NER), Stanza (biomedical models), NLTK.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Natural language processing. It works with Transformers and PubMed. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
piphuggingface-cliFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
download.pytorch.orgAlso links to:
doi.orghuggingface.cobiocreative.bioinformatics.udel.eduFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Transformers Bio NLP loads about 4.9k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 873 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 873 words, ~4,935 tokens.
.claude/skills/transformers-bio-nlp/SKILL.md (or your agent's skills folder).HuggingFace Transformers provides a unified API to load, run, and fine-tune 500+ biomedical language models. The key biomedical models — BioBERT (trained on PubMed abstracts + PMC full text), PubMedBERT (trained from scratch on PubMed), BioGPT (generative, trained on PubMed), and BioMedLM — significantly outperform general-purpose BERT on biomedical NER, relation extraction, and question answering. The pipeline() abstraction handles tokenization, inference, and postprocessing in one call. Fine-tuning on task-specific labeled data (e.g., BC5CDR for chemical/disease NER) takes under an hour on a single GPU. The datasets library provides direct access to standard biomedical benchmarks.
transformers, torch, datasets, accelerate, sentencepiecepip install transformers torch datasets accelerate sentencepiece
# For GPU (CUDA 11.8)
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118from transformers import pipeline
# Named entity recognition with BioBERT
ner = pipeline("ner", model="allenai/scibert_scivocab_cased",
aggregation_strategy="simple")
text = "BRCA1 mutations are associated with increased risk of breast cancer and ovarian cancer."
entities = ner(text)
for ent in entities:
print(f" {ent['word']:20s} {ent['entity_group']:10s} score={ent['score']:.3f}")Extract biomedical entities using pre-trained NER models.
from transformers import pipeline, AutoTokenizer, AutoModelForTokenClassification
# BioBERT fine-tuned for NER (genes, diseases, chemicals)
# Common choices:
# "allenai/scibert_scivocab_cased" — scientific NER
# "d4data/biomedical-ner-all" — multi-entity biomedical NER
# "pruas/BENT-PubMedBERT-NER-Gene" — gene-specific NER
ner_pipe = pipeline(
"ner",
model="d4data/biomedical-ner-all",
aggregation_strategy="simple", # merge subword tokens into words
device=-1 # -1=CPU, 0=GPU
)
abstracts = [
"Imatinib inhibits the BCR-ABL1 tyrosine kinase and is first-line treatment for CML.",
"EGFR mutations in non-small cell lung cancer predict response to erlotinib.",
]
for text in abstracts:
entities = ner_pipe(text)
print(f"\nText: {text[:60]}...")
for e in entities:
print(f" [{e['entity_group']}] '{e['word']}' (score={e['score']:.2f})")# Manual tokenization + inference for batch processing
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch
model_name = "allenai/scibert_scivocab_cased"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
model.eval()
text = "Metformin activates AMPK and reduces hepatic glucose production in type 2 diabetes."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits # shape: (1, seq_len, n_labels)
predictions = logits.argmax(dim=-1)[0]
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
labels = [model.config.id2label[p.item()] for p in predictions]
for token, label in zip(tokens[1:-1], labels[1:-1]): # skip [CLS] and [SEP]
if label != "O":
print(f" {token:20s} {label}")Classify biomedical abstracts or sentences.
from transformers import pipeline
# Zero-shot classification — no fine-tuning needed
zs_clf = pipeline("zero-shot-classification",
model="facebook/bart-large-mnli",
device=-1)
abstract = """
This randomized controlled trial evaluated the efficacy of pembrolizumab versus
chemotherapy in patients with advanced non-small-cell lung cancer. Overall survival
was significantly improved in the pembrolizumab arm (HR=0.60, 95% CI 0.41-0.89).
"""
candidate_labels = ["clinical trial", "basic research", "meta-analysis", "review"]
result = zs_clf(abstract, candidate_labels)
print("Zero-shot classification:")
for label, score in zip(result["labels"], result["scores"]):
print(f" {label:20s}: {score:.3f}")# Fine-tuned sentiment/outcome classification
from transformers import pipeline
# Example: classify clinical outcome sentiment
clf = pipeline("text-classification",
model="pruas/BENT-PubMedBERT-NER-Gene", # use appropriate task-specific model
device=-1)
sentences = [
"Treatment significantly improved overall survival (p<0.001).",
"No statistically significant difference was observed between groups.",
]
results = clf(sentences)
for sent, result in zip(sentences, results):
print(f" [{result['label']} | {result['score']:.2f}] {sent[:50]}...")Extract answers from biomedical text passages.
from transformers import pipeline
# Extractive QA: find answer span within context
qa_pipe = pipeline(
"question-answering",
model="sultan/BioM-ELECTRA-Large-SQuAD2", # biomedical QA model
device=-1
)
context = """
BRCA1 is a tumor suppressor gene located on chromosome 17q21. Pathogenic variants
in BRCA1 confer a lifetime breast cancer risk of 50-72% and ovarian cancer risk
of 44-46%. BRCA1 protein functions in DNA double-strand break repair via
homologous recombination.
"""
questions = [
"What chromosome is BRCA1 located on?",
"What is the lifetime breast cancer risk from BRCA1 variants?",
"What DNA repair pathway does BRCA1 participate in?",
]
for q in questions:
result = qa_pipe(question=q, context=context)
print(f"Q: {q}")
print(f"A: {result['answer']} (score={result['score']:.3f})\n")Generate biomedical text, hypotheses, and summaries.
from transformers import AutoTokenizer, BioGptForCausalLM
import torch
model_name = "microsoft/biogpt"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = BioGptForCausalLM.from_pretrained(model_name)
model.eval()
prompt = "The role of VEGF in tumor angiogenesis"
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=100,
num_beams=5,
early_stopping=True,
no_repeat_ngram_size=3,
pad_token_id=tokenizer.eos_token_id,
)
generated = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Generated:\n{generated}")Embed biomedical text for similarity search and clustering.
from transformers import AutoTokenizer, AutoModel
import torch
import numpy as np
def mean_pooling(model_output, attention_mask):
"""Mean pooling across token embeddings."""
token_embeddings = model_output.last_hidden_state
input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
return (token_embeddings * input_mask_expanded).sum(1) / input_mask_expanded.sum(1)
# PubMedBERT for biomedical sentence embeddings
model_name = "microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
model.eval()
sentences = [
"BRCA1 is involved in DNA double-strand break repair.",
"Homologous recombination requires BRCA1 and BRCA2.",
"Metformin inhibits hepatic gluconeogenesis via AMPK.",
]
inputs = tokenizer(sentences, padding=True, truncation=True,
max_length=512, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
embeddings = mean_pooling(outputs, inputs["attention_mask"])
embeddings = torch.nn.functional.normalize(embeddings, p=2, dim=1).numpy()
# Compute cosine similarity
from numpy.linalg import norm
sim_01 = np.dot(embeddings[0], embeddings[1])
sim_02 = np.dot(embeddings[0], embeddings[2])
print(f"Similarity (BRCA1 repair vs. HR): {sim_01:.3f}")
print(f"Similarity (BRCA1 repair vs. Metformin): {sim_02:.3f}")Fine-tune a biomedical model on a labeled NER dataset.
from transformers import (AutoTokenizer, AutoModelForTokenClassification,
TrainingArguments, Trainer, DataCollatorForTokenClassification)
from datasets import Dataset
import numpy as np
# Example: minimal NER fine-tuning setup
model_name = "microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract"
label_list = ["O", "B-GENE", "I-GENE", "B-DISEASE", "I-DISEASE"]
id2label = {i: l for i, l in enumerate(label_list)}
label2id = {l: i for i, l in enumerate(label_list)}
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(
model_name, num_labels=len(label_list), id2label=id2label, label2id=label2id
)
# Training arguments
training_args = TrainingArguments(
output_dir="./biomed_ner_finetuned",
num_train_epochs=3,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
warmup_steps=100,
weight_decay=0.01,
logging_dir="./logs",
evaluation_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
)
print(f"Model ready for fine-tuning: {model_name}")
print(f"Labels: {label_list}")
# trainer = Trainer(model=model, args=training_args, ...)
# trainer.train()Biomedical text contains special tokens (gene symbols, drug names, chemical SMILES, numeric values) that WordPiece and BPE tokenizers split unexpectedly. For example, "BRCA1" → ["BR", "##CA", "##1"]. This subword splitting does not affect classification tasks but does affect NER — use aggregation_strategy="simple" or "first" in pipeline() to merge subword predictions back to word level.
NER uses BIO (Begin-Inside-Outside) tagging: B-GENE marks the first token of a gene name, I-GENE marks continuation tokens, O marks non-entity tokens. During fine-tuning, align labels to subword tokens by setting non-first subword labels to -100 (ignored by the loss function).
from transformers import pipeline
import pandas as pd
ner_pipe = pipeline("ner", model="d4data/biomedical-ner-all",
aggregation_strategy="simple", device=-1)
abstracts = [
"Pembrolizumab combined with chemotherapy significantly improved progression-free survival in HER2-positive breast cancer.",
"Inhibition of EGFR by gefitinib is effective in patients with activating EGFR mutations in exons 19 and 21.",
"CRISPR-Cas9 editing of the PCSK9 gene in hepatocytes reduces LDL cholesterol in murine models.",
]
records = []
for i, text in enumerate(abstracts):
entities = ner_pipe(text)
for e in entities:
records.append({
"abstract_id": i,
"entity": e["word"],
"type": e["entity_group"],
"score": round(e["score"], 3),
})
df = pd.DataFrame(records)
print(df.groupby("type")["entity"].apply(list).to_string())
df.to_csv("extracted_entities.csv", index=False)
print(f"\nExtracted {len(df)} entity mentions across {len(abstracts)} abstracts")from transformers import AutoTokenizer, AutoModel
import torch
import numpy as np
model_name = "microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
model.eval()
def embed(texts):
enc = tokenizer(texts, padding=True, truncation=True,
max_length=512, return_tensors="pt")
with torch.no_grad():
out = model(**enc)
vecs = out.last_hidden_state[:, 0, :] # [CLS] token
return torch.nn.functional.normalize(vecs, dim=1).numpy()
query = "CRISPR base editing for correction of point mutations in genetic disease"
corpus = [
"Base editing enables precise single-base changes in genomic DNA without double-strand breaks.",
"CAR-T cell therapy targets CD19 in B-cell acute lymphoblastic leukemia.",
"Prime editing uses reverse transcriptase to install targeted edits at specific loci.",
"RNA interference silences gene expression via RISC-mediated mRNA cleavage.",
]
q_emb = embed([query])
c_emb = embed(corpus)
scores = (q_emb @ c_emb.T).flatten()
ranked = sorted(zip(scores, corpus), reverse=True)
print("Top results:")
for score, text in ranked:
print(f" [{score:.3f}] {text[:70]}...")| Parameter | Module/Function | Default | Range / Options | Effect |
|---|---|---|---|---|
model | pipeline() | — | HuggingFace model ID string | Pre-trained model to load; must match task |
aggregation_strategy | NER pipeline | "none" | "none", "simple", "first", "average" | Merge subword NER predictions; use "simple" for word-level output |
device | pipeline() | -1 | -1 (CPU), 0 (GPU 0), 1 (GPU 1) | Inference device |
max_length | tokenizer | 512 | 128–2048 (model-dependent) | Max token length; truncates longer inputs |
max_new_tokens | model.generate() | 20 | 1–1000 | Tokens to generate for text generation models |
num_beams | model.generate() | 1 | 1–10 | Beam search width; larger = better quality, slower |
num_train_epochs | TrainingArguments | 3 | 1–10 | Fine-tuning epochs |
per_device_train_batch_size | TrainingArguments | 8 | 4–32 | Batch size per GPU; reduce if OOM |
weight_decay | TrainingArguments | 0.0 | 0.01–0.1 | L2 regularization for fine-tuning |
Use domain-specific models, not general BERT: PubMedBERT trained from scratch on PubMed outperforms BERT-base by 5–15% on biomedical NER. Always start with biomedical pre-training before fine-tuning on task-specific data.
Verify model licenses before production use: Some models (BioGPT, BioMedLM) have research-only licenses. Check the HuggingFace model card's license field before deploying in commercial applications.
Use aggregation_strategy="simple" for word-level NER output: The default "none" returns subword tokens, making post-processing difficult. "simple" merges subword tokens using the first-token strategy.
Truncate at sentence boundaries, not mid-sentence: Long biomedical abstracts that exceed 512 tokens should be split at sentence boundaries before encoding. Mid-sentence truncation degrades NER accuracy for entities near the cutoff.
from transformers import pipeline
from itertools import product
ner = pipeline("ner", model="d4data/biomedical-ner-all",
aggregation_strategy="simple", device=-1)
def extract_drug_disease_pairs(text):
entities = ner(text)
drugs = [e["word"] for e in entities if e["entity_group"] in ("DRUG", "CHEMICAL")]
diseases = [e["word"] for e in entities if e["entity_group"] in ("DISEASE", "CONDITION")]
return list(product(drugs, diseases))
text = "Imatinib and nilotinib both target BCR-ABL1 in chronic myeloid leukemia and Philadelphia chromosome-positive ALL."
pairs = extract_drug_disease_pairs(text)
print("Drug-Disease pairs:")
for drug, disease in pairs:
print(f" {drug} → {disease}")from transformers import pipeline
clf = pipeline("zero-shot-classification",
model="facebook/bart-large-mnli", device=-1)
abstracts = [
"We present a phase 3 randomized controlled trial of semaglutide in type 2 diabetes.",
"Structural analysis of the SARS-CoV-2 spike protein RBD domain by cryo-EM.",
"A retrospective cohort study of 1,200 ICU patients during the COVID-19 pandemic.",
]
label_options = ["randomized controlled trial", "observational study", "structural biology", "computational study"]
for abstract in abstracts:
result = clf(abstract, label_options)
print(f"Type: {result['labels'][0]} ({result['scores'][0]:.2f})")
print(f" {abstract[:70]}...\n")| Problem | Cause | Solution |
|---|---|---|
CUDA out of memory during inference | Batch too large for GPU VRAM | Reduce batch size; use device=-1 for CPU; use model.half() for FP16 |
NER returns subword tokens (##CA) | aggregation_strategy not set | Set aggregation_strategy="simple" in pipeline() |
| Model download times out | Large model files (1–10 GB); slow connection | Set HF_HUB_OFFLINE=1 and download manually with huggingface-cli download |
| NER misses entities at end of long abstracts | Input truncated at 512 tokens | Split abstracts into sentences; process each separately |
Fine-tuning loss is NaN | Learning rate too high or gradient explosion | Reduce learning_rate to 2e-5; enable gradient clipping max_grad_norm=1.0 |
| Wrong entities for specialized domain | Generic biomedical model not suited to subdomain | Fine-tune on domain-labeled data; use more specific model (e.g., gene-only NER) |
| BioGPT generates repetitive text | no_repeat_ngram_size too small | Set no_repeat_ngram_size=3 or 4; increase num_beams |
pubmed-database — retrieve PubMed abstracts that serve as input to biomedical NLP pipelinesbiorxiv-database — retrieve preprints for NLP analysis before peer reviewscientific-critical-thinking — evaluate quality of NLP-extracted evidence before using for research conclusions© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/scientific-computing/transformers-bio-nlp of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Transformers Bio NLP next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Transformers Bio NLP this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Transformers Usagedavila7/claude-code-templates | 33k | 11 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Transformersynulihao/AgentSkillOS | 618 | — | ~2.9k | Automated safety check: Pass | None | |
| Transformers.jshuggingface/skills | 11k | 1 repos | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| TransformersK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~2.8k | Automated safety check: Notes | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
davila7/claude-code-templates
Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.
ynulihao/AgentSkillOS
Work with state-of-the-art machine learning models for NLP, computer vision, audio, and multimodal tasks using HuggingFace Transformers.
huggingface/skills
Runs pre-trained Hugging Face models in JavaScript or TypeScript with Transformers.js, in browsers or Node.js, Bun and Deno, for text, vision, audio and multimodal tasks.
K-Dense-AI/scientific-agent-skills
Hugging Face Transformers for loading Hub models, running pipeline inference, text generation, and Trainer fine-tuning on NLP, vision, audio, and multimodal tasks.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Orchestra-Research/AI-Research-SKILLs
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
HuggingFace Transformers with biomedical LMs (BioBERT, PubMedBERT, BioGPT, BioMedLM) for scientific NLP: NER (genes, diseases, chemicals), relation extraction, QA, text classification, abstract…. Transformers Bio NLP is an agent skill from jaechang-hits/SciAgent-Skills. HuggingFace Transformers with biomedical LMs (BioBERT, PubMedBERT, BioGPT, BioMedLM) for scientific NLP: NER (genes, diseases, chemicals), relation extraction, QA, text classification, abstract summarization.
Transformers Bio NLP fits situations like: tasks that involve Natural language processing.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a claude-code`. Or copy the skill folder (skills/scientific-computing/transformers-bio-nlp in jaechang-hits/SciAgent-Skills) into .claude/skills/transformers-bio-nlp in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a codex`. Or copy the skill folder (skills/scientific-computing/transformers-bio-nlp in jaechang-hits/SciAgent-Skills) into .agents/skills/transformers-bio-nlp in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill transformers-bio-nlp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/transformers-bio-nlp, .gemini/skills/transformers-bio-nlp, .github/skills/transformers-bio-nlp and .opencode/skills/transformers-bio-nlp in your project.
Going by SKILL.md and its folder, Transformers Bio NLP needs the command-line tools its instructions call (pip and huggingface-cli). Our summary lists: Python 3.
SKILL.md names 4 domains. In commands or code: download.pytorch.org; the agent is likely to contact it when it follows the instructions. As links in the text: doi.org, huggingface.co and biocreative.bioinformatics.udel.edu. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Transformers Bio NLP is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Transformers Bio NLP: Hugging Face Transformers Usage (davila7/claude-code-templates, 33k stars), Transformers (ynulihao/AgentSkillOS, 618 stars), Transformers.js (huggingface/skills, 11k stars) and Transformers (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.