Agent skill

Train Sentence Transformers

by sickn33 in sickn33/agentic-awesome-skills

Train or fine-tune SentenceTransformer, CrossEncoder, and SparseEncoder models for retrieval, similarity, clustering, classification, reranking, and related embedding tasks.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Train Sentence Transformers

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill train-sentence-transformers -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills train-sentence-transformers --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/train-sentence-transformers .claude/skills/train-sentence-transformers && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
train-sentence-transformers
GitHub stars
47k
Used in
1 other repo
Token cost
~2.4k tokens
SKILL.md length
896 words
Files
28 (incl. scripts, references)
Skills in repo
1,493
Repo updated
First seen
Licence
Apache-2.0

At a glance

Train or fine-tune SentenceTransformer, CrossEncoder, and SparseEncoder models for retrieval, similarity, clustering, classification, reranking, and related embedding tasks.

  • Works in 5 steps: Identify the model type → Required reading → Defaults → …
  • Tasks that involve Embeddings
  • SKILL.md covers When to Use, 1. Identify the model type, 2. Required reading and 3. Defaults, plus 4 more sections
  • Runs Python scripts from its folder; calls pip and hf; needs HF_TOKEN

What it does

Train Sentence Transformers is an agent skill from sickn33/agentic-awesome-skills. Train or fine-tune SentenceTransformer, CrossEncoder, and SparseEncoder models for retrieval, similarity, clustering, classification, reranking, and related embedding tasks.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 29 other files, including scripts and reference files (for example `references/base_model_selection.md`, `references/dataset_formats.md` and `references/evaluators_cross_encoder.md`).

It sits in AI & LLM Engineering, covering Embeddings and Retrieval-augmented generation. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Embeddings
  • Tasks that involve Retrieval-augmented generation

Example prompts

  • “/train-sentence-transformers”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Identify the model type
  2. Required reading
  3. Defaults
  4. Constraints the produced script must satisfy
  5. Workflow

What it can do on your machine

Read from SKILL.md and the folder at commit 680176d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Train Sentence Transformers loads about 2.4k tokens when it runs, and up to ~31k if it reads all its reference files. Until then it costs about 50 tokens; SKILL.md has 896 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~31k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 680176d, republished under its Apache-2.0 licence (© sickn33). 896 words, ~2,387 tokens.

Download SKILL.mdSave it as .claude/skills/train-sentence-transformers/SKILL.md (or your agent's skills folder). This skill also uses 27 other files; get the full folder from GitHub.
name
train-sentence-transformers
description
Train or fine-tune SentenceTransformer, CrossEncoder, and SparseEncoder models for retrieval, similarity, clustering, classification, reranking, and related embedding tasks.
risk
critical
source
https://github.com/huggingface/skills/tree/main/skills/train-sentence-transformers
source_repo
huggingface/skills
source_type
official
date_added
2026-07-01
license
Apache-2.0
license_source
https://github.com/huggingface/skills/blob/main/LICENSE

Train a sentence-transformers Model

When to Use

Use this skill when you need train or fine-tune sentence-transformers models across SentenceTransformer (bi-encoder; dense or static embedding model; for retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal), CrossEncoder (reranker; pair scoring for two-stage retrieval / pair...

This SKILL.md is a router, not a manual. It tells you which references and example scripts to load for your task. The actual content — recommended losses, evaluators, training-script structure, model selection, training-arg knobs, troubleshooting — lives in references/ and scripts/.

Do not synthesize a training script from this file alone. Open the per-type production template (scripts/train_<type>_example.py) and copy it as your starting point. The templates contain load-bearing scaffolding (autocast helper, model-card class, logger silencing list, force=True, seed, TF32, version-compatible imports, named-evaluator metric handling) that prior agent runs have repeatedly missed when rolling their own from a synthesized snippet.

1. Identify the model type

TagClassWhat it doesWhen to pick
[SentenceTransformer]SentenceTransformer (bi-encoder)Maps each input to a fixed-dim dense vectorRetrieval, similarity, clustering, classification, paraphrase mining, dedup
[CrossEncoder]CrossEncoder (reranker)Scores (query, passage) pairs jointlyTwo-stage retrieval (rerank top-100 from bi-encoder), pair classification
[SparseEncoder]SparseEncoder (SPLADE)Sparse vectors over the vocabularyLearned-sparse retrieval, inverted-index backends (Elasticsearch / OpenSearch / Lucene)

Tiebreakers when the request is ambiguous: "embedding model" / "vector search" / "similarity" → [SentenceTransformer]. "rerank" / "ranker" / "two-stage" → [CrossEncoder]. "SPLADE" / "sparse" / "inverted index" → [SparseEncoder]. If still unclear, ask.

2. Required reading

Read these in full before writing any code. Do not triage by perceived relevance.

Per-type — always required

[SentenceTransformer]

  • references/losses_sentence_transformer.md — loss-to-data-shape mapping; BatchSamplers.NO_DUPLICATES requirement for MNRL-family; Cached* ↔ gradient_checkpointing incompatibility.
  • references/evaluators_sentence_transformer.md — evaluator-to-task mapping; metric_for_best_model key construction (named vs unnamed); per-evaluator primary_metric values.
  • references/model_architectures.md — encoder vs decoder vs static vs Router pipelines; pooling rules (mean / cls / lasttoken); auto-mean-pooling behavior for fresh-start MLM bases.
  • scripts/train_sentence_transformer_example.py — production template; copy this as your starting point.

[CrossEncoder]

  • references/losses_cross_encoder.md — pointwise / pairwise / listwise / distillation; pos_weight derivation; activation_fn=Identity() mandatory for non-BCE losses (silent eval-rank collapse otherwise).
  • references/evaluators_cross_encoder.md — CrossEncoderRerankingEvaluator recipe; named-evaluator key format eval_{name}_{primary_metric}.
  • scripts/train_cross_encoder_example.py — production template; copy this as your starting point.

[SparseEncoder]

  • references/losses_sparse_encoder.md — SpladeLoss wrapper requirement; FLOPS regularizer weights; smoke-test active-dim ramp behavior.
  • references/evaluators_sparse_encoder.md — SparseNanoBEIREvaluator (English-only) and the in-domain alternative; eval_{name}_{primary_metric} key format.
  • scripts/train_sparse_encoder_example.py — production template; copy this as your starting point.
Cross-cutting — always required (regardless of task)
  • references/training_args.md — TrainingArguments knobs, precision rules (load fp32 + autocast bf16/fp16; never torch_dtype=bfloat16), warmup_steps (float) vs deprecated warmup_ratio, save_steps must be a multiple of eval_steps for load_best_model_at_end, schedulers, HPO, tracker, resume, hub-push variants.
  • references/dataset_formats.md — column-matching rules (label name auto-detection; column-order-not-name); reshaping recipes; hard-negative mining options.
  • references/base_model_selection.md — discovery commands; per-type model namespaces; ModernBERT-family max_seq_length=8192 trap; datasets >= 4 script-loader rejection; non-English starting-point shortcuts.
  • references/troubleshooting.md — symptom-indexed failure recipes. Skim the section headings on every run, even a healthy one; the "Metrics don't improve" and "Hub push fails" entries cover bugs that bite frequently and are cheaper to recognize before they fire than to debug after.
Cross-cutting — load when applicable
  • references/hardware_guide.md — VRAM sizing, multi-GPU, FSDP / DeepSpeed, HF Jobs flavors. Required for >24GB models, multi-GPU, or HF Jobs runs.
  • references/hf_jobs_execution.md — required when running on HF Jobs.
  • references/prompts_and_instructions.md — required when using prompt-tuned bases (E5, BGE, GTE, Qwen3-Embedding, Instructor, Nomic, etc.) or adding query: / passage: style prefixes.
Variant scripts (open when the task matches)
  • [SentenceTransformer] scripts/train_sentence_transformer_<matryoshka|multi_dataset|with_lora|distillation|make_multilingual|static_embedding>_example.py.
  • [CrossEncoder] scripts/train_cross_encoder_<distillation|listwise>_example.py.
  • [SparseEncoder] scripts/train_sparse_encoder_distillation_example.py.
  • Hard-negative mining CLI — scripts/mine_hard_negatives.py.
Show full SKILL.md (362 more words)Show less

3. Defaults

Override only if the user specifies otherwise:

  • Local execution. Pitch HF Jobs only if local hardware can't fit the job.
  • Single run. After it completes, propose experimentation if the user would benefit (weak/marginal verdict, "see how high you can push it" framing, etc.). Iteration rules in references/training_args.md (Experimentation section).
  • Public Hub push at end-of-run, wrapped in try-except. On HF Jobs (ephemeral env) ALSO enable in-trainer push (push_to_hub=True + hub_strategy="every_save"); details in references/hf_jobs_execution.md.

4. Constraints the produced script must satisfy

These are non-negotiable contracts. Implementation lives in the production templates and references — do not reinvent.

  • Capture the pre-training evaluator score as baseline_eval before trainer.train().
  • Emit a single end-of-run line: VERDICT: WIN|MARGINAL|REGRESSION | score=... | baseline=... | delta=.... A monitor scrapes for this.
  • Silence httpx, httpcore, huggingface_hub, urllib3, filelock, fsspec to WARNING (otherwise HF download URLs flood the agent's context).
  • Tee logs to logs/{RUN_NAME}.log.
  • End with model.push_to_hub(...) wrapped in try/except.
  • Smoke-test before any long run (max_steps=1 + tiny dataset slice). The production templates show one common pattern (SMOKE_TEST env var).
  • [CrossEncoder] Include EarlyStoppingCallback(patience>=3) — CE rerankers often peak mid-training and regress.
  • [SparseEncoder] Log query_active_dims / corpus_active_dims on the verdict line; high nDCG with collapsed sparsity is not a win. The keys come back name-prefixed (e.g. ..._query_active_dims); use suffix matching to pluck them — see the SPARSE production template for the exact pattern.

5. Workflow

  1. Identify the model type (§1). Ask if ambiguous.
  2. Load the §2 required-reading files for that type.
  3. Open scripts/train_<type>_example.py and copy it as your starting point.
  4. Replace MODEL_NAME, DATASET_NAME, RUN_NAME, the loss, and the evaluator with the user's task. Cross-check loss/data-shape match against references/losses_<type>.md; cross-check the metric_for_best_model key against references/evaluators_<type>.md (named evaluators format the key as eval_{name}_{primary_metric}).
  5. Smoke-test (max_steps=1).
  6. Run.
  7. After the run, append to logs/experiments.md and propose iteration if the verdict is weak/marginal.

Prerequisites

bash
pip install "sentence-transformers[train]>=5.0"        # add [train,image] / [audio] / [video] for [SentenceTransformer] multimodal
pip install trackio                                    # optional tracker; or wandb / tensorboard / mlflow
hf auth login                                          # or set HF_TOKEN with write scope (for Hub push)

GPU strongly recommended. CPU works only for demos and [SentenceTransformer] StaticEmbedding.

Limitations

  • Use this skill only when the task clearly matches its upstream product or API scope.
  • Verify commands, API behavior, pricing, quotas, credentials, and deployment effects against current official documentation before making changes.
  • Do not treat generated examples as a substitute for environment-specific tests, security review, or user approval for destructive or costly actions.

© sickn33, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 27 other files (scripts, references) in skills/train-sentence-transformers of sickn33/agentic-awesome-skills.

  • SKILL.md
  • references/base_model_selection.md
  • references/dataset_formats.md
  • references/evaluators_cross_encoder.md
  • references/evaluators_sentence_transformer.md
  • references/evaluators_sparse_encoder.md
  • references/hardware_guide.md
  • references/hf_jobs_execution.md
  • references/losses_cross_encoder.md
  • references/losses_sentence_transformer.md
  • references/losses_sparse_encoder.md
  • references/model_architectures.md
  • references/prompts_and_instructions.md
  • references/training_args.md
  • references/troubleshooting.md
  • scripts/mine_hard_negatives.py
  • scripts/train_cross_encoder_distillation_example.py
  • scripts/train_cross_encoder_example.py
  • scripts/train_cross_encoder_listwise_example.py
  • … and 9 more

Open the folder on GitHubat commit 680176d

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Train Sentence Transformers next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Train Sentence Transformers compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Train Sentence Transformers this skillsickn33/agentic-awesome-skills47k1 repos~2.4kAutomated safety check: PassApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Ms Agent Framework RAGshuyu-labs/WebCode278—~1.1kAutomated safety check: PassCustom licence
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0
Memory Upgradeprofbernardoj/everclaw-community-branches112—~574Automated safety check: PassMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 15 days ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Memory Upgrade

    profbernardoj/everclaw-community-branches

    Diagnose and fix broken memory search in OpenClaw. An agent skill from profbernardoj/everclaw-community-branches.

    112 GitHub stars~574 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Embedding Strategies

    wshobson/agents

    Helps choose and tune embedding models for semantic search and RAG: model comparison, chunking, preprocessing, normalization and caching.

    40k GitHub starsUsed in 10 repos~710 tokens
    AI & LLM EngineeringAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,493 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Train Sentence Transformers

What does Train Sentence Transformers do?

Train or fine-tune SentenceTransformer, CrossEncoder, and SparseEncoder models for retrieval, similarity, clustering, classification, reranking, and related embedding tasks. Train Sentence Transformers is an agent skill from sickn33/agentic-awesome-skills. Train or fine-tune SentenceTransformer, CrossEncoder, and SparseEncoder models for retrieval, similarity, clustering, classification, reranking, and related embedding tasks.

When should I use Train Sentence Transformers?

Train Sentence Transformers fits situations like: tasks that involve Embeddings; tasks that involve Retrieval-augmented generation.

How do I install Train Sentence Transformers in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill train-sentence-transformers -a claude-code`. Or copy the skill folder (skills/train-sentence-transformers in sickn33/agentic-awesome-skills) into .claude/skills/train-sentence-transformers in your project. Claude Code loads it when a task matches its description.

How do I install Train Sentence Transformers in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill train-sentence-transformers -a codex`. Or copy the skill folder (skills/train-sentence-transformers in sickn33/agentic-awesome-skills) into .agents/skills/train-sentence-transformers in your project. Codex loads it when a task matches its description.

Can I use Train Sentence Transformers in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill train-sentence-transformers -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/train-sentence-transformers, .gemini/skills/train-sentence-transformers, .github/skills/train-sentence-transformers and .opencode/skills/train-sentence-transformers in your project.

What does Train Sentence Transformers need to run?

Going by SKILL.md and its folder, Train Sentence Transformers needs Python for the scripts in its folder, the command-line tools its instructions call (pip and hf) and credentials named HF_TOKEN. Our summary lists: Python 3.

Does Train Sentence Transformers access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Train Sentence Transformers safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Train Sentence Transformers use?

Train Sentence Transformers is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Train Sentence Transformers use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 29k tokens, read only when the agent opens those files.

What are the alternatives to Train Sentence Transformers?

Skills that share tags, products or a category with Train Sentence Transformers: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ms Agent Framework RAG (shuyu-labs/WebCode, 278 stars), Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars) and Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Train Sentence Transformers?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,379 GitHub stars. The repository holds 1,493 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.