Official agent skill

Sentence-Transformers Training Router

by huggingface in huggingface/skills

Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Sentence-Transformers Training Router

skills CLI
$ npx skills add huggingface/skills --skill train-sentence-transformers -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huggingface/skills train-sentence-transformers --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/train-sentence-transformers .claude/skills/train-sentence-transformers && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
train-sentence-transformers
GitHub stars
11k
Used in
1 other repo
Token cost
~2.6k tokens
SKILL.md length
944 words
Files
31 (incl. scripts, references)
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

  • Works in 5 steps: Identify the model type → Required reading → Defaults → …
  • Fine-tuning an embedding model for retrieval or similarity search
  • SKILL.md covers 1. Identify the model type, 2. Required reading, 3. Defaults and 4. Constraints the produced…, plus 2 more sections
  • Runs Python scripts from its folder; calls pip and hf; needs HF_TOKEN

What it does

The skill describes itself as a router rather than a manual: it identifies which of four model types fits the request, SentenceTransformer bi-encoders for retrieval and similarity, CrossEncoder rerankers for scoring query-passage pairs, SparseEncoder for SPLADE-style learned-sparse retrieval, and MultiVectorEncoder for ColBERT-style late interaction, using keyword tiebreakers such as rerank, SPLADE or MaxSim when the request is ambiguous, and asking when it still is not clear.

For the chosen type it requires reading specific reference files in full, such as the loss-to-data-shape mapping for SentenceTransformer, rather than skimming by perceived relevance. It warns against writing a training script from the SKILL.md alone, instead directing the agent to copy a per-type production template from scripts/, because those templates carry load-bearing details like an autocast helper, model-card generation and logger silencing that prior ad hoc attempts have missed. Further references cover hard-negative mining, evaluators per type, distillation, LoRA, Matryoshka embeddings and publishing to the Hub.

When your agent uses it

  • Fine-tuning an embedding model for retrieval or similarity search
  • Training a reranker to score query and passage pairs
  • Training a SPLADE sparse retrieval model or a ColBERT-style multi-vector model

Example prompts

  • “Fine-tune a SentenceTransformer bi-encoder for product search with hard negatives.”
  • “Train a CrossEncoder reranker to rerank the top 100 results from our retriever.”
  • “Set up Matryoshka embedding training and publish the result to the Hub.”

Requirements

  • Python with sentence-transformers installed
  • A Hugging Face account, to publish a trained model

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Identify the model type
  2. Required reading
  3. Defaults
  4. Constraints the produced script must satisfy
  5. Workflow

What it can do on your machine

Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sentence-Transformers Training Router loads about 2.6k tokens when it runs, and up to ~36k if it reads all its reference files. Until then it costs about 169 tokens; SKILL.md has 944 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~169
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~36k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 944 words, ~2,622 tokens.

Download SKILL.mdSave it as .claude/skills/train-sentence-transformers/SKILL.md (or your agent's skills folder). This skill also uses 30 other files; get the full folder from GitHub.
name
train-sentence-transformers
description
Train or fine-tune sentence-transformers models across `SentenceTransformer` (bi-encoder, dense or static embedding model for retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal), `CrossEncoder` (reranker, pair scoring for two-stage retrieval / pair classification), `SparseEncoder` (SPLADE, sparse embedding model for learned-sparse retrieval), and `MultiVectorEncoder` (ColBERT / late-interaction, per-token embeddings scored with MaxSim). Covers loss selection, hard-negative mining, evaluators, distillation, LoRA, Matryoshka, and Hugging Face Hub publishing. Use for any sentence-transformers training task.

Train a sentence-transformers Model

This SKILL.md is a router, not a manual. It tells you which references and example scripts to load for your task. The actual content (recommended losses, evaluators, training-script structure, model selection, training-arg knobs, troubleshooting) lives in references/ and scripts/.

Do not synthesize a training script from this file alone. Open the per-type production template (scripts/train_<type>_example.py) and copy it as your starting point. The templates contain load-bearing scaffolding (autocast helper, model-card class, logger silencing list, force=True, seed, TF32, version-compatible imports, named-evaluator metric handling) that prior agent runs have repeatedly missed when rolling their own from a synthesized snippet.

1. Identify the model type

TagClassWhat it doesWhen to pick
[SentenceTransformer]SentenceTransformer (bi-encoder)Maps each input to a fixed-dim dense vectorRetrieval, similarity, clustering, classification, paraphrase mining, dedup
[CrossEncoder]CrossEncoder (reranker)Scores (query, passage) pairs jointlyTwo-stage retrieval (rerank top-100 from bi-encoder), pair classification
[SparseEncoder]SparseEncoder (SPLADE)Sparse vectors over the vocabularyLearned-sparse retrieval, inverted-index backends (Elasticsearch / OpenSearch / Lucene)
[MultiVectorEncoder]MultiVectorEncoder (ColBERT)One embedding per token, scored with MaxSimLate-interaction retrieval, recall gains over bi-encoders at higher storage cost, multimodal (ColPali / ColQwen2)

Tiebreakers when the request is ambiguous: "embedding model" / "vector search" / "similarity" → [SentenceTransformer]. "rerank" / "ranker" / "two-stage" → [CrossEncoder]. "SPLADE" / "sparse" / "inverted index" → [SparseEncoder]. "ColBERT" / "late interaction" / "multi-vector" / "MaxSim" / "ColPali" / "ColQwen" → [MultiVectorEncoder]. If still unclear, ask.

2. Required reading

Read these in full before writing any code. Do not triage by perceived relevance.

Per-type: always required

[SentenceTransformer]

  • references/losses_sentence_transformer.md: loss-to-data-shape mapping, BatchSamplers.NO_DUPLICATES requirement for MNRL-family, Cached* ↔ gradient_checkpointing incompatibility.
  • references/evaluators_sentence_transformer.md: evaluator-to-task mapping, metric_for_best_model key construction (named vs unnamed), per-evaluator primary_metric values.
  • references/model_architectures.md: encoder vs decoder vs static vs Router pipelines, pooling rules (mean / cls / lasttoken), auto-mean-pooling behavior for fresh-start MLM bases.
  • scripts/train_sentence_transformer_example.py: production template. Copy this as your starting point.

[CrossEncoder]

  • references/losses_cross_encoder.md: pointwise / pairwise / listwise / distillation, pos_weight derivation, activation_fn=Identity() mandatory for non-BCE losses (silent eval-rank collapse otherwise).
  • references/evaluators_cross_encoder.md: CrossEncoderRerankingEvaluator recipe, named-evaluator key format eval_{name}_{primary_metric}.
  • scripts/train_cross_encoder_example.py: production template. Copy this as your starting point.

[SparseEncoder]

  • references/losses_sparse_encoder.md: SpladeLoss wrapper requirement, FLOPS regularizer weights, smoke-test active-dim ramp behavior.
  • references/evaluators_sparse_encoder.md: SparseNanoBEIREvaluator (English-only) and the in-domain alternative, eval_{name}_{primary_metric} key format.
  • scripts/train_sparse_encoder_example.py: production template. Copy this as your starting point.

[MultiVectorEncoder]

  • references/losses_multi_vector_encoder.md: MaxSim scoring, scale choice per scoring mode (scale=1.0 for MaxSim, roughly the average query length for MeanMaxSim), MNRL / CachedMNRL / MarginMSE / DistillKLDiv, XTR-vs-ColBERT scoring, CachedMNRL ↔ gradient_checkpointing incompatibility.
  • references/evaluators_multi_vector_encoder.md: MultiVectorNanoBEIREvaluator (English-only) and the in-domain alternative, eval_NanoBEIR_mean_maxsim_ndcg@10 key format, distillation-eval spearman variant.
  • scripts/train_multi_vector_encoder_example.py: production template. Copy this as your starting point.
Cross-cutting: always required (regardless of task)
  • references/training_args.md: TrainingArguments knobs, precision rules (load fp32 + autocast bf16/fp16, never torch_dtype=bfloat16), warmup_steps (float) vs deprecated warmup_ratio, save_steps must be a multiple of eval_steps for load_best_model_at_end, schedulers, HPO, tracker, resume, hub-push variants.
  • references/dataset_formats.md: column-matching rules (label name auto-detection, column-order-not-name), reshaping recipes, hard-negative mining options.
  • references/base_model_selection.md: discovery commands, per-type model namespaces, ModernBERT-family max_seq_length=8192 trap, datasets >= 4 script-loader rejection, non-English starting-point shortcuts.
  • references/troubleshooting.md: symptom-indexed failure recipes. Skim the section headings on every run, even a healthy one. The "Metrics don't improve" and "Hub push fails" entries cover bugs that bite frequently and are cheaper to recognize before they fire than to debug after.
Cross-cutting: load when applicable
  • references/hardware_guide.md: VRAM sizing, multi-GPU, FSDP / DeepSpeed, HF Jobs flavors. Required for >24GB models, multi-GPU, or HF Jobs runs.
  • references/hf_jobs_execution.md: required when running on HF Jobs.
  • references/prompts_and_instructions.md: required when using prompt-tuned bases (E5, BGE, GTE, Qwen3-Embedding, Instructor, Nomic, etc.) or adding query: / passage: style prefixes.
Show full SKILL.md (385 more words)Show less
Variant scripts (open when the task matches)
  • [SentenceTransformer] scripts/train_sentence_transformer_<matryoshka|multi_dataset|with_lora|distillation|make_multilingual|static_embedding>_example.py.
  • [CrossEncoder] scripts/train_cross_encoder_<distillation|listwise>_example.py.
  • [SparseEncoder] scripts/train_sparse_encoder_distillation_example.py.
  • Hard-negative mining CLI: scripts/mine_hard_negatives.py.

3. Defaults

Override only if the user specifies otherwise:

  • Local execution. Pitch HF Jobs only if local hardware can't fit the job.
  • Single run. After it completes, propose experimentation if the user would benefit (weak/marginal verdict, "see how high you can push it" framing, etc.). Iteration rules in references/training_args.md (Experimentation section).
  • Public Hub push at end-of-run, wrapped in try-except. On HF Jobs (ephemeral env) ALSO enable in-trainer push (push_to_hub=True + hub_strategy="every_save"). Details in references/hf_jobs_execution.md.

4. Constraints the produced script must satisfy

These are non-negotiable contracts. Implementation lives in the production templates and references. Do not reinvent.

  • Capture the pre-training evaluator score as baseline_eval before trainer.train().
  • Emit a single end-of-run line: VERDICT: WIN|MARGINAL|REGRESSION | score=... | baseline=... | delta=.... A monitor scrapes for this.
  • Silence httpx, httpcore, huggingface_hub, urllib3, filelock, fsspec to WARNING (otherwise HF download URLs flood the agent's context).
  • Tee logs to logs/{RUN_NAME}.log.
  • End with model.push_to_hub(...) wrapped in try/except.
  • Smoke-test before any long run (max_steps=1 + tiny dataset slice). The production templates show one common pattern (SMOKE_TEST env var).
  • [CrossEncoder] Include EarlyStoppingCallback(patience>=3). CE rerankers often peak mid-training and regress.
  • [SparseEncoder] Log query_active_dims / corpus_active_dims on the verdict line. High nDCG with collapsed sparsity is not a win. The keys come back name-prefixed (e.g. ..._query_active_dims). Use suffix matching to pluck them. See the SPARSE production template for the exact pattern.
  • [MultiVectorEncoder] Match scale to the scoring mode on any MNRL-family loss: near 1.0 for unnormalized MaxSim (do not copy scale=20.0 from bi-encoder MNRL), roughly the average query length with length-normalized MeanMaxSim, since each score is divided by its query's token count. XTRScores is a train-only similarity_fct: the evaluators reject it, so evaluation always scores with MaxSim, including for XTR-trained models.

5. Workflow

  1. Identify the model type (§1). Ask if ambiguous.
  2. Load the §2 required-reading files for that type.
  3. Open scripts/train_<type>_example.py and copy it as your starting point.
  4. Replace MODEL_NAME, DATASET_NAME, RUN_NAME, the loss, and the evaluator with the user's task. Cross-check loss/data-shape match against references/losses_<type>.md. Cross-check the metric_for_best_model key against references/evaluators_<type>.md (named evaluators format the key as eval_{name}_{primary_metric}).
  5. Smoke-test (max_steps=1).
  6. Run.
  7. After the run, append to logs/experiments.md and propose iteration if the verdict is weak/marginal.

Prerequisites

bash
pip install "sentence-transformers[train]>=5.0"        # add [train,image] / [audio] / [video] for [SentenceTransformer] multimodal
                                                       # [MultiVectorEncoder] requires >=6.0
pip install trackio                                    # optional tracker (or wandb / tensorboard / mlflow)
hf auth login                                          # or set HF_TOKEN with write scope (for Hub push)

GPU strongly recommended. CPU works only for demos and [SentenceTransformer] StaticEmbedding.

© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 30 other files (scripts, references) in skills/train-sentence-transformers of huggingface/skills.

  • SKILL.md
  • references/base_model_selection.md
  • references/dataset_formats.md
  • references/evaluators_cross_encoder.md
  • references/evaluators_multi_vector_encoder.md
  • references/evaluators_sentence_transformer.md
  • references/evaluators_sparse_encoder.md
  • references/hardware_guide.md
  • references/hf_jobs_execution.md
  • references/losses_cross_encoder.md
  • references/losses_multi_vector_encoder.md
  • references/losses_sentence_transformer.md
  • references/losses_sparse_encoder.md
  • references/model_architectures.md
  • references/prompts_and_instructions.md
  • references/training_args.md
  • references/troubleshooting.md
  • scripts/mine_hard_negatives.py
  • scripts/train_cross_encoder_distillation_example.py
  • … and 12 more

Open the folder on GitHubat commit ca0325b

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sentence-Transformers Training Router next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sentence-Transformers Training Router compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sentence-Transformers Training Router this skillhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Discover MLrand/cc-polymath1811 repos~574Automated safety check: PassMIT
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Discover ML

    rand/cc-polymath

    Automatically discover machine learning and AI skills when working with machine learning, PyTorch, training, inference, RAG, embeddings, fine-tuning, LLM, DSPy, HuggingFace, or diffusion models.

    181 GitHub starsUsed in 1 repo~574 tokens
    AI & LLM EngineeringAuto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface Vision Trainer

    waybarrios/opencode-power-pack

    Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

    533 GitHub stars~2.7k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from huggingface/skills

All 25 skills in this repo
  • Official

    Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

    11k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    Auto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed
  • Official

    Indexes research papers on the Hugging Face Hub from arXiv, links them to models and datasets, claims authorship and generates markdown research articles from templates.

    11k GitHub starsUsed in 5 repos~4.2k tokens
    Auto-check passed

Works with

Questions about Sentence-Transformers Training Router

What does Sentence-Transformers Training Router do?

Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models. The skill describes itself as a router rather than a manual: it identifies which of four model types fits the request, SentenceTransformer bi-encoders for retrieval and similarity, CrossEncoder rerankers for scoring query-passage pairs, SparseEncoder for SPLADE-style learned-sparse retrieval, and MultiVectorEncoder for ColBERT-style late interaction, using keyword tiebreakers such as rerank, SPLADE or MaxSim when the request is ambiguous, and asking when it still is not clear.

When should I use Sentence-Transformers Training Router?

Sentence-Transformers Training Router fits situations like: fine-tuning an embedding model for retrieval or similarity search; training a reranker to score query and passage pairs; training a SPLADE sparse retrieval model or a ColBERT-style multi-vector model.

How do I install Sentence-Transformers Training Router in Claude Code?

Run `npx skills add huggingface/skills --skill train-sentence-transformers -a claude-code`. Or copy the skill folder (skills/train-sentence-transformers in huggingface/skills) into .claude/skills/train-sentence-transformers in your project. Claude Code loads it when a task matches its description.

How do I install Sentence-Transformers Training Router in Codex?

Run `npx skills add huggingface/skills --skill train-sentence-transformers -a codex`. Or copy the skill folder (skills/train-sentence-transformers in huggingface/skills) into .agents/skills/train-sentence-transformers in your project. Codex loads it when a task matches its description.

Can I use Sentence-Transformers Training Router in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill train-sentence-transformers -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/train-sentence-transformers, .gemini/skills/train-sentence-transformers, .github/skills/train-sentence-transformers and .opencode/skills/train-sentence-transformers in your project.

What does Sentence-Transformers Training Router need to run?

Going by SKILL.md and its folder, Sentence-Transformers Training Router needs Python for the scripts in its folder, the command-line tools its instructions call (pip and hf) and credentials named HF_TOKEN. Our summary lists: Python with sentence-transformers installed; A Hugging Face account, to publish a trained model.

Does Sentence-Transformers Training Router access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Sentence-Transformers Training Router safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Sentence-Transformers Training Router use?

Sentence-Transformers Training Router is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sentence-Transformers Training Router use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 33k tokens, read only when the agent opens those files.

What are the alternatives to Sentence-Transformers Training Router?

Skills that share tags, products or a category with Sentence-Transformers Training Router: Discover ML (rand/cc-polymath, 181 stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars), Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sentence-Transformers Training Router?

huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,148 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.

Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.