Sentence-Transformers Training Router
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR).
$ npx skills add wshobson/agents --skill grpo-rlvr-training -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents grpo-rlvr-training --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/grpo-rlvr-training .claude/skills/grpo-rlvr-training && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "grpo-rlvr-training" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/grpo-rlvr-training into .claude/skills/grpo-rlvr-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grpo-rlvr-training", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/grpo-rlvr-trainingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill grpo-rlvr-training -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents grpo-rlvr-training --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/llm-finetuning/skills/grpo-rlvr-training .agents/skills/grpo-rlvr-training && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "grpo-rlvr-training" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/grpo-rlvr-training into .agents/skills/grpo-rlvr-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grpo-rlvr-training", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill grpo-rlvr-training -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents grpo-rlvr-training --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/llm-finetuning/skills/grpo-rlvr-training .cursor/skills/grpo-rlvr-training && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "grpo-rlvr-training" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/grpo-rlvr-training into .cursor/skills/grpo-rlvr-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grpo-rlvr-training", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/llm-finetuning/skills/grpo-rlvr-training--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill grpo-rlvr-training -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents grpo-rlvr-training --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/llm-finetuning/skills/grpo-rlvr-training .gemini/skills/grpo-rlvr-training && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "grpo-rlvr-training" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/grpo-rlvr-training into .gemini/skills/grpo-rlvr-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grpo-rlvr-training", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents grpo-rlvr-trainingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill grpo-rlvr-training -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/llm-finetuning/skills/grpo-rlvr-training .github/skills/grpo-rlvr-training && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "grpo-rlvr-training" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/grpo-rlvr-training into .github/skills/grpo-rlvr-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grpo-rlvr-training", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill grpo-rlvr-training -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents grpo-rlvr-training --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/llm-finetuning/skills/grpo-rlvr-training .opencode/skills/grpo-rlvr-training && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "grpo-rlvr-training" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/grpo-rlvr-training into .opencode/skills/grpo-rlvr-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grpo-rlvr-training", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
grpo-rlvr-trainingTrain reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR).
Grpo Rlvr Training is an agent skill from wshobson/agents. Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/grpo-memory.md` and `references/reward-functions.md`).
It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Grpo Rlvr Training loads about 1.9k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 910 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 910 words, ~1,944 tokens.
.claude/skills/grpo-rlvr-training/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.This skill assumes finetuning-method-selection
already routed here because the target behavior
has a verifiable pass/fail signal — not
demonstrations (lora-qlora-recipes) or
preference pairs (preference-optimization).
What follows is when RL is the right tool, the
reference recipe, the mandatory reward-inspection
gate, and how to pick a GRPO variant when the
base recipe misbehaves.
Input: a routing decision (RLVR via GRPO)
plus a verifier (code executor, test suite,
schema checker, or grader) for the target task.
Output format: a validated GRPO config — the
kwarg values in references/grpo-memory.md and
the reward functions in
references/reward-functions.md, not free-form
advice — that llm-finetuning-training-engineer
consumes directly.
GRPO+RLVR only pays off when task success is
algorithmically checkable — a unit test
passes, a parser accepts the output, a tool call
matches an expected schema, a math answer matches
a ground truth. If grading the output requires
human judgment or a subjective rubric, that's an
eval-harness and judge-calibration problem first
— see eval-harness-first — not a reason to skip
straight to RL.
Before opening a GRPO run, confirm the model can sometimes succeed on the target task already. RL sharpens an existing capability by reweighting toward the samples that already work; it does not install a capability from zero.
lora-qlora-recipes) and only return to this
skill once the base success rate is nonzero.The standing rule for the whole plugin: DPO for
taste, GRPO for reasoning. If the signal is a
preference between two acceptable outputs, that's
preference-optimization, not this skill.
The reference recipe is TRL's GRPOTrainer with
vLLM-backed generation:
from trl import GRPOConfig, GRPOTrainer
grpo_args = GRPOConfig(
output_dir="./outputs-grpo",
use_vllm=True,
vllm_mode="colocate", # single GPU; "server" for multi-GPU
num_generations=8, # floor — fewer starves the group-relative baseline
learning_rate=5e-7, # settled range for GRPO
beta=0.01, # KL coefficient vs the reference policy
per_device_train_batch_size=8,
gradient_accumulation_steps=4,
bf16=True,
logging_steps=10,
seed=3407,
)
trainer = GRPOTrainer(
model=SFT_CHECKPOINT,
args=grpo_args,
reward_funcs=[format_reward, correctness_reward], # references/reward-functions.md
train_dataset=prompts, # prompt-only — GRPO generates its own completions
processing_class=tokenizer,
)
trainer.train()vllm_mode="colocate" runs generation and
training on the same GPU — the default for a
single-GPU box.vllm_mode="server" points at a separate
vLLM server process and is the multi-GPU path —
generation and training don't compete for the
same device.num_generations ≥ 8 is a floor, not a
suggestion: GRPO's advantage estimate is
relative to the group mean, and fewer than 8
samples per prompt produces a noisy baseline.learning_rate=5e-7 and beta=0.01 are
the settled starting point; deviate only after
the base run is stable and reward-inspected
(below).Memory sizing for this recipe by target size
class: references/grpo-memory.md.
Run the reward function against 50–100 sampled outputs and manually read the results before starting the actual training run. This is a gate, not a one-time sanity check.
If the reward function's judgment disagrees with a human reading of that sample, fix the reward function first. Training against an uninspected reward, or tuning hyperparameters to compensate for one silently scoring the wrong thing, is how a run reward-hacks: the model optimizes cleanly toward the wrong target, and that doesn't surface as a training-loop bug.
This inspection is a Phase 1 gate input for
/finetune — the same 50–100-sample read that
catches a broken reward function here is what that
command checks for before it lets a GRPO brief
proceed.
Complete reward function implementations to
inspect against — exact-match, schema-validation,
unit-test-execution, a length-penalty wrapper, and
a rubric-as-reward judge pattern:
references/reward-functions.md.
The base recipe above is the default. Reach for a variant only when a specific failure mode shows up, not preemptively:
| Failure mode | Variant | Why |
|---|---|---|
| Entropy collapse / degenerate long chain-of-thought | DAPO | Decouples clip bounds and relaxes the KL penalty that over-regularizes exploration on long reasoning traces |
| Reward or output length trends up regardless of quality | Dr.GRPO | Removes GRPO's length-normalization bias so reward tracks correctness, not completion length |
| Training a mixture-of-experts model | GSPO | Moves the importance-sampling ratio to the sequence level instead of per-token — per-token ratios are unstable on MoE routing, so GSPO is required here, not optional |
Start with plain GRPO. Watch for the specific symptom — collapsing entropy on long CoT, a length-reward correlation, or MoE instability — and only then swap in the matching variant above. Don't pre-select a variant before the base recipe has actually shown the failure mode.
Vision-language RL is not executed by this plugin in v1 — it's documented here for context, not as a runnable path. Tooling is fragmented across ms-swift and EasyR1-derived forks with no one-line TRL command yet, and naive text-only GRPO applied to a VLM tends to reward-hack by optimizing the text-reasoning trace while ignoring the image — the model learns to sound right without looking at the input. A VLM RL run is a research spike outside this skill's supported recipe, not a variant of The Recipe above.
references/reward-functions.md — complete
Python reward functions (exact-match
correctness, schema validation, unit-test
execution, a length-penalty wrapper, and a
rubric-as-reward judge pattern) to inspect under
The Inspection Rule before any training run.references/grpo-memory.md — memory sizing by
target size class, vLLM sleep-mode and
optimizer-state tactics, Unsloth's long-context
RL chunking, and the DGX Spark bandwidth caveat
for decode-heavy rollouts.Related skills: finetuning-method-selection
routes here once a verifiable pass/fail signal
exists; preference-optimization is the sibling
skill for preference pairs rather than verifiable
rewards; eval-harness-first covers judge
calibration for any reward that isn't purely
code-checkable. On DGX Spark, defer to the
dgx-spark-ops plugin's skills, when installed,
for the memory/thermal remediation ladder this
skill's memory table doesn't cover.
© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in plugins/llm-finetuning/skills/grpo-rlvr-training of wshobson/agents.
Open the folder on GitHubat commit 46891e7
Grpo Rlvr Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Grpo Rlvr Training this skillwshobson/agents | 40k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Sentence-Transformers Training Routerhuggingface/skills | 11k | 1 repos | ~2.6k | Automated safety check: Pass | Apache-2.0 | |
| Train RlOpenPipe/ART | 11k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~830 | Automated safety check: Pass | Apache-2.0 | |
| Dataset Evaluationawslabs/agent-plugins | 916 | 1 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Train SftOpenPipe/ART | 11k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
OpenPipe/ART
RL training reference for the ART framework. An agent skill from OpenPipe/ART.
R6410418/Jackrong-llm-finetuning-guide
Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.
awslabs/agent-plugins
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).
OpenPipe/ART
SFT training reference for the ART framework. An agent skill from OpenPipe/ART.
Orchestra-Research/AI-Research-SKILLs
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
Categories
Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Grpo Rlvr Training is an agent skill from wshobson/agents. Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR).
Grpo Rlvr Training fits situations like: task success is algorithmically checkable (math; structured output); designing GRPO reward functions; A GRPO run diverges.
Run `npx skills add wshobson/agents --skill grpo-rlvr-training -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/grpo-rlvr-training in wshobson/agents) into .claude/skills/grpo-rlvr-training in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill grpo-rlvr-training -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/grpo-rlvr-training in wshobson/agents) into .agents/skills/grpo-rlvr-training in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill grpo-rlvr-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grpo-rlvr-training, .gemini/skills/grpo-rlvr-training, .github/skills/grpo-rlvr-training and .opencode/skills/grpo-rlvr-training in your project.
SKILL.md names no scripts, command-line tools or credentials: Grpo Rlvr Training is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Grpo Rlvr Training is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Grpo Rlvr Training: Sentence-Transformers Training Router (huggingface/skills, 11k stars), Train Rl (OpenPipe/ART, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Dataset Evaluation (awslabs/agent-plugins, 916 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.