Hugging Face LLM Trainer
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
A skill your agent uses for reinforcement learning tasks including training RL agents (PPO, SAC, DQN, TD3, DDPG, A2C, etc.), creating custom Gym environments, implementing callbacks for monitoring…
$ npx skills add davila7/claude-code-templates --skill stable-baselines3 -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install davila7/claude-code-templates stable-baselines3 --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/scientific/stable-baselines3 .claude/skills/stable-baselines3 && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "stable-baselines3" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/scientific/stable-baselines3 into .claude/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/scientific/stable-baselines3Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add davila7/claude-code-templates --skill stable-baselines3 -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install davila7/claude-code-templates stable-baselines3 --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .agents/skills && cp -r skills-src/cli-tool/components/skills/scientific/stable-baselines3 .agents/skills/stable-baselines3 && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "stable-baselines3" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/scientific/stable-baselines3 into .agents/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davila7/claude-code-templates --skill stable-baselines3 -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install davila7/claude-code-templates stable-baselines3 --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/cli-tool/components/skills/scientific/stable-baselines3 .cursor/skills/stable-baselines3 && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "stable-baselines3" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/scientific/stable-baselines3 into .cursor/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/davila7/claude-code-templates.git --path cli-tool/components/skills/scientific/stable-baselines3--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add davila7/claude-code-templates --skill stable-baselines3 -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install davila7/claude-code-templates stable-baselines3 --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/cli-tool/components/skills/scientific/stable-baselines3 .gemini/skills/stable-baselines3 && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "stable-baselines3" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/scientific/stable-baselines3 into .gemini/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install davila7/claude-code-templates stable-baselines3Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add davila7/claude-code-templates --skill stable-baselines3 -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .github/skills && cp -r skills-src/cli-tool/components/skills/scientific/stable-baselines3 .github/skills/stable-baselines3 && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "stable-baselines3" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/scientific/stable-baselines3 into .github/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add davila7/claude-code-templates --skill stable-baselines3 -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install davila7/claude-code-templates stable-baselines3 --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/cli-tool/components/skills/scientific/stable-baselines3 .opencode/skills/stable-baselines3 && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "stable-baselines3" agent skill from https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/scientific/stable-baselines3 into .opencode/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
stable-baselines3A skill your agent uses for reinforcement learning tasks including training RL agents (PPO, SAC, DQN, TD3, DDPG, A2C, etc.), creating custom Gym environments, implementing callbacks for monitoring…
Stable Baselines3 is an agent skill from davila7/claude-code-templates. Use this skill for reinforcement learning tasks including training RL agents (PPO, SAC, DQN, TD3, DDPG, A2C, etc.), creating custom Gym environments, implementing callbacks for monitoring and control, using vectorized environments for parallel training, and integrating with deep RL workflows. This skill should be used when users request RL algorithm implementation, agent training, environment design, or RL experimentation.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/algorithms.md`, `references/callbacks.md` and `references/custom_environments.md`).
It sits in AI & LLM Engineering, covering Reinforcement learning. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Stable Baselines3 loads about 2.4k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 628 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 628 words, ~2,372 tokens.
.claude/skills/stable-baselines3/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Stable Baselines3 (SB3) is a PyTorch-based library providing reliable implementations of reinforcement learning algorithms. This skill provides comprehensive guidance for training RL agents, creating custom environments, implementing callbacks, and optimizing training workflows using SB3's unified API.
Basic Training Pattern:
import gymnasium as gym
from stable_baselines3 import PPO
# Create environment
env = gym.make("CartPole-v1")
# Initialize agent
model = PPO("MlpPolicy", env, verbose=1)
# Train the agent
model.learn(total_timesteps=10000)
# Save the model
model.save("ppo_cartpole")
# Load the model (without prior instantiation)
model = PPO.load("ppo_cartpole", env=env)Important Notes:
total_timesteps is a lower bound; actual training may exceed this due to batch collectionmodel.load() as a static method, not on an existing instanceAlgorithm Selection:
Use references/algorithms.md for detailed algorithm characteristics and selection guidance. Quick reference:
See scripts/train_rl_agent.py for a complete training template with best practices.
Requirements:
Custom environments must inherit from gymnasium.Env and implement:
__init__(): Define action_space and observation_spacereset(seed, options): Return initial observation and info dictstep(action): Return observation, reward, terminated, truncated, inforender(): Visualization (optional)close(): Cleanup resourcesKey Constraints:
np.uint8 in range [0, 255]normalize_images=False in policy_kwargs if pre-normalizedDiscrete or MultiDiscrete spaces with start!=0Validation:
from stable_baselines3.common.env_checker import check_env
check_env(env, warn=True)See scripts/custom_env_template.py for a complete custom environment template and references/custom_environments.md for comprehensive guidance.
Purpose: Vectorized environments run multiple environment instances in parallel, accelerating training and enabling certain wrappers (frame-stacking, normalization).
Types:
Quick Setup:
from stable_baselines3.common.env_util import make_vec_env
# Create 4 parallel environments
env = make_vec_env("CartPole-v1", n_envs=4, vec_env_cls=SubprocVecEnv)
model = PPO("MlpPolicy", env, verbose=1)
model.learn(total_timesteps=25000)Off-Policy Optimization:
When using multiple environments with off-policy algorithms (SAC, TD3, DQN), set gradient_steps=-1 to perform one gradient update per environment step, balancing wall-clock time and sample efficiency.
API Differences:
reset() returns only observations (info available in vec_env.reset_infos)step() returns 4-tuple: (obs, rewards, dones, infos) not 5-tupleinfos[env_idx]["terminal_observation"]See references/vectorized_envs.md for detailed information on wrappers and advanced usage.
Purpose: Callbacks enable monitoring metrics, saving checkpoints, implementing early stopping, and custom training logic without modifying core algorithms.
Common Callbacks:
Custom Callback Structure:
from stable_baselines3.common.callbacks import BaseCallback
class CustomCallback(BaseCallback):
def _on_training_start(self):
# Called before first rollout
pass
def _on_step(self):
# Called after each environment step
# Return False to stop training
return True
def _on_rollout_end(self):
# Called at end of rollout
passAvailable Attributes:
self.model: The RL algorithm instanceself.num_timesteps: Total environment stepsself.training_env: The training environmentChaining Callbacks:
from stable_baselines3.common.callbacks import CallbackList
callback = CallbackList([eval_callback, checkpoint_callback, custom_callback])
model.learn(total_timesteps=10000, callback=callback)See references/callbacks.md for comprehensive callback documentation.
Saving and Loading:
# Save model
model.save("model_name")
# Save normalization statistics (if using VecNormalize)
vec_env.save("vec_normalize.pkl")
# Load model
model = PPO.load("model_name", env=env)
# Load normalization statistics
vec_env = VecNormalize.load("vec_normalize.pkl", vec_env)Parameter Access:
# Get parameters
params = model.get_parameters()
# Set parameters
model.set_parameters(params)
# Access PyTorch state dict
state_dict = model.policy.state_dict()Evaluation:
from stable_baselines3.common.evaluation import evaluate_policy
mean_reward, std_reward = evaluate_policy(
model,
env,
n_eval_episodes=10,
deterministic=True
)Video Recording:
from stable_baselines3.common.vec_env import VecVideoRecorder
# Wrap environment with video recorder
env = VecVideoRecorder(
env,
"videos/",
record_video_trigger=lambda x: x % 2000 == 0,
video_length=200
)See scripts/evaluate_agent.py for a complete evaluation and recording template.
Learning Rate Schedules:
def linear_schedule(initial_value):
def func(progress_remaining):
# progress_remaining goes from 1 to 0
return progress_remaining * initial_value
return func
model = PPO("MlpPolicy", env, learning_rate=linear_schedule(0.001))Multi-Input Policies (Dict Observations):
model = PPO("MultiInputPolicy", env, verbose=1)Use when observations are dictionaries (e.g., combining images with sensor data).
Hindsight Experience Replay:
from stable_baselines3 import SAC, HerReplayBuffer
model = SAC(
"MultiInputPolicy",
env,
replay_buffer_class=HerReplayBuffer,
replay_buffer_kwargs=dict(
n_sampled_goal=4,
goal_selection_strategy="future",
),
)TensorBoard Integration:
model = PPO("MlpPolicy", env, tensorboard_log="./tensorboard/")
model.learn(total_timesteps=10000)Starting a New RL Project:
references/algorithms.md for selection guidancescripts/custom_env_template.py if neededcheck_env() before trainingscripts/train_rl_agent.py as starting templatescripts/evaluate_agent.py for assessmentCommon Issues:
buffer_size for off-policy algorithms or use fewer parallel environmentsstable_baselines3 is installed: uv pip install stable-baselines3[extra]train_rl_agent.py: Complete training script template with best practicesevaluate_agent.py: Agent evaluation and video recording templatecustom_env_template.py: Custom Gym environment templatealgorithms.md: Detailed algorithm comparison and selection guidecustom_environments.md: Comprehensive custom environment creation guidecallbacks.md: Complete callback system referencevectorized_envs.md: Vectorized environment usage and wrappers# Basic installation
uv pip install stable-baselines3
# With extra dependencies (Tensorboard, etc.)
uv pip install stable-baselines3[extra]© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in cli-tool/components/skills/scientific/stable-baselines3 of davila7/claude-code-templates.
Open the folder on GitHubat commit 14680ec
We found 20 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 9 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.
Stable Baselines3 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Stable Baselines3 this skilldavila7/claude-code-templates | 32k | 9 repos | ~2.4k | Automated safety check: Pass | MIT | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 3 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| Train RlOpenPipe/ART | 11k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~830 | Automated safety check: Pass | Apache-2.0 | |
| Fine Tuning With TrlOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Optim AgentOptim-Agent/optim-agent | 801 | — | ~1.3k | Automated safety check: Pass | MIT |
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
OpenPipe/ART
RL training reference for the ART framework. An agent skill from OpenPipe/ART.
R6410418/Jackrong-llm-finetuning-guide
Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.
Orchestra-Research/AI-Research-SKILLs
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.
Optim-Agent/optim-agent
A skill your agent uses when the user wants to optimize configurable system parameters against a measurable scalar objective, especially for model training, inference, quantitative strategies…
AI45Lab/SAfactory
Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.
davila7/claude-code-templates
Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.
davila7/claude-code-templates
Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.
davila7/claude-code-templates
Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.
davila7/claude-code-templates
Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.
davila7/claude-code-templates
Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.
davila7/claude-code-templates
Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.
Categories
A skill your agent uses for reinforcement learning tasks including training RL agents (PPO, SAC, DQN, TD3, DDPG, A2C, etc.), creating custom Gym environments, implementing callbacks for monitoring…. Stable Baselines3 is an agent skill from davila7/claude-code-templates.), creating custom Gym environments, implementing callbacks for monitoring and control, using vectorized environments for parallel training, and integrating with deep RL workflows.
Stable Baselines3 fits situations like: reinforcement learning tasks including training RL agents (PPO; creating custom Gym environments; implementing callbacks for monitoring and control; using vectorized environments for parallel training.
Run `npx skills add davila7/claude-code-templates --skill stable-baselines3 -a claude-code`. Or copy the skill folder (cli-tool/components/skills/scientific/stable-baselines3 in davila7/claude-code-templates) into .claude/skills/stable-baselines3 in your project. Claude Code loads it when a task matches its description.
Run `npx skills add davila7/claude-code-templates --skill stable-baselines3 -a codex`. Or copy the skill folder (cli-tool/components/skills/scientific/stable-baselines3 in davila7/claude-code-templates) into .agents/skills/stable-baselines3 in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill stable-baselines3 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stable-baselines3, .gemini/skills/stable-baselines3, .github/skills/stable-baselines3 and .opencode/skills/stable-baselines3 in your project.
Going by SKILL.md and its folder, Stable Baselines3 needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Stable Baselines3 is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Stable Baselines3: Hugging Face LLM Trainer (huggingface/skills, 11k stars), Train Rl (OpenPipe/ART, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Fine Tuning With Trl (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.
Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.