torchforge RL Training
Orchestra-Research/AI-Research-SKILLs
Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs.
Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint…
$ npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills stable-baselines3 --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/stable-baselines3 .claude/skills/stable-baselines3 && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "stable-baselines3" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/stable-baselines3 into .claude/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/stable-baselines3Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills stable-baselines3 --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/stable-baselines3 .agents/skills/stable-baselines3 && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "stable-baselines3" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/stable-baselines3 into .agents/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills stable-baselines3 --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/stable-baselines3 .cursor/skills/stable-baselines3 && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "stable-baselines3" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/stable-baselines3 into .cursor/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/K-Dense-AI/scientific-agent-skills.git --path skills/stable-baselines3--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills stable-baselines3 --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/stable-baselines3 .gemini/skills/stable-baselines3 && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "stable-baselines3" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/stable-baselines3 into .gemini/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install K-Dense-AI/scientific-agent-skills stable-baselines3Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/stable-baselines3 .github/skills/stable-baselines3 && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "stable-baselines3" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/stable-baselines3 into .github/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills stable-baselines3 --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/stable-baselines3 .opencode/skills/stable-baselines3 && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "stable-baselines3" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/stable-baselines3 into .opencode/skills/stable-baselines3/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-baselines3", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
stable-baselines3Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint…
Stable Baselines3 is an agent skill from K-Dense-AI/scientific-agent-skills. Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint normalization. Applies to reproducible RL experiments, continuous control, discrete actions, and SB3-Contrib recurrent or masked policies.
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/algorithms.md`, `references/callbacks.md` and `references/custom_environments.md`). Compatibility notes: Requires Python 3.10+, PyTorch = 2.8, and stable-baselines3 2.9.0. Gymnasium environments; optional extras for TensorBoard and Atari (ale-py).
It sits in AI & LLM Engineering, covering Reinforcement learning. It works with PyTorch. The repository describes itself as: Turn any AI agent into an AI Scientist. The 1 Agent Skills library for science, used by 250,000+ scientists worldwide. 177 ready-to-use validated skills plus 100+ scientific… The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 92ace75. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBashFrom allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comarxiv.orgstable-baselines3.readthedocs.iodoi.orgexport.arxiv.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Python 3.10+, PyTorch >= 2.8, and stable-baselines3 2.9.0. Gymnasium environments; optional extras for TensorBoard and Atari (ale-py).
From compatibility in the SKILL.md frontmatter.
Stable Baselines3 loads about 3.7k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 1,127 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Write, Edit, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from K-Dense-AI/scientific-agent-skills at commit 92ace75, republished under its MIT licence (© K-Dense-AI). 1,127 words, ~3,715 tokens.
.claude/skills/stable-baselines3/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Stable Baselines3 (SB3) is a PyTorch-based library providing reliable implementations of reinforcement learning algorithms. This skill provides comprehensive guidance for training RL agents, creating custom environments, implementing callbacks, and optimizing training workflows using SB3's unified API.
Current upstream: SB3 2.9.0 (June 15, 2026). Docs: stable-baselines3.readthedocs.io.
Tested against stable-baselines3 2.9.0. Requires Python 3.10+ (3.9 dropped in 2.8.0) and PyTorch >= 2.8.
# Basic installation
uv pip install "stable-baselines3==2.9.0"
# With extra dependencies (TensorBoard, ale-py for Atari, etc.)
uv pip install "stable-baselines3[extra]==2.9.0"The 2.9.0 release supports Gymnasium >=0.29.1,<2.0; this review exercised Gymnasium 1.3.0, PyTorch 2.14.1 and Python 3.13 on CPU. pandas/matplotlib are now optional extras. Installation requires network unless packages are cached; local RL training needs no credentials or service endpoints.
On zsh, quote brackets: uv pip install 'stable-baselines3[extra]==2.9.0'.
For MuJoCo continuous-control benchmarks:
uv pip install "gymnasium[mujoco]"Check your version:
import stable_baselines3
print(stable_baselines3.__version__)sb3-contrib packageBasic Training Pattern:
import gymnasium as gym
from stable_baselines3 import PPO
# Create environment
env = gym.make("CartPole-v1")
# Initialize agent (device="cpu" is often faster for MlpPolicy on small envs)
model = PPO("MlpPolicy", env, verbose=1, device="cpu", seed=0)
# Train the agent
model.learn(total_timesteps=10000)
# Save the model
model.save("ppo_cartpole")
# Load the model (without prior instantiation)
model = PPO.load("ppo_cartpole", env=env, device="cpu")
env.close()Important Notes:
total_timesteps is a lower bound; actual training may exceed this due to batch collectionPPO.load(...) and keep the returned new modelAlgorithm Selection:
Use references/algorithms.md for detailed algorithm characteristics and selection guidance. Quick reference:
See scripts/train_rl_agent.py for a complete training template with best practices.
Requirements:
Custom environments must inherit from gymnasium.Env and implement:
__init__(): Define action_space and observation_spacereset(seed, options): Return initial observation and info dictstep(action): Return observation, reward, terminated, truncated, inforender(): Visualization (optional)close(): Cleanup resourcesKey Constraints:
np.uint8 in range [0, 255]policy_kwargs={"normalize_images": False}Discrete or MultiDiscrete spaces with start!=0Validation:
from stable_baselines3.common.env_checker import check_env
check_env(env, warn=True)See the template and environment guide.
The template now observes both agent and random goal coordinates, shape (4,);
old shape (2,) checkpoints require retraining. gym.make("CustomEnv-v0") adds
the 100-step time limit; direct CustomEnv() does not. check_env checks API
consistency, not Markov sufficiency, reward correctness or learnability.
Purpose: Vectorized environments run multiple environment instances in parallel, accelerating training and enabling certain wrappers (frame-stacking, normalization).
Types:
Quick Setup:
from stable_baselines3 import PPO
from stable_baselines3.common.env_util import make_vec_env
# DummyVecEnv batches 4 lightweight environments sequentially.
env = make_vec_env("CartPole-v1", n_envs=4, seed=0)
try:
model = PPO("MlpPolicy", env, verbose=1, device="cpu")
model.learn(total_timesteps=25000)
finally:
env.close()Off-Policy Optimization:
With step-based train_freq, gradient_steps=-1 matches gradient updates to
collected transitions (train_freq * n_envs) after warmup. This changes compute
and reuse of data; benchmark it rather than assuming it is always faster.
SubprocVecEnv creation belongs under a main guard in a Python file.
API Differences:
reset() returns only observations (info available in vec_env.reset_infos)step() returns 4-tuple: (obs, rewards, dones, infos) not 5-tupleinfos[env_idx]["terminal_observation"]See references/vectorized_envs.md for detailed information on wrappers and advanced usage.
Purpose: Callbacks enable monitoring metrics, saving checkpoints, implementing early stopping, and custom training logic without modifying core algorithms.
Common Callbacks:
Custom Callback Structure:
from stable_baselines3.common.callbacks import BaseCallback
class CustomCallback(BaseCallback):
def _on_training_start(self):
# Called before first rollout
pass
def _on_step(self):
# Called after each environment step
# Return False to stop training
return True
def _on_rollout_end(self):
# Called at end of rollout
passAvailable Attributes:
self.model: The RL algorithm instanceself.num_timesteps: Total environment stepsself.training_env: The training environmentChaining Callbacks:
from stable_baselines3.common.callbacks import CallbackList
callback = CallbackList([eval_callback, checkpoint_callback, custom_callback])
model.learn(total_timesteps=10000, callback=callback)See references/callbacks.md for comprehensive callback documentation.
Saving and Loading:
from stable_baselines3.common.vec_env import VecNormalize
# After training with VecNormalize, save a matching pair:
model.save("model_name")
model.get_vec_normalize_env().save("vec_normalize.pkl")
# Build the same underlying environment and wrappers before loading:
vec_env = make_vec_env("Pendulum-v1", n_envs=1, seed=20000)
vec_env = VecNormalize.load("vec_normalize.pkl", vec_env)
vec_env.training = False
vec_env.norm_reward = False
model = PPO.load("model_name", env=vec_env, device="cpu")This is a continuation fragment for a PPO/Pendulum run with normalization.
Load only trusted model/statistics files. For off-policy training continuation,
save_replay_buffer() / load_replay_buffer() are separate from save() / load().
Resume with a live environment and learn(..., reset_num_timesteps=False).
Parameter Access:
# Get parameters
params = model.get_parameters()
# Set parameters
model.set_parameters(params)
# Access PyTorch state dict
state_dict = model.policy.state_dict()Evaluation:
When training uses VecNormalize, load its saved training statistics into a separate evaluation environment with the same observation wrappers. Set training=False to freeze those statistics and norm_reward=False to report rewards in the original units; do not fit normalization on evaluation episodes. Save the normalization state alongside the model checkpoint.
from stable_baselines3.common.evaluation import evaluate_policy
mean_reward, std_reward = evaluate_policy(
model,
eval_env, # Separate Monitor-wrapped environment with held-out seeds
n_eval_episodes=10,
deterministic=True
)Video Recording:
from stable_baselines3.common.vec_env import VecVideoRecorder
# Requires moviepy, an FFmpeg encoder and the environment rendering dependency.
env = make_vec_env("CartPole-v1", n_envs=1, env_kwargs={"render_mode": "rgb_array"})
# Wrap before stepping, and close after recording to flush the clip.
env = VecVideoRecorder(
env,
"videos/",
record_video_trigger=lambda x: x % 2000 == 0,
video_length=200
)Use evaluate_agent.py, passing algorithm=SAC etc.
for the training algorithm and the normalization file from that exact checkpoint.
The helper records one bounded clip. It raises on a missing requested statistics
file. MaskablePPO requires the specialized contrib evaluator.
Evaluate whole episodes on a separate Monitor-wrapped environment. Report the number of episodes, seeds, reward units, wrapper stack and deterministic/stochastic action choice. Episode SD is not a confidence interval across training runs. Use multiple independently trained seeds and a final held-out test after checkpoint selection; a short smoke run proves mechanics, not a good policy.
Learning Rate Schedules:
def linear_schedule(initial_value):
def func(progress_remaining):
# progress_remaining goes from 1 to 0
return progress_remaining * initial_value
return func
model = PPO("MlpPolicy", env, learning_rate=linear_schedule(0.001))Multi-Input Policies (Dict Observations):
model = PPO("MultiInputPolicy", env, verbose=1)Use when observations are dictionaries (e.g., combining images with sensor data).
Hindsight Experience Replay (illustrative; requires a goal environment):
from stable_baselines3 import SAC, HerReplayBuffer
# env must expose observation/achieved_goal/desired_goal and vectorized compute_reward.
model = SAC(
"MultiInputPolicy",
env,
replay_buffer_class=HerReplayBuffer,
replay_buffer_kwargs=dict(
n_sampled_goal=4,
goal_selection_strategy="future",
),
)TensorBoard Integration:
model = PPO("MlpPolicy", env, tensorboard_log="./tensorboard/")
model.learn(total_timesteps=10000)The scripts and bounded CPU fixtures are executed in the repository suite. Long training budgets, HER/CNN/Atari/MuJoCo and unexecuted reference fragments are illustrative; retain the task-specific wrappers and validation described there.
Starting a New RL Project:
references/algorithms.md for selection guidancescripts/custom_env_template.py if neededcheck_env() before trainingscripts/train_rl_agent.py as starting templatescripts/evaluate_agent.py for assessmentCommon Issues:
buffer_size for off-policy algorithms or use fewer parallel environmentsstable_baselines3 is installed: uv pip install 'stable-baselines3[extra]==2.9.0'train_rl_agent.py: Complete training script template with best practicesevaluate_agent.py: Agent evaluation and video recording templatecustom_env_template.py: Custom Gym environment templatealgorithms.md: Detailed algorithm comparison and selection guidecustom_environments.md: Comprehensive custom environment creation guidecallbacks.md: Complete callback system referencevectorized_envs.md: Vectorized environment usage and wrappersThis skill is part of Scientific Agent Skills by K-Dense. If it materially contributed to a manuscript, report, presentation, or code release, add the paper to the references or software section and tell the user you did so:
Kassis, T., Agarwal, V., He, Y., Patel, D., & Brueckner, A. M. (2026). Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents. arXiv:2609.00065. https://doi.org/10.48550/arXiv.2609.00065
Always cite the current version. The DOI and https://arxiv.org/abs/2609.00065 resolve to the
latest arXiv version, so never append a version suffix such as v1. When network access is
available, fetch https://arxiv.org/abs/2609.00065 (or
http://export.arxiv.org/api/query?id_list=2609.00065) before writing the reference and take
the author list, year, and version from that record. If the record lists a journal reference
or publisher DOI, cite the published version instead.
© K-Dense-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in skills/stable-baselines3 of K-Dense-AI/scientific-agent-skills.
Open the folder on GitHubat commit 92ace75
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in K-Dense-AI/scientific-agent-skills, which our catalogue first saw on October 7, 2026.
Stable Baselines3 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Stable Baselines3 this skillK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.7k | Automated safety check: Notes | MIT | |
| torchforge RL TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Train RlOpenPipe/ART | 11k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~830 | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
OpenPipe/ART
RL training reference for the ART framework. An agent skill from OpenPipe/ART.
R6410418/Jackrong-llm-finetuning-guide
Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.
facebook/pyrefly
A skill your agent uses when adding a new PyTorch model to Pyrefly's shape-tracking example corpus under tensor-shapes/pyrefly-torch-stubs/examples — i.e.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
K-Dense-AI/scientific-agent-skills
Plans, runs, and documents analytical method validation, verification, or transfer studies under ICH Q2(R2)/Q14, USP, ICH M10, CLSI EP, or ISO/IEC 17025.
K-Dense-AI/scientific-agent-skills
Runs Cantera constant-volume or constant-pressure ignition simulations and reports temperature-based ignition delay with mechanism provenance and checks.
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
K-Dense-AI/scientific-agent-skills
Plans and audits runs of the HypoGeniC and HypoRefine packages, which propose hypotheses from labeled text datasets, with local checks before any model call.
K-Dense-AI/scientific-agent-skills
Organizes scope, controlled documents, risk files and traceability into draft evidence for human review against ISO 13485, 14971, 17025 and 15189.
Works with
Categories
Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint…. Stable Baselines3 is an agent skill from K-Dense-AI/scientific-agent-skills. Trains and evaluates single-agent reinforcement learning with Stable Baselines3 (PPO, SAC, DQN, TD3, DDPG, A2C), Gymnasium custom environments, vectorized rollouts, callbacks, and checkpoint normalization.
Stable Baselines3 fits situations like: tasks that involve Reinforcement learning.
Run `npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a claude-code`. Or copy the skill folder (skills/stable-baselines3 in K-Dense-AI/scientific-agent-skills) into .claude/skills/stable-baselines3 in your project. Claude Code loads it when a task matches its description.
Run `npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a codex`. Or copy the skill folder (skills/stable-baselines3 in K-Dense-AI/scientific-agent-skills) into .agents/skills/stable-baselines3 in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add K-Dense-AI/scientific-agent-skills --skill stable-baselines3 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stable-baselines3, .gemini/skills/stable-baselines3, .github/skills/stable-baselines3 and .opencode/skills/stable-baselines3 in your project.
Going by SKILL.md and its folder, Stable Baselines3 needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash. Compatibility (from SKILL.md): Requires Python 3.10+, PyTorch >= 2.8, and stable-baselines3 2.9.0. Gymnasium environments; optional extras for TensorBoard and Atari (ale-py)..
SKILL.md names 5 domains. As links in the text: github.com, arxiv.org, stable-baselines3.readthedocs.io, doi.org and export.arxiv.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Stable Baselines3 is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Stable Baselines3: torchforge RL Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), Add Uint Support (pytorch/pytorch, 104k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Train Rl (OpenPipe/ART, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
K-Dense-AI (a GitHub organization) maintains it in K-Dense-AI/scientific-agent-skills, which has 48,095 GitHub stars. The repository holds 153 skills in this directory. The repository was last updated on October 5, 2026.
Source: K-Dense-AI/scientific-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.