Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Training Hub Guide

skills CLI
$ npx skills add Red-Hat-AI-Innovation-Team/training_hub --skill training-hub-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Red-Hat-AI-Innovation-Team/training_hub training-hub-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Red-Hat-AI-Innovation-Team/training_hub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/training-hub-guide .claude/skills/training-hub-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
training-hub-guide
GitHub stars
100
Token cost
~2.8k tokens
SKILL.md length
1,100 words
Files
4
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves…

  • Works in 5 steps: uv cache clean → Remove GPU-related caches from ~/.cache/… → Remove ~/.triton/ if it exists (triton… → …
  • The user is working with traininghub
  • SKILL.md covers Installation, Algorithm selection, Hyperparameter configuration and Memory model and OOM…, plus 5 more sections
  • Calls uv

What it does

Training Hub Guide is an agent skill from Red-Hat-AI-Innovation-Team/training_hub. Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves, and leveraging backend-specific features. Use when the user is working with traininghub, fine-tuning language models, asking about SFT/OSFT/LoRA training, or debugging GPU/CUDA training issues.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `backend-kwargs.md`, `hyperparameter-guide.md` and `installation-troubleshooting.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. It works with CUDA. The repository describes itself as: An algorithm-focused interface for common llm training, continual learning, and reinforcement learning techniques. The licence is Apache-2.0.

When your agent uses it

  • The user is working with traininghub
  • Fine-tuning language models
  • Asking about SFT/OSFT/LoRA training
  • Debugging GPU/CUDA training issues

Example prompts

  • “Use the training-hub-guide skill to guide users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT…”
  • “/training-hub-guide”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. uv cache clean
  2. Remove GPU-related caches from ~/.cache/ (torch, triton, flash_attn, vllm, and similar)
  3. Remove ~/.triton/ if it exists (triton kernel cache)
  4. Delete the current venv and recreate it fresh
  5. Reinstall with the two-step process above

What it can do on your machine

Read from SKILL.md and the folder at commit 511a905. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ai-innovation.team

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Training Hub Guide loads about 2.8k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 1,100 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Red-Hat-AI-Innovation-Team/training_hub at commit 511a905, republished under its Apache-2.0 licence (© Red-Hat-AI-Innovation-Team). 1,100 words, ~2,780 tokens.

Download SKILL.mdSave it as .claude/skills/training-hub-guide/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
training-hub-guide
description
Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves, and leveraging backend-specific features. Use when the user is working with training_hub, fine-tuning language models, asking about SFT/OSFT/LoRA training, or debugging GPU/CUDA training issues.

Training Hub Guide

Training Hub is an abstraction layer for LLM post-training algorithms. It packages SFT, OSFT, and LoRA behind a unified interface so users do not need to learn multiple backend APIs. Backends are wired together internally; users interact with a single API surface.

For API reference and conceptual overviews, consult the live documentation at https://ai-innovation.team/training_hub/#/ and the docs/ directory in the repo root. This skill covers practical knowledge, decision frameworks, and troubleshooting that supplements the official docs.

Installation

Install targets
bash
# Minimal (no backends, no GPU training)
uv pip install training_hub

# SFT + OSFT (high-scale distributed fine-tuning via CUDA backends)
# IMPORTANT: base install MUST come first, then [cuda] with --no-build-isolation
uv pip install training_hub && uv pip install training_hub[cuda] --no-build-isolation

# LoRA (budget-friendly, single/few-GPU via Unsloth — does NOT require [cuda])
uv pip install training_hub[lora]

The [cuda] extra is only needed for SFT and OSFT algorithms. LoRA uses the Unsloth backend which handles its own CUDA dependencies through [lora].

The two-step install for [cuda] is required because flash-attn and other CUDA packages need torch and packaging to already be present at build time.

Third-party loggers

Loggers are not bundled. Install separately as needed:

bash
uv pip install wandb       # Weights & Biases
uv pip install mlflow      # MLflow
uv pip install tensorboard # TensorBoard
Fixing CUDA/kernel import errors

When users encounter errors like cannot import from flash_attn: unknown symbol or similar issues with optimized kernels (flash attention, liger, causal-conv1d, mamba-ssm), the root cause is usually stale cached builds. Fix with:

  1. uv cache clean
  2. Remove GPU-related caches from ~/.cache/ (torch, triton, flash_attn, vllm, and similar)
  3. Remove ~/.triton/ if it exists (triton kernel cache)
  4. Delete the current venv and recreate it fresh
  5. Reinstall with the two-step process above

See installation-troubleshooting.md for the full cleanup procedure.

Algorithm selection

Read the algorithm guides at https://ai-innovation.team/training_hub/#/algorithms/sft, https://ai-innovation.team/training_hub/#/algorithms/osft, and https://ai-innovation.team/training_hub/#/algorithms/lora for conceptual overviews. The decision framework:

NeedAlgorithmWhy
Compute-constrained or simple task fine-tuningLoRALow VRAM, fast iteration, but lower capacity and higher forgetting
Maximum capacity, forgetting is acceptableSFTFull-parameter training, distributed multi-node support
New knowledge while preserving existing capabilitiesOSFTOrthogonal subspace prevents catastrophic forgetting

Always try multiple algorithms and pick the one that performs best on your evaluation. See the example notebooks in examples/notebooks/ that compare SFT vs OSFT for continual learning scenarios.

Hyperparameter configuration

The three hyperparameters that matter most for any algorithm are learning rate, effective batch size, and number of epochs. Detailed guidance including dataset-size-dependent recommendations lives in hyperparameter-guide.md.

Quick reference

SFT / OSFT:

  • Learning rate: start at 5e-6 for smaller datasets, 1e-6 for larger datasets
  • Epochs: 2-3 is typical
  • Batch size: 32-64 for <1k samples, 128 for 1k-10k, 256+ beyond 10k

LoRA:

  • Learning rate: start at 1e-5, adjust up toward 1e-4 only if needed
  • Rank (lora_r): start at 16, increase if underfitting
  • Alpha (lora_alpha): typically 2x rank

OSFT-specific:

  • unfreeze_rank_ratio: recommended default 0.5. Rarely need above 0.5 for models around 8B parameters. Larger models generally need less; smaller models may need more (see hyperparameter-guide.md for why).
  • target_patterns: optionally restrict OSFT to specific modules (e.g., only MLPs or only attention)
Key parameters
  • effective_batch_size: Taken as the exact minibatch size on any backend. The algorithm translates this into whatever gradient accumulation is needed internally.
  • max_seq_len: Samples exceeding this length are dropped. Important for long-context data and affects training speed and memory.
  • unmask_messages (OSFT) / unmask field (SFT): Unmasks all messages except the system message for loss computation. When training on knowledge data where documents are embedded in user messages (e.g., from sdg_hub), this significantly boosts knowledge ingestion.
  • is_pretraining: Enables pretraining mode for document-style data. Uses block_size to pack documents into fixed-length sample blocks. Start with block_size=2048, or 512 for short/numerous documents.
  • accelerate_full_state_at_epoch (SFT only): Saves FP32 full-state checkpoints at every epoch. Very expensive (an 8B model checkpoint is ~108GB).

Memory model and OOM troubleshooting

SFT and OSFT train in FP32 + mixed precision. Memory requirement to load a model: 16 bytes per parameter (4 bytes x 4 copies: parameter, gradient, 2x AdamW optimizer states).

When you hit OOM, these are the available knobs:

  • nproc_per_node: Use all available GPUs to distribute the workload
  • max_seq_len: Reduce if sequences are longer than what fits in a single forward/backward pass
  • max_tokens_per_gpu: Reduce to fit fewer tokens per GPU per step
  • Liger kernels: Enable use_liger=True to reduce memory via fused kernels
  • Flash attention: Ensure flash-attn is installed and importable for memory-efficient attention
  • LoRA-specific: Decrease rank or target fewer modules
  • OSFT-specific: Decrease unfreeze_rank_ratio to reduce SVD memory overhead

If all knobs are exhausted, choose a smaller model or a node with more GPU memory.

Use from training_hub import estimate for upfront memory estimation. See examples/notebooks/memory_estimator_example.ipynb.

Show full SKILL.md (404 more words)Show less

Experiment tracking

All three algorithms (sft(), osft(), lora_sft()) expose logging configuration as first-class parameters. Loggers are auto-detected: they are automatically enabled when their configuration parameters are set.

Logging parameters

All algorithms accept the same logging parameters:

ParameterLoggerDescription
wandb_projectW&BProject name (enables W&B logging)
wandb_entityW&BTeam or user entity
wandb_run_nameW&BRun display name
mlflow_tracking_uriMLflowTracking server URI (enables MLflow logging)
mlflow_experiment_nameMLflowExperiment name
mlflow_run_nameMLflowRun name
tensorboard_log_dirTensorBoardLog directory (enables TensorBoard logging)
Example
python
from training_hub import sft

sft(
    model_path="my-model",
    data_path="data.jsonl",
    ckpt_output_dir="./checkpoints",
    # W&B logging — enabled automatically because wandb_project is set
    wandb_project="my-finetune",
    wandb_entity="my-team",
    wandb_run_name="sft-run-1",
    # MLflow — also enabled, multiple loggers can run simultaneously
    mlflow_tracking_uri="http://localhost:5000",
    mlflow_experiment_name="sft-experiment",
)
Environment variable fallback

If logging parameters are not passed explicitly, backends will check these environment variables as fallback:

ParameterEnvironment variable
wandb_projectWANDB_PROJECT
wandb_entityWANDB_ENTITY
wandb_run_nameWANDB_RUN_NAME
mlflow_tracking_uriMLFLOW_TRACKING_URI
mlflow_experiment_nameMLFLOW_EXPERIMENT_NAME
mlflow_run_nameMLFLOW_RUN_NAME

Explicit kwargs always take precedence over environment variables.

Logger support matrix
LoggerSFTOSFTLoRA
W&BYesYesYes
MLflowYesYesYes
TensorBoardYesLimitedYes

Loss monitoring and convergence

Each backend emits loss in a different format. plot_loss() auto-detects all of them:

python
from training_hub import plot_loss
plot_loss(["./run1", "./run2"], labels=["baseline", "tuned"], ema=True)
Log formats
BackendFormatFileLoss key
SFT (instructlab-training)JSONLtraining_log.jsonlavg_loss
OSFT (mini-trainer)JSONLtraining_log.jsonlloss
LoRA (Unsloth/TRL)JSONcheckpoint-*/trainer_state.jsonloss
Interpreting convergence
  • Train loss alone is insufficient. A model can converge on train loss forever while overfitting badly.
  • Validation loss is the real signal. Look for where validation loss stops decreasing or starts increasing (optimal training regime vs overfitting).
  • SFT and OSFT backends support validation loss via validation_split and validation_frequency kwargs. OSFT also supports save_best_val_loss. See backend-kwargs.md and the Validation Loss Guide.
  • When validation loss is unavailable (LoRA, GRPO), rely on downstream evaluation to gauge actual model quality.

Backend kwargs passthrough

Every algorithm exposes a curated parameter set, but backends support many more options. Any parameter not directly exposed can be passed as a kwarg to the algorithm function, and it will be forwarded to the backend.

This is covered in detail in backend-kwargs.md, including links to each backend's full parameter definitions and practical examples like running plain SFT through the OSFT backend with osft=False.

Evaluation

Validation loss is a useful proxy but not always sufficient. Ideally, evaluate the model on a downstream benchmark or task-specific eval harness both before and after training to measure the actual impact. Evaluation itself is outside Training Hub's scope.

Additional resources

© Red-Hat-AI-Innovation-Team, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in .claude/skills/training-hub-guide of Red-Hat-AI-Innovation-Team/training_hub.

  • SKILL.md
  • backend-kwargs.md
  • hyperparameter-guide.md
  • installation-troubleshooting.md

Open the folder on GitHubat commit 511a905

Compare with similar skills

Training Hub Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Training Hub Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Training Hub Guide this skillRed-Hat-AI-Innovation-Team/training_hub100—~2.8kAutomated safety check: PassApache-2.0
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0
ML Generative Mattergenlearningmatter-mit/AtomisticSkills176—~1.8kAutomated safety check: PassMIT
Cosmos3 Post TrainingNVIDIA/cosmos-framework560—~2.7kAutomated safety check: PassCustom licence
Spark Environment Setupwshobson/agents40k—~2kAutomated safety check: PassMIT
Kermt FinetuneNVIDIA/skills3.6k1 repos~4.1kAutomated safety check: PassApache-2.0

Similar skills

  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • ML Generative Mattergen

    learningmatter-mit/AtomisticSkills

    Generate inorganic material structures using MatterGen, a diffusion-based generative model.

    176 GitHub stars~1.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Cosmos3 Post Training

    NVIDIA/cosmos-framework

    Official

    Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired…

    560 GitHub stars~2.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Kermt Finetune

    NVIDIA/skills

    Official

    Finetune a pretrained KERMT encoder on a labeled CSV. An agent skill from NVIDIA/skills.

    3.6k GitHub starsUsed in 1 repo~4.1k tokens
    AI & LLM EngineeringAuto-check passed
  • NV-Tesseract AD Diffusion — diffusion-based anomaly detection and fine-tuning for multivariate time series.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Data & AnalyticsAuto-check: notes

More from Red-Hat-AI-Innovation-Team/training_hub

  • Setup Guide

    Red-Hat-AI-Innovation-Team/training_hub

    A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment.

    100 GitHub stars~959 tokensUpdated 3 days ago
    Auto-check passed
  • Memory Estimation

    Red-Hat-AI-Innovation-Team/training_hub

    A skill your agent uses when the user wants to estimate GPU memory (VRAM) requirements for a training configuration, check if a model will fit on their GPUs, or plan GPU allocation for training.

    100 GitHub stars~393 tokensUpdated 3 days ago
    Auto-check passed
  • Training Guide

    Red-Hat-AI-Innovation-Team/training_hub

    A skill your agent uses when the user wants to run a training job using a saved configuration.

    100 GitHub stars~346 tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Training Hub Guide

What does Training Hub Guide do?

Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves…. Training Hub Guide is an agent skill from Red-Hat-AI-Innovation-Team/training_hub. Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves, and leveraging backend-specific features.

When should I use Training Hub Guide?

Training Hub Guide fits situations like: the user is working with traininghub; fine-tuning language models; asking about SFT/OSFT/LoRA training; debugging GPU/CUDA training issues.

How do I install Training Hub Guide in Claude Code?

Run `npx skills add Red-Hat-AI-Innovation-Team/training_hub --skill training-hub-guide -a claude-code`. Or copy the skill folder (.claude/skills/training-hub-guide in Red-Hat-AI-Innovation-Team/training_hub) into .claude/skills/training-hub-guide in your project. Claude Code loads it when a task matches its description.

How do I install Training Hub Guide in Codex?

Run `npx skills add Red-Hat-AI-Innovation-Team/training_hub --skill training-hub-guide -a codex`. Or copy the skill folder (.claude/skills/training-hub-guide in Red-Hat-AI-Innovation-Team/training_hub) into .agents/skills/training-hub-guide in your project. Codex loads it when a task matches its description.

Can I use Training Hub Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Red-Hat-AI-Innovation-Team/training_hub --skill training-hub-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/training-hub-guide, .gemini/skills/training-hub-guide, .github/skills/training-hub-guide and .opencode/skills/training-hub-guide in your project.

What does Training Hub Guide need to run?

Going by SKILL.md and its folder, Training Hub Guide needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Training Hub Guide access the network?

SKILL.md names 1 domain. As links in the text: ai-innovation.team. This is read from the text; nothing was executed.

Is Training Hub Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Training Hub Guide use?

Training Hub Guide is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Training Hub Guide use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Training Hub Guide?

Skills that share tags, products or a category with Training Hub Guide: Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 146 stars), ML Generative Mattergen (learningmatter-mit/AtomisticSkills, 176 stars), Cosmos3 Post Training (NVIDIA/cosmos-framework, 560 stars) and Spark Environment Setup (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Training Hub Guide?

Red-Hat-AI-Innovation-Team (a GitHub organization) maintains it in Red-Hat-AI-Innovation-Team/training_hub, which has 100 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.

Source: Red-Hat-AI-Innovation-Team/training_hub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.