Adapting Transfer Learning Models
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
Fine-tune vision-language models (VLMs) with supervised learning on image+text data.
$ npx skills add wshobson/agents --skill vision-sft -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents vision-sft --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/vision-sft .claude/skills/vision-sft && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vision-sft" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/vision-sft into .claude/skills/vision-sft/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-sft", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/vision-sftType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill vision-sft -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents vision-sft --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/llm-finetuning/skills/vision-sft .agents/skills/vision-sft && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vision-sft" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/vision-sft into .agents/skills/vision-sft/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-sft", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill vision-sft -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents vision-sft --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/llm-finetuning/skills/vision-sft .cursor/skills/vision-sft && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vision-sft" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/vision-sft into .cursor/skills/vision-sft/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-sft", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/llm-finetuning/skills/vision-sft--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill vision-sft -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents vision-sft --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/llm-finetuning/skills/vision-sft .gemini/skills/vision-sft && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vision-sft" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/vision-sft into .gemini/skills/vision-sft/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-sft", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents vision-sftInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill vision-sft -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/llm-finetuning/skills/vision-sft .github/skills/vision-sft && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vision-sft" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/vision-sft into .github/skills/vision-sft/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-sft", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill vision-sft -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents vision-sft --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/llm-finetuning/skills/vision-sft .opencode/skills/vision-sft && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vision-sft" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/vision-sft into .opencode/skills/vision-sft/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vision-sft", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vision-sftFine-tune vision-language models (VLMs) with supervised learning on image+text data.
Vision Sft is an agent skill from wshobson/agents. Fine-tune vision-language models (VLMs) with supervised learning on image+text data. Use when adapting a VLM to a visual domain or task, configuring frozen-vision-tower LoRA, or debugging a VLM fine-tune that trains without learning.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/collators-and-pitfalls.md`).
It sits in AI & LLM Engineering, covering Fine-tuning, Computer vision and Machine learning. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vision Sft loads about 2k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 989 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 989 words, ~1,954 tokens.
.claude/skills/vision-sft/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.This skill assumes finetuning-method-selection
already routed here: the data shape is
image+text demonstrations, not preference pairs
or a verifiable reward signal, and the base is a
vision-language model rather than a text-only
one. lora-qlora-recipes covers the text-only
LoRA/QLoRA recipe this skill specializes for the
vision tower and projector; read that skill first
if the LoRA fundamentals (rank, alpha, target
modules) aren't already familiar.
Input: an image+text dataset and a VLM base
model already picked from the model catalog.
Output format: a validated adapter config —
which components are frozen, LoRA target modules,
and a min_pixels/max_pixels budget — that
llm-finetuning-training-engineer consumes
directly when it generates a runnable script.
| Situation | Default |
|---|---|
| Adapting behavior on familiar images | Frozen tower+projector, LoRA r=8–16, α=16–32 |
| Visual domain shift | Unfreeze last-6 ViT layers, vision LR 5–10x lower |
| Doesn't fit in bf16 at target rank | QLoRA — frozen vision tower only |
fast_inference=True | finetune_vision_layers=False |
| Loss normal, eval not improving | Check the Two Silent Killers below first |
Freeze the vision tower and the projector. Put
LoRA on the LLM only, all-linear (the same
attention + MLP target list as text-only SFT —
see lora-qlora-recipes), at r=8–16,
α=16–32. This is the settled default for
adapting a VLM's behavior without disturbing how
it sees.
# freeze tower + projector; LoRA on LLM only
for name, param in model.named_parameters():
if "vision_tower" in name or "projector" in name:
param.requires_grad = False
target_modules = [
"q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj",
] # LLM-only, all-linear — r=8-16, alpha=16-32Unfreezing vision layers is a deliberate escalation, not a default decision — reach for it only when the domain shift is visual, not textual.
Both produce a run that trains without error and without learning: the loss curve looks normal, the model doesn't improve, and neither throws an exception — both need an explicit pre-training check, not just a clean training log.
references/collators-and-pitfalls.md.min_pixels/max_pixels resolution budget.
This pair is the single most consequential
hyperparameter for quality and memory in VLM
SFT — more than rank, alpha, or LR. Too low
silently downsamples images below what the task
needs (small document text becomes unreadable
even though training "succeeds"); too high blows
the activation memory budget or forces too small
a batch to train stably. Set it deliberately per
dataset, don't leave it at a framework default.UnslothVisionDataCollator is the collator
Unsloth expects for VLM SFT — it handles the
image-tag alignment and per-architecture
processor contract described in
references/collators-and-pitfalls.md. Don't
substitute a text-only collator for VLM data.finetune_vision_layers=False is required
when fast_inference=True. vLLM cannot serve
LoRA adapters on vision layers, so a fast-
inference setup that also unfreezes vision
layers fails at serve time even if training
succeeds. If the recipe calls for unfreezing the
last-6 ViT layers (see When to Unfreeze above),
fast inference is off the table for that run —
choose one or the other, not both.Base VLM choice is out of scope for this skill —
it lives in one place, the model catalog at
finetuning-method-selection's
references/model-catalog.md. This skill and its
references describe recipes by architecture
family only, never by recommending one model over
another.
VLM reinforcement learning (VLM-GRPO) is
reference-only in this plugin — the fragmented
tooling and reward-hacking failure modes specific
to VLM-RL are covered in grpo-rlvr-training,
not here. This skill's scope stops at supervised
fine-tuning.
The recurring mistake across every section above
is treating a clean loss curve as proof the run
is healthy. A normal-looking curve is consistent
with both a working run and either silent
killer, since the model trains on something
either way — just not the aligned image-text
signal when a killer is present. A flat eval score
next to a normal loss curve means re-run the
checklist in references/collators-and-pitfalls.md
before touching any hyperparameter.
references/collators-and-pitfalls.md — per-
architecture collator table, dataset-format
examples with image placeholders, a pre-
training validation checklist, and the two-
stage projector-alignment recipe as an advanced
pattern.Related skills: finetuning-method-selection
routes here; lora-qlora-recipes covers the
text-only LoRA fundamentals this skill
specializes; grpo-rlvr-training covers VLM-RL
(reference-only); dataset-curation covers
image+text dataset preparation this skill doesn't.
© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/llm-finetuning/skills/vision-sft of wshobson/agents.
Open the folder on GitHubat commit 46891e7
Vision Sft next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vision Sft this skillwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Adapting Transfer Learning Modelsjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Transformersynulihao/AgentSkillOS | 617 | — | ~2.9k | Automated safety check: Pass | None | |
| AI ML Skillswentorai/research-plugins | 298 | 1 repos | ~993 | Automated safety check: Pass | MIT | |
| RuView Model Trainingruvnet/RuView | 97k | — | ~1.3k | Automated safety check: Notes | MIT | |
| Hugging Face Vision Trainerhuggingface/skills | 11k | 1 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 |
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
ynulihao/AgentSkillOS
Work with state-of-the-art machine learning models for NLP, computer vision, audio, and multimodal tasks using HuggingFace Transformers.
wentorai/research-plugins
27 ai & machine learning skills. An agent skill from wentorai/research-plugins.
ruvnet/RuView
Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing.
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
waybarrios/opencode-power-pack
Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
Categories
Fine-tune vision-language models (VLMs) with supervised learning on image+text data. Vision Sft is an agent skill from wshobson/agents. Fine-tune vision-language models (VLMs) with supervised learning on image+text data.
Vision Sft fits situations like: adapting a VLM to a visual domain; configuring frozen-vision-tower LoRA; debugging a VLM fine-tune that trains without learning.
Run `npx skills add wshobson/agents --skill vision-sft -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/vision-sft in wshobson/agents) into .claude/skills/vision-sft in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill vision-sft -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/vision-sft in wshobson/agents) into .agents/skills/vision-sft in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill vision-sft -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vision-sft, .gemini/skills/vision-sft, .github/skills/vision-sft and .opencode/skills/vision-sft in your project.
SKILL.md names no scripts, command-line tools or credentials: Vision Sft is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vision Sft is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Vision Sft: Adapting Transfer Learning Models (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Transformers (ynulihao/AgentSkillOS, 617 stars), AI ML Skills (wentorai/research-plugins, 298 stars) and RuView Model Training (ruvnet/RuView, 97k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,305 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.