Agent skill

Vision Sft

by wshobson in wshobson/agents

Fine-tune vision-language models (VLMs) with supervised learning on image+text data.

MITAuto-check passedAI & LLM Engineering

Install Vision Sft

skills CLI
$ npx skills add wshobson/agents --skill vision-sft -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents vision-sft --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/vision-sft .claude/skills/vision-sft && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vision-sft
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
989 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Fine-tune vision-language models (VLMs) with supervised learning on image+text data.

  • Adapting a VLM to a visual domain
  • SKILL.md covers Quick Reference, The Consensus Recipe, When to Unfreeze and The Two Silent Killers, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Configuring frozen-vision-tower LoRA

What it does

Vision Sft is an agent skill from wshobson/agents. Fine-tune vision-language models (VLMs) with supervised learning on image+text data. Use when adapting a VLM to a visual domain or task, configuring frozen-vision-tower LoRA, or debugging a VLM fine-tune that trains without learning.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/collators-and-pitfalls.md`).

It sits in AI & LLM Engineering, covering Fine-tuning, Computer vision and Machine learning. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Adapting a VLM to a visual domain
  • Configuring frozen-vision-tower LoRA
  • Debugging a VLM fine-tune that trains without learning

Example prompts

  • “/vision-sft”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vision Sft loads about 2k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 989 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 989 words, ~1,954 tokens.

Download SKILL.mdSave it as .claude/skills/vision-sft/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vision-sft
description
Fine-tune vision-language models (VLMs) with supervised learning on image+text data. Use when adapting a VLM to a visual domain or task, configuring frozen-vision-tower LoRA, or debugging a VLM fine-tune that trains without learning.

Vision-Language SFT

This skill assumes finetuning-method-selection already routed here: the data shape is image+text demonstrations, not preference pairs or a verifiable reward signal, and the base is a vision-language model rather than a text-only one. lora-qlora-recipes covers the text-only LoRA/QLoRA recipe this skill specializes for the vision tower and projector; read that skill first if the LoRA fundamentals (rank, alpha, target modules) aren't already familiar.

Input: an image+text dataset and a VLM base model already picked from the model catalog. Output format: a validated adapter config — which components are frozen, LoRA target modules, and a min_pixels/max_pixels budget — that llm-finetuning-training-engineer consumes directly when it generates a runnable script.

Quick Reference

SituationDefault
Adapting behavior on familiar imagesFrozen tower+projector, LoRA r=8–16, α=16–32
Visual domain shiftUnfreeze last-6 ViT layers, vision LR 5–10x lower
Doesn't fit in bf16 at target rankQLoRA — frozen vision tower only
fast_inference=Truefinetune_vision_layers=False
Loss normal, eval not improvingCheck the Two Silent Killers below first

The Consensus Recipe

Freeze the vision tower and the projector. Put LoRA on the LLM only, all-linear (the same attention + MLP target list as text-only SFT — see lora-qlora-recipes), at r=8–16, α=16–32. This is the settled default for adapting a VLM's behavior without disturbing how it sees.

  • The vision tower and projector stay frozen by default. They already encode a general visual representation; retraining them is rarely necessary and adds risk without adding capability for most tasks.
  • LoRA rank runs lower than the text-only general default (r=8–16 here vs r=16–32 for text-only SFT) because the LLM-only adapter is adapting behavior, not injecting new visual knowledge.
  • QLoRA is permitted only with a frozen vision tower. Quantizing the base while also unfreezing and training vision layers is unsupported and unstable — treat this as a hard pairing rule, not a tunable. If the vision tower needs to unfreeze, drop QLoRA and use bf16 LoRA instead.
python
# freeze tower + projector; LoRA on LLM only
for name, param in model.named_parameters():
    if "vision_tower" in name or "projector" in name:
        param.requires_grad = False

target_modules = [
    "q_proj", "k_proj", "v_proj", "o_proj",
    "gate_proj", "up_proj", "down_proj",
]  # LLM-only, all-linear — r=8-16, alpha=16-32

When to Unfreeze

Unfreezing vision layers is a deliberate escalation, not a default decision — reach for it only when the domain shift is visual, not textual.

  • Unfreeze only for visual domain shift. If the task is teaching new behavior on images the tower already understands (charts, everyday photos), the frozen-tower recipe above is sufficient. Unfreeze when the visual domain itself is unfamiliar to the tower — satellite imagery, medical scans, dense technical diagrams — and the frozen-tower recipe plateaus.
  • Last-6 ViT layers is the sweet spot. Unfreezing the final six vision-transformer layers (not the whole tower) measured +1.7pt DocVQA at ~1.75x training cost over the frozen baseline. Treat six layers as the ceiling worth paying for; going further spends compute without a matched result.
  • Vision LR must run 5–10x lower than the LLM LR when unfrozen. The vision tower's pretrained representation is more fragile than the LLM's adapter; the same LR for both risks overwriting the visual representation faster than the LLM adapter can compensate.
  • High LoRA rank on the patch- embedding layer risks NaN. If patch embedding is in the unfrozen set, keep its rank low and watch early-step loss closely — one of the most fragile places to apply LoRA in a VLM.
Show full SKILL.md (478 more words)Show less

The Two Silent Killers

Both produce a run that trains without error and without learning: the loss curve looks normal, the model doesn't improve, and neither throws an exception — both need an explicit pre-training check, not just a clean training log.

  • Image-tag/count mismatch. Every image placeholder token in the templated text must map 1:1 to a media item actually passed to the collator. A mismatch (one placeholder, zero or two images attached; or an image with no placeholder) doesn't error in most collators — it silently misaligns image and text, and the model "trains but learns nothing." Validate the 1:1 placeholder-to-media mapping before training starts, on every example, not just a sample. Full validation-checklist detail: references/collators-and-pitfalls.md.
  • min_pixels/max_pixels resolution budget. This pair is the single most consequential hyperparameter for quality and memory in VLM SFT — more than rank, alpha, or LR. Too low silently downsamples images below what the task needs (small document text becomes unreadable even though training "succeeds"); too high blows the activation memory budget or forces too small a batch to train stably. Set it deliberately per dataset, don't leave it at a framework default.

Unsloth Specifics

  • UnslothVisionDataCollator is the collator Unsloth expects for VLM SFT — it handles the image-tag alignment and per-architecture processor contract described in references/collators-and-pitfalls.md. Don't substitute a text-only collator for VLM data.
  • finetune_vision_layers=False is required when fast_inference=True. vLLM cannot serve LoRA adapters on vision layers, so a fast- inference setup that also unfreezes vision layers fails at serve time even if training succeeds. If the recipe calls for unfreezing the last-6 ViT layers (see When to Unfreeze above), fast inference is off the table for that run — choose one or the other, not both.

Model Choice

Base VLM choice is out of scope for this skill — it lives in one place, the model catalog at finetuning-method-selection's references/model-catalog.md. This skill and its references describe recipes by architecture family only, never by recommending one model over another.

VLM reinforcement learning (VLM-GRPO) is reference-only in this plugin — the fragmented tooling and reward-hacking failure modes specific to VLM-RL are covered in grpo-rlvr-training, not here. This skill's scope stops at supervised fine-tuning.

Failure Modes

The recurring mistake across every section above is treating a clean loss curve as proof the run is healthy. A normal-looking curve is consistent with both a working run and either silent killer, since the model trains on something either way — just not the aligned image-text signal when a killer is present. A flat eval score next to a normal loss curve means re-run the checklist in references/collators-and-pitfalls.md before touching any hyperparameter.

References

  • references/collators-and-pitfalls.md — per- architecture collator table, dataset-format examples with image placeholders, a pre- training validation checklist, and the two- stage projector-alignment recipe as an advanced pattern.

Related skills: finetuning-method-selection routes here; lora-qlora-recipes covers the text-only LoRA fundamentals this skill specializes; grpo-rlvr-training covers VLM-RL (reference-only); dataset-curation covers image+text dataset preparation this skill doesn't.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/llm-finetuning/skills/vision-sft of wshobson/agents.

  • SKILL.md
  • references/collators-and-pitfalls.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Vision Sft next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vision Sft compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vision Sft this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Adapting Transfer Learning Modelsjeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMIT
Transformersynulihao/AgentSkillOS617—~2.9kAutomated safety check: PassNone
AI ML Skillswentorai/research-plugins2981 repos~993Automated safety check: PassMIT
RuView Model Trainingruvnet/RuView97k—~1.3kAutomated safety check: NotesMIT
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0

Similar skills

  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Transformers

    ynulihao/AgentSkillOS

    Work with state-of-the-art machine learning models for NLP, computer vision, audio, and multimodal tasks using HuggingFace Transformers.

    617 GitHub stars~2.9k tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • AI ML Skills

    wentorai/research-plugins

    27 ai & machine learning skills. An agent skill from wentorai/research-plugins.

    298 GitHub starsUsed in 1 repo~993 tokens
    AI & LLM EngineeringAuto-check passed
  • Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing.

    97k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface Vision Trainer

    waybarrios/opencode-power-pack

    Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

    533 GitHub stars~2.7k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed

Questions about Vision Sft

What does Vision Sft do?

Fine-tune vision-language models (VLMs) with supervised learning on image+text data. Vision Sft is an agent skill from wshobson/agents. Fine-tune vision-language models (VLMs) with supervised learning on image+text data.

When should I use Vision Sft?

Vision Sft fits situations like: adapting a VLM to a visual domain; configuring frozen-vision-tower LoRA; debugging a VLM fine-tune that trains without learning.

How do I install Vision Sft in Claude Code?

Run `npx skills add wshobson/agents --skill vision-sft -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/vision-sft in wshobson/agents) into .claude/skills/vision-sft in your project. Claude Code loads it when a task matches its description.

How do I install Vision Sft in Codex?

Run `npx skills add wshobson/agents --skill vision-sft -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/vision-sft in wshobson/agents) into .agents/skills/vision-sft in your project. Codex loads it when a task matches its description.

Can I use Vision Sft in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill vision-sft -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vision-sft, .gemini/skills/vision-sft, .github/skills/vision-sft and .opencode/skills/vision-sft in your project.

What does Vision Sft need to run?

SKILL.md names no scripts, command-line tools or credentials: Vision Sft is instructions for the agent only. Our summary lists: Python 3.

Does Vision Sft access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vision Sft safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vision Sft use?

Vision Sft is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vision Sft use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Vision Sft?

Skills that share tags, products or a category with Vision Sft: Adapting Transfer Learning Models (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Transformers (ynulihao/AgentSkillOS, 617 stars), AI ML Skills (wentorai/research-plugins, 298 stars) and RuView Model Training (ruvnet/RuView, 97k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vision Sft?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,305 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.