Hugging Face Vision Trainer
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
Agent skill
by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs
Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs llava --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/18-multimodal/llava .claude/skills/llava && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "llava" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/llava into .claude/skills/llava/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llava", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/llavaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs llava --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/18-multimodal/llava .agents/skills/llava && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "llava" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/llava into .agents/skills/llava/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llava", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs llava --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/18-multimodal/llava .cursor/skills/llava && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "llava" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/llava into .cursor/skills/llava/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llava", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 18-multimodal/llava--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs llava --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/18-multimodal/llava .gemini/skills/llava && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "llava" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/llava into .gemini/skills/llava/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llava", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs llavaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/18-multimodal/llava .github/skills/llava && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "llava" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/llava into .github/skills/llava/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llava", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs llava --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/18-multimodal/llava .opencode/skills/llava && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "llava" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/18-multimodal/llava into .opencode/skills/llava/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llava", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
llavaGuide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.
LLaVA is an open-source vision-language model that pairs a CLIP vision encoder with Vicuna or LLaMA language models. The skill covers installing from the repository, loading a pretrained model in Python, running single-image queries from the CLI, launching a Gradio web interface and holding multi-turn conversations with a conversation template.
A table lists the 7B, 13B and 34B variants with approximate memory needs of roughly 14, 28 and 70 GB. Common tasks include captioning, visual question answering, textual object listing, scene understanding and document understanding with images. The skill also names alternatives: GPT-4V for top API quality, CLIP for zero-shot classification and BLIP-2 for captioning alone, and a training reference file is included.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonbashgitpipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comAlso links to:
arxiv.orgllava.hliu.cchuggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DEFAULT_IMAGE_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
LLaVA Vision-Language Model loads about 2k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 290 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 290 words, ~1,957 tokens.
.claude/skills/llava/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Open-source vision-language model for conversational image understanding.
Use when:
Metrics:
Use alternatives instead:
# Clone repository
git clone https://github.com/haotian-liu/LLaVA
cd LLaVA
# Install
pip install -e .from llava.model.builder import load_pretrained_model
from llava.mm_utils import get_model_name_from_path, process_images, tokenizer_image_token
from llava.constants import IMAGE_TOKEN_INDEX, DEFAULT_IMAGE_TOKEN
from llava.conversation import conv_templates
from PIL import Image
import torch
# Load model
model_path = "liuhaotian/llava-v1.5-7b"
tokenizer, model, image_processor, context_len = load_pretrained_model(
model_path=model_path,
model_base=None,
model_name=get_model_name_from_path(model_path)
)
# Load image
image = Image.open("image.jpg")
image_tensor = process_images([image], image_processor, model.config)
image_tensor = image_tensor.to(model.device, dtype=torch.float16)
# Create conversation
conv = conv_templates["llava_v1"].copy()
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?")
conv.append_message(conv.roles[1], None)
prompt = conv.get_prompt()
# Generate response
input_ids = tokenizer_image_token(prompt, tokenizer, IMAGE_TOKEN_INDEX, return_tensors='pt').unsqueeze(0).to(model.device)
with torch.inference_mode():
output_ids = model.generate(
input_ids,
images=image_tensor,
do_sample=True,
temperature=0.2,
max_new_tokens=512
)
response = tokenizer.decode(output_ids[0], skip_special_tokens=True).strip()
print(response)| Model | Parameters | VRAM | Quality |
|---|---|---|---|
| LLaVA-v1.5-7B | 7B | ~14 GB | Good |
| LLaVA-v1.5-13B | 13B | ~28 GB | Better |
| LLaVA-v1.6-34B | 34B | ~70 GB | Best |
# Load different models
model_7b = "liuhaotian/llava-v1.5-7b"
model_13b = "liuhaotian/llava-v1.5-13b"
model_34b = "liuhaotian/llava-v1.6-34b"
# 4-bit quantization for lower VRAM
load_4bit = True # Reduces VRAM by ~4×# Single image query
python -m llava.serve.cli \
--model-path liuhaotian/llava-v1.5-7b \
--image-file image.jpg \
--query "What is in this image?"
# Multi-turn conversation
python -m llava.serve.cli \
--model-path liuhaotian/llava-v1.5-7b \
--image-file image.jpg
# Then type questions interactively# Launch Gradio interface
python -m llava.serve.gradio_web_server \
--model-path liuhaotian/llava-v1.5-7b \
--load-4bit # Optional: reduce VRAM
# Access at http://localhost:7860# Initialize conversation
conv = conv_templates["llava_v1"].copy()
# Turn 1
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?")
conv.append_message(conv.roles[1], None)
response1 = generate(conv, model, image) # "A dog playing in a park"
# Turn 2
conv.messages[-1][1] = response1 # Add previous response
conv.append_message(conv.roles[0], "What breed is the dog?")
conv.append_message(conv.roles[1], None)
response2 = generate(conv, model, image) # "Golden Retriever"
# Turn 3
conv.messages[-1][1] = response2
conv.append_message(conv.roles[0], "What time of day is it?")
conv.append_message(conv.roles[1], None)
response3 = generate(conv, model, image)question = "Describe this image in detail."
response = ask(model, image, question)question = "How many people are in the image?"
response = ask(model, image, question)question = "List all the objects you can see in this image."
response = ask(model, image, question)question = "What is happening in this scene?"
response = ask(model, image, question)question = "What is the main topic of this document?"
response = ask(model, document_image, question)# Stage 1: Feature alignment (558K image-caption pairs)
bash scripts/v1_5/pretrain.sh
# Stage 2: Visual instruction tuning (150K instruction data)
bash scripts/v1_5/finetune.sh# 4-bit quantization
tokenizer, model, image_processor, context_len = load_pretrained_model(
model_path="liuhaotian/llava-v1.5-13b",
model_base=None,
model_name=get_model_name_from_path("liuhaotian/llava-v1.5-13b"),
load_4bit=True # Reduces VRAM ~4×
)
# 8-bit quantization
load_8bit=True # Reduces VRAM ~2×| Model | VRAM (FP16) | VRAM (4-bit) | Speed (tokens/s) |
|---|---|---|---|
| 7B | ~14 GB | ~4 GB | ~20 |
| 13B | ~28 GB | ~8 GB | ~12 |
| 34B | ~70 GB | ~18 GB | ~5 |
On A100 GPU
LLaVA achieves competitive scores on:
from langchain.llms.base import LLM
class LLaVALLM(LLM):
def _call(self, prompt, stop=None):
# Custom LLaVA inference
return response
llm = LLaVALLM()import gradio as gr
def chat(image, text, history):
response = ask_llava(model, image, text)
return response
demo = gr.ChatInterface(
chat,
additional_inputs=[gr.Image(type="pil")],
title="LLaVA Chat"
)
demo.launch()© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in 18-multimodal/llava of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 6 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
LLaVA Vision-Language Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| LLaVA Vision-Language Model this skillOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2k | Automated safety check: Pass | MIT | |
| Hugging Face Vision Trainerhuggingface/skills | 11k | 1 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Segmentation Sam2SharpAI/DeepCamera | 3.1k | — | ~594 | Automated safety check: Pass | MIT | |
| Yolo Detection 2026 Coral Tpu Win WslSharpAI/DeepCamera | 3.1k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Gradio App Builderhuggingface/skills | 11k | 4 repos | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Transformers Usagedavila7/claude-code-templates | 33k | 11 repos | ~1.2k | Automated safety check: Pass | MIT |
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
SharpAI/DeepCamera
Interactive click-to-segment using Segment Anything 2 — AI-assisted labeling for Annotation Studio
SharpAI/DeepCamera
Google Coral Edge TPU — real-time object detection natively via Windows WSL
huggingface/skills
Guides building Gradio web UIs and ML demos in Python, covering Interface, Blocks, ChatInterface, event listeners, layouts and chatbots.
davila7/claude-code-templates
Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.
Hermes-brasil/hermes-brasil
Portuguese guide to building a retrieval-augmented generation assistant over a company's documents, with embeddings, section-based chunking, retrieval and a client workflow.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
Categories
Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code. LLaVA is an open-source vision-language model that pairs a CLIP vision encoder with Vicuna or LLaMA language models. The skill covers installing from the repository, loading a pretrained model in Python, running single-image queries from the CLI, launching a Gradio web interface and holding multi-turn conversations with a conversation template.
LLaVA Vision-Language Model fits situations like: building a chatbot that can discuss uploaded images; answering questions about the contents of a picture; captioning images or describing scenes in detail; running a local web demo for image conversations.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a claude-code`. Or copy the skill folder (18-multimodal/llava in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/llava in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a codex`. Or copy the skill folder (18-multimodal/llava in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/llava in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llava, .gemini/skills/llava, .github/skills/llava and .opencode/skills/llava in your project.
Going by SKILL.md and its folder, LLaVA Vision-Language Model needs the command-line tools its instructions call (python, bash, git and pip) and credentials named DEFAULT_IMAGE_TOKEN. Our summary lists: Python with the LLaVA repository installed; A GPU with enough VRAM for the chosen model size, about 14 GB for the 7B model.
SKILL.md names 4 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: arxiv.org, llava.hliu.cc and huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
LLaVA Vision-Language Model is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with LLaVA Vision-Language Model: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Segmentation Sam2 (SharpAI/DeepCamera, 3.1k stars), Yolo Detection 2026 Coral Tpu Win Wsl (SharpAI/DeepCamera, 3.1k stars) and Gradio App Builder (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.