Agent skill

LLaVA Vision-Language Model

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.

MITAuto-check passedAI & LLM Engineering

Install LLaVA Vision-Language Model

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs llava --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/18-multimodal/llava .claude/skills/llava && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llava
GitHub stars
13k
Used in
6 other repos
Token cost
~2k tokens
SKILL.md length
290 words
Files
2 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.

  • Works in 8 steps: Start with 7B model - Good quality,… → Use 4-bit quantization - Reduces VRAM… → GPU required - CPU inference extremely… → …
  • Building a chatbot that can discuss uploaded images
  • SKILL.md covers When to use LLaVA, Quick start, Available models and CLI usage, plus 11 more sections
  • Calls python, bash and git; reaches github.com; needs DEFAULT_IMAGE_TOKEN

What it does

LLaVA is an open-source vision-language model that pairs a CLIP vision encoder with Vicuna or LLaMA language models. The skill covers installing from the repository, loading a pretrained model in Python, running single-image queries from the CLI, launching a Gradio web interface and holding multi-turn conversations with a conversation template.

A table lists the 7B, 13B and 34B variants with approximate memory needs of roughly 14, 28 and 70 GB. Common tasks include captioning, visual question answering, textual object listing, scene understanding and document understanding with images. The skill also names alternatives: GPT-4V for top API quality, CLIP for zero-shot classification and BLIP-2 for captioning alone, and a training reference file is included.

When your agent uses it

  • Building a chatbot that can discuss uploaded images
  • Answering questions about the contents of a picture
  • Captioning images or describing scenes in detail
  • Running a local web demo for image conversations

Example prompts

  • “Load LLaVA 7B and describe the photo at ./images/street.jpg.”
  • “Set up a multi-turn chat where I can ask follow-up questions about an uploaded chart.”
  • “Launch the LLaVA Gradio demo with the 13B model.”

Requirements

  • Python with the LLaVA repository installed
  • A GPU with enough VRAM for the chosen model size, about 14 GB for the 7B model

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Start with 7B model - Good quality, manageable VRAM
  2. Use 4-bit quantization - Reduces VRAM significantly
  3. GPU required - CPU inference extremely slow
  4. Clear prompts - Specific questions get better answers
  5. Multi-turn conversations - Maintain conversation context
  6. Temperature 0.2-0.7 - Balance creativity/consistency
  7. max_new_tokens 512-1024 - For detailed responses
  8. Batch processing - Process multiple images sequentially

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • bash
    • git
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • arxiv.org
    • llava.hliu.cc
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DEFAULT_IMAGE_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLaVA Vision-Language Model loads about 2k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 290 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 290 words, ~1,957 tokens.

Download SKILL.mdSave it as .claude/skills/llava/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
llava
description
Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
version
1.0.0
author
Orchestra Research
license
MIT
tags
LLaVA, Vision-Language, Multimodal, Visual Question Answering, Image Chat, CLIP, Vicuna, Conversational AI, Instruction Tuning, VQA
dependencies
transformers, torch, pillow

LLaVA - Large Language and Vision Assistant

Open-source vision-language model for conversational image understanding.

When to use LLaVA

Use when:

  • Building vision-language chatbots
  • Visual question answering (VQA)
  • Image description and captioning
  • Multi-turn image conversations
  • Visual instruction following
  • Document understanding with images

Metrics:

  • 23,000+ GitHub stars
  • GPT-4V level capabilities (targeted)
  • Apache 2.0 License
  • Multiple model sizes (7B-34B params)

Use alternatives instead:

  • GPT-4V: Highest quality, API-based
  • CLIP: Simple zero-shot classification
  • BLIP-2: Better for captioning only
  • Flamingo: Research, not open-source

Quick start

Installation
bash
# Clone repository
git clone https://github.com/haotian-liu/LLaVA
cd LLaVA

# Install
pip install -e .
Basic usage
python
from llava.model.builder import load_pretrained_model
from llava.mm_utils import get_model_name_from_path, process_images, tokenizer_image_token
from llava.constants import IMAGE_TOKEN_INDEX, DEFAULT_IMAGE_TOKEN
from llava.conversation import conv_templates
from PIL import Image
import torch

# Load model
model_path = "liuhaotian/llava-v1.5-7b"
tokenizer, model, image_processor, context_len = load_pretrained_model(
    model_path=model_path,
    model_base=None,
    model_name=get_model_name_from_path(model_path)
)

# Load image
image = Image.open("image.jpg")
image_tensor = process_images([image], image_processor, model.config)
image_tensor = image_tensor.to(model.device, dtype=torch.float16)

# Create conversation
conv = conv_templates["llava_v1"].copy()
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?")
conv.append_message(conv.roles[1], None)
prompt = conv.get_prompt()

# Generate response
input_ids = tokenizer_image_token(prompt, tokenizer, IMAGE_TOKEN_INDEX, return_tensors='pt').unsqueeze(0).to(model.device)

with torch.inference_mode():
    output_ids = model.generate(
        input_ids,
        images=image_tensor,
        do_sample=True,
        temperature=0.2,
        max_new_tokens=512
    )

response = tokenizer.decode(output_ids[0], skip_special_tokens=True).strip()
print(response)

Available models

ModelParametersVRAMQuality
LLaVA-v1.5-7B7B~14 GBGood
LLaVA-v1.5-13B13B~28 GBBetter
LLaVA-v1.6-34B34B~70 GBBest
python
# Load different models
model_7b = "liuhaotian/llava-v1.5-7b"
model_13b = "liuhaotian/llava-v1.5-13b"
model_34b = "liuhaotian/llava-v1.6-34b"

# 4-bit quantization for lower VRAM
load_4bit = True  # Reduces VRAM by ~4×

CLI usage

bash
# Single image query
python -m llava.serve.cli \
    --model-path liuhaotian/llava-v1.5-7b \
    --image-file image.jpg \
    --query "What is in this image?"

# Multi-turn conversation
python -m llava.serve.cli \
    --model-path liuhaotian/llava-v1.5-7b \
    --image-file image.jpg
# Then type questions interactively

Web UI (Gradio)

bash
# Launch Gradio interface
python -m llava.serve.gradio_web_server \
    --model-path liuhaotian/llava-v1.5-7b \
    --load-4bit  # Optional: reduce VRAM

# Access at http://localhost:7860

Multi-turn conversations

python
# Initialize conversation
conv = conv_templates["llava_v1"].copy()

# Turn 1
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?")
conv.append_message(conv.roles[1], None)
response1 = generate(conv, model, image)  # "A dog playing in a park"

# Turn 2
conv.messages[-1][1] = response1  # Add previous response
conv.append_message(conv.roles[0], "What breed is the dog?")
conv.append_message(conv.roles[1], None)
response2 = generate(conv, model, image)  # "Golden Retriever"

# Turn 3
conv.messages[-1][1] = response2
conv.append_message(conv.roles[0], "What time of day is it?")
conv.append_message(conv.roles[1], None)
response3 = generate(conv, model, image)

Common tasks

Image captioning
python
question = "Describe this image in detail."
response = ask(model, image, question)
Visual question answering
python
question = "How many people are in the image?"
response = ask(model, image, question)
Object detection (textual)
python
question = "List all the objects you can see in this image."
response = ask(model, image, question)
Scene understanding
python
question = "What is happening in this scene?"
response = ask(model, image, question)
Document understanding
python
question = "What is the main topic of this document?"
response = ask(model, document_image, question)

Training custom model

bash
# Stage 1: Feature alignment (558K image-caption pairs)
bash scripts/v1_5/pretrain.sh

# Stage 2: Visual instruction tuning (150K instruction data)
bash scripts/v1_5/finetune.sh

Quantization (reduce VRAM)

python
# 4-bit quantization
tokenizer, model, image_processor, context_len = load_pretrained_model(
    model_path="liuhaotian/llava-v1.5-13b",
    model_base=None,
    model_name=get_model_name_from_path("liuhaotian/llava-v1.5-13b"),
    load_4bit=True  # Reduces VRAM ~4×
)

# 8-bit quantization
load_8bit=True  # Reduces VRAM ~2×

Best practices

  1. Start with 7B model - Good quality, manageable VRAM
  2. Use 4-bit quantization - Reduces VRAM significantly
  3. GPU required - CPU inference extremely slow
  4. Clear prompts - Specific questions get better answers
  5. Multi-turn conversations - Maintain conversation context
  6. Temperature 0.2-0.7 - Balance creativity/consistency
  7. max_new_tokens 512-1024 - For detailed responses
  8. Batch processing - Process multiple images sequentially

Performance

ModelVRAM (FP16)VRAM (4-bit)Speed (tokens/s)
7B~14 GB~4 GB~20
13B~28 GB~8 GB~12
34B~70 GB~18 GB~5

On A100 GPU

Benchmarks

LLaVA achieves competitive scores on:

  • VQAv2: 78.5%
  • GQA: 62.0%
  • MM-Vet: 35.4%
  • MMBench: 64.3%

Limitations

  1. Hallucinations - May describe things not in image
  2. Spatial reasoning - Struggles with precise locations
  3. Small text - Difficulty reading fine print
  4. Object counting - Imprecise for many objects
  5. VRAM requirements - Need powerful GPU
  6. Inference speed - Slower than CLIP

Integration with frameworks

LangChain
python
from langchain.llms.base import LLM

class LLaVALLM(LLM):
    def _call(self, prompt, stop=None):
        # Custom LLaVA inference
        return response

llm = LLaVALLM()
Gradio App
python
import gradio as gr

def chat(image, text, history):
    response = ask_llava(model, image, text)
    return response

demo = gr.ChatInterface(
    chat,
    additional_inputs=[gr.Image(type="pil")],
    title="LLaVA Chat"
)
demo.launch()

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in 18-multimodal/llava of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/training.md

Open the folder on GitHubat commit 773a529

Used in 6 other repositories

We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 6 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLaVA Vision-Language Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLaVA Vision-Language Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLaVA Vision-Language Model this skillOrchestra-Research/AI-Research-SKILLs13k6 repos~2kAutomated safety check: PassMIT
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0
Segmentation Sam2SharpAI/DeepCamera3.1k—~594Automated safety check: PassMIT
Yolo Detection 2026 Coral Tpu Win WslSharpAI/DeepCamera3.1k—~1.1kAutomated safety check: PassMIT
Gradio App Builderhuggingface/skills11k4 repos~6.2kAutomated safety check: PassApache-2.0
Hugging Face Transformers Usagedavila7/claude-code-templates33k11 repos~1.2kAutomated safety check: PassMIT

Similar skills

  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Segmentation Sam2

    SharpAI/DeepCamera

    Interactive click-to-segment using Segment Anything 2 — AI-assisted labeling for Annotation Studio

    3.1k GitHub stars~594 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Google Coral Edge TPU — real-time object detection natively via Windows WSL

    3.1k GitHub stars~1.1k tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Gradio App Builder

    huggingface/skills

    Official

    Guides building Gradio web UIs and ML demos in Python, covering Interface, Blocks, ChatInterface, event listeners, layouts and chatbots.

    11k GitHub starsUsed in 4 repos~6.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Transformers Usage

    davila7/claude-code-templates

    Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

    33k GitHub starsUsed in 11 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Company Knowledge Assistant

    Hermes-brasil/hermes-brasil

    Portuguese guide to building a retrieval-augmented generation assistant over a company's documents, with embeddings, section-based chunking, retrieval and a client workflow.

    154 GitHub stars~1.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Works with

Questions about LLaVA Vision-Language Model

What does LLaVA Vision-Language Model do?

Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code. LLaVA is an open-source vision-language model that pairs a CLIP vision encoder with Vicuna or LLaMA language models. The skill covers installing from the repository, loading a pretrained model in Python, running single-image queries from the CLI, launching a Gradio web interface and holding multi-turn conversations with a conversation template.

When should I use LLaVA Vision-Language Model?

LLaVA Vision-Language Model fits situations like: building a chatbot that can discuss uploaded images; answering questions about the contents of a picture; captioning images or describing scenes in detail; running a local web demo for image conversations.

How do I install LLaVA Vision-Language Model in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a claude-code`. Or copy the skill folder (18-multimodal/llava in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/llava in your project. Claude Code loads it when a task matches its description.

How do I install LLaVA Vision-Language Model in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a codex`. Or copy the skill folder (18-multimodal/llava in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/llava in your project. Codex loads it when a task matches its description.

Can I use LLaVA Vision-Language Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llava -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llava, .gemini/skills/llava, .github/skills/llava and .opencode/skills/llava in your project.

What does LLaVA Vision-Language Model need to run?

Going by SKILL.md and its folder, LLaVA Vision-Language Model needs the command-line tools its instructions call (python, bash, git and pip) and credentials named DEFAULT_IMAGE_TOKEN. Our summary lists: Python with the LLaVA repository installed; A GPU with enough VRAM for the chosen model size, about 14 GB for the 7B model.

Does LLaVA Vision-Language Model access the network?

SKILL.md names 4 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: arxiv.org, llava.hliu.cc and huggingface.co. This is read from the text; nothing was executed.

Is LLaVA Vision-Language Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLaVA Vision-Language Model use?

LLaVA Vision-Language Model is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLaVA Vision-Language Model use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to LLaVA Vision-Language Model?

Skills that share tags, products or a category with LLaVA Vision-Language Model: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Segmentation Sam2 (SharpAI/DeepCamera, 3.1k stars), Yolo Detection 2026 Coral Tpu Win Wsl (SharpAI/DeepCamera, 3.1k stars) and Gradio App Builder (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLaVA Vision-Language Model?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.