CLIP Image-Text Matching
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Analyze images and multi-frame sequences using OpenAI GPT vision models
$ npx skills add benchflow-ai/skillsbench --skill openai-vision -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install benchflow-ai/skillsbench openai-vision --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/openai-vision .claude/skills/openai-vision && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "openai-vision" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/openai-vision into .claude/skills/openai-vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "openai-vision", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/openai-visionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add benchflow-ai/skillsbench --skill openai-vision -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install benchflow-ai/skillsbench openai-vision --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .agents/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/openai-vision .agents/skills/openai-vision && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "openai-vision" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/openai-vision into .agents/skills/openai-vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "openai-vision", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/skillsbench --skill openai-vision -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install benchflow-ai/skillsbench openai-vision --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/openai-vision .cursor/skills/openai-vision && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "openai-vision" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/openai-vision into .cursor/skills/openai-vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "openai-vision", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/benchflow-ai/skillsbench.git --path tasks/jpg-ocr-stat/environment/skills/openai-vision--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add benchflow-ai/skillsbench --skill openai-vision -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install benchflow-ai/skillsbench openai-vision --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/openai-vision .gemini/skills/openai-vision && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "openai-vision" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/openai-vision into .gemini/skills/openai-vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "openai-vision", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install benchflow-ai/skillsbench openai-visionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add benchflow-ai/skillsbench --skill openai-vision -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .github/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/openai-vision .github/skills/openai-vision && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "openai-vision" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/openai-vision into .github/skills/openai-vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "openai-vision", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add benchflow-ai/skillsbench --skill openai-vision -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install benchflow-ai/skillsbench openai-vision --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/tasks/jpg-ocr-stat/environment/skills/openai-vision .opencode/skills/openai-vision && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "openai-vision" agent skill from https://github.com/benchflow-ai/skillsbench/tree/main/tasks/jpg-ocr-stat/environment/skills/openai-vision into .opencode/skills/openai-vision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "openai-vision", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
openai-visionAnalyze images and multi-frame sequences using OpenAI GPT vision models
Openai Vision is an agent skill from benchflow-ai/skillsbench. Analyze images and multi-frame sequences using OpenAI GPT vision models
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Computer vision. It works with OpenAI. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Openai Vision loads about 5k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 522 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 522 words, ~5,001 tokens.
.claude/skills/openai-vision/SKILL.md (or your agent's skills folder).This skill enables image analysis, scene understanding, text extraction, and multi-frame comparison using OpenAI's vision-capable GPT models (e.g., gpt-4o, gpt-4o-mini). It supports single images, multiple images for comparison, and sequential frames for temporal analysis.
The following Python libraries are required:
from openai import OpenAI
import base64
import json
import os
from pathlib import PathAnalysis results should be returned as valid JSON conforming to this schema:
{
"success": true,
"images_analyzed": 1,
"analysis": {
"description": "A detailed scene description...",
"objects": [
{"name": "car", "color": "red", "position": "foreground center"},
{"name": "tree", "count": 3, "position": "background"}
],
"text_content": "Any text visible in the image...",
"colors": ["blue", "green", "white"],
"scene_type": "outdoor/urban"
},
"comparison": {
"differences": ["Object X appeared", "Color changed from A to B"],
"similarities": ["Background unchanged", "Layout consistent"]
},
"metadata": {
"model_used": "gpt-4o",
"detail_level": "high",
"token_usage": {"prompt": 1500, "completion": 200}
},
"warnings": []
}success: Boolean indicating whether analysis completedimages_analyzed: Number of images processed in the requestanalysis.description: Natural language description of the image contentanalysis.objects: Array of detected objects with attributesanalysis.text_content: Any text extracted from the imageanalysis.colors: Dominant colors identifiedanalysis.scene_type: Classification of the scenecomparison: Present when multiple images are analyzed; describes differences and similaritiesmetadata.model_used: The GPT model used for analysismetadata.detail_level: Resolution level used (low, high, or auto)metadata.token_usage: Token consumption for cost trackingwarnings: Array of any issues or limitations encounteredfrom openai import OpenAI
client = OpenAI()
def analyze_image_url(image_url, prompt="Describe this image in detail."):
"""Analyze an image from a URL using GPT-4o vision."""
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {
"url": image_url,
"detail": "high"
}
}
]
}
],
max_tokens=1000
)
return response.choices[0].message.contentfrom openai import OpenAI
import base64
client = OpenAI()
def encode_image_to_base64(image_path):
"""Encode a local image file to base64."""
with open(image_path, "rb") as image_file:
return base64.standard_b64encode(image_file.read()).decode("utf-8")
def get_image_media_type(image_path):
"""Determine the media type based on file extension."""
ext = image_path.lower().split('.')[-1]
media_types = {
'jpg': 'image/jpeg',
'jpeg': 'image/jpeg',
'png': 'image/png',
'gif': 'image/gif',
'webp': 'image/webp'
}
return media_types.get(ext, 'image/jpeg')
def analyze_local_image(image_path, prompt="Describe this image in detail."):
"""Analyze a local image file using GPT-4o vision."""
base64_image = encode_image_to_base64(image_path)
media_type = get_image_media_type(image_path)
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {
"url": f"data:{media_type};base64,{base64_image}",
"detail": "high"
}
}
]
}
],
max_tokens=1000
)
return response.choices[0].message.contentfrom openai import OpenAI
import base64
client = OpenAI()
def compare_images(image_paths, comparison_prompt=None):
"""Compare multiple images and identify differences."""
if comparison_prompt is None:
comparison_prompt = (
"Compare these images carefully. "
"List all differences and similarities you observe. "
"Describe any changes in objects, colors, positions, or text."
)
content = [{"type": "text", "text": comparison_prompt}]
for i, image_path in enumerate(image_paths):
base64_image = encode_image_to_base64(image_path)
media_type = get_image_media_type(image_path)
# Add label for each image
content.append({
"type": "text",
"text": f"Image {i + 1}:"
})
content.append({
"type": "image_url",
"image_url": {
"url": f"data:{media_type};base64,{base64_image}",
"detail": "high"
}
})
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": content}],
max_tokens=2000
)
return response.choices[0].message.contentfrom openai import OpenAI
import base64
from pathlib import Path
client = OpenAI()
def analyze_video_frames(frame_paths, analysis_prompt=None):
"""Analyze a sequence of video frames for temporal understanding."""
if analysis_prompt is None:
analysis_prompt = (
"These are sequential frames from a video. "
"Describe what is happening over time. "
"Identify any motion, changes, or events that occur across the frames."
)
content = [{"type": "text", "text": analysis_prompt}]
for i, frame_path in enumerate(frame_paths):
base64_image = encode_image_to_base64(frame_path)
media_type = get_image_media_type(frame_path)
content.append({
"type": "text",
"text": f"Frame {i + 1}:"
})
content.append({
"type": "image_url",
"image_url": {
"url": f"data:{media_type};base64,{base64_image}",
"detail": "auto" # Use auto for frames to balance cost
}
})
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": content}],
max_tokens=2000
)
return response.choices[0].message.contentfrom openai import OpenAI
import base64
import json
import os
client = OpenAI()
def analyze_image_to_json(image_path, extract_text=True):
"""Perform comprehensive image analysis and return structured JSON."""
filename = os.path.basename(image_path)
prompt = """Analyze this image and return a JSON object with the following structure:
{
"description": "detailed scene description",
"objects": [{"name": "object name", "attributes": "color, size, position"}],
"text_content": "any visible text or null if none",
"colors": ["dominant", "colors"],
"scene_type": "indoor/outdoor/abstract/etc",
"people_count": 0,
"notable_features": ["list of notable visual elements"]
}
Return ONLY valid JSON, no other text."""
try:
base64_image = encode_image_to_base64(image_path)
media_type = get_image_media_type(image_path)
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {
"url": f"data:{media_type};base64,{base64_image}",
"detail": "high"
}
}
]
}
],
max_tokens=1500
)
# Parse the response as JSON
analysis_text = response.choices[0].message.content
# Remove markdown code blocks if present
if analysis_text.startswith("```"):
analysis_text = analysis_text.split("```")[1]
if analysis_text.startswith("json"):
analysis_text = analysis_text[4:]
analysis = json.loads(analysis_text.strip())
result = {
"success": True,
"filename": filename,
"analysis": analysis,
"metadata": {
"model_used": "gpt-4o",
"detail_level": "high",
"token_usage": {
"prompt": response.usage.prompt_tokens,
"completion": response.usage.completion_tokens
}
},
"warnings": []
}
except json.JSONDecodeError as e:
result = {
"success": False,
"filename": filename,
"analysis": {"raw_response": response.choices[0].message.content},
"metadata": {"model_used": "gpt-4o"},
"warnings": [f"Failed to parse JSON: {str(e)}"]
}
except Exception as e:
result = {
"success": False,
"filename": filename,
"analysis": {},
"metadata": {},
"warnings": [f"Analysis failed: {str(e)}"]
}
return result
# Usage
result = analyze_image_to_json("photo.jpg")
print(json.dumps(result, indent=2))from openai import OpenAI
import base64
import json
from pathlib import Path
client = OpenAI()
def process_image_directory(directory_path, output_file, prompt=None):
"""Process all images in a directory and save results."""
if prompt is None:
prompt = "Describe this image briefly, including any visible text."
image_extensions = {'.jpg', '.jpeg', '.png', '.webp', '.gif'}
results = []
for file_path in sorted(Path(directory_path).iterdir()):
if file_path.suffix.lower() in image_extensions:
print(f"Processing: {file_path.name}")
try:
analysis = analyze_local_image(str(file_path), prompt)
results.append({
"filename": file_path.name,
"success": True,
"analysis": analysis
})
except Exception as e:
results.append({
"filename": file_path.name,
"success": False,
"error": str(e)
})
# Save results
with open(output_file, 'w') as f:
json.dump(results, f, indent=2)
return resultsThe detail parameter controls image resolution and token usage:
# Low detail: 512x512 fixed, ~85 tokens per image
# Best for: Quick summaries, dominant colors, general scene type
{"detail": "low"}
# High detail: Full resolution processing
# Best for: Reading text, detecting small objects, detailed analysis
{"detail": "high"}
# Auto: Model decides based on image size
# Best for: General use when cost vs quality tradeoff is acceptable
{"detail": "auto"}def get_recommended_detail(task_type):
"""Recommend detail level based on task type."""
high_detail_tasks = {
'ocr', 'text_extraction', 'document_analysis',
'small_object_detection', 'detailed_comparison',
'fine_grained_analysis'
}
low_detail_tasks = {
'scene_classification', 'dominant_colors',
'general_description', 'thumbnail_preview'
}
if task_type.lower() in high_detail_tasks:
return "high"
elif task_type.lower() in low_detail_tasks:
return "low"
else:
return "auto"For extracting text from images using vision models:
def extract_text_from_image(image_path, preserve_layout=False):
"""Extract text from an image using GPT-4o vision."""
if preserve_layout:
prompt = (
"Extract ALL text visible in this image. "
"Preserve the original layout and formatting as much as possible. "
"Include headers, paragraphs, captions, and any other text. "
"Return only the extracted text, nothing else."
)
else:
prompt = (
"Extract all text visible in this image. "
"Return the text in reading order (top to bottom, left to right). "
"Return only the extracted text, nothing else."
)
return analyze_local_image(image_path, prompt)
def extract_structured_text(image_path):
"""Extract text with structure information as JSON."""
prompt = """Extract all text from this image and return as JSON:
{
"headers": ["list of headers/titles"],
"paragraphs": ["list of paragraph texts"],
"labels": ["list of labels or captions"],
"other_text": ["any other text elements"],
"reading_order": ["all text in reading order"]
}
Return ONLY valid JSON."""
response = analyze_local_image(image_path, prompt)
try:
# Clean and parse JSON
if response.startswith("```"):
response = response.split("```")[1]
if response.startswith("json"):
response = response[4:]
return json.loads(response.strip())
except json.JSONDecodeError:
return {"raw_text": response, "parse_error": True}Issue: API rate limits exceeded
import time
from openai import RateLimitError
def analyze_with_retry(image_path, prompt, max_retries=3):
"""Analyze image with exponential backoff retry."""
for attempt in range(max_retries):
try:
return analyze_local_image(image_path, prompt)
except RateLimitError:
if attempt < max_retries - 1:
wait_time = 2 ** attempt # Exponential backoff
print(f"Rate limited, waiting {wait_time}s...")
time.sleep(wait_time)
else:
raiseIssue: Image too large
from PIL import Image
import io
def resize_image_if_needed(image_path, max_size_mb=15):
"""Resize image if it exceeds size limit."""
file_size_mb = os.path.getsize(image_path) / (1024 * 1024)
if file_size_mb <= max_size_mb:
return encode_image_to_base64(image_path)
# Resize the image
img = Image.open(image_path)
# Calculate new dimensions (reduce by 50% iteratively)
while file_size_mb > max_size_mb:
new_width = img.width // 2
new_height = img.height // 2
img = img.resize((new_width, new_height), Image.Resampling.LANCZOS)
# Check new size
buffer = io.BytesIO()
img.save(buffer, format='JPEG', quality=85)
file_size_mb = len(buffer.getvalue()) / (1024 * 1024)
# Encode resized image
buffer = io.BytesIO()
img.save(buffer, format='JPEG', quality=85)
return base64.standard_b64encode(buffer.getvalue()).decode("utf-8")Issue: Invalid image format
def validate_image(image_path):
"""Validate image before processing."""
valid_extensions = {'.jpg', '.jpeg', '.png', '.gif', '.webp'}
path = Path(image_path)
if not path.exists():
return False, "File does not exist"
if path.suffix.lower() not in valid_extensions:
return False, f"Unsupported format: {path.suffix}"
try:
with Image.open(image_path) as img:
img.verify()
return True, "Valid image"
except Exception as e:
return False, f"Invalid image: {str(e)}"Before returning results, verify:
detail: highApproximate token costs for image inputs:
| Detail Level | Tokens per Image | Best For |
|---|---|---|
| low | ~85 tokens (fixed) | Quick classification, color detection |
| high | 85 + 170 per 512x512 tile | OCR, detailed analysis, small objects |
| auto | Variable | General use |
def estimate_image_tokens(image_path, detail="high"):
"""Estimate token usage for an image."""
if detail == "low":
return 85
with Image.open(image_path) as img:
width, height = img.size
# High detail: image is scaled to fit in 2048x2048, then tiled at 512x512
scale = min(2048 / max(width, height), 1.0)
scaled_width = int(width * scale)
scaled_height = int(height * scale)
# Ensure minimum 768 on shortest side
if min(scaled_width, scaled_height) < 768:
scale = 768 / min(scaled_width, scaled_height)
scaled_width = int(scaled_width * scale)
scaled_height = int(scaled_height * scale)
# Calculate tiles
tiles_x = (scaled_width + 511) // 512
tiles_y = (scaled_height + 511) // 512
total_tiles = tiles_x * tiles_y
return 85 + (170 * total_tiles)© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in tasks/jpg-ocr-stat/environment/skills/openai-vision of benchflow-ai/skillsbench.
Open the folder on GitHubat commit 9a1f4dd
Openai Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Openai Vision this skillbenchflow-ai/skillsbench | 1.8k | — | ~5k | Automated safety check: Pass | Apache-2.0 | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~1.7k | Automated safety check: Pass | MIT | |
| I4h Workflow Dataset AnnotateNVIDIA/skills | 3.5k | 1 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Visionaiskillstore/marketplace | 430 | — | ~1.1k | Automated safety check: Pass | None | |
| ModLens Image Vision Bridgeliustack/modlens | 4.2k | — | ~1.3k | Automated safety check: Notes | MIT | |
| Visionxiincs/claude-code-vision-skill | 170 | — | ~1.2k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
NVIDIA/skills
Grade or filter workflow HDF5 episodes with an OpenAI-compatible vision model.
aiskillstore/marketplace
See and understand images when you (the current model) have no native vision.
liustack/modlens
Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.
xiincs/claude-code-vision-skill
Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.
QianWen-AI/qianwen-ai
Generate text, have conversations, write code, reason, and call functions with Qwen models.
benchflow-ai/skillsbench
This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
benchflow-ai/skillsbench
AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.
benchflow-ai/skillsbench
Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.
benchflow-ai/skillsbench
Build deterministic, verifiable data visualizations with D3.js (v6).
benchflow-ai/skillsbench
DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.
Works with
Categories
Analyze images and multi-frame sequences using OpenAI GPT vision models. Openai Vision is an agent skill from benchflow-ai/skillsbench.
Openai Vision fits situations like: tasks that involve Computer vision.
Run `npx skills add benchflow-ai/skillsbench --skill openai-vision -a claude-code`. Or copy the skill folder (tasks/jpg-ocr-stat/environment/skills/openai-vision in benchflow-ai/skillsbench) into .claude/skills/openai-vision in your project. Claude Code loads it when a task matches its description.
Run `npx skills add benchflow-ai/skillsbench --skill openai-vision -a codex`. Or copy the skill folder (tasks/jpg-ocr-stat/environment/skills/openai-vision in benchflow-ai/skillsbench) into .agents/skills/openai-vision in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill openai-vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/openai-vision, .gemini/skills/openai-vision, .github/skills/openai-vision and .opencode/skills/openai-vision in your project.
SKILL.md names no scripts, command-line tools or credentials: Openai Vision is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Openai Vision is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Openai Vision: CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), I4h Workflow Dataset Annotate (NVIDIA/skills, 3.5k stars), Vision (aiskillstore/marketplace, 430 stars) and ModLens Image Vision Bridge (liustack/modlens, 4.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,834 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.
Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.