Agent skill

Gpt Multimodal

by benchflow-ai in benchflow-ai/skillsbench

Analyze images and multi-frame sequences using OpenAI GPT series

Apache-2.0Auto-check passed

Install Gpt Multimodal

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill gpt-multimodal -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench gpt-multimodal --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks-extra/pedestrian-traffic-counting/environment/skills/gpt-multimodal .claude/skills/gpt-multimodal && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gpt-multimodal
GitHub stars
1.8k
Token cost
~4.9k tokens
SKILL.md length
765 words
Files
1
Skills in repo
189
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze images and multi-frame sequences using OpenAI GPT series

  • Works in 4 steps: Use low detail for video frames -… → Resize large images before uploading -… → Batch related questions - Analyze… → …
  • SKILL.md covers Purpose, When to Use, Required Libraries and Input Requirements, plus 7 more sections
  • Needs OPENAI_API_KEY

What it does

Gpt Multimodal is an agent skill from benchflow-ai/skillsbench. Analyze images and multi-frame sequences using OpenAI GPT series

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with OpenAI. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

Example prompts

  • “/gpt-multimodal”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Use low detail for video frames - Temporal analysis doesn't need high resolution
  2. Resize large images before uploading - Reduce dimensions to 1024×1024 if high detail not needed
  3. Batch related questions - Analyze multiple aspects in one API call
  4. Cache analysis results - Store results for repeated processing

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gpt Multimodal loads about 4.9k tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 765 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 765 words, ~4,873 tokens.

Download SKILL.mdSave it as .claude/skills/gpt-multimodal/SKILL.md (or your agent's skills folder).
name
gpt-multimodal
description
Analyze images and multi-frame sequences using OpenAI GPT series

OpenAI Vision Analysis Skill

Purpose

This skill enables image analysis, scene understanding, text extraction, and multi-frame comparison using OpenAI's vision-capable GPT models (e.g., gpt-4o, gpt-5). It supports single and multiple images analysis and sequential frames for temporal analysis.

When to Use

  • Analyzing image content (objects, scenes, colors, spatial relationships)
  • Extracting and reading text from images (OCR via vision models)
  • Comparing multiple images to detect differences or changes
  • Processing video frames to understand temporal progression
  • Generating detailed image descriptions or captions
  • Answering questions about visual content

Required Libraries

The following Python libraries are required:

python
from openai import OpenAI
import base64
import json
import os
from pathlib import Path

Input Requirements

  • File formats: JPG, JPEG, PNG, WEBP, non-animated GIF
  • Image quality: Clear and legible; minimum 512×512px recommended
  • File size: Under 20MB per image recommended
  • Maximum per request: Up to 500 images, 50MB total payload
  • URL or Base64: Images can be provided as URLs or base64-encoded data

Output Schema

All analysis results should be returned as valid JSON conforming to this schema:

json
{
  "success": true,
  "model": "gpt-5",
  "analysis": "Detailed description or analysis of the image content...",
  "metadata": {
    "image_count": 1,
    "detail_level": "high",
    "tokens_used": 850,
    "processing_time_ms": 1234
  },
  "extracted_data": {
    "objects": ["car", "person", "building"],
    "text_found": "Sample text from image",
    "colors": ["blue", "white", "gray"],
    "scene_type": "urban street"
  },
  "warnings": []
}
Field Descriptions
  • success: Boolean indicating whether the API call succeeded
  • model: The GPT model used for analysis (e.g., "gpt-4o", "gpt-5")
  • analysis: Complete textual analysis or description from the model
  • metadata.image_count: Number of images analyzed in this request
  • metadata.detail_level: Detail parameter used ("low", "high", or "auto")
  • metadata.tokens_used: Approximate token count for the request
  • metadata.processing_time_ms: Time taken to process the request
  • extracted_data: Structured information extracted from the image(s)
  • warnings: Array of issues or limitations encountered

Code Examples

Basic Image Analysis
python
from openai import OpenAI
import base64

def analyze_image(image_path, prompt="What's in this image?"):
    """Analyze a single image using GPT-5 Vision."""
    client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
    
    # Read and encode image
    with open(image_path, "rb") as image_file:
        base64_image = base64.b64encode(image_file.read()).decode('utf-8')
    
    response = client.chat.completions.create(
        model="gpt-5",
        messages=[
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": prompt},
                    {
                        "type": "image_url",
                        "image_url": {
                            "url": f"data:image/jpeg;base64,{base64_image}"
                        }
                    }
                ]
            }
        ],
        max_tokens=300
    )
    
    return response.choices[0].message.content
Using Image URLs
python
from openai import OpenAI

def analyze_image_url(image_url, prompt="Describe this image"):
    """Analyze an image from a URL."""
    client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
    
    response = client.chat.completions.create(
        model="gpt-5",
        messages=[
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": prompt},
                    {
                        "type": "image_url",
                        "image_url": {"url": image_url}
                    }
                ]
            }
        ]
    )
    
    return response.choices[0].message.content
Multiple Images Analysis
python
from openai import OpenAI
import base64

def analyze_multiple_images(image_paths, prompt="Compare these images"):
    """Analyze multiple images in a single request."""
    client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
    
    # Build content array with text and all images
    content = [{"type": "text", "text": prompt}]
    
    for image_path in image_paths:
        with open(image_path, "rb") as image_file:
            base64_image = base64.b64encode(image_file.read()).decode('utf-8')
        
        content.append({
            "type": "image_url",
            "image_url": {
                "url": f"data:image/jpeg;base64,{base64_image}"
            }
        })
    
    response = client.chat.completions.create(
        model="gpt-5",
        messages=[{"role": "user", "content": content}],
        max_tokens=500
    )
    
    return response.choices[0].message.content
Full Analysis with JSON Output
python
from openai import OpenAI
import base64
import json
import time

def analyze_image_to_json(image_path, prompt="Analyze this image"):
    """Analyze image and return structured JSON output."""
    client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
    
    start_time = time.time()
    warnings = []
    
    try:
        # Read and encode image
        with open(image_path, "rb") as image_file:
            base64_image = base64.b64encode(image_file.read()).decode('utf-8')
        
        # Make API call
        response = client.chat.completions.create(
            model="gpt-5",
            messages=[
                {
                    "role": "user",
                    "content": [
                        {"type": "text", "text": prompt},
                        {
                            "type": "image_url",
                            "image_url": {
                                "url": f"data:image/jpeg;base64,{base64_image}",
                                "detail": "high"
                            }
                        }
                    ]
                }
            ],
            max_tokens=500
        )
        
        analysis = response.choices[0].message.content
        tokens_used = response.usage.total_tokens
        processing_time = int((time.time() - start_time) * 1000)
        
        result = {
            "success": True,
            "model": "gpt-5",
            "analysis": analysis,
            "metadata": {
                "image_count": 1,
                "detail_level": "high",
                "tokens_used": tokens_used,
                "processing_time_ms": processing_time
            },
            "extracted_data": {},
            "warnings": warnings
        }
        
    except Exception as e:
        result = {
            "success": False,
            "model": "gpt-5",
            "analysis": "",
            "metadata": {
                "image_count": 0,
                "detail_level": "high",
                "tokens_used": 0,
                "processing_time_ms": 0
            },
            "extracted_data": {},
            "warnings": [f"API call failed: {str(e)}"]
        }
    
    return result

# Usage
result = analyze_image_to_json("photo.jpg", "Describe what you see in detail")
print(json.dumps(result, indent=2))
Batch Processing with Sequential Frames
python
from openai import OpenAI
import base64
from pathlib import Path

def process_video_frames(frames_directory, analysis_prompt):
    """Process sequential video frames for temporal analysis."""
    client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
    
    image_extensions = {'.jpg', '.jpeg', '.png', '.webp'}
    frame_paths = sorted([
        f for f in Path(frames_directory).iterdir() 
        if f.suffix.lower() in image_extensions
    ])
    
    # Analyze frames in groups (e.g., 5 frames at a time)
    batch_size = 5
    results = []
    
    for i in range(0, len(frame_paths), batch_size):
        batch = frame_paths[i:i+batch_size]
        
        # Build content with all frames in batch
        content = [{"type": "text", "text": analysis_prompt}]
        
        for frame_path in batch:
            with open(frame_path, "rb") as f:
                base64_image = base64.b64encode(f.read()).decode('utf-8')
            
            content.append({
                "type": "image_url",
                "image_url": {
                    "url": f"data:image/jpeg;base64,{base64_image}",
                    "detail": "low"  # Use low detail for video frames to save tokens
                }
            })
        
        response = client.chat.completions.create(
            model="gpt-5",
            messages=[{"role": "user", "content": content}],
            max_tokens=800
        )
        
        results.append({
            "batch_index": i // batch_size,
            "frame_range": f"{batch[0].name} to {batch[-1].name}",
            "analysis": response.choices[0].message.content
        })
    
    return results
Text Extraction from Images (OCR Alternative)
python
from openai import OpenAI
import base64

def extract_text_with_gpt(image_path):
    """Extract text from image using GPT Vision as OCR alternative."""
    client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
    
    with open(image_path, "rb") as image_file:
        base64_image = base64.b64encode(image_file.read()).decode('utf-8')
    
    response = client.chat.completions.create(
        model="gpt-5",
        messages=[
            {
                "role": "user",
                "content": [
                    {
                        "type": "text", 
                        "text": "Extract all text from this image. Return only the text content, preserving the layout and structure."
                    },
                    {
                        "type": "image_url",
                        "image_url": {
                            "url": f"data:image/jpeg;base64,{base64_image}",
                            "detail": "high"
                        }
                    }
                ]
            }
        ],
        max_tokens=1000
    )
    
    return response.choices[0].message.content

Model Selection and Configuration

Available Models
python
# GPT-4o - Best for general vision tasks, fast and cost-effective
model = "gpt-4o"

# GPT-5-nano - Faster and cheaper for simple vision tasks
model = "gpt-5-nano"

# GPT-5 - More capable for complex reasoning
model = "gpt-5"
Detail Level Configuration

Control how much visual detail the model processes:

python
# Low detail - 512×512px resolution, fewer tokens, faster
"image_url": {
    "url": image_url,
    "detail": "low"
}

# High detail - Full resolution with tiling, more tokens, better accuracy
"image_url": {
    "url": image_url,
    "detail": "high"
}

# Auto - Model chooses appropriate detail level
"image_url": {
    "url": image_url,
    "detail": "auto"
}

When to use each detail level:

  • Low: Video frames, simple scene classification, color/shape detection
  • High: Text extraction, detailed object detection, fine-grained analysis
  • Auto: General purpose when unsure; model optimizes cost vs. quality

Token Cost Management

Understanding Image Tokens

Image tokens count toward your request limits and costs:

  • Low detail: Fixed ~85 tokens per image (gpt-5)
  • High detail: Base tokens + tile tokens based on image dimensions
    • Images are scaled to fit within 2048×2048px
    • Divided into 512×512px tiles
    • Each tile costs additional tokens
Cost Calculation Examples
python
# For gpt-5 with high detail:
# - Base: 85 tokens
# - Per tile: 170 tokens
# - Example: 1024×1024 image = 85 + (2×2 tiles × 170) = 765 tokens

def estimate_tokens_high_detail(width, height):
    """Estimate token cost for high-detail image (gpt-5)."""
    # Scale to fit within 2048×2048
    scale = min(2048 / width, 2048 / height, 1.0)
    scaled_w = int(width * scale)
    scaled_h = int(height * scale)
    
    # Calculate tiles (512×512)
    tiles_x = (scaled_w + 511) // 512
    tiles_y = (scaled_h + 511) // 512
    total_tiles = tiles_x * tiles_y
    
    # Token calculation
    base_tokens = 85
    tile_tokens = total_tiles * 170
    
    return base_tokens + tile_tokens
Cost Optimization Strategies
  1. Use low detail for video frames - Temporal analysis doesn't need high resolution
  2. Resize large images before uploading - Reduce dimensions to 1024×1024 if high detail not needed
  3. Batch related questions - Analyze multiple aspects in one API call
  4. Cache analysis results - Store results for repeated processing

Advanced Use Cases

Image Comparison
python
def compare_images(image1_path, image2_path):
    """Compare two images and identify differences."""
    client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
    
    images = []
    for path in [image1_path, image2_path]:
        with open(path, "rb") as f:
            base64_image = base64.b64encode(f.read()).decode('utf-8')
            images.append(base64_image)
    
    response = client.chat.completions.create(
        model="gpt-5",
        messages=[
            {
                "role": "user",
                "content": [
                    {
                        "type": "text",
                        "text": "Compare these two images. List all differences you observe, including changes in objects, colors, positions, or any other visual elements."
                    },
                    {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{images[0]}"}},
                    {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{images[1]}"}}
                ]
            }
        ]
    )
    
    return response.choices[0].message.content
Structured Data Extraction
python
def extract_structured_data(image_path, schema_description):
    """Extract structured information from image based on schema."""
    client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
    
    with open(image_path, "rb") as f:
        base64_image = base64.b64encode(f.read()).decode('utf-8')
    
    prompt = f"""Analyze this image and extract information in JSON format following this schema:
{schema_description}

Return only valid JSON, no additional text."""
    
    response = client.chat.completions.create(
        model="gpt-5",
        messages=[
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": prompt},
                    {
                        "type": "image_url",
                        "image_url": {
                            "url": f"data:image/jpeg;base64,{base64_image}",
                            "detail": "high"
                        }
                    }
                ]
            }
        ],
        response_format={"type": "json_object"}
    )
    
    return json.loads(response.choices[0].message.content)

# Usage example
schema = """
{
  "products": [{"name": string, "price": number, "quantity": number}],
  "total": number,
  "date": string
}
"""
data = extract_structured_data("receipt.jpg", schema)

Error Handling

Common Issues and Solutions

Issue: API authentication failed

python
# Verify API key is set
import os
api_key = os.environ.get("OPENAI_API_KEY")
if not api_key:
    raise ValueError("OPENAI_API_KEY environment variable not set")

Issue: Image too large

python
from PIL import Image

def resize_if_needed(image_path, max_size=2048):
    """Resize image if dimensions exceed maximum."""
    img = Image.open(image_path)
    
    if max(img.size) > max_size:
        img.thumbnail((max_size, max_size), Image.Resampling.LANCZOS)
        resized_path = image_path.replace('.', '_resized.')
        img.save(resized_path, quality=95)
        return resized_path
    
    return image_path

Issue: Token limit exceeded

python
# Reduce max_tokens or use low detail mode
response = client.chat.completions.create(
    model="gpt-5",
    messages=[...],
    max_tokens=300  # Reduce from default
)

Issue: Rate limit errors

python
import time
from openai import RateLimitError

def analyze_with_retry(image_path, max_retries=3):
    """Analyze image with exponential backoff on rate limits."""
    for attempt in range(max_retries):
        try:
            return analyze_image(image_path)
        except RateLimitError:
            if attempt < max_retries - 1:
                wait_time = 2 ** attempt  # Exponential backoff
                print(f"Rate limit hit, waiting {wait_time}s...")
                time.sleep(wait_time)
            else:
                raise

Best Practices

Show full SKILL.md (326 more words)Show less
Prompt Engineering for Vision
  1. Be specific: "Count the number of people wearing red shirts" vs "Analyze this image"
  2. Request structured output: Ask for JSON, lists, or tables when appropriate
  3. Provide context: "This is a medical diagram showing..." helps the model understand
  4. Use examples: Show the format you want in your prompt
Image Quality Guidelines
  • Use clear, well-lit images
  • Ensure text is readable at original size
  • Avoid extreme angles or distortions
  • Crop to relevant content to save tokens
  • Use standard orientations (avoid rotated images)
Multi-Image Analysis
  • Order matters: Present images in logical sequence
  • Reference images explicitly: "In the first image..."
  • Limit to 10-20 images per request for best results
  • Use low detail for large batches of similar images

Quality Self-Check

Before returning results, verify:

  • Output is valid JSON (use json.loads() to validate)
  • All required fields are present and properly typed
  • API errors are caught and handled gracefully
  • Token usage is tracked and within limits
  • Image formats are supported (JPG, PNG, WEBP, GIF)
  • Base64 encoding is correct (no corruption)
  • Model name is valid and available
  • Results are consistent with the prompt request

Limitations

Vision Model Limitations
  • Not for medical diagnosis: Cannot interpret CT scans, X-rays, or provide medical advice
  • Poor text recognition: Struggles with rotated, upside-down, or very small text (< 10pt)
  • Non-Latin scripts: Reduced accuracy for non-Latin alphabets and special characters
  • Spatial reasoning: Weak at precise localization tasks (e.g., chess positions, exact coordinates)
  • Graph interpretation: Cannot reliably distinguish line styles (solid vs dashed) or precise color gradients
  • Object counting: Provides approximate counts; may miss or double-count objects
  • Panoramic/fisheye: Distorted perspectives reduce accuracy
  • Metadata loss: Original EXIF data and exact dimensions are not preserved
Performance Considerations
  • Latency: High-detail mode takes longer to process
  • Token costs: Multiple images can quickly consume token budgets
  • Rate limits: Vision requests count toward TPM (tokens per minute) limits
  • File size: 20MB per image practical limit; 50MB total per request
  • Batch size: Over 20 images may degrade quality or hit timeouts

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks-extra/pedestrian-traffic-counting/environment/skills/gpt-multimodal of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Gpt Multimodal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gpt Multimodal compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gpt Multimodal this skillbenchflow-ai/skillsbench1.8k—~4.9kAutomated safety check: PassApache-2.0
Geo Fundamentalswasp-lang/wasp19k9 repos~861Automated safety check: PassMIT
AI SDKvercel-labs/ai-facts16820 repos~1.2kAutomated safety check: PassNone
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
PR Design DocOpenHands/OpenHands90k—~2.4kAutomated safety check: PassMIT
SEO GeoReScienceLab/opc-skills1.8k4 repos~2.1kAutomated safety check: PassApache-2.0

Similar skills

  • Geo Fundamentals

    wasp-lang/wasp

    Generative Engine Optimization for AI search engines (ChatGPT, Claude, Perplexity).

    19k GitHub starsUsed in 9 repos~861 tokens
    Marketing & SEOAuto-check passed
  • AI SDK

    vercel-labs/ai-facts

    Official

    Answer questions about the AI SDK and help build AI-powered features.

    168 GitHub starsUsed in 20 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • SEO Geo

    ReScienceLab/opc-skills

    SEO & GEO (Generative Engine Optimization) for websites. An agent skill from ReScienceLab/opc-skills.

    1.8k GitHub starsUsed in 4 repos~2.1k tokens
    Marketing & SEOAuto-check passed
  • Open Code Review CLI

    alibaba/open-code-review

    Runs the ocr command-line tool to review Git changes, a commit or a branch comparison with an AI model, returning line-level comments and optionally applying fixes.

    45k GitHub stars~3.1k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from benchflow-ai/skillsbench

All 189 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Gpt Multimodal

What does Gpt Multimodal do?

Analyze images and multi-frame sequences using OpenAI GPT series. Gpt Multimodal is an agent skill from benchflow-ai/skillsbench.

How do I install Gpt Multimodal in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill gpt-multimodal -a claude-code`. Or copy the skill folder (tasks-extra/pedestrian-traffic-counting/environment/skills/gpt-multimodal in benchflow-ai/skillsbench) into .claude/skills/gpt-multimodal in your project. Claude Code loads it when a task matches its description.

How do I install Gpt Multimodal in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill gpt-multimodal -a codex`. Or copy the skill folder (tasks-extra/pedestrian-traffic-counting/environment/skills/gpt-multimodal in benchflow-ai/skillsbench) into .agents/skills/gpt-multimodal in your project. Codex loads it when a task matches its description.

Can I use Gpt Multimodal in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill gpt-multimodal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpt-multimodal, .gemini/skills/gpt-multimodal, .github/skills/gpt-multimodal and .opencode/skills/gpt-multimodal in your project.

What does Gpt Multimodal need to run?

Going by SKILL.md and its folder, Gpt Multimodal needs credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Gpt Multimodal access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gpt Multimodal safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gpt Multimodal use?

Gpt Multimodal is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gpt Multimodal use?

About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gpt Multimodal?

Skills that share tags, products or a category with Gpt Multimodal: Geo Fundamentals (wasp-lang/wasp, 19k stars), AI SDK (vercel-labs/ai-facts, 168 stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars) and PR Design Doc (OpenHands/OpenHands, 90k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gpt Multimodal?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,834 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.