Agent skill

Image Understand

by Ali-Marandi in Ali-Marandi/Web-Scraper-Framework

Implement specialized image understanding capabilities using the z-ai-web-dev-sdk.

MITAuto-check passed

Install Image Understand

skills CLI
$ npx skills add Ali-Marandi/Web-Scraper-Framework --skill image-understand -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Ali-Marandi/Web-Scraper-Framework image-understand --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Ali-Marandi/Web-Scraper-Framework.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/image-understand .claude/skills/image-understand && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-understand
GitHub stars
107
Token cost
~5.6k tokens
SKILL.md length
854 words
Files
3 (incl. scripts)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Implement specialized image understanding capabilities using the z-ai-web-dev-sdk.

  • Works in 5 steps: Image Quality and Preparation → Prompt Engineering for Images → Error Handling → …
  • The user needs to analyze static images
  • SKILL.md covers Skills Path, Overview, Prerequisites and CLI Usage (For Simple Tasks), plus 5 more sections
  • Runs TypeScript scripts from its folder

What it does

Image Understand is an agent skill from Ali-Marandi/Web-Scraper-Framework. Implement specialized image understanding capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to analyze static images, extract visual information, perform OCR, detect objects, classify images, or understand visual content. Optimized for PNG, JPEG, GIF, WebP, and BMP formats.

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/image-understand.ts`).

It works with Zhipu GLM. The repository describes itself as: Flexible and Scalable Web Scraping Framework. The licence is MIT.

When your agent uses it

  • The user needs to analyze static images
  • Extract visual information
  • Classify images
  • Understand visual content

Example prompts

  • “/image-understand”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Image Quality and Preparation
  2. Prompt Engineering for Images
  3. Error Handling
  4. Performance Optimization
  5. Security Considerations

What it can do on your machine

Read from SKILL.md and the folder at commit f8af4cd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (TypeScript), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image Understand loads about 5.6k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 854 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Ali-Marandi/Web-Scraper-Framework at commit f8af4cd, republished under its MIT licence (© Ali-Marandi). 854 words, ~5,610 tokens.

Download SKILL.mdSave it as .claude/skills/image-understand/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
image-understand
description
Implement specialized image understanding capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to analyze static images, extract visual information, perform OCR, detect objects, classify images, or understand visual content. Optimized for PNG, JPEG, GIF, WebP, and BMP formats.
license
MIT

Image Understanding Skill

This skill provides specialized image understanding functionality using the z-ai-web-dev-sdk package, enabling AI models to analyze, describe, and extract information from static images.

Skills Path

Skill Location: {project_path}/skills/image-understand

this skill is located at above path in your project.

Reference Scripts: Example test scripts are available in the {Skill Location}/scripts/ directory for quick testing and reference. See {Skill Location}/scripts/image-understand.ts for a working example.

Overview

Image Understanding focuses specifically on static image analysis, providing capabilities for:

  • Image description and scene understanding
  • Object detection and recognition
  • OCR (Optical Character Recognition) and text extraction
  • Image classification and categorization
  • Visual content analysis
  • Quality assessment
  • Accessibility (alt text generation)

IMPORTANT: z-ai-web-dev-sdk MUST be used in backend code only. Never use it in client-side code.

Prerequisites

The z-ai-web-dev-sdk package is already installed. Import it as shown in the examples below.

CLI Usage (For Simple Tasks)

For quick image analysis tasks, you can use the z-ai CLI instead of writing code. This is ideal for simple image descriptions, testing, or automation.

Basic Image Analysis
bash
# Describe an image from URL
z-ai vision --prompt "What's in this image?" --image "https://example.com/photo.jpg"

# Using short options
z-ai vision -p "Describe this image" -i "https://example.com/image.png"
Analyze Local Images
bash
# Analyze a local image file
z-ai vision -p "What objects are in this photo?" -i "./photo.jpg"

# Save response to file
z-ai vision -p "Describe the scene" -i "./landscape.png" -o description.json
Multiple Images Comparison
bash
# Compare multiple images
z-ai vision \
  -p "Compare these two images and highlight the differences" \
  -i "./photo1.jpg" \
  -i "./photo2.jpg" \
  -o comparison.json

# Analyze a series of images
z-ai vision \
  --prompt "What patterns do you see across these images?" \
  --image "https://example.com/img1.jpg" \
  --image "https://example.com/img2.jpg" \
  --image "https://example.com/img3.jpg"
Advanced Analysis with Thinking
bash
# Enable chain-of-thought reasoning for complex tasks
z-ai vision \
  -p "Count all people in this image and describe what each person is doing" \
  -i "./crowd.jpg" \
  --thinking \
  -o analysis.json

# Complex object detection with reasoning
z-ai vision \
  -p "Identify all safety hazards in this workplace image" \
  -i "./workplace.jpg" \
  --thinking
Streaming Output
bash
# Stream the analysis in real-time
z-ai vision -p "Provide a detailed description" -i "./photo.jpg" --stream
CLI Parameters
  • --prompt, -p <text>: Required - Question or instruction about the image(s)
  • --image, -i <URL or path>: Optional - Image URL or local file path (can be used multiple times)
  • --thinking, -t: Optional - Enable chain-of-thought reasoning (default: disabled)
  • --output, -o <path>: Optional - Output file path (JSON format)
  • --stream: Optional - Stream the response in real-time
Supported Image Formats
  • PNG (.png) - Best for diagrams, screenshots, graphics with transparency
  • JPEG (.jpg, .jpeg) - Best for photos and complex images
  • GIF (.gif) - Supports both static and animated images
  • WebP (.webp) - Modern format with good compression
  • BMP (.bmp) - Uncompressed bitmap format
When to Use CLI vs SDK

Use CLI for:

  • Quick image analysis or descriptions
  • One-off OCR tasks
  • Testing image understanding capabilities
  • Simple batch processing scripts
  • Generating alt text for accessibility

Use SDK for:

  • Multi-turn conversations about images
  • Complex image processing pipelines
  • Production applications with error handling
  • Custom integration with your application logic
  • Batch processing with custom business logic

For better performance and reliability, use base64 encoding to pass images to the model instead of image URLs.

Basic Image Understanding Implementation

Single Image Analysis
javascript
import ZAI from 'z-ai-web-dev-sdk';

async function analyzeImage(imageUrl, prompt) {
  const zai = await ZAI.create();

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: [
          {
            type: 'text',
            text: prompt
          },
          {
            type: 'image_url',
            image_url: {
              url: imageUrl
            }
          }
        ]
      }
    ],
    thinking: { type: 'disabled' }
  });

  return response.choices[0]?.message?.content;
}

// Usage examples
const description = await analyzeImage(
  'https://example.com/landscape.jpg',
  'Describe this landscape in detail, including colors, lighting, and mood'
);

const objectDetection = await analyzeImage(
  'https://example.com/room.jpg',
  'List all objects visible in this room'
);
Multiple Images Comparison
javascript
import ZAI from 'z-ai-web-dev-sdk';

async function compareImages(imageUrls, question) {
  const zai = await ZAI.create();

  const content = [
    {
      type: 'text',
      text: question
    },
    ...imageUrls.map(url => ({
      type: 'image_url',
      image_url: { url }
    }))
  ];

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: content
      }
    ],
    thinking: { type: 'disabled' }
  });

  return response.choices[0]?.message?.content;
}

// Usage
const comparison = await compareImages(
  [
    'https://example.com/before.jpg',
    'https://example.com/after.jpg'
  ],
  'What are the key differences between these before and after images?'
);
javascript
import ZAI from 'z-ai-web-dev-sdk';
import fs from 'fs';
import path from 'path';

async function analyzeLocalImage(imagePath, prompt) {
  const zai = await ZAI.create();

  // Read image file and convert to base64
  const imageBuffer = fs.readFileSync(imagePath);
  const base64Image = imageBuffer.toString('base64');
  
  // Determine MIME type based on file extension
  const ext = path.extname(imagePath).toLowerCase();
  const mimeTypes = {
    '.png': 'image/png',
    '.jpg': 'image/jpeg',
    '.jpeg': 'image/jpeg',
    '.gif': 'image/gif',
    '.webp': 'image/webp',
    '.bmp': 'image/bmp'
  };
  const mimeType = mimeTypes[ext] || 'image/jpeg';

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: [
          {
            type: 'text',
            text: prompt
          },
          {
            type: 'image_url',
            image_url: {
              url: `data:${mimeType};base64,${base64Image}`
            }
          }
        ]
      }
    ],
    thinking: { type: 'disabled' }
  });

  return response.choices[0]?.message?.content;
}

// Usage
const result = await analyzeLocalImage(
  './product-photo.jpg',
  'Analyze this product image for e-commerce listing'
);

Advanced Use Cases

OCR and Text Extraction
javascript
import ZAI from 'z-ai-web-dev-sdk';

async function extractText(imageUrl, options = {}) {
  const zai = await ZAI.create();

  const prompt = options.preserveLayout 
    ? 'Extract all text from this image. Preserve the exact layout, formatting, and structure.'
    : 'Extract all visible text from this image.';

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: [
          { type: 'text', text: prompt },
          { type: 'image_url', image_url: { url: imageUrl } }
        ]
      }
    ],
    thinking: { type: 'disabled' }
  });

  return response.choices[0]?.message?.content;
}

// Usage examples
const receiptText = await extractText(
  'https://example.com/receipt.jpg',
  { preserveLayout: true }
);

const businessCardInfo = await extractText(
  'https://example.com/business-card.jpg'
);
Object Detection and Counting
javascript
import ZAI from 'z-ai-web-dev-sdk';

async function detectObjects(imageUrl, objectType) {
  const zai = await ZAI.create();

  const prompt = objectType 
    ? `Count and locate all ${objectType} in this image. Provide their positions and describe each one.`
    : 'Detect and list all objects in this image with their approximate locations.';

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: [
          { type: 'text', text: prompt },
          { type: 'image_url', image_url: { url: imageUrl } }
        ]
      }
    ],
    thinking: { type: 'enabled' } // Enable thinking for complex counting
  });

  return response.choices[0]?.message?.content;
}

// Usage
const peopleCount = await detectObjects(
  'https://example.com/crowd.jpg',
  'people'
);

const allObjects = await detectObjects(
  'https://example.com/room.jpg'
);
Image Classification and Tagging
javascript
import ZAI from 'z-ai-web-dev-sdk';

async function classifyAndTag(imageUrl) {
  const zai = await ZAI.create();

  const prompt = `Analyze this image and provide a comprehensive classification:
1. Primary category (e.g., nature, urban, portrait, product)
2. Subject matter (main focus of the image)
3. Style or mood (e.g., professional, casual, artistic, vintage)
4. Color palette description
5. Suggested tags (10-15 keywords, comma-separated)

Format your response as structured JSON.`;

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: [
          { type: 'text', text: prompt },
          { type: 'image_url', image_url: { url: imageUrl } }
        ]
      }
    ],
    thinking: { type: 'disabled' }
  });

  const content = response.choices[0]?.message?.content;
  
  try {
    return JSON.parse(content);
  } catch (e) {
    return { rawResponse: content };
  }
}

// Usage
const classification = await classifyAndTag(
  'https://example.com/photo.jpg'
);
console.log('Tags:', classification.tags);
Quality Assessment
javascript
import ZAI from 'z-ai-web-dev-sdk';

async function assessImageQuality(imageUrl) {
  const zai = await ZAI.create();

  const prompt = `Assess the technical quality of this image:
1. Sharpness and focus (1-10)
2. Exposure and brightness (1-10)
3. Color balance (1-10)
4. Composition (1-10)
5. Any technical issues (blur, noise, artifacts, etc.)
6. Overall quality rating (1-10)
7. Suggestions for improvement

Provide specific feedback for each criterion.`;

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: [
          { type: 'text', text: prompt },
          { type: 'image_url', image_url: { url: imageUrl } }
        ]
      }
    ],
    thinking: { type: 'disabled' }
  });

  return response.choices[0]?.message?.content;
}
Accessibility - Alt Text Generation
javascript
import ZAI from 'z-ai-web-dev-sdk';

async function generateAltText(imageUrl, context = '') {
  const zai = await ZAI.create();

  const prompt = context
    ? `Generate concise, descriptive alt text for this image. Context: ${context}. Focus on the most important visual elements that convey the image's purpose.`
    : 'Generate concise, descriptive alt text for this image suitable for screen readers. Focus on key visual elements.';

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: [
          { type: 'text', text: prompt },
          { type: 'image_url', image_url: { url: imageUrl } }
        ]
      }
    ],
    thinking: { type: 'disabled' }
  });

  return response.choices[0]?.message?.content;
}

// Usage
const altText = await generateAltText(
  'https://example.com/hero-image.jpg',
  'Website hero section for a tech startup'
);
Scene Understanding
javascript
import ZAI from 'z-ai-web-dev-sdk';

async function understandScene(imageUrl) {
  const zai = await ZAI.create();

  const prompt = `Provide a comprehensive scene analysis:
1. Setting/location type (indoor/outdoor, specific place)
2. Time of day and lighting conditions
3. Weather (if applicable)
4. People present (number, activities, interactions)
5. Key objects and their arrangement
6. Overall atmosphere and mood
7. Notable details or interesting elements`;

  const response = await zai.chat.completions.createVision({
    messages: [
      {
        role: 'user',
        content: [
          { type: 'text', text: prompt },
          { type: 'image_url', image_url: { url: imageUrl } }
        ]
      }
    ],
    thinking: { type: 'disabled' }
  });

  return response.choices[0]?.message?.content;
}

Batch Processing

Process Multiple Images
javascript
import ZAI from 'z-ai-web-dev-sdk';

class ImageBatchProcessor {
  constructor() {
    this.zai = null;
  }

  async initialize() {
    this.zai = await ZAI.create();
  }

  async processImage(imageUrl, prompt) {
    const response = await this.zai.chat.completions.createVision({
      messages: [
        {
          role: 'user',
          content: [
            { type: 'text', text: prompt },
            { type: 'image_url', image_url: { url: imageUrl } }
          ]
        }
      ],
      thinking: { type: 'disabled' }
    });

    return response.choices[0]?.message?.content;
  }

  async processBatch(imageUrls, prompt) {
    const results = [];
    
    for (const imageUrl of imageUrls) {
      try {
        const result = await this.processImage(imageUrl, prompt);
        results.push({ imageUrl, success: true, result });
      } catch (error) {
        results.push({ 
          imageUrl, 
          success: false, 
          error: error.message 
        });
      }
    }

    return results;
  }
}

// Usage
const processor = new ImageBatchProcessor();
await processor.initialize();

const images = [
  'https://example.com/img1.jpg',
  'https://example.com/img2.jpg',
  'https://example.com/img3.jpg'
];

const results = await processor.processBatch(
  images,
  'Generate a short description suitable for social media'
);

Best Practices

1. Image Quality and Preparation
  • Use high-resolution images for better analysis accuracy
  • Ensure images are well-lit and properly exposed
  • For OCR, ensure text is clear and readable
  • Optimize file size to balance quality and performance
  • Supported formats: PNG (best for text/diagrams), JPEG (best for photos), WebP, GIF, BMP
2. Prompt Engineering for Images
  • Be specific about what information you need
  • Mention the type of image (photo, diagram, screenshot, etc.)
  • For complex tasks, break down into specific questions
  • Use structured prompts for JSON output
  • Include context when relevant
3. Error Handling
javascript
async function safeImageAnalysis(imageUrl, prompt) {
  try {
    const zai = await ZAI.create();
    
    const response = await zai.chat.completions.createVision({
      messages: [
        {
          role: 'user',
          content: [
            { type: 'text', text: prompt },
            { type: 'image_url', image_url: { url: imageUrl } }
          ]
        }
      ],
      thinking: { type: 'disabled' }
    });

    return {
      success: true,
      content: response.choices[0]?.message?.content
    };
  } catch (error) {
    console.error('Image analysis error:', error);
    return {
      success: false,
      error: error.message
    };
  }
}
4. Performance Optimization
  • Cache SDK instance for batch processing
  • Use base64 encoding for local images
  • Implement request throttling for large batches
  • Consider image preprocessing (resize, compress) for large files
  • Use appropriate thinking mode (disabled for simple tasks, enabled for complex reasoning)
Show full SKILL.md (326 more words)Show less
5. Security Considerations
  • Validate image URLs before processing
  • Implement rate limiting for public APIs
  • Sanitize user-provided image data
  • Never expose SDK credentials in client-side code
  • Implement content moderation for user-uploaded images

Common Use Cases

  1. E-commerce Product Analysis: Analyze product images, extract features, generate descriptions
  2. Document Processing: Extract text from receipts, invoices, forms, business cards
  3. Content Moderation: Detect inappropriate content, verify image compliance
  4. Quality Control: Identify defects, assess product quality in manufacturing
  5. Accessibility: Generate alt text for images automatically
  6. Image Cataloging: Auto-tag and categorize image libraries
  7. Visual Search: Understand and index images for search functionality
  8. Medical Imaging: Preliminary analysis with appropriate disclaimers
  9. Real Estate: Analyze property photos, extract features
  10. Social Media: Generate captions, hashtags, and descriptions

Integration Examples

Express.js API Endpoint
javascript
import express from 'express';
import ZAI from 'z-ai-web-dev-sdk';
import multer from 'multer';

const app = express();
const upload = multer({ storage: multer.memoryStorage() });

let zaiInstance;

async function initZAI() {
  zaiInstance = await ZAI.create();
}

// Analyze image from URL
app.post('/api/analyze-image', express.json(), async (req, res) => {
  try {
    const { imageUrl, prompt } = req.body;

    if (!imageUrl || !prompt) {
      return res.status(400).json({ 
        error: 'imageUrl and prompt are required' 
      });
    }

    const response = await zaiInstance.chat.completions.createVision({
      messages: [
        {
          role: 'user',
          content: [
            { type: 'text', text: prompt },
            { type: 'image_url', image_url: { url: imageUrl } }
          ]
        }
      ],
      thinking: { type: 'disabled' }
    });

    res.json({
      success: true,
      analysis: response.choices[0]?.message?.content
    });
  } catch (error) {
    res.status(500).json({
      success: false,
      error: error.message
    });
  }
});

// Analyze uploaded image file
app.post('/api/analyze-upload', upload.single('image'), async (req, res) => {
  try {
    const { prompt } = req.body;
    const imageFile = req.file;

    if (!imageFile || !prompt) {
      return res.status(400).json({ 
        error: 'image file and prompt are required' 
      });
    }

    // Convert to base64
    const base64Image = imageFile.buffer.toString('base64');
    const mimeType = imageFile.mimetype;

    const response = await zaiInstance.chat.completions.createVision({
      messages: [
        {
          role: 'user',
          content: [
            { type: 'text', text: prompt },
            { 
              type: 'image_url', 
              image_url: { 
                url: `data:${mimeType};base64,${base64Image}` 
              } 
            }
          ]
        }
      ],
      thinking: { type: 'disabled' }
    });

    res.json({
      success: true,
      analysis: response.choices[0]?.message?.content
    });
  } catch (error) {
    res.status(500).json({
      success: false,
      error: error.message
    });
  }
});

initZAI().then(() => {
  app.listen(3000, () => {
    console.log('Image understanding API running on port 3000');
  });
});
Next.js API Route
javascript
// pages/api/image-understand.js
import ZAI from 'z-ai-web-dev-sdk';

let zaiInstance = null;

async function getZAI() {
  if (!zaiInstance) {
    zaiInstance = await ZAI.create();
  }
  return zaiInstance;
}

export default async function handler(req, res) {
  if (req.method !== 'POST') {
    return res.status(405).json({ error: 'Method not allowed' });
  }

  try {
    const { imageUrl, prompt } = req.body;

    if (!imageUrl || !prompt) {
      return res.status(400).json({ 
        error: 'imageUrl and prompt are required' 
      });
    }

    const zai = await getZAI();

    const response = await zai.chat.completions.createVision({
      messages: [
        {
          role: 'user',
          content: [
            { type: 'text', text: prompt },
            { type: 'image_url', image_url: { url: imageUrl } }
          ]
        }
      ],
      thinking: { type: 'disabled' }
    });

    res.status(200).json({
      success: true,
      analysis: response.choices[0]?.message?.content
    });
  } catch (error) {
    console.error('Error:', error);
    res.status(500).json({
      success: false,
      error: error.message
    });
  }
}

Troubleshooting

Issue: "SDK must be used in backend"

  • Solution: Ensure z-ai-web-dev-sdk is only imported and used in server-side code, never in client/browser code

Issue: Image not loading or being analyzed

  • Solution: Verify the image URL is accessible, returns correct MIME type, and is in a supported format

Issue: Poor OCR accuracy

  • Solution: Ensure text is clear and readable, increase image resolution, ensure proper lighting and contrast

Issue: Inaccurate object detection or counting

  • Solution: Enable thinking mode for complex counting tasks, use high-resolution images, provide specific prompts

Issue: Slow response times

  • Solution: Optimize image size (resize before upload), use base64 for local images, cache SDK instance for batch processing

Issue: Base64 encoding fails

  • Solution: Verify file path is correct, check file permissions, ensure MIME type matches file extension

Remember

  • Always use z-ai-web-dev-sdk in backend code only
  • The SDK is already installed - import as shown in examples
  • Use image_url content type for static images
  • Base64 encoding is recommended for better performance
  • Structure prompts clearly for best results
  • Enable thinking mode for complex reasoning tasks (counting, detailed analysis)
  • Handle errors gracefully in production
  • Validate and sanitize user inputs
  • Consider privacy and security when processing user images

© Ali-Marandi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/image-understand of Ali-Marandi/Web-Scraper-Framework.

  • SKILL.md
  • LICENSE.txt
  • scripts/image-understand.ts

Open the folder on GitHubat commit f8af4cd

Compare with similar skills

Image Understand next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image Understand compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image Understand this skillAli-Marandi/Web-Scraper-Framework107—~5.6kAutomated safety check: PassMIT
Higress Openclaw Integrationhigress-group/higress9.5k—~2.5kAutomated safety check: PassApache-2.0
Image Generationjjyaoao/HelloAgents3.2k1 repos~3.8kAutomated safety check: PassMIT
Weave Router Local Testingweave-os/router5.6k—~3.1kAutomated safety check: NotesApache-2.0
Video Understandjjyaoao/HelloAgents3.2k1 repos~6.2kAutomated safety check: PassMIT
Web Readerjjyaoao/HelloAgents3.2k1 repos~7.1kAutomated safety check: PassMIT

Similar skills

  • Higress Openclaw Integration

    higress-group/higress

    Deploy and configure Higress AI Gateway for OpenClaw integration.

    9.5k GitHub stars~2.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Image Generation

    jjyaoao/HelloAgents

    Implement AI image generation capabilities using the z-ai-web-dev-sdk.

    3.2k GitHub starsUsed in 1 repo~3.8k tokens
    Media & CreativeAuto-check passed
  • Stands up the Weave model router in Docker Compose and drives it with claude -p against a real or mocked upstream to reproduce and verify routing and streaming behavior.

    5.6k GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Video Understand

    jjyaoao/HelloAgents

    Implement specialized video understanding capabilities using the z-ai-web-dev-sdk.

    3.2k GitHub starsUsed in 1 repo~6.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Web Reader

    jjyaoao/HelloAgents

    Implement web page content extraction capabilities using the z-ai-web-dev-sdk.

    3.2k GitHub starsUsed in 1 repo~7.1k tokens
    Data & AnalyticsAuto-check passed
  • AI Model Nodejs

    TencentCloudBase/CloudBase-AI-Toolkit

    A skill your agent uses for Node.js backend AI via @cloudbase/node-sdk (=3.16.0) — cloud functions, CloudRun, Express/Koa/NestJS, serverless APIs, scheduled jobs, LLM proxies, agent orchestration.

    1.1k GitHub starsUsed in 2 repos~5k tokens
    Backend & APIsAuto-check passed

More from Ali-Marandi/Web-Scraper-Framework

All 12 skills in this repo
  • Aminer Deep Search

    Ali-Marandi/Web-Scraper-Framework

    Activate this skill when the user wants deep, multi-round academic paper collection for a survey or literature review.

    107 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Interview Prep

    Ali-Marandi/Web-Scraper-Framework

    帮用户准备面试。基于目标 JD、公司、岗位方向,生成"高频面试题 + 参考回答 + 行为面 / 技术面 / Case 面分类题库",并产出可打印的『面试备战手册』。当用户说"帮我准备面试""明天有面试 / 后天面试""面试题""面经""模拟面试""我要面 X 公司 Y 岗位""帮我准备 STAR 故事""怎么回答这道面试题""自我介绍 / 离职原因 / 优缺点 怎么答",必须触发本…

    107 GitHub stars~735 tokensUpdated 1 mo ago
    Auto-check passed
  • Jd Resume Tailor

    Ali-Marandi/Web-Scraper-Framework

    给定一份 JD 和一份现有简历,做"JD 拆解 + 简历定向改写"。拆 JD 抽出硬技能、软技能、加分项;对照简历做 gap 分析;产出针对该岗位重写后的简历,突出相关经验、补齐关键词缺口、并保留候选人真实经历不编造。当用户说"针对这个岗位 / 这家公司改简历""帮我对一下这个 JD""我想投这个职位你看怎么改""把这份简历针对 X 公司优化""做一份定向版简历",或同时给出 JD 文本 +…

    107 GitHub stars~714 tokensUpdated 1 mo ago
    Auto-check passed
  • Job Intent Tracker

    Ali-Marandi/Web-Scraper-Framework

    帮助用户梳理求职意向、生成目标岗位画像,并维护一份结构化的"岗位投递追踪表"。当用户说"我想换工作 / 不知道投什么岗 / 帮我看看我适合什么岗位 / 帮我管理投递进度 / 我投了好几家但记不住状态了 / 想做一个求职 OKR / 整理一下求职方向",或上传简历但没说要改简历时,应该主动触发本 skill。本 skill 也适用于实习生、应届生、转行候选人在求职启动阶段做"自我盘点 +…

    107 GitHub stars~819 tokensUpdated 1 mo ago
    Auto-check passed
  • Resume Builder

    Ali-Marandi/Web-Scraper-Framework

    从零生成或全面优化一份中文简历,并导出 docx / pdf / markdown 多种格式。用 STAR 法则改写经历、做 ATS 关键词覆盖率检查、根据行业(互联网产品 / 技术 / 金融 / 通用)选模板。当用户说"帮我写简历 / 优化简历 / 简历不会写 / 我的简历太弱了 / 简历看起来不专业 / 简历改一改 / 给我做个简历模板 / 简历导出 / 简历加点关键词",或者上传…

    107 GitHub stars~740 tokensUpdated 1 mo ago
    Auto-check passed
  • Dream Interpreter

    Ali-Marandi/Web-Scraper-Framework

    AI 解梦大师。用户描述梦境,智能追问关键细节后,从三个视角(周公解梦/心理分析/赛博神棍)生成解读,输出结构化 JSON 供前端渲染"梦境解析卡"。

    107 GitHub stars~406 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Image Understand

What does Image Understand do?

Implement specialized image understanding capabilities using the z-ai-web-dev-sdk. Image Understand is an agent skill from Ali-Marandi/Web-Scraper-Framework. Implement specialized image understanding capabilities using the z-ai-web-dev-sdk.

When should I use Image Understand?

Image Understand fits situations like: the user needs to analyze static images; extract visual information; classify images; understand visual content.

How do I install Image Understand in Claude Code?

Run `npx skills add Ali-Marandi/Web-Scraper-Framework --skill image-understand -a claude-code`. Or copy the skill folder (skills/image-understand in Ali-Marandi/Web-Scraper-Framework) into .claude/skills/image-understand in your project. Claude Code loads it when a task matches its description.

How do I install Image Understand in Codex?

Run `npx skills add Ali-Marandi/Web-Scraper-Framework --skill image-understand -a codex`. Or copy the skill folder (skills/image-understand in Ali-Marandi/Web-Scraper-Framework) into .agents/skills/image-understand in your project. Codex loads it when a task matches its description.

Can I use Image Understand in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Ali-Marandi/Web-Scraper-Framework --skill image-understand -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-understand, .gemini/skills/image-understand, .github/skills/image-understand and .opencode/skills/image-understand in your project.

What does Image Understand need to run?

Going by SKILL.md and its folder, Image Understand needs TypeScript for the scripts in its folder. Our summary lists: Node.js.

Does Image Understand access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image Understand safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Image Understand use?

Image Understand is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image Understand use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image Understand?

Skills that share tags, products or a category with Image Understand: Higress Openclaw Integration (higress-group/higress, 9.5k stars), Image Generation (jjyaoao/HelloAgents, 3.2k stars), Weave Router Local Testing (weave-os/router, 5.6k stars) and Video Understand (jjyaoao/HelloAgents, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image Understand?

Ali-Marandi (a GitHub user) maintains it in Ali-Marandi/Web-Scraper-Framework, which has 107 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on August 31, 2026.

Source: Ali-Marandi/Web-Scraper-Framework on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.