Official agent skill

Azure AI Vision Imageanalysis Java

by microsoft in microsoft/skills

Build image analysis applications with Azure AI Vision SDK for Java.

OfficialMITAuto-check passedAI & LLM Engineering

Install Azure AI Vision Imageanalysis Java

skills CLI
$ npx skills add microsoft/skills --skill azure-ai-vision-imageanalysis-java -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/skills azure-ai-vision-imageanalysis-java --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/plugins/azure-sdk-java/skills/azure-ai-vision-imageanalysis-java .claude/skills/azure-ai-vision-imageanalysis-java && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
azure-ai-vision-imageanalysis-java
GitHub stars
3.1k
Used in
6 other repos
Token cost
~2.2k tokens
SKILL.md length
161 words
Files
2 (incl. references)
Skills in repo
150
Repo updated
First seen
Licence
MIT

At a glance

Build image analysis applications with Azure AI Vision SDK for Java.

  • Implementing image captioning
  • SKILL.md covers Installation, Client Creation, Visual Features and Core Patterns, plus 5 more sections
  • Reaches learn.microsoft.com; needs VISION_KEY and AZURE_TOKEN_CREDENTIALS
  • OCR text extraction

What it does

Azure AI Vision Imageanalysis Java is an agent skill from microsoft/skills, published by the product's own GitHub organization. Build image analysis applications with Azure AI Vision SDK for Java. Use when implementing image captioning, OCR text extraction, object detection, tagging, or smart cropping.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/examples.md`).

It sits in AI & LLM Engineering, covering Computer vision. It works with Azure AI Vision and Java. The repository describes itself as: Skills, MCP servers, Custom Agents, Agents.md for SDKs to ground Coding Agents. The licence is MIT.

When your agent uses it

  • Implementing image captioning
  • OCR text extraction
  • Object detection

Example prompts

  • “/azure-ai-vision-imageanalysis-java”

Requirements

  • A credential in VISION_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 354361d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are java, xml and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • learn.microsoft.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VISION_KEY
    • AZURE_TOKEN_CREDENTIALS

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Azure AI Vision Imageanalysis Java loads about 2.2k tokens when it runs, and up to ~7.7k if it reads all its reference files. Until then it costs about 53 tokens; SKILL.md has 161 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/skills at commit 354361d, republished under its MIT licence (© microsoft). 161 words, ~2,183 tokens.

Download SKILL.mdSave it as .claude/skills/azure-ai-vision-imageanalysis-java/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
azure-ai-vision-imageanalysis-java
description
Build image analysis applications with Azure AI Vision SDK for Java. Use when implementing image captioning, OCR text extraction, object detection, tagging, or smart cropping.
license
MIT
metadata.author
Microsoft
metadata.version
1.0.0
metadata.package
com.azure:azure-ai-vision-imageanalysis

Azure AI Vision Image Analysis SDK for Java

Build image analysis applications using the Azure AI Vision Image Analysis SDK for Java.

Installation

xml
<dependency>
    <groupId>com.azure</groupId>
    <artifactId>azure-ai-vision-imageanalysis</artifactId>
    <version>1.1.0-beta.1</version>
</dependency>

Client Creation

With API Key
java
import com.azure.ai.vision.imageanalysis.ImageAnalysisClient;
import com.azure.ai.vision.imageanalysis.ImageAnalysisClientBuilder;
import com.azure.core.credential.KeyCredential;

String endpoint = System.getenv("VISION_ENDPOINT");
String key = System.getenv("VISION_KEY");

ImageAnalysisClient client = new ImageAnalysisClientBuilder()
    .endpoint(endpoint)
    .credential(new KeyCredential(key))
    .buildClient();
Async Client
java
import com.azure.ai.vision.imageanalysis.ImageAnalysisAsyncClient;

ImageAnalysisAsyncClient asyncClient = new ImageAnalysisClientBuilder()
    .endpoint(endpoint)
    .credential(new KeyCredential(key))
    .buildAsyncClient();
With DefaultAzureCredential
java
import com.azure.core.credential.TokenCredential;
import com.azure.identity.AzureIdentityEnvVars;
import com.azure.identity.DefaultAzureCredentialBuilder;
import com.azure.identity.ManagedIdentityCredentialBuilder;

TokenCredential credential = new DefaultAzureCredentialBuilder()
    .requireEnvVars(AzureIdentityEnvVars.AZURE_TOKEN_CREDENTIALS)
    .build();
// Or use a specific credential directly in production:
// See https://learn.microsoft.com/java/api/overview/azure/identity-readme?view=azure-java-stable#credential-classes
// TokenCredential credential = new ManagedIdentityCredentialBuilder().build();

ImageAnalysisClient client = new ImageAnalysisClientBuilder()
    .endpoint(endpoint)
    .credential(credential)
    .buildClient();

Visual Features

FeatureDescription
CAPTIONGenerate human-readable image description
DENSE_CAPTIONSCaptions for up to 10 regions
READOCR - Extract text from images
TAGSContent tags for objects, scenes, actions
OBJECTSDetect objects with bounding boxes
SMART_CROPSSmart thumbnail regions
PEOPLEDetect people with locations

Core Patterns

Generate Caption
java
import com.azure.ai.vision.imageanalysis.models.*;
import com.azure.core.util.BinaryData;
import java.io.File;
import java.util.Arrays;

// From file
BinaryData imageData = BinaryData.fromFile(new File("image.jpg").toPath());

ImageAnalysisResult result = client.analyze(
    imageData,
    Arrays.asList(VisualFeatures.CAPTION),
    new ImageAnalysisOptions().setGenderNeutralCaption(true));

System.out.printf("Caption: \"%s\" (confidence: %.4f)%n",
    result.getCaption().getText(),
    result.getCaption().getConfidence());
Generate Caption from URL
java
ImageAnalysisResult result = client.analyzeFromUrl(
    "https://example.com/image.jpg",
    Arrays.asList(VisualFeatures.CAPTION),
    new ImageAnalysisOptions().setGenderNeutralCaption(true));

System.out.printf("Caption: \"%s\"%n", result.getCaption().getText());
Extract Text (OCR)
java
ImageAnalysisResult result = client.analyze(
    BinaryData.fromFile(new File("document.jpg").toPath()),
    Arrays.asList(VisualFeatures.READ),
    null);

for (DetectedTextBlock block : result.getRead().getBlocks()) {
    for (DetectedTextLine line : block.getLines()) {
        System.out.printf("Line: '%s'%n", line.getText());
        System.out.printf("  Bounding polygon: %s%n", line.getBoundingPolygon());
        
        for (DetectedTextWord word : line.getWords()) {
            System.out.printf("  Word: '%s' (confidence: %.4f)%n",
                word.getText(),
                word.getConfidence());
        }
    }
}
Detect Objects
java
ImageAnalysisResult result = client.analyzeFromUrl(
    imageUrl,
    Arrays.asList(VisualFeatures.OBJECTS),
    null);

for (DetectedObject obj : result.getObjects()) {
    System.out.printf("Object: %s (confidence: %.4f)%n",
        obj.getTags().get(0).getName(),
        obj.getTags().get(0).getConfidence());
    
    ImageBoundingBox box = obj.getBoundingBox();
    System.out.printf("  Location: x=%d, y=%d, w=%d, h=%d%n",
        box.getX(), box.getY(), box.getWidth(), box.getHeight());
}
Get Tags
java
ImageAnalysisResult result = client.analyzeFromUrl(
    imageUrl,
    Arrays.asList(VisualFeatures.TAGS),
    null);

for (DetectedTag tag : result.getTags()) {
    System.out.printf("Tag: %s (confidence: %.4f)%n",
        tag.getName(),
        tag.getConfidence());
}
Detect People
java
ImageAnalysisResult result = client.analyzeFromUrl(
    imageUrl,
    Arrays.asList(VisualFeatures.PEOPLE),
    null);

for (DetectedPerson person : result.getPeople()) {
    ImageBoundingBox box = person.getBoundingBox();
    System.out.printf("Person at x=%d, y=%d (confidence: %.4f)%n",
        box.getX(), box.getY(), person.getConfidence());
}
Smart Cropping
java
ImageAnalysisResult result = client.analyzeFromUrl(
    imageUrl,
    Arrays.asList(VisualFeatures.SMART_CROPS),
    new ImageAnalysisOptions().setSmartCropsAspectRatios(Arrays.asList(1.0, 1.5)));

for (CropRegion crop : result.getSmartCrops()) {
    System.out.printf("Crop region: aspect=%.2f, x=%d, y=%d, w=%d, h=%d%n",
        crop.getAspectRatio(),
        crop.getBoundingBox().getX(),
        crop.getBoundingBox().getY(),
        crop.getBoundingBox().getWidth(),
        crop.getBoundingBox().getHeight());
}
Dense Captions
java
ImageAnalysisResult result = client.analyzeFromUrl(
    imageUrl,
    Arrays.asList(VisualFeatures.DENSE_CAPTIONS),
    new ImageAnalysisOptions().setGenderNeutralCaption(true));

for (DenseCaption caption : result.getDenseCaptions()) {
    System.out.printf("Caption: \"%s\" (confidence: %.4f)%n",
        caption.getText(),
        caption.getConfidence());
    System.out.printf("  Region: x=%d, y=%d, w=%d, h=%d%n",
        caption.getBoundingBox().getX(),
        caption.getBoundingBox().getY(),
        caption.getBoundingBox().getWidth(),
        caption.getBoundingBox().getHeight());
}
Multiple Features
java
ImageAnalysisResult result = client.analyzeFromUrl(
    imageUrl,
    Arrays.asList(
        VisualFeatures.CAPTION,
        VisualFeatures.TAGS,
        VisualFeatures.OBJECTS,
        VisualFeatures.READ),
    new ImageAnalysisOptions()
        .setGenderNeutralCaption(true)
        .setLanguage("en"));

// Access all results
System.out.println("Caption: " + result.getCaption().getText());
System.out.println("Tags: " + result.getTags().size());
System.out.println("Objects: " + result.getObjects().size());
System.out.println("Text blocks: " + result.getRead().getBlocks().size());
Async Analysis
java
asyncClient.analyzeFromUrl(
    imageUrl,
    Arrays.asList(VisualFeatures.CAPTION),
    null)
    .subscribe(
        result -> System.out.println("Caption: " + result.getCaption().getText()),
        error -> System.err.println("Error: " + error.getMessage()),
        () -> System.out.println("Complete")
    );

Error Handling

java
import com.azure.core.exception.HttpResponseException;

try {
    client.analyzeFromUrl(imageUrl, Arrays.asList(VisualFeatures.CAPTION), null);
} catch (HttpResponseException e) {
    System.out.println("Status: " + e.getResponse().getStatusCode());
    System.out.println("Error: " + e.getMessage());
}

Environment Variables

bash
VISION_ENDPOINT=https://<resource>.cognitiveservices.azure.com/ # Required for all auth methods
VISION_KEY=<your-api-key> # Only required for AzureKeyCredential auth
AZURE_TOKEN_CREDENTIALS=prod  # Required only if DefaultAzureCredential is used in production

Image Requirements

  • Formats: JPEG, PNG, GIF, BMP, WEBP, ICO, TIFF, MPO
  • Size: < 20 MB
  • Dimensions: 50x50 to 16000x16000 pixels

Regional Availability

Caption and Dense Captions require GPU-supported regions. Check supported regions before deployment.

Trigger Phrases

  • "image analysis Java"
  • "Azure Vision SDK"
  • "image captioning"
  • "OCR image text extraction"
  • "object detection image"
  • "smart crop thumbnail"
  • "detect people image"

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .github/plugins/azure-sdk-java/skills/azure-ai-vision-imageanalysis-java of microsoft/skills.

  • SKILL.md
  • references/examples.md

Open the folder on GitHubat commit 354361d

Used in 6 other repositories

We found 17 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 6 other GitHub owners. This page covers the copy in microsoft/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Azure AI Vision Imageanalysis Java next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azure AI Vision Imageanalysis Java compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azure AI Vision Imageanalysis Java this skillmicrosoft/skills3.1k6 repos~2.2kAutomated safety check: PassMIT
Azure AI Vision ReferenceMicrosoftDocs/Agent-Skills775—~1.6kAutomated safety check: PassCC-BY-4.0
Azure Custom VisionMicrosoftDocs/Agent-Skills775—~1.6kAutomated safety check: PassCC-BY-4.0
ModLens Image Vision Bridgeliustack/modlens4.1k—~1.3kAutomated safety check: NotesMIT
Claude APIKocoro-lab/Kocoro4148 repos~4.5kAutomated safety check: PassApache-2.0
Agent AI Codingw8123/EnterpriseAgentFramework8301 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Azure AI Vision Reference

    MicrosoftDocs/Agent-Skills

    Official

    Looks up Microsoft Learn guidance for Azure AI Vision: Image Analysis, Read OCR containers, smart-crop thumbnails, background removal and video frame analysis, plus limits and deployment.

    775 GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Azure Custom Vision

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure AI Custom Vision development including best practices, decision making, limits & quotas, security, integrations & coding patterns, and deployment.

    775 GitHub stars~1.6k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.1k GitHub stars~1.3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes
  • Claude API

    Kocoro-lab/Kocoro

    Build apps with the Claude API or Anthropic SDK. An agent skill from Kocoro-lab/Kocoro.

    414 GitHub starsUsed in 8 repos~4.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent AI Coding

    w8123/EnterpriseAgentFramework

    Create, inspect, and safely update project-scoped ReachAI Agents; edit and publish Supervisor config drafts; discover published bindable Skills; and attach or detach exact Skill versions through the…

    830 GitHub starsUsed in 1 repo~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Workflow AI Coding

    w8123/EnterpriseAgentFramework

    Edit, validate, debug, publish, and inspect ReachAI Workflow drafts through the Workflow AI Coding REST API.

    830 GitHub stars~3.7k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from microsoft/skills

All 150 skills in this repo
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    Auto-check passed
  • Official

    Python guidance for the Azure AI Search SDK covering vector, hybrid and semantic search, index management and indexers, with Entra ID authentication preferred over keys.

    3.1k GitHub starsUsed in 6 repos~4.4k tokens
    Auto-check passed
  • Official

    Covers producer, consumer, and checkpoint-store setup for Azure Event Hubs streaming in Python, with Entra ID auth and partition targeting.

    3.1k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Pydantic Models Py

    microsoft/skills

    Official

    Create Pydantic models following the multi-model pattern with Base, Create, Update, Response, and InDB variants.

    3.1k GitHub starsUsed in 6 repos~496 tokens
    Auto-check passed
  • Official

    Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

    3.1k GitHub starsUsed in 1 repo~947 tokens
    Auto-check passed
  • Frontend UI Dark TS

    microsoft/skills

    Official

    Build dark-themed React applications using Tailwind CSS with custom theming, glassmorphism effects, and Framer Motion animations.

    3.1k GitHub starsUsed in 5 repos~3.6k tokens
    Auto-check passed

Questions about Azure AI Vision Imageanalysis Java

What does Azure AI Vision Imageanalysis Java do?

Build image analysis applications with Azure AI Vision SDK for Java. Azure AI Vision Imageanalysis Java is an agent skill from microsoft/skills, published by the product's own GitHub organization. Build image analysis applications with Azure AI Vision SDK for Java.

When should I use Azure AI Vision Imageanalysis Java?

Azure AI Vision Imageanalysis Java fits situations like: implementing image captioning; OCR text extraction; object detection.

How do I install Azure AI Vision Imageanalysis Java in Claude Code?

Run `npx skills add microsoft/skills --skill azure-ai-vision-imageanalysis-java -a claude-code`. Or copy the skill folder (.github/plugins/azure-sdk-java/skills/azure-ai-vision-imageanalysis-java in microsoft/skills) into .claude/skills/azure-ai-vision-imageanalysis-java in your project. Claude Code loads it when a task matches its description.

How do I install Azure AI Vision Imageanalysis Java in Codex?

Run `npx skills add microsoft/skills --skill azure-ai-vision-imageanalysis-java -a codex`. Or copy the skill folder (.github/plugins/azure-sdk-java/skills/azure-ai-vision-imageanalysis-java in microsoft/skills) into .agents/skills/azure-ai-vision-imageanalysis-java in your project. Codex loads it when a task matches its description.

Can I use Azure AI Vision Imageanalysis Java in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/skills --skill azure-ai-vision-imageanalysis-java -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azure-ai-vision-imageanalysis-java, .gemini/skills/azure-ai-vision-imageanalysis-java, .github/skills/azure-ai-vision-imageanalysis-java and .opencode/skills/azure-ai-vision-imageanalysis-java in your project.

What does Azure AI Vision Imageanalysis Java need to run?

Going by SKILL.md and its folder, Azure AI Vision Imageanalysis Java needs credentials named VISION_KEY and AZURE_TOKEN_CREDENTIALS. Our summary lists: A credential in VISION_KEY.

Does Azure AI Vision Imageanalysis Java access the network?

SKILL.md names 1 domain. In commands or code: learn.microsoft.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Azure AI Vision Imageanalysis Java safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Azure AI Vision Imageanalysis Java use?

Azure AI Vision Imageanalysis Java is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azure AI Vision Imageanalysis Java use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.

What are the alternatives to Azure AI Vision Imageanalysis Java?

Skills that share tags, products or a category with Azure AI Vision Imageanalysis Java: Azure AI Vision Reference (MicrosoftDocs/Agent-Skills, 775 stars), Azure Custom Vision (MicrosoftDocs/Agent-Skills, 775 stars), ModLens Image Vision Bridge (liustack/modlens, 4.1k stars) and Claude API (Kocoro-lab/Kocoro, 414 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azure AI Vision Imageanalysis Java?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/skills, which has 3,086 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 6, 2026.

Source: microsoft/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.