Agent skill

Multimodal Embedding Serving User

by open-edge-platform in open-edge-platform/edge-ai-libraries

Deploy and consume the Multimodal Embedding Serving microservice — bring it up with setup.sh + docker compose (from a repo clone, or by fetching those same files from GitHub when no clone exists)…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Multimodal Embedding Serving User

skills CLI
$ npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-embedding-serving-user -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-edge-platform/edge-ai-libraries multimodal-embedding-serving-user --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-edge-platform/edge-ai-libraries.git skills-src && mkdir -p .claude/skills && cp -r skills-src/microservices/multimodal-embedding-serving/.github/skills/multimodal-embedding-serving-user .claude/skills/multimodal-embedding-serving-user && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multimodal-embedding-serving-user
GitHub stars
168
Token cost
~1.7k tokens
SKILL.md length
556 words
Files
6
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploy and consume the Multimodal Embedding Serving microservice — bring it up with setup.sh + docker compose (from a repo clone, or by fetching those same files from GitHub when no clone exists)…

  • Works in 5 steps: Context routing — repo clone or… → Bring-up (identical in both contexts) → Capability check (mandatory before… → …
  • An app needs embeddings for similarity search
  • SKILL.md covers When to Use, Example Prompts, Docs & deploy files — with or… and 1. Context routing — repo…, plus 5 more sections
  • Calls curl, docker and bash

What it does

Multimodal Embedding Serving User is an agent skill from open-edge-platform/edge-ai-libraries. Deploy and consume the Multimodal Embedding Serving microservice — bring it up with setup.sh + docker compose (from a repo clone, or by fetching those same files from GitHub when no clone exists) using the prebuilt intel/multimodal-embedding-serving image, embed text/images/videos over REST on port 9777, choose among 19 models (CLIP/SigLIP/MobileCLIP/CN-CLIP/Blip2/ QwenText), or integrate in-process via the Python SDK wheel. Use when an app needs embeddings for similarity search or retrieval. Not for modifying…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `benchmark/benchmark.md`, `evals/evals.json` and `example-prompts/image-duplicate-detector.md`).

It sits in AI & LLM Engineering, covering Embeddings and Microservices. It works with Python, Docker and GitHub. The repository describes itself as: Libraries, microservices, tools, and other reference software, supporting development of performance-optimized Edge AI applications. The licence is Apache-2.0.

When your agent uses it

  • An app needs embeddings for similarity search
  • Tasks that involve Embeddings
  • Tasks that involve Microservices

Example prompts

  • “/multimodal-embedding-serving-user”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Context routing — repo clone or standalone? REST or SDK?
  2. Bring-up (identical in both contexts)
  3. Capability check (mandatory before non-text inputs)
  4. Embed
  5. Stop / clean

What it can do on your machine

Read from SKILL.md and the folder at commit 960d2e4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • docker
    • bash
    • poetry

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multimodal Embedding Serving User loads about 1.7k tokens when it runs. Until then it costs about 154 tokens; SKILL.md has 556 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~154
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-edge-platform/edge-ai-libraries at commit 960d2e4, republished under its Apache-2.0 licence (© open-edge-platform). 556 words, ~1,717 tokens.

Download SKILL.mdSave it as .claude/skills/multimodal-embedding-serving-user/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
multimodal-embedding-serving-user
description
Deploy and consume the Multimodal Embedding Serving microservice — bring it up with setup.sh + docker compose (from a repo clone, or by fetching those same files from GitHub when no clone exists) using the prebuilt intel/multimodal-embedding-serving image, embed text/images/videos over REST on port 9777, choose among 19 models (CLIP/SigLIP/MobileCLIP/CN-CLIP/Blip2/ QwenText), or integrate in-process via the Python SDK wheel. Use when an app needs embeddings for similarity search or retrieval. Not for modifying the service's source — that is multimodal-embedding-serving-dev.

Multimodal Embedding Serving — User

Run and call the embedding service. Run commands yourself and relay output. REST base URL: http://localhost:9777 (host port hardcoded by setup.sh; container 8000).

When to Use

  • Deploy the embedding service (Docker Compose or standalone) and confirm health
  • Embed text, images, or videos over REST on port 9777
  • Choose or switch among the CLIP/SigLIP/MobileCLIP/CN-CLIP/Blip2/QwenText models
  • Integrate embeddings in-process via the Python SDK wheel
  • Diagnose 400/422 errors or model-capability mismatches

Example Prompts

Sample Problem-solving scenarios this skill handles end-to-end:

ExampleProblem it solves
image-similarity-finder.mdFind visually/semantically similar images in a local folder
text-to-image-search.mdSearch an image folder with natural-language queries ("Google Lens" for local media)
image-duplicate-detector.mdFlag duplicate / near-duplicate images in a folder for cleanup, QC, and dataset deduplication

Docs & deploy files — with or without a clone

All paths below are relative to microservices/multimodal-embedding-serving/ in the edge-ai-libraries repo. No clone? Fetch any of them from GitHub raw:

https://raw.githubusercontent.com/open-edge-platform/edge-ai-libraries/main/microservices/multimodal-embedding-serving/<path>

Load these existing docs only when needed:

ResourceLoad when…
docs/user-guide/api-reference.md + docs/user-guide/api-docs/openapi.yamlbuilding non-text payloads (image/video, base64, segment_config) or parsing responses/errors
docs/user-guide/supported-models.mdchoosing or switching models (dimensions, modalities, language, size)
docs/user-guide/sdk-usage.md + docs/user-guide/wheel-installation.mdintegrating in-process via the Python SDK wheel
docs/user-guide/get-started.mdmore curl examples and env-var tables
setup.sh, docker/compose.yamlthe deploy artifacts used below

1. Context routing — repo clone or standalone? REST or SDK?

  • In-process Python integration wanted (no separate server): the service doubles as an SDK — build the wheel with poetry build (needs the repo) and use get_model_handler(...) + EmbeddingModel; see docs/user-guide/sdk-usage.md. Rule of thumb: default to REST; pick the SDK when a Python process embeds heavily and an HTTP hop per item would dominate. Note: MobileCLIP/Blip2 extras exist only in the Docker image, and the wheel is not on PyPI.
  • Otherwise detect a clone:
    bash
    [ -f setup.sh ] && grep -q 'name = "multimodal-embedding-serving"' pyproject.toml 2>/dev/null \
      && echo REPO || echo STANDALONE
    REPO → Step 2 from the microservice root. STANDALONE → fetch the two deploy files, then the exact same Step 2:
    bash
    RAW=https://raw.githubusercontent.com/open-edge-platform/edge-ai-libraries/main/microservices/multimodal-embedding-serving
    mkdir -p embedding-serving/docker && cd embedding-serving
    curl -fsSL $RAW/setup.sh -o setup.sh
    curl -fsSL $RAW/docker/compose.yaml -o docker/compose.yaml
    Already running (curl -sf localhost:9777/health) → Step 3.
Show full SKILL.md (272 more words)Show less

2. Bring-up (identical in both contexts)

  1. Pick a model — default CLIP/clip-vit-b-32 (512-dim, text+image+video). Trade-offs: docs/user-guide/supported-models.md.
  2. setup.sh must be sourced and hard-fails without EMBEDDING_MODEL_NAME. REGISTRY_URL=intel selects the prebuilt image; --no-build prevents a source build. Run in the background — first start downloads the model:
    bash
    bash -c 'export EMBEDDING_MODEL_NAME="CLIP/clip-vit-b-32" REGISTRY_URL=intel TAG=latest \
      && source setup.sh && docker compose -f docker/compose.yaml up -d --no-build'
    Intel GPU: also export EMBEDDING_DEVICE=GPU (setup.sh then auto-enables OpenVINO + THROUGHPUT mode).
  3. Wait for readiness:
    bash
    until curl -sf http://localhost:9777/health; do sleep 5; done

3. Capability check (mandatory before non-text inputs)

bash
curl -s http://localhost:9777/model/capabilities

QwenText models are text-only — image/video requests return 400. GET /model/current shows the exact loaded model id to use in requests.

4. Embed

Text (single string or list of strings):

bash
curl -s http://localhost:9777/embeddings -H 'Content-Type: application/json' -d '{
  "model": "CLIP/clip-vit-b-32",
  "input": {"type": "text", "text": "a red truck at a loading dock"},
  "encoding_format": "float"
}'

Response: {"embedding": [...]} — a flat vector for text/image; a list of per-frame vectors for video inputs.

  • model must equal the loaded model (else 400).
  • Image: {"type":"image_url","image_url":"https://…"} (plain string, not a nested object) or image_base64. Video: video_url/video_base64/ video_frames with segment_config (num_frames default 64, extraction_fps, frame_indexes) — full shapes and examples: docs/user-guide/api-reference.md.

5. Stop / clean

  • docker compose -f docker/compose.yaml down
  • Volumes ov-models (model caches) and data-prep persist; removing them forces re-downloads — confirm with the user first.

Troubleshooting

SymptomLikely cause → action
source setup.sh prints ERROR and stopsEMBEDDING_MODEL_NAME not exported → export it first
No response on 9777still starting/downloading → docker logs -f multimodal-embedding-serving
400 "model mismatch"request model ≠ loaded model → GET /model/current
400 unsupported modality on image/videotext-only model (QwenText) → switch model or send text
First non-text request slowlazy OpenVINO conversion/compile → expected once
422 on /embeddingsmalformed input union → check shapes in docs/user-guide/api-reference.md
Port 9777 busystop the conflicting service (EMBEDDING_SERVER_PORT is hardcoded by setup.sh)

© open-edge-platform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in microservices/multimodal-embedding-serving/.github/skills/multimodal-embedding-serving-user of open-edge-platform/edge-ai-libraries.

  • SKILL.md
  • benchmark/benchmark.md
  • evals/evals.json
  • example-prompts/image-duplicate-detector.md
  • example-prompts/image-similarity-finder.md
  • example-prompts/text-to-image-search.md

Open the folder on GitHubat commit 960d2e4

Compare with similar skills

Multimodal Embedding Serving User next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multimodal Embedding Serving User compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multimodal Embedding Serving User this skillopen-edge-platform/edge-ai-libraries168—~1.7kAutomated safety check: PassApache-2.0
Copilot SDKgithub/awesome-copilot40k5 repos~6.3kAutomated safety check: PassMIT
Creating Kboaustegard/claude-skills150—~1.5kAutomated safety check: PassMIT
Molmim NimNVIDIA/skills3.5k1 repos~1.9kAutomated safety check: NotesApache-2.0
Vss Deploy Detection Tracking 2DNVIDIA/skills3.5k1 repos~4.5kAutomated safety check: PassApache-2.0
Awf Debug Toolsgithub/gh-aw-firewall148—~2.6kAutomated safety check: NotesMIT

Similar skills

  • Copilot SDK

    github/awesome-copilot

    Official

    Build agentic applications with GitHub Copilot SDK. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 5 repos~6.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Creating Kb

    oaustegard/claude-skills

    Builds a portable, embedding-free knowledgebase from a set of files and delivers it as a self-contained .skill bundle (BM25 index + bundled searcher + query protocol).

    150 GitHub stars~1.5k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Molmim Nim

    NVIDIA/skills

    Official

    A skill your agent uses for MolMIM, NVIDIA's BioNeMo NIM microservice for small-molecule latent-space generation and optimization.

    3.5k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check: notes
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Backend & APIsAuto-check passed
  • Awf Debug Tools

    github/gh-aw-firewall

    Official

    Practical Python scripts for debugging awf - parse logs, diagnose issues, inspect containers, test domains

    148 GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check: notes
  • Take Tour

    microbus-io/fabric

    Runs the agent-guided tour of Microbus using examples. An agent skill from microbus-io/fabric.

    172 GitHub stars~7.6k tokensUpdated 10 days ago
    Backend & APIsAuto-check passed

More from open-edge-platform/edge-ai-libraries

All 29 skills in this repo
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    168 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Vss Add Nest Module

    open-edge-platform/edge-ai-libraries

    Scaffolds and wires a new NestJS service/module for the Video Search & Summarization sample app's pipeline-manager using the repo's real conventions.

    168 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Chatqna Helm Deploy

    open-edge-platform/edge-ai-libraries

    Deploy Chat Question-and-Answer Core to Kubernetes using Helm (OpenVINO CPU, OpenVINO GPU, or Ollama), including values.yaml configuration, helm install/upgrade, deployment verification, uninstall…

    168 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Generate Changelog

    open-edge-platform/edge-ai-libraries

    Generates or updates CHANGELOG.md by analyzing git commit history between two branches, tags, or revisions in ANY git repository or folder.

    168 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Vss Deploy

    open-edge-platform/edge-ai-libraries

    Deploys and manages VSS through setup.sh and its Docker Compose overlays.

    168 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Vss Deploy Helm

    open-edge-platform/edge-ai-libraries

    A skill your agent uses whenever a developer needs to deploy VSS to Kubernetes, helm install VSS, configure values.yaml for VSS, or run VSS on k8s with GPU/vLLM for the…

    168 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed

Questions about Multimodal Embedding Serving User

What does Multimodal Embedding Serving User do?

Deploy and consume the Multimodal Embedding Serving microservice — bring it up with setup.sh + docker compose (from a repo clone, or by fetching those same files from GitHub when no clone exists)…. Multimodal Embedding Serving User is an agent skill from open-edge-platform/edge-ai-libraries.sh + docker compose (from a repo clone, or by fetching those same files from GitHub when no clone exists) using the prebuilt intel/multimodal-embedding-serving image, embed text/images/videos over REST on port 9777, choose among 19 models (CLIP/SigLIP/MobileCLIP/CN-CLIP/Blip2/ QwenText), or integrate in-process via the Python SDK wheel.

When should I use Multimodal Embedding Serving User?

Multimodal Embedding Serving User fits situations like: an app needs embeddings for similarity search; tasks that involve Embeddings; tasks that involve Microservices.

How do I install Multimodal Embedding Serving User in Claude Code?

Run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-embedding-serving-user -a claude-code`. Or copy the skill folder (microservices/multimodal-embedding-serving/.github/skills/multimodal-embedding-serving-user in open-edge-platform/edge-ai-libraries) into .claude/skills/multimodal-embedding-serving-user in your project. Claude Code loads it when a task matches its description.

How do I install Multimodal Embedding Serving User in Codex?

Run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-embedding-serving-user -a codex`. Or copy the skill folder (microservices/multimodal-embedding-serving/.github/skills/multimodal-embedding-serving-user in open-edge-platform/edge-ai-libraries) into .agents/skills/multimodal-embedding-serving-user in your project. Codex loads it when a task matches its description.

Can I use Multimodal Embedding Serving User in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-embedding-serving-user -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multimodal-embedding-serving-user, .gemini/skills/multimodal-embedding-serving-user, .github/skills/multimodal-embedding-serving-user and .opencode/skills/multimodal-embedding-serving-user in your project.

What does Multimodal Embedding Serving User need to run?

Going by SKILL.md and its folder, Multimodal Embedding Serving User needs the command-line tools its instructions call (curl, docker, bash and poetry). Our summary lists: Python 3; Docker.

Does Multimodal Embedding Serving User access the network?

SKILL.md contains no URLs. Its commands use curl and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Multimodal Embedding Serving User safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Multimodal Embedding Serving User use?

Multimodal Embedding Serving User is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multimodal Embedding Serving User use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Multimodal Embedding Serving User?

Skills that share tags, products or a category with Multimodal Embedding Serving User: Copilot SDK (github/awesome-copilot, 40k stars), Creating Kb (oaustegard/claude-skills, 150 stars), Molmim Nim (NVIDIA/skills, 3.5k stars) and Vss Deploy Detection Tracking 2D (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multimodal Embedding Serving User?

open-edge-platform (a GitHub organization) maintains it in open-edge-platform/edge-ai-libraries, which has 168 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.

Source: open-edge-platform/edge-ai-libraries on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.