Agent skill

Multimodal Dataprep User

by open-edge-platform in open-edge-platform/edge-ai-libraries

Deploy and consume Intel Multimodal DataPrep from prebuilt images or a repository checkout.

Apache-2.0Auto-check passedBackend & APIs

Install Multimodal Dataprep User

skills CLI
$ npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-dataprep-user -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-edge-platform/edge-ai-libraries multimodal-dataprep-user --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-edge-platform/edge-ai-libraries.git skills-src && mkdir -p .claude/skills && cp -r skills-src/microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.github/skills/multimodal-dataprep-user .claude/skills/multimodal-dataprep-user && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multimodal-dataprep-user
GitHub stars
169
Token cost
~2.2k tokens
SKILL.md length
631 words
Files
6
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploy and consume Intel Multimodal DataPrep from prebuilt images or a repository checkout.

  • Works in 7 steps: Select deployment backends → Obtain deployment files → Start the prebuilt stack → …
  • Configuring VDMS
  • SKILL.md covers Load project resources as needed, 1. Select deployment backends, 2. Obtain deployment files and 3. Start the prebuilt stack, plus 5 more sections
  • Calls curl, docker and python3; needs MINIO_ROOT_PASSWORD

What it does

Multimodal Dataprep User is an agent skill from open-edge-platform/edge-ai-libraries. Deploy and consume Intel Multimodal DataPrep from prebuilt images or a repository checkout. Use for configuring VDMS or Milvus vector storage, MinIO or local media storage, checking service dependencies, and ingesting, listing, streaming, or deleting videos and images; submitting batch jobs; adding text-summary embeddings; and inspecting telemetry. This service prepares retrieval data but does not execute semantic search. Use multimodal-dataprep-dev for source changes, tests, or image builds.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `benchmark/benchmark.md`, `evals/evals.json` and `example-prompts/edge-video-preprocessing-box.md`).

It sits in Backend & APIs, covering Embeddings, File uploads and storage and Background jobs. It works with Milvus and Docker. The repository describes itself as: Libraries, microservices, tools, and other reference software, supporting development of performance-optimized Edge AI applications. The licence is Apache-2.0.

When your agent uses it

  • Configuring VDMS
  • Milvus vector storage
  • Local media storage
  • Checking service dependencies

Example prompts

  • “/multimodal-dataprep-user”

Requirements

  • Python 3
  • Docker

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Select deployment backends
  2. Obtain deployment files
  3. Start the prebuilt stack
  4. Verify readiness and dependencies
  5. Ingest supported media
  6. Add a text summary
  7. Manage and observe

What it can do on your machine

Read from SKILL.md and the folder at commit 3084578. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • docker
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MINIO_ROOT_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multimodal Dataprep User loads about 2.2k tokens when it runs. Until then it costs about 131 tokens; SKILL.md has 631 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-edge-platform/edge-ai-libraries at commit 3084578, republished under its Apache-2.0 licence (© open-edge-platform). 631 words, ~2,152 tokens.

Download SKILL.mdSave it as .claude/skills/multimodal-dataprep-user/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
multimodal-dataprep-user
description
Deploy and consume Intel Multimodal DataPrep from prebuilt images or a repository checkout. Use for configuring VDMS or Milvus vector storage, MinIO or local media storage, checking service dependencies, and ingesting, listing, streaming, or deleting videos and images; submitting batch jobs; adding text-summary embeddings; and inspecting telemetry. This service prepares retrieval data but does not execute semantic search. Use multimodal-dataprep-dev for source changes, tests, or image builds.

Multimodal DataPrep — User

Run deployment and API commands when authorized, then report their actual output. API base: http://localhost:6007/v1/dataprep.

Multimodal DataPrep creates embeddings and metadata for retrieval. It does not offer a vector-query/search endpoint.

Load project resources as needed

Paths are relative to microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.

ResourceRead when
docs/user-guide/api-reference.md and docs/user-guide/api-docs/openapi.yamlConstructing media, image, batch, summary, download, delete, or telemetry requests
docs/user-guide/get-started.mdConfiguring devices, detection, batching, duplicate policy, or environment variables
docs/user-guide/pluggable-backends.mdSelecting VDMS/Milvus or MinIO/local and diagnosing backend behavior
docs/user-guide/telemetry-metrics.mdReading ingestion telemetry or configuring Metrics Manager
setup.sh and docker/compose*.yamlDeploying a stack

Example scenarios:

ExamplePurpose
manufacturing-inspection-archive.mdIngest object-aware production media
object-aware-video-catalog.mdBuild retrieval-ready frame and crop records
edge-video-preprocessing-box.mdBatch-ingest a mounted edge directory

1. Select deployment backends

Vector backendMedia storageCompose files
VDMS (default)MinIO (default)docker/compose.yaml
MilvusMinIOdocker/compose-milvus.yaml
VDMSLocal filesystemdocker/compose.yaml then docker/compose.storage-local.yaml

Use MM_DATAPREP_VECTORDB_BACKEND (vdms or milvus) and MM_DATAPREP_STORAGE_BACKEND (minio or local) for custom deployments. Keep the service, retriever, and collection naming consistent.

2. Obtain deployment files

In a repository checkout, run from the microservice root.

Without a checkout, fetch the setup script and the compose file(s) for the selected backend:

bash
RAW='https://raw.githubusercontent.com/open-edge-platform/edge-ai-libraries/main/microservices/visual-data-preparation-for-retrieval/multimodal-dataprep'
mkdir -p multimodal-dataprep/docker
cd multimodal-dataprep
curl -fsSLo setup.sh "$RAW/setup.sh"
curl -fsSLo docker/compose.yaml "$RAW/docker/compose.yaml"

Also fetch docker/compose-milvus.yaml or docker/compose.storage-local.yaml when selected.

3. Start the prebuilt stack

Never commit credentials. For the default VDMS + MinIO deployment:

bash
export MINIO_ROOT_USER='<user>'
export MINIO_ROOT_PASSWORD='<strong-password>'
export EMBEDDING_MODEL_NAME='CLIP/clip-vit-b-32'
export REGISTRY_URL='docker.io/intel'
export TAG='latest'

source ./setup.sh --nosetup
docker compose -f docker/compose.yaml up -d --no-build

For Milvus, replace the compose file with docker/compose-milvus.yaml. For local media storage, layer the storage override after docker/compose.yaml.

setup.sh must be sourced because it exports Compose variables. With no argument it only exports defaults; it does not start containers.

4. Verify readiness and dependencies

Wait for the in-process embedding client:

bash
until curl -fsS http://localhost:6007/v1/dataprep/health \
  | python3 -c 'import json,sys; d=json.load(sys.stdin); raise SystemExit(0 if d.get("status") == "ok" and d.get("embedding_client_status") == "preloaded" else 1)'
do
  sleep 10
done

The current fields are embedding_client_status, model_name, embedding_device, use_openvino, detection_model, and detection_device.

Check dependencies separately:

bash
docker compose -f docker/compose.yaml ps
curl -fsS http://localhost:6010/minio/health/live
timeout 2 bash -c '</dev/tcp/localhost/6020'   # default VDMS

For Milvus, use the Milvus compose file and probe localhost:19530. DataPrep's health response is not a substitute for checking every dependency.

5. Ingest supported media

Upload a video:

bash
curl -fsS -X POST \
  'http://localhost:6007/v1/dataprep/media/upload?frame_interval=15&enable_object_detection=true&tags=camera-1' \
  -F 'file=@/path/to/video.mp4;type=video/mp4'

Upload an image through the same multipart endpoint:

bash
curl -fsS -X POST \
  'http://localhost:6007/v1/dataprep/media/upload?enable_object_detection=true&tags=inspection' \
  -F 'file=@/path/to/image.jpg;type=image/jpeg'

For inline base64 or remote HTTP(S) images, use POST /media/ingest. For media already in the selected storage backend, use POST /media/process.

Asynchronous workflows:

  • POST /media/upload/batch
  • POST /media/process/batch
  • POST /media/ingest/batch
  • POST /media/ingest-dir (also store_copy: false to embed files in place without copying them into storage — such media is still listed by GET /media with "stored": false and streamable via GET /media/download — and metadata / meta/<basename>.json sidecars for user-defined filterable fields)
  • GET /media/jobs/{job_id}
  • DELETE /media/jobs/{job_id} to request cancellation

Clean-up: DELETE /media/{bucket_name}/{video_id} removes one item, DELETE /media/{bucket_name} clears a whole bucket (storage + embeddings).

Read the API reference for exact request schemas and configured batch limits.

Show full SKILL.md (222 more words)Show less

6. Add a text summary

Use the video_id returned by media listing/job results and its bucket:

bash
curl -fsS -X POST 'http://localhost:6007/v1/dataprep/summary' \
  -H 'Content-Type: application/json' \
  -d '{
    "bucket_name": "video-summary",
    "video_id": "dp_video_1730000000",
    "video_summary": "forklift narrowly misses pedestrian",
    "video_start_time": 33,
    "video_end_time": 41,
    "tags": ["safety"]
  }'

The selected model must support text embeddings.

7. Manage and observe

bash
curl -fsS 'http://localhost:6007/v1/dataprep/media'
curl -fsS 'http://localhost:6007/v1/dataprep/telemetry?limit=5'
curl -L 'http://localhost:6007/v1/dataprep/media/download?video_id=dp_video_1730000000' \
  -o media.bin

GET /media/download supports HTTP Range requests for seeking.

Deletion removes the entire media directory and its matching vectors:

bash
curl -fsS -X DELETE \
  'http://localhost:6007/v1/dataprep/media/video-summary/dp_video_1730000000'

There is no current single-file video_name deletion option. Deletion is destructive, so obtain explicit confirmation before running it.

Set MM_DATAPREP_METRICS_MANAGER_URL to publish completed-pipeline dataprep_embeddings_per_second values asynchronously. /telemetry remains the direct source for detailed per-ingestion stage timings.

Troubleshooting

SymptomAction
API is unavailableInspect docker compose ... ps and DataPrep logs; initial model/YOLOX downloads can take time
embedding_client_status is not_loaded or errorVerify EMBEDDING_MODEL_NAME, model compatibility, device access, and startup logs
Dimension mismatchUse a fresh collection compatible with the selected model; never wipe an existing collection without confirmation
Milvus connection failure behind a proxyUse docker/compose-milvus.yaml and ensure Milvus/etcd/in-cluster addresses bypass proxies
MinIO authentication failureVerify the effective MM_DATAPREP_MINIO_* values and reuse the same credentials across restarts
Large upload returns 413Identify the rejecting proxy/server from headers and logs; stage media in storage and call /media/process
Duplicate upload returns 409MM_DATAPREP_ALLOW_DUPLICATE_UPLOADS=false is enforcing content-hash deduplication
Object crops are absentCheck whether YOLOX downloaded and whether detection is enabled on a supported device
Local directory ingest is rejectedKeep dir_path beneath MM_DATAPREP_INGEST_DATA_ROOT; traversal outside that root is intentionally blocked

© open-edge-platform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.github/skills/multimodal-dataprep-user of open-edge-platform/edge-ai-libraries.

  • SKILL.md
  • benchmark/benchmark.md
  • evals/evals.json
  • example-prompts/edge-video-preprocessing-box.md
  • example-prompts/manufacturing-inspection-archive.md
  • example-prompts/object-aware-video-catalog.md

Open the folder on GitHubat commit 3084578

Compare with similar skills

Multimodal Dataprep User next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multimodal Dataprep User compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multimodal Dataprep User this skillopen-edge-platform/edge-ai-libraries169—~2.2kAutomated safety check: PassApache-2.0
Vss Deploy Detection Tracking 2DNVIDIA/skills3.5k1 repos~4.5kAutomated safety check: PassApache-2.0
Molmim NimNVIDIA/skills3.5k1 repos~1.9kAutomated safety check: NotesApache-2.0
Debugging Signals PipelinePostHog/posthog40k—~2.4kAutomated safety check: NotesCustom licence
LLM Gatewaysickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
Deepstream SopNVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0

Similar skills

  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Backend & APIsAuto-check passed
  • Molmim Nim

    NVIDIA/skills

    Official

    A skill your agent uses for MolMIM, NVIDIA's BioNeMo NIM microservice for small-molecule latent-space generation and optimization.

    3.5k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check: notes
  • Official

    Debug the signals pipeline locally end-to-end. An agent skill from PostHog/posthog.

    40k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • LLM Gateway

    sickn33/agentic-awesome-skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed
  • Deepstream Sop

    NVIDIA/skills

    Official

    A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

    3.5k GitHub stars~4.7k tokensUpdated today
    Backend & APIsAuto-check: notes
  • Convex Runtime

    IgorWarzocha/Opencode-Workflows

    Implement Convex runtime features: HTTP actions, file storage, search (full text + vector), scheduling (crons + scheduled functions), and RAG patterns.

    122 GitHub stars~1.1k tokensUpdated 8 mo ago
    Backend & APIsAuto-check passed

More from open-edge-platform/edge-ai-libraries

All 29 skills in this repo
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    169 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Vss Add Nest Module

    open-edge-platform/edge-ai-libraries

    Scaffolds and wires a new NestJS service/module for the Video Search & Summarization sample app's pipeline-manager using the repo's real conventions.

    169 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Chatqna Helm Deploy

    open-edge-platform/edge-ai-libraries

    Deploy Chat Question-and-Answer Core to Kubernetes using Helm (OpenVINO CPU, OpenVINO GPU, or Ollama), including values.yaml configuration, helm install/upgrade, deployment verification, uninstall…

    169 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Generate Changelog

    open-edge-platform/edge-ai-libraries

    Generates or updates CHANGELOG.md by analyzing git commit history between two branches, tags, or revisions in ANY git repository or folder.

    169 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Vss Deploy

    open-edge-platform/edge-ai-libraries

    Deploys and manages VSS through setup.sh and its Docker Compose overlays.

    169 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Vss Deploy Helm

    open-edge-platform/edge-ai-libraries

    A skill your agent uses whenever a developer needs to deploy VSS to Kubernetes, helm install VSS, configure values.yaml for VSS, or run VSS on k8s with GPU/vLLM for the…

    169 GitHub stars~3.8k tokensUpdated today
    Auto-check passed

Works with

Questions about Multimodal Dataprep User

What does Multimodal Dataprep User do?

Deploy and consume Intel Multimodal DataPrep from prebuilt images or a repository checkout. Multimodal Dataprep User is an agent skill from open-edge-platform/edge-ai-libraries. Deploy and consume Intel Multimodal DataPrep from prebuilt images or a repository checkout.

When should I use Multimodal Dataprep User?

Multimodal Dataprep User fits situations like: configuring VDMS; milvus vector storage; local media storage; checking service dependencies.

How do I install Multimodal Dataprep User in Claude Code?

Run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-dataprep-user -a claude-code`. Or copy the skill folder (microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.github/skills/multimodal-dataprep-user in open-edge-platform/edge-ai-libraries) into .claude/skills/multimodal-dataprep-user in your project. Claude Code loads it when a task matches its description.

How do I install Multimodal Dataprep User in Codex?

Run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-dataprep-user -a codex`. Or copy the skill folder (microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.github/skills/multimodal-dataprep-user in open-edge-platform/edge-ai-libraries) into .agents/skills/multimodal-dataprep-user in your project. Codex loads it when a task matches its description.

Can I use Multimodal Dataprep User in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-dataprep-user -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multimodal-dataprep-user, .gemini/skills/multimodal-dataprep-user, .github/skills/multimodal-dataprep-user and .opencode/skills/multimodal-dataprep-user in your project.

What does Multimodal Dataprep User need to run?

Going by SKILL.md and its folder, Multimodal Dataprep User needs the command-line tools its instructions call (curl, docker and python3) and credentials named MINIO_ROOT_PASSWORD. Our summary lists: Python 3; Docker.

Does Multimodal Dataprep User access the network?

SKILL.md contains no URLs. Its commands use curl and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Multimodal Dataprep User safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Multimodal Dataprep User use?

Multimodal Dataprep User is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multimodal Dataprep User use?

About 2.2k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Multimodal Dataprep User?

Skills that share tags, products or a category with Multimodal Dataprep User: Vss Deploy Detection Tracking 2D (NVIDIA/skills, 3.5k stars), Molmim Nim (NVIDIA/skills, 3.5k stars), Debugging Signals Pipeline (PostHog/posthog, 40k stars) and LLM Gateway (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multimodal Dataprep User?

open-edge-platform (a GitHub organization) maintains it in open-edge-platform/edge-ai-libraries, which has 169 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.

Source: open-edge-platform/edge-ai-libraries on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.