Agent skill

Multimodal Dataprep Dev

by open-edge-platform in open-edge-platform/edge-ai-libraries

Develop and debug the Multimodal DataPrep microservice: its FastAPI media endpoints, in-process embedding pipeline, batch jobs, object detection, telemetry and Metrics Manager publishing, and…

Apache-2.0Auto-check passedBackend & APIs

Install Multimodal Dataprep Dev

skills CLI
$ npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-dataprep-dev -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-edge-platform/edge-ai-libraries multimodal-dataprep-dev --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-edge-platform/edge-ai-libraries.git skills-src && mkdir -p .claude/skills && cp -r skills-src/microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.github/skills/multimodal-dataprep-dev .claude/skills/multimodal-dataprep-dev && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multimodal-dataprep-dev
GitHub stars
168
Token cost
~1.3k tokens
SKILL.md length
462 words
Files
6 (incl. references)
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

Develop and debug the Multimodal DataPrep microservice: its FastAPI media endpoints, in-process embedding pipeline, batch jobs, object detection, telemetry and Metrics Manager publishing, and…

  • Changing source
  • SKILL.md covers Start with the relevant…, Build-context rule, Local development loop and Build and run, plus 3 more sections
  • Calls poetry and docker; needs MINIO_ROOT_PASSWORD
  • Adding a backend

What it does

Multimodal Dataprep Dev is an agent skill from open-edge-platform/edge-ai-libraries. Develop and debug the Multimodal DataPrep microservice: its FastAPI media endpoints, in-process embedding pipeline, batch jobs, object detection, telemetry and Metrics Manager publishing, and pluggable VDMS/Milvus vector stores plus MinIO/local storage. Use when changing source, adding a backend, running pytest/coverage/format checks, or building the service image. Use multimodal-dataprep-user for deployment and API-consumer workflows.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `evals/evals.json`, `example-prompts/onboard-embedding-model.md` and `example-prompts/update-test-cases.md`).

It sits in Backend & APIs, covering Microservices, Embeddings and Backend development. It works with Milvus, FastAPI, pytest and Docker. The repository describes itself as: Libraries, microservices, tools, and other reference software, supporting development of performance-optimized Edge AI applications. The licence is Apache-2.0.

When your agent uses it

  • Changing source
  • Adding a backend
  • Running pytest/coverage/format checks
  • Building the service image

Example prompts

  • “/multimodal-dataprep-dev”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 960d2e4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • poetry
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MINIO_ROOT_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multimodal Dataprep Dev loads about 1.3k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 462 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-edge-platform/edge-ai-libraries at commit 960d2e4, republished under its Apache-2.0 licence (© open-edge-platform). 462 words, ~1,335 tokens.

Download SKILL.mdSave it as .claude/skills/multimodal-dataprep-dev/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
multimodal-dataprep-dev
description
Develop and debug the Multimodal DataPrep microservice: its FastAPI media endpoints, in-process embedding pipeline, batch jobs, object detection, telemetry and Metrics Manager publishing, and pluggable VDMS/Milvus vector stores plus MinIO/local storage. Use when changing source, adding a backend, running pytest/coverage/format checks, or building the service image. Use multimodal-dataprep-user for deployment and API-consumer workflows.

Multimodal DataPrep — Dev

Work from microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/ inside an edge-ai-libraries checkout. If the request only deploys or consumes the service, use ../multimodal-dataprep-user/SKILL.md.

Start with the relevant reference

ReferenceRead when
references/source-map.mdLocating endpoints, pipeline code, backend abstractions, or configuration
references/testing-and-build.mdInstalling dependencies, testing, formatting, building, or debugging containers

Example tasks:

ExamplePurpose
onboard-embedding-model.mdExercise a different embedding model safely
update-test-cases.mdAdd coverage for a source change

Build-context rule

Use ./build.sh; do not run docker build ... . from this directory. The Dockerfile copies both this service and the sibling multimodal-embedding-serving source, so its context must be microservices/. build.sh supplies that context.

Local development loop

bash
poetry install --with dev
poetry run python -m pytest tests
poetry run coverage run --rcfile ./pyproject.toml -m pytest tests
poetry run coverage report -m
poetry run black --check src tests
poetry run isort --check-only src tests

Run a focused test file while iterating:

bash
poetry run python -m pytest tests/test_vectorstores.py

Do not conceal collection failures or attribute unrelated failures to the current change. See references/testing-and-build.md.

Build and run

setup.sh must be sourced. With no argument it exports defaults and creates the YOLOX model volume; it does not start the stack.

bash
export MINIO_ROOT_USER='<user>'
export MINIO_ROOT_PASSWORD='<strong-password>'
export EMBEDDING_MODEL_NAME='CLIP/clip-vit-b-32'
source ./setup.sh --nosetup

./build.sh
docker compose -f docker/compose.yaml up -d --build

Other supported setup actions are --conf, --down, --build [custom-tag], and --nd (foreground docker compose ... up --build).

For Milvus, use docker/compose-milvus.yaml. For local media storage, layer docker/compose.storage-local.yaml after the default compose file.

Current architecture

src/main.py creates the FastAPI app at /v1/dataprep, starts the optional Metrics Manager publisher, preloads the embedding client and YOLOX detector, and asks the active vector store to update its index during shutdown.

Requests flow through src/endpoints/ into src/core/embedding/embedding_orchestrator.py. Video work is executed by the threaded/shared-memory pipeline in embedding_helper.py; client.py wraps the in-process model from the sibling embedding package and persists vectors through src/core/vectorstores/. Media bytes and metadata go through src/core/storage/.

Show full SKILL.md (211 more words)Show less

Change rules

  • Preserve backend neutrality. Use get_vector_store() and get_storage() rather than importing a concrete backend in endpoint or orchestration code.
  • Keep search/query behavior out of this service; BaseVectorStore covers ingestion-time add, delete, health, and index-update operations.
  • Add endpoint schemas in src/common/schema.py and include routers in src/main.py.
  • Keep media routes under /media; supported inputs include MP4 video and common image formats, plus text summaries at /summary.
  • Mock external storage, model, and vector-store calls in unit tests. The Milvus integration test is opt-in through MILVUS_IT_URI.
  • Add the repository SPDX header to every new source, config, test, or documentation file.
  • Never commit credentials.

Important configuration facts

FactImpact
All application settings use Pydantic's MM_DATAPREP_ prefixSet container variables such as MM_DATAPREP_VECTORDB_BACKEND, MM_DATAPREP_STORAGE_BACKEND, and MM_DATAPREP_EMBEDDING_MODEL_NAME
setup.sh sets INDEX_NAME=video-rag; default compose maps it to MM_DATAPREP_DB_COLLECTIONFor a one-off VDMS collection, set INDEX_NAME after sourcing and before Compose
Milvus collection names cannot contain hyphenscompose-milvus.yaml uses MILVUS_INDEX_NAME with default video_rag
Changing embedding dimensions is incompatible with an existing collectionChoose a fresh collection or obtain confirmation before deleting data
YOLOX weights are downloaded on first useAn offline first run can leave object detection unavailable while other ingestion continues
Metrics Manager publishing is optionalIt is enabled only when MM_DATAPREP_METRICS_MANAGER_URL is non-empty and must not delay ingestion

© open-edge-platform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.github/skills/multimodal-dataprep-dev of open-edge-platform/edge-ai-libraries.

  • SKILL.md
  • evals/evals.json
  • example-prompts/onboard-embedding-model.md
  • example-prompts/update-test-cases.md
  • references/source-map.md
  • references/testing-and-build.md

Open the folder on GitHubat commit 960d2e4

Compare with similar skills

Multimodal Dataprep Dev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multimodal Dataprep Dev compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multimodal Dataprep Dev this skillopen-edge-platform/edge-ai-libraries168—~1.3kAutomated safety check: PassApache-2.0
Deepstream SopNVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0
Flowfile Debugging PlaybookEdwardvaneechoud/Flowfile370—~6.3kAutomated safety check: PassMIT
Python API Designjohnku2011/boilerplates-with-ai-skills240—~449Automated safety check: PassMIT
Fastapi Patternsaffaan-m/ECC274k1 repos~3.9kAutomated safety check: NotesMIT
Python Devdoccker/cc-use-exp1.1k—~790Automated safety check: PassCustom licence

Similar skills

  • Deepstream Sop

    NVIDIA/skills

    Official

    A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Backend & APIsAuto-check: notes
  • Flowfile Debugging Playbook

    Edwardvaneechoud/Flowfile

    Symptom-to-cause triage playbook for Flowfile (core/worker/kernel/frontend/AI) — covers "no such table" DB cascades (two distinct causes), import-time Alembic migration corruption, silent…

    370 GitHub stars~6.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Python API Design

    johnku2011/boilerplates-with-ai-skills

    A skill your agent uses when adding or changing FastAPI routes, dependencies, or tests in this Python service — keep endpoints typed, validated, and covered by pytest.

    240 GitHub stars~449 tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Fastapi Patterns

    affaan-m/ECC

    FastAPI best practices covering project structure, Pydantic v2 schemas, dependency injection, async handlers, authentication, authorization, transactional service layers, and testing with httpx and…

    274k GitHub starsUsed in 1 repo~3.9k tokens
    Backend & APIsAuto-check: notes
  • Python Dev

    doccker/cc-use-exp

    Python 开发规范。当用户操作 .py、pyproject.toml、requirements.txt、setup.py 文件, 或涉及 FastAPI、Django、Flask、pytest、asyncio 开发时触发。

    1.1k GitHub stars~790 tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Fastapi Endpoint

    davila7/claude-code-templates

    Plan and build production-ready FastAPI endpoints with async SQLAlchemy, Pydantic v2 models, dependency injection for auth, and pytest tests.

    32k GitHub stars~3.9k tokensUpdated today
    Backend & APIsAuto-check passed

More from open-edge-platform/edge-ai-libraries

All 29 skills in this repo
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    168 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Vss Add Nest Module

    open-edge-platform/edge-ai-libraries

    Scaffolds and wires a new NestJS service/module for the Video Search & Summarization sample app's pipeline-manager using the repo's real conventions.

    168 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Chatqna Helm Deploy

    open-edge-platform/edge-ai-libraries

    Deploy Chat Question-and-Answer Core to Kubernetes using Helm (OpenVINO CPU, OpenVINO GPU, or Ollama), including values.yaml configuration, helm install/upgrade, deployment verification, uninstall…

    168 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Generate Changelog

    open-edge-platform/edge-ai-libraries

    Generates or updates CHANGELOG.md by analyzing git commit history between two branches, tags, or revisions in ANY git repository or folder.

    168 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Vss Deploy

    open-edge-platform/edge-ai-libraries

    Deploys and manages VSS through setup.sh and its Docker Compose overlays.

    168 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Vss Deploy Helm

    open-edge-platform/edge-ai-libraries

    A skill your agent uses whenever a developer needs to deploy VSS to Kubernetes, helm install VSS, configure values.yaml for VSS, or run VSS on k8s with GPU/vLLM for the…

    168 GitHub stars~3.8k tokensUpdated today
    Auto-check passed

Questions about Multimodal Dataprep Dev

What does Multimodal Dataprep Dev do?

Develop and debug the Multimodal DataPrep microservice: its FastAPI media endpoints, in-process embedding pipeline, batch jobs, object detection, telemetry and Metrics Manager publishing, and…. Multimodal Dataprep Dev is an agent skill from open-edge-platform/edge-ai-libraries. Develop and debug the Multimodal DataPrep microservice: its FastAPI media endpoints, in-process embedding pipeline, batch jobs, object detection, telemetry and Metrics Manager publishing, and pluggable VDMS/Milvus vector stores plus MinIO/local storage.

When should I use Multimodal Dataprep Dev?

Multimodal Dataprep Dev fits situations like: changing source; adding a backend; running pytest/coverage/format checks; building the service image.

How do I install Multimodal Dataprep Dev in Claude Code?

Run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-dataprep-dev -a claude-code`. Or copy the skill folder (microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.github/skills/multimodal-dataprep-dev in open-edge-platform/edge-ai-libraries) into .claude/skills/multimodal-dataprep-dev in your project. Claude Code loads it when a task matches its description.

How do I install Multimodal Dataprep Dev in Codex?

Run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-dataprep-dev -a codex`. Or copy the skill folder (microservices/visual-data-preparation-for-retrieval/multimodal-dataprep/.github/skills/multimodal-dataprep-dev in open-edge-platform/edge-ai-libraries) into .agents/skills/multimodal-dataprep-dev in your project. Codex loads it when a task matches its description.

Can I use Multimodal Dataprep Dev in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-edge-platform/edge-ai-libraries --skill multimodal-dataprep-dev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multimodal-dataprep-dev, .gemini/skills/multimodal-dataprep-dev, .github/skills/multimodal-dataprep-dev and .opencode/skills/multimodal-dataprep-dev in your project.

What does Multimodal Dataprep Dev need to run?

Going by SKILL.md and its folder, Multimodal Dataprep Dev needs the command-line tools its instructions call (poetry and docker) and credentials named MINIO_ROOT_PASSWORD. Our summary lists: Python 3; Docker.

Does Multimodal Dataprep Dev access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Multimodal Dataprep Dev safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Multimodal Dataprep Dev use?

Multimodal Dataprep Dev is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multimodal Dataprep Dev use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Multimodal Dataprep Dev?

Skills that share tags, products or a category with Multimodal Dataprep Dev: Deepstream Sop (NVIDIA/skills, 3.5k stars), Flowfile Debugging Playbook (Edwardvaneechoud/Flowfile, 370 stars), Python API Design (johnku2011/boilerplates-with-ai-skills, 240 stars) and Fastapi Patterns (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multimodal Dataprep Dev?

open-edge-platform (a GitHub organization) maintains it in open-edge-platform/edge-ai-libraries, which has 168 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 7, 2026.

Source: open-edge-platform/edge-ai-libraries on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.