Dr Jskill
jdubois/dr-jskill
Creates Java + Spring Boot projects: Web applications, full-stack apps with Vue.js or Angular or React or vanilla JS, PostgreSQL, REST APIs, and Docker.
Start, query, and stop a network-specific TAO inference microservice ({networkarch}-inference-microservice) by delegating container execution to the appropriate platform skill.
$ npx skills add NVIDIA/skills --skill tao-run-inference-service -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-run-inference-service --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-run-inference-service .claude/skills/tao-run-inference-service && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-run-inference-service" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service into .claude/skills/tao-run-inference-service/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-inference-service", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-serviceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-run-inference-service -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-run-inference-service --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-run-inference-service .agents/skills/tao-run-inference-service && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-run-inference-service" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service into .agents/skills/tao-run-inference-service/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-inference-service", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-run-inference-service -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-run-inference-service --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-run-inference-service .cursor/skills/tao-run-inference-service && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-run-inference-service" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service into .cursor/skills/tao-run-inference-service/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-inference-service", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-run-inference-service--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-run-inference-service -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-run-inference-service --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-run-inference-service .gemini/skills/tao-run-inference-service && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-run-inference-service" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service into .gemini/skills/tao-run-inference-service/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-inference-service", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-run-inference-serviceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-run-inference-service -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-run-inference-service .github/skills/tao-run-inference-service && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-run-inference-service" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service into .github/skills/tao-run-inference-service/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-inference-service", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-run-inference-service -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-run-inference-service --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-run-inference-service .opencode/skills/tao-run-inference-service && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-run-inference-service" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service into .opencode/skills/tao-run-inference-service/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-inference-service", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-run-inference-serviceStart, query, and stop a network-specific TAO inference microservice ({networkarch}-inference-microservice) by delegating container execution to the appropriate platform skill.
Tao Run Inference Service is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Start, query, and stop a network-specific TAO inference microservice ({networkarch}-inference-microservice) by delegating container execution to the appropriate platform skill. Handles container image resolution, job-payload JSON construction, and the service registry. Use when the user wants to run inference on a TAO model checkpoint using a microservice container, deploy a TAO inference endpoint, or stop a running inference container.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: The inference service has no cloud-storage dependency — model weights come from the HuggingFace Hub (HFTOKEN env var for gated models) or a local container…
It sits in Backend & APIs, covering Microservices and Containers. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashWriteFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENWANDB_API_KEYCLEARML_API_ACCESS_KEYCLEARML_API_SECRET_KEYTAO_API_KEYTAO_USER_KEYTAO_ADMIN_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
The inference service has no cloud-storage dependency — model weights come from the HuggingFace Hub (HF_TOKEN env var for gated models) or a local container path. Platform prerequisites are checked by each platform skill.
From compatibility in the SKILL.md frontmatter.
Tao Run Inference Service loads about 4.6k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 117 tokens; SKILL.md has 2,048 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
ile loaded with `set -a; source /path/to/.env; set +a`.allowed-tools: Read, Bash, WriteAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 2,048 words, ~4,629 tokens.
.claude/skills/tao-run-inference-service/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
To start an inference service:
references/code-templates.yaml → job_payload_builder.skills/platform/<platform>/SKILL.md and start the container (Section 4.2).references/code-templates.yaml → registry_write.<platform> and readiness_check.To send an inference request:
job_id, by network_arch, or by explicit user choice when multiple services run — never silently default to "latest" when more than one service exists), then read the endpoint from references/code-templates.yaml → request.registry_read with the resolved job_id.max_tokens, top_p, temperature (and any per-arch extras) with their defaults; let the user override or skip each one to accept the default. Never silently use defaults.To stop a service: Read references/code-templates.yaml → stop.registry_read to resolve the job_id, read skills/platform/<platform>/SKILL.md, then follow Section 5.
Reference data (schemas, mappings, valid values — no instructions):
references/service.yaml — image mappings, valid network_arch names, job payload schema, env var names, secrets classification.references/request.yaml — endpoint definition, request field schema, response shapes, code examples.references/code-templates.yaml — Python templates for payload building, registry writes, readiness checks, and stop/request flows.Never ask the user to type a secret value into a prompt. For every secret value:
export HF_TOKEN=... in their shell, or a KEY=value line in a user-approved env file loaded with set -a; source /path/to/.env; set +a.cat, or Read such a file — verify presence only: [ -n "$HF_TOKEN" ] && echo SET || echo UNSET.os.environ["VAR_NAME"] — never hard-code, interpolate, or prompt for the value.Secret env vars (full list in references/service.yaml → secrets_handling):
HF_TOKEN, WANDB_API_KEY, CLEARML_API_ACCESS_KEY, CLEARML_API_SECRET_KEY, TAO_API_KEY, TAO_USER_KEY.
Safe to collect in the prompt: network_arch, model_path, num_gpus, prompt text, WANDB_* config URLs, CLEARML_*_HOST URLs.
| Input | Role |
|---|---|
network_arch | Chooses container image, the per-arch inner command shape (references/service.yaml → container_commands.<network_arch>), and neural_network_name in the job JSON when applicable. Must match a basename in valid_network_arch_config_basenames in references/service.yaml (e.g. cosmos-rl, cosmos-predict2.5). |
model_path | The trained model checkpoint. Valid forms: hf_model://<org>/<model> (HuggingFace Hub — set HF_TOKEN for gated models) or a local container filesystem path. Cloud URIs (s3://, gs://, az://) are NOT supported — the inference service has no cloud-storage dependency. Always ask the user; never substitute a placeholder. See references/service.yaml → model_path_protocols. |
platform | Compute platform: local-docker, brev, slurm, or kubernetes. |
num_gpus | Defaults to 1; minimum 1 for inference. |
Each network_arch has a sidecar config file named {network_arch}.config.json. Resolve the container image as follows:
{network_arch}.config.json and take api_params.image (e.g. COSMOS_RL). This selects an entry from docker_image_defaults in references/service.yaml.IMAGE_<KEY> is set (e.g. IMAGE_COSMOS_RL), it overrides the packaged default.model_skill_defaults, call scripts/resolve_tao_image.py with its model, action, and backend. This keeps the exact Cosmos image in the Cosmos model skill's references/skill_info.yaml.mapping, resolve the dotted value against the repo-root versions.yaml with scripts/resolve_versions_key.py. Absolute environment overrides pass through unchanged. The Python examples live in references/code-templates.yaml.api_params.image is empty, fall back to the COSMOS_RL key.The config file also has spec_params.inference.model_path which drives folder vs file path semantics: if the value contains the substring folder, the container treats the path as a directory.
Set these in env_payload before encoding env_json. Do not set TAO_LOGGING_SERVER_URL or TAO_ADMIN_KEY.
TAO_EXECUTION_BACKEND — must match the platform:
| Platform | TAO_EXECUTION_BACKEND value |
|---|---|
| local-docker | local-docker |
| brev | local-docker |
| slurm | slurm |
| kubernetes | local-k8s |
CLOUD_BASED — always "False" for this skill (disables callback posting to TAO_LOGGING_SERVER_URL).
GPU env vars — only needed when the platform skill does not handle GPU injection automatically:
--runtime=nvidia with NVIDIA_DRIVER_CAPABILITIES=all and NVIDIA_VISIBLE_DEVICES=<ids>.device_requests. The platform skill handles this.The job payload and inner command (Sections 1–3) are platform-agnostic. For each platform, read skills/platform/<name>/SKILL.md for preflight checks and credentials before generating any execution code.
The inner-command shape is per network_arch — there is no uniform template. Look up the per-arch entry in references/service.yaml → container_commands.<network_arch>; if not present, the arch is unsupported — stop and ask. Pick the matching sub-block in references/code-templates.yaml → job_payload_builder.<network_arch>. Prefix the command with umask 0 && and keep it identical across platforms (local-docker, brev, slurm, kubernetes).
Common across arches:
job_id: fresh uuid.uuid4() — becomes the container name and registry key.image: resolve per Section 2.access_key, secret_key, HF_TOKEN, etc.) are read from env vars at runtime — never hard-code, never log or print.Arch-specific notes (full details in references/service.yaml → container_commands):
cosmos-rl — single --job '<JOB_JSON>' --docker_env_vars '<ENV_JSON>' blob; json.dumps(...) + shlex.quote(...). env_payload carries TAO_EXECUTION_BACKEND (per Section 3 table), TAO_API_JOB_ID, CLOUD_BASED=False. The inference service has no cloud-storage dependency; HF_TOKEN is the only cred env var that ever applies (for gated HuggingFace models).cosmos-predict2.5 — flag-style cosmos_predict inference_microservice start ... --port 8080 (no setup. prefix; uses tyro.conf.OmitArgPrefixes). --job/--docker_env_vars are not accepted. Translate model_path to --checkpoint-path (local path) or --model <registered_key> (hf_model://); cloud URIs are rejected. The only cred env var that ever applies is HF_TOKEN for gated HuggingFace models. Per-request params (prompt, inference_type, num_output_frames, guidance, seed, num_steps, negative_prompt) go in the request body, not at startup. TAO_EXECUTION_BACKEND/TAO_API_JOB_ID/CLOUD_BASED are unused and may be omitted.Read skills/platform/<platform>/SKILL.md and follow it to start the container.
Base parameters (all platforms):
| Parameter | Value |
|---|---|
image | resolved container image (Section 2) |
command | inner — the shell string built in Section 4.1 |
gpu_count | num_gpus |
env_vars | env_payload |
| job / container name | job_id — must equal the UUID from 4.1 so the registry can reference it |
host_port (local-docker, brev) | host-side port to bind to container port 8080. Default 8080, but must be unique per concurrent service — see the port-allocation rule below. |
Platform-specific additional inputs:
| Platform | Additional inputs |
|---|---|
| local-docker | None beyond base |
| brev | instance_id (optional — reuse an existing instance); on multi-credential / multi-workspace accounts also cloud_cred_id and workspace_group_id for first-create — see skills/platform/tao-run-on-brev/SKILL.md |
| slurm | partition and account — check SLURM_PARTITION/SLURM_ACCOUNT env vars; ask user if unset |
| kubernetes | namespace (default: default); image_pull_secret (required for nvcr.io images) |
Port binding (local-docker and brev): use direct docker run so that -p <host_port>:8080 can be passed and the container name equals job_id exactly.
Port allocation rule (local-docker and brev, REQUIRED for concurrent services): Before starting a service, read the registry (/tmp/tao-inf-ms-state.json) and collect the set of host_port values from every existing entry on the same platform (and, for brev, the same instance_id). Pick the lowest free port starting from 8080 that is not in that set — e.g. host_port = next(p for p in range(8080, 8200) if p not in used_ports). The default 8080 only applies when no other service is running. This is what makes "start 3 services, each reachable at a distinct host_url" work; without it, services 2 and 3 fail with bind: address already in use. SLURM and kubernetes get distinct endpoints from their own platform mechanisms and do not need this step.
Write the service registry immediately after the platform confirms the container is running. The registry (/tmp/tao-inf-ms-state.json) is keyed by job_id; "latest" always points to the most recently started service.
See references/code-templates.yaml → registry_write.<platform> for the Python template.
| Platform | host_url | platform_job_id | Extra step before writing |
|---|---|---|---|
| local-docker | http://localhost:{host_port} | — | None |
| brev | http://{brev_ip}:{host_port} | — | brev ls → get instance IP (localhost is invalid on remote VM) |
| slurm | http://localhost:{host_port} | SLURM scheduler job ID | Wait until Running; SSH port-forward localhost:{host_port}→{node}:8080 |
| kubernetes | http://{external_ip}:8080 | k8s job name | kubectl expose job … --type=LoadBalancer; wait for external IP |
After writing the registry, print the job_id and URL:
print(f"Inference service started.")
print(f" Job ID : {job_id}")
print(f" Arch : {network_arch}")
print(f" URL : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")Then poll for readiness — see references/code-templates.yaml → readiness_check. The container loads the model in the background; do not send requests before it returns 200.
Ask the user for the job_id to stop. If they don't provide one, default to state["latest"] and confirm which job_id is being stopped. Read the registry using references/code-templates.yaml → stop.registry_read, then read skills/platform/<platform>/SKILL.md and use its cancellation / stop mechanism.
| Platform | Identifier to pass | Extra cleanup |
|---|---|---|
| local-docker | job_id_to_stop — container name | None |
| brev | job_id_to_stop — container name | None |
| slurm | entry["platform_job_id"] — SLURM job ID | pkill -f "ssh.*-L.*{entry['host_port']}" |
| kubernetes | entry["platform_job_id"] — k8s job name | kubectl delete svc {entry["platform_job_id"]} -n <namespace> |
where entry = state[job_id_to_stop]. After stopping, clean up the registry: references/code-templates.yaml → stop.registry_cleanup.
Each request must be routed to the specific service that runs the matching model. Routing happens by job_id — the registry stores network_arch per entry, so you can resolve a target by arch when the user names a model instead of a job_id. Apply these rules in order:
job_id → use it. Verify it exists in state.network_arch (e.g. "send this to the cosmos-rl service") → look up matching entries: candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch].job_ids and their started_at; do not auto-pick.job_id and no network_arch → count non-"latest" entries in state:state["latest"]. Prompt the user with the full list (job_id, network_arch, host_url) and require an explicit choice. The "latest" pointer is a convenience for single-service workflows, not a routing fallback when multiple services coexist.After resolving, read the endpoint from the registry (references/code-templates.yaml → request.registry_read), passing the resolved job_id as user_provided_job_id. Confirm to the user: "Sending to job_id=… arch=… url=…". If the service may still be loading, poll readiness first (references/code-templates.yaml → readiness_check).
Cross-check before sending: if the user-supplied request body contains arch-specific fields (e.g. guidance / num_steps / seed / negative_prompt → cosmos-predict2.5; required image_url/video_url content items → cosmos-rl), verify they are consistent with state[job_id]["network_arch"]. On mismatch, stop and ask — sending a cosmos-predict2.5 body to a cosmos-rl service will fail at the container with a 4xx/5xx that is harder to diagnose than catching it here.
Before constructing the request body, you MUST explicitly prompt the user for the vLLM-style sampling parameters. Do not silently apply defaults. Use a structured prompt, one question per field, that:
After the prompt, apply each user-entered value verbatim and substitute the default for any skipped field. Do not invent values or silently clamp.
Field list, defaults, and per-arch applicability: references/request.yaml → chat_completions_request_body (base sampling fields: max_tokens, top_p, temperature) and network_arch_constraints.<network_arch> (per-arch overrides and extras such as guidance/num_steps/seed/negative_prompt for cosmos-predict2.5). If a field is marked unsupported for the active arch, do not prompt for it and do not include it in the body.
Send a POST to {BASE_URL}/v1/chat/completions with Content-Type: application/json and a timeout of at least 300 s. The body is OpenAI-compatible (vLLM chat completions); see references/request.yaml → chat_completions_request_body for the full field schema and content-item shapes (text / image_url / video_url), and code_examples for ready-to-run Python and curl samples.
Constraints: only the first user message is processed. No secret values in request bodies. Per-network constraints (e.g. cosmos-rl requires every request to include an image or video; cosmos-rl rejects data: URIs) are in references/request.yaml → network_arch_constraints.
| HTTP status | Meaning | Action |
|---|---|---|
| 200 | Success — choices[0].message.content has the generated text | Read result |
| 202 | Server still initializing or model still loading | Retry after a delay |
| 503 | Initialization failed, model load failed, or model not yet ready | Inspect error.type: model_not_ready → retry; initialization_error / model_load_error → give up and check logs |
| 400 | Missing or empty JSON body | Fix request |
| 500 | Unhandled exception during inference | Check container logs |
For 202 and 503, the body contains {"error": {"type": "<error_type>", "message": "<reason>"}}. See container_response_shapes in references/request.yaml for error type strings.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (references) in skills/tao-run-inference-service of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Tao Run Inference Service next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Run Inference Service this skillNVIDIA/skills | 3.6k | — | ~4.6k | Automated safety check: Notes | Apache-2.0 | |
| Dr Jskilljdubois/dr-jskill | 342 | — | ~4.6k | Automated safety check: Notes | Apache-2.0 | |
| Polylith Project ManagementDavidVujic/python-polylith | 554 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Vss Buildopen-edge-platform/edge-ai-libraries | 171 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | |
| Dlsps Useropen-edge-platform/edge-ai-libraries | 171 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Spring Boot Project Creatorgiuseppe-trisciuoglio/developer-kit | 357 | — | ~3.4k | Automated safety check: Notes | MIT |
jdubois/dr-jskill
Creates Java + Spring Boot projects: Web applications, full-stack apps with Vue.js or Angular or React or vanilla JS, PostgreSQL, REST APIs, and Docker.
DavidVujic/python-polylith
Create a deployable Polylith project with poly create project — a lightweight pyproject.toml under projects/<name/ that references bricks for deployment as a Docker image, wheel, AWS Lambda, GCP…
open-edge-platform/edge-ai-libraries
Build (and optionally push) the VSS Docker images from source with make build, make build-deps, and make push - the application services, the dependency microservices, or both, with registry/tag…
open-edge-platform/edge-ai-libraries
Deploy and operate DL Streamer Pipeline Server — a microservice that wraps DL Streamer pipelines behind a REST API for containerized, no-code operation.
giuseppe-trisciuoglio/developer-kit
Creates and scaffolds a new Spring Boot project (3.x or 4.x) by downloading from Spring Initializr, generating package structure (DDD or Layered architecture), configuring JPA, SpringDoc OpenAPI…
open-edge-platform/edge-ai-libraries
Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Categories
Start, query, and stop a network-specific TAO inference microservice ({networkarch}-inference-microservice) by delegating container execution to the appropriate platform skill. Tao Run Inference Service is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Start, query, and stop a network-specific TAO inference microservice ({networkarch}-inference-microservice) by delegating container execution to the appropriate platform skill.
Tao Run Inference Service fits situations like: the user wants to run inference on a TAO model checkpoint using a microservice container; deploy a TAO inference endpoint; stop a running inference container.
Run `npx skills add NVIDIA/skills --skill tao-run-inference-service -a claude-code`. Or copy the skill folder (skills/tao-run-inference-service in NVIDIA/skills) into .claude/skills/tao-run-inference-service in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-run-inference-service -a codex`. Or copy the skill folder (skills/tao-run-inference-service in NVIDIA/skills) into .agents/skills/tao-run-inference-service in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-run-inference-service -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-run-inference-service, .gemini/skills/tao-run-inference-service, .github/skills/tao-run-inference-service and .opencode/skills/tao-run-inference-service in your project.
Going by SKILL.md and its folder, Tao Run Inference Service needs the command-line tools its instructions call (kubectl) and credentials named HF_TOKEN, WANDB_API_KEY, CLEARML_API_ACCESS_KEY and CLEARML_API_SECRET_KEY. Our summary lists: Python 3; Docker; A credential in WANDB_API_KEY; A credential in CLEARML_API_ACCESS_KEY. Its frontmatter pre-approves these tools: Read, Bash, Write. Compatibility (from SKILL.md): The inference service has no cloud-storage dependency — model weights come from the HuggingFace Hub (HF_TOKEN env var for gated models) or a local container path. Platform prerequisites are checked by each platform skill..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Tao Run Inference Service is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 10k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Run Inference Service: Dr Jskill (jdubois/dr-jskill, 342 stars), Polylith Project Management (DavidVujic/python-polylith, 554 stars), Vss Build (open-edge-platform/edge-ai-libraries, 171 stars) and Dlsps User (open-edge-platform/edge-ai-libraries, 171 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.