Agent skill

Find Serving Recipe

by ai-dynamo in ai-dynamo/dynamo

Answers "for this model, this hardware, this GPU budget and this workload, what is the best known serving configuration, and how much do I trust it?" by walking an ordered set of recipe catalogs…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Find Serving Recipe

skills CLI
$ npx skills add ai-dynamo/dynamo --skill find-serving-recipe -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-dynamo/dynamo find-serving-recipe --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/find-serving-recipe .claude/skills/find-serving-recipe && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
find-serving-recipe
GitHub stars
8.3k
Token cost
~4.3k tokens
SKILL.md length
2,218 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
Apache-2.0

At a glance

Answers "for this model, this hardware, this GPU budget and this workload, what is the best known serving configuration, and how much do I trust it?" by walking an ordered set of recipe catalogs…

  • Works in 6 steps: Never answer from memory. Every claim in… → Refuse to guess a hardware match. If no… → The tag-resolvability gate, applied to… → …
  • Deploy anything
  • SKILL.md covers Ground rules, Tier 0: this repository, Tier 1: authoritative external… and Tier 2: engine-native catalogs, plus 5 more sections
  • Reaches hub.docker.com and recipes.vllm.ai

What it does

Find Serving Recipe is an agent skill from ai-dynamo/dynamo. Answers "for this model, this hardware, this GPU budget and this workload, what is the best known serving configuration, and how much do I trust it?" by walking an ordered set of recipe catalogs with provenance gates, and writes a recipe dossier recording what was found, where, and at what confidence. Use at baseline-selection time to perform the ladder's recipe-catalog scan, and during optimization whenever a performance question may already be answered by a published recipe (before spending GPU time deriving a…

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. The repository describes itself as: A Datacenter Scale Distributed Inference Serving Framework. The licence is Apache-2.0.

When your agent uses it

  • Deploy anything
  • Replace an established baseline

Example prompts

  • “Use the find-serving-recipe skill to answer "for this model, this hardware, this GPU budget and this workload, what is the best known serving…”
  • “/find-serving-recipe”

Requirements

  • Docker

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Never answer from memory. Every claim in the dossier comes from a source fetched or read
  2. Refuse to guess a hardware match. If no source names the target GPU or a documented
  3. The tag-resolvability gate, applied to every candidate at every tier. Resolve the recipe's
  4. Carry the model card forward. Before any tier, read the Hugging Face model card for the
  5. Stop at the first tier that yields a candidate meeting the confidence bar for the
  6. Freshness is part of the verdict, never a reason to discard. Serving recipes rot

What it can do on your machine

Read from SKILL.md and the folder at commit f54f2a4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • hub.docker.com
    • recipes.vllm.ai
    • nvidia.github.io
    • docs.sglang.io
    • developer.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Find Serving Recipe loads about 4.3k tokens when it runs. Until then it costs about 157 tokens; SKILL.md has 2,218 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~157
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-dynamo/dynamo at commit f54f2a4, republished under its Apache-2.0 licence (© ai-dynamo). 2,218 words, ~4,284 tokens.

Download SKILL.mdSave it as .claude/skills/find-serving-recipe/SKILL.md (or your agent's skills folder).
name
find-serving-recipe
description
Answers "for this model, this hardware, this GPU budget and this workload, what is the best known serving configuration, and how much do I trust it?" by walking an ordered set of recipe catalogs with provenance gates, and writes a recipe dossier recording what was found, where, and at what confidence. Use at baseline-selection time to perform the ladder's recipe-catalog scan, and during optimization whenever a performance question may already be answered by a published recipe (before spending GPU time deriving a config from scratch). Do not use to deploy anything or to replace an established baseline.
license
Apache-2.0
user-invocable
true
metadata.author
NVIDIA
metadata.tags
dynamo, recipes, optimization, provenance

Find a Serving Recipe

<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->

Known-good serving configurations exist for far more model and hardware combinations than this repository's recipes/ tree carries, and deriving one from scratch on GPU time is the most expensive way to obtain one. This skill defines "the catalog": an ordered set of sources, each with a trust level, so that a recipe scan finds what exists and never presents an unverifiable config as deployable.

Two invocation contexts, with different outputs:

  • Baseline selection (the interviewer's recipe-catalog scan): the dossier feeds the proposal-and-confirmation flow. This skill only finds and ranks; proposing and capturing confirmation stay with the interviewer.
  • Mid-engagement candidate hunting: before deriving a candidate configuration from first principles, check whether a published recipe already answers the performance question. A found recipe becomes a HYPOTHESIS for the normal candidate pipeline (consult, materialize, adversarial review); it never bypasses that pipeline and it NEVER re-selects the engagement baseline.

Ground rules

  1. Never answer from memory. Every claim in the dossier comes from a source fetched or read during this invocation, cited with its path or URL and, where available, a commit or date. If a source cannot be reached and no usable local copy exists, say so explicitly and move to the next tier; do not reconstruct what the catalog "probably" contains.
  2. Refuse to guess a hardware match. If no source names the target GPU or a documented equivalent, report no match for that source. Do not silently map one SKU onto another.
  3. The tag-resolvability gate, applied to every candidate at every tier. Resolve the recipe's container image before ranking it:
    • nvcr.io/... images: query the NGC registry API for the tag.
    • Docker Hub images: https://hub.docker.com/v2/repositories/<org>/<repo>/tags/<tag>. A candidate whose image tag does not resolve, or resolves only to a mutable tag (latest, dev, nightly without a digest), is ceiling-only: it may inform expectations but must never be handed to a deploy step. A digest-pinned image (@sha256:...) passes regardless of tag mutability.
  4. Carry the model card forward. Before any tier, read the Hugging Face model card for the target model. It frequently points at the authoritative recipe for that model, states a minimum engine version, and is the only source of the model author's sampling and context correctness settings. Record those settings in the dossier; no catalog carries them.
  5. Stop at the first tier that yields a candidate meeting the confidence bar for the invocation context (deployable for baseline selection; hypothesis-grade for candidate hunting). Later tiers may still be consulted for expected performance.
  6. Freshness is part of the verdict, never a reason to discard. Serving recipes rot unevenly. Split every candidate into its DURABLE content (topology, parallelism dimensions, precision and quantization, KV-cache dtype, memory fractions, sizing, workload fit, and the reasoning behind those choices) and its VERSION-BOUND content (exact flag names and defaults, the image tag, kernel and backend selections, measured performance). Durable content is first-class evidence regardless of the recipe's age; carry it into the dossier and into any candidate's rationale. Version-bound content ages: record the engine version the recipe pins (image tag, min_*_version, or commit) and its verification or publication date, and compare against the engine's CURRENT stable release (registry tag list, engine release page, or the catalog's latest entry). When the pinned engine is behind current, verify each flag against the current CLI and port renamed ones, treat the recipe's measured performance as historical (a shape hint, not a target), pin the current image, and re-verify before anything is graded deployable. A recipe verified on the current or immediately previous minor release with unchanged flags may be graded deployable; anything older is hypothesis until re-verified. A version label alone is evidence of relabeling, not revalidation; only a recorded re-verification resets the clock. Prefer the fresher of otherwise comparable candidates, but never let a fresher, thinner recipe displace the durable lessons of an older, richer one.

Tier 0: this repository

  • docs/fern/pages/recipes/_catalog/: the schema-validated machine-readable index (index.yaml, recipes/*.yaml, validated by validate.py against schema.json). Match on model.hf_id, targets[].hardware, targets[].runtime.framework, and targets[].topology. Two fields live at different levels: status (validated or experimental) is an ENTRY-level property and gates every target under it, while recommended is a TARGET-level property. Prefer targets with recommended: true under entries with status: validated; an experimental entry caps all of its targets at hypothesis. Surface the entry's gaps list in the dossier rather than hiding it.
  • The matched entry's deploy.asset points at the recipes/<model>/.../deploy.yaml DynamoGraphDeployment and its benchmark manifests; expected_performance.summary carries the measured numbers when available: true.

A Tier 0 match with a resolvable image is the best possible outcome: zero translation, attached perf, and a validated flag. Use it and stop. A Tier 0 entry that structurally matches but fails the verdict bar (mutable or unresolved image, experimental status, unsatisfiable prerequisites, stale pin) is recorded with its lower grade and the walk CONTINUES: a structural match is not a usable match.

Tier 1: authoritative external catalogs

Consult in this order when Tier 0 yields no candidate meeting the required verdict.

  1. vllm-project/recipes (vLLM engine). Consume the JSON API, not the YAML files: https://recipes.vllm.ai/models.json, then /<hf_org>/<hf_repo>.json, then /<hf_org>/<hf_repo>/hw/<hardware>.json. Read the MODEL-LEVEL JSON first and take three things from it, because they do not appear on the per-hardware response: the version floor at $.model.min_vllm_version (variants may carry their own under $.variants.<name>.min_vllm_version), the verification signal at $.meta.hardware, a hardware-to-status map (a hardware key absent from that map is unverified for this model even if /hw/<hardware>.json renders a config), and the freshness date at $.meta.date_updated. Then fetch the per-hardware response and branch on its deploy_type: for single_node it returns argv, docker_image, env, and a hardware_profile whose gpu_count is the total; for multi_node it returns head_argv, worker_argvs, node_count, and a hardware_profile whose gpu_count is PER NODE, so the total requirement is node_count times gpu_count (the strategy_spec names the interconnect assumption, for example InfiniBand for multi-node TP). Never read the flat single-node fields from a multi-node response; a missing argv is a schema branch, not an absent recipe. Two caveats: most recipes fall back to a mutable latest image (the tag gate then classifies them ceiling-only unless the engagement pins its own image), and the exported JSON drops the prefill/decode strategy_overrides; for disaggregated topology, read the recipe's YAML from the git repo and compose against its strategies.json.
  2. NVIDIA/srt-slurm-recipes for frontier models on Blackwell-class hardware, especially disaggregated and multi-node. Recipes are SLURM-shaped but carry the full engine configuration, worker split, and image. Exclude **/agentic/ and *-sa/ paths from deployable candidates: those port externally tuned benchmark configs and are quarantined to ceiling-only (see Tier 4). Expect a meaningful fraction of container tags to fail the gate.
  3. llm-d/llm-d guides/ ("well-lit paths"). Tested, benchmarked, Kubernetes-native recipes with sized prefill/decode Deployments; when the target model is covered, the best public source of sized disaggregation topology. Two checks before any llm-d candidate is graded above hypothesis: (a) IMAGE CAPABILITY: the guide's image component may select a stock upstream image or an llm-d-hosted variant (ghcr.io/llm-d/llm-d-*) that carries patches upstream lacks, such as the NVSHMEM fix for RoCE; record which, and treat a patched-variant dependency as a prerequisite the target must satisfy, not a stock image; (b) COORDINATION TRANSLATION: guides that use LeaderWorkerSet (LWS_WORKER_INDEX, LWS_GROUP_SIZE, LWS_LEADER_ADDRESS) compute ranks and addresses in shell at start-up, which has no literal DGD equivalent; those manifests are not a mechanical translation and stay hypothesis until the multi-node coordination is re-expressed in Dynamo's own terms and verified. Only single-pod-per-worker guides with literal args: arrays on stock images translate near-mechanically.
Show full SKILL.md (979 more words)Show less

Tier 2: engine-native catalogs

For flag detail, and for models Tier 1 misses.

  • TensorRT-LLM: fetch https://nvidia.github.io/TensorRT-LLM/_static/config_db.json (versioned snapshots live under /<version>/_static/... for drift checks). The published JSON covers aggregated serving only; for disaggregated entries read examples/configs/curated/lookup.yaml in the TensorRT-LLM repo. Treat validated_trtllm_commit as a floor, not a freshness signal. Never harvest from examples/models/, which is legacy.
  • SGLang cookbook (https://docs.sglang.io/cookbook/, source under docs/cookbook/ in the sglang repo). Treat cells as FLAG SOURCES, not deployable recipes: prefer verified: true cells, skip in-progress ones, and note that the cookbook's benchmark records share the cell's match key, so a verified cell often carries measured TTFT/TPOT/throughput for the same config. Its NVIDIA-side images are mutable tags and fail the gate; take the flags, not the image.
  • NIM support matrix as a sanity check on what TP and precision NVIDIA ships for a model on a given GPU. It will not give engine flags. Scrape the version-pinned docs URL, not latest.

Conflict precedence. First-party sources contradict each other. When two in-tier sources disagree, prefer in order: (1) a config with an attached measured result on the target hardware, (2) a config with a resolvable image digest, (3) a config with a validated commit pin, (4) recency. Never silently pick one; record the conflict in the dossier.

Tier 3: narrative sources

Engine release blogs and LMSYS posts. Consult only for a model that just released and appears in no catalog. Harvested flags are hypothesis-grade at best; blog-pinned nightly images rot within weeks, so everything from this tier fails the tag gate by construction. Both blogs increasingly point at their catalogs; follow the pointer instead of scraping the post.

Tier 4: quarantined, ceiling reference only

SemiAnalysis / InferenceX configs (SemiAnalysisAI/InferenceX, and their ports under agentic/ and *-sa/ in NVIDIA repos) are state-of-the-art but tuned for a benchmark leaderboard: a large fraction depend on containers that no longer exist, on release candidates, or on feature-branch builds, and their speculative-decoding results use SIMULATED acceptance lengths, not measured ones. Use them for exactly one thing: headroom. The dossier may say "an externally published result reports X tok/s/GPU for this model on this hardware; the current configuration achieves 0.7X", with the simulation caveat attached when the ceiling involves speculative decoding. A Tier 4 config must never be handed to a deploy step, regardless of whether its image resolves.

MLPerf inference results (mlcommons/inference_results_v*) sit here too: frozen, high-integrity, wrong harness shape. Submitter scripts/slurm_llm/ directories occasionally carry real disaggregated topology worth reading for sizing, with measured results attached.

Expected performance is a separate axis

Most catalogs carry configs without numbers. When the dossier's config comes from a source without measured performance, fill the expectation from:

  • Tier 0's expected_performance fields, when a related target exists; and
  • https://developer.nvidia.com/search-data/nv_inference_benchmark.json, a public index of measured records (TTFT, TPOT, per-GPU throughput, prefill/decode split) including Dynamo entries. Note in the dossier when the record does not name the image it ran on.

Label every expectation with its source and hardware; an expectation from different hardware is a shape hint, not a target.

Verdicts

Every candidate gets exactly one grade, decided by these conditions in order:

  • ceiling-only if ANY of: the image tag does not resolve or is mutable without a digest; the source is Tier 4 (quarantined) or a Tier 3 narrative with no pinned image; the config could not be fetched or read during this invocation.
  • deployable if ALL of: the source is Tier 0 or Tier 1 (or a Tier 2 config with a stock release image); the model, hardware class, and GPU count match the engagement exactly (no inferred SKU mapping); every infrastructure prerequisite the recipe declares is satisfiable on the stated target or has a documented adaptation; the image resolves to an immutable release tag or digest; the freshness check passes (current or immediately previous minor with unchanged flags); the source is not marked experimental or unverified (for Tier 0 that is the ENTRY-level status; for vLLM recipes it is the model-level $.meta.hardware map naming the target hardware); and the translation into a DynamoGraphDeployment is mechanical (no guessed fields).
  • hypothesis otherwise: a real, fetched, resolvable config that needs porting, re-verification, adaptation, or a close-but-not-exact hardware or topology match before it could be proposed.

Record the failed conditions next to any grade below deployable, so a reader knows what would promote it.

Output: the recipe dossier

Write the dossier as an IMMUTABLE snapshot at <EXP_ROOT>/analysis/recipe-dossier/<NNN>-<UTC timestamp>.md (NNN zero-padded, increasing per invocation), in both invocation contexts (the interviewer establishes EXP_ROOT before the ladder runs). When invoked standalone with no EXP_ROOT, use ./recipe-dossier/ under the current working directory with the same layout and tell the operator where it landed. Never modify a snapshot after writing it; a later invocation writes the next snapshot and may summarize deltas against the previous one. Maintain <EXP_ROOT>/analysis/recipe-dossier/index.md listing every snapshot with its SHA256. Return the snapshot path and its SHA256 to the caller; callers cite exactly that pair, so an evidence record written against snapshot 001 still verifies after snapshot 002 exists. Later iterations reuse the latest snapshot by reading the index. Each snapshot contains:

  • the question (model, hardware, GPU budget, workload, topology preference);
  • the model card's correctness settings (sampling, context, parsers) and minimum engine version;
  • every tier consulted, what was searched, and what it returned (including "no match");
  • for each candidate: source with path or URL and commit or date, image and its gate result (resolvable, digest-pinned, mutable, missing), pinned engine version versus current stable and the verification date (the freshness check), flags or manifest, measured performance with its source, the verdict (deployable, hypothesis, or ceiling-only) and, below deployable, the conditions that failed;
  • conflicts encountered and how precedence resolved them;
  • the selected candidate and why, or an explicit statement that nothing met the bar.

The dossier is evidence for the interviewer or the hypothesis pipeline. This skill does not deploy, does not edit tracked recipes, and does not modify the engagement baseline.

© ai-dynamo, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/find-serving-recipe of ai-dynamo/dynamo.

Open the folder on GitHubat commit f54f2a4

Compare with similar skills

Find Serving Recipe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Find Serving Recipe compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Find Serving Recipe this skillai-dynamo/dynamo8.3k—~4.3kAutomated safety check: PassApache-2.0
Hexstellarbrayonpi/hexstellar1.4k—~5kAutomated safety check: PassProprietary
Safactory WorkflowsAI45Lab/SAfactory236—~1.8kAutomated safety check: PassNone
Langfuselangfuse/skills300—~2.1kAutomated safety check: NotesMIT
Areno Debug RuntimeinclusionAI/AReno323—~486Automated safety check: PassApache-2.0
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Hexstellar

    brayonpi/hexstellar

    Decide, don't guess — trigger on ANY combinatorial or ground-state decision where a plausible guess is worse than silence: rosters and on-call schedules, packing and placement, RAG passage…

    1.4k GitHub stars~5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Safactory Workflows

    AI45Lab/SAfactory

    Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.

    236 GitHub stars~1.8k tokensUpdated 15 days ago
    AI & LLM EngineeringAuto-check passed
  • Langfuse

    langfuse/skills

    Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications.

    300 GitHub stars~2.1k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check: notes
  • Areno Debug Runtime

    inclusionAI/AReno

    Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

    323 GitHub stars~486 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Eval

    agentevals-dev/agentevals

    Evaluate and score agent behavior against a golden reference.

    162 GitHub stars~904 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from ai-dynamo/dynamo

All 27 skills in this repo
  • Visual Review

    ai-dynamo/dynamo

    Create self-contained interactive HTML code-review dashboards from GitHub or GitLab pull requests, checked-out branch diffs, or supplied unified diffs, with correctness and safe-to-merge scores…

    8.3k GitHub stars~4.5k tokensUpdated today
    Auto-check passed
  • Fern Components

    ai-dynamo/dynamo

    Knowledge of Fern's built-in MDX component library (accordions, callouts, cards, steps, tabs, code blocks, API-reference snippets, and more) for authoring docs pages.

    8.3k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Fern Navigation

    ai-dynamo/dynamo

    Knowledge of Fern's site-level navigation and structure configuration — how a docs site is organized in docs.yml (and product/version .yml files) using sections, pages, folders, tabs, tab variants…

    8.3k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Dynamo Agent Harness

    ai-dynamo/dynamo

    Drives persistent Claude Code, Codex, or OpenCode agent sessions through a Dynamo OpenAI/Anthropic-compatible endpoint over Agent Client Protocol (ACP).

    8.3k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Benchmark and profile the Dynamo frontend (dynamo.frontend HTTP + tokenizer + KV router) against mock workers (dynamo.mocker).

    8.3k GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Selects and freezes a question-driven AIPerf workload, objective, load policy, and Kubernetes execution manifest for a successfully deployed Dynamo candidate.

    8.3k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Questions about Find Serving Recipe

What does Find Serving Recipe do?

Answers "for this model, this hardware, this GPU budget and this workload, what is the best known serving configuration, and how much do I trust it?" by walking an ordered set of recipe catalogs…. Find Serving Recipe is an agent skill from ai-dynamo/dynamo." by walking an ordered set of recipe catalogs with provenance gates, and writes a recipe dossier recording what was found, where, and at what confidence.

When should I use Find Serving Recipe?

Find Serving Recipe fits situations like: deploy anything; replace an established baseline.

How do I install Find Serving Recipe in Claude Code?

Run `npx skills add ai-dynamo/dynamo --skill find-serving-recipe -a claude-code`. Or copy the skill folder (.agents/skills/find-serving-recipe in ai-dynamo/dynamo) into .claude/skills/find-serving-recipe in your project. Claude Code loads it when a task matches its description.

How do I install Find Serving Recipe in Codex?

Run `npx skills add ai-dynamo/dynamo --skill find-serving-recipe -a codex`. Or copy the skill folder (.agents/skills/find-serving-recipe in ai-dynamo/dynamo) into .agents/skills/find-serving-recipe in your project. Codex loads it when a task matches its description.

Can I use Find Serving Recipe in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-dynamo/dynamo --skill find-serving-recipe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/find-serving-recipe, .gemini/skills/find-serving-recipe, .github/skills/find-serving-recipe and .opencode/skills/find-serving-recipe in your project.

What does Find Serving Recipe need to run?

SKILL.md names no scripts, command-line tools or credentials: Find Serving Recipe is instructions for the agent only. Our summary lists: Docker.

Does Find Serving Recipe access the network?

SKILL.md names 5 domains. In commands or code: hub.docker.com, recipes.vllm.ai, nvidia.github.io, docs.sglang.io and developer.nvidia.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Find Serving Recipe safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Find Serving Recipe use?

Find Serving Recipe is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Find Serving Recipe use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Find Serving Recipe?

Skills that share tags, products or a category with Find Serving Recipe: Hexstellar (brayonpi/hexstellar, 1.4k stars), Safactory Workflows (AI45Lab/SAfactory, 236 stars), Langfuse (langfuse/skills, 300 stars) and Areno Debug Runtime (inclusionAI/AReno, 323 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Find Serving Recipe?

ai-dynamo (a GitHub organization) maintains it in ai-dynamo/dynamo, which has 8,250 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 9, 2026.

Source: ai-dynamo/dynamo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.