Agent skill

Managed Model Endpoints

by JimLiu in JimLiu/science-skills

Register a model service in the managed family — a local model server container the daemon starts/stops on demand, or a remote upstream model API (https).

Apache-2.0Auto-check passedDevOps & Cloud

Install Managed Model Endpoints

skills CLI
$ npx skills add JimLiu/science-skills --skill managed-model-endpoints -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JimLiu/science-skills managed-model-endpoints --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/managed-model-endpoints .claude/skills/managed-model-endpoints && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
managed-model-endpoints
GitHub stars
227
Used in
2 other repos
Token cost
~2.7k tokens
SKILL.md length
1,012 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
Apache-2.0

At a glance

Register a model service in the managed family — a local model server container the daemon starts/stops on demand, or a remote upstream model API (https).

  • Wants a model service available for inference
  • SKILL.md covers Calling a registered endpoint…, Enablement — once per machine, Register (repl kernel) and Remote endpoints — upstream…, plus 2 more sections
  • Calls docker; needs NVIDIA_API_KEY and INFER_API_KEY
  • Listcompute shows managed endpoints

What it does

Managed Model Endpoints is an agent skill from JimLiu/science-skills. Register a model service in the managed family — a local model server container the daemon starts/stops on demand, or a remote upstream model API (https). Read the runbook, allocate a port (local only), compose idempotent start/stop scripts (local only), register once. Load when the user wants a model service available for inference, or when listcompute shows managed endpoints.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Runbooks and postmortems. The licence is Apache-2.0.

When your agent uses it

  • Wants a model service available for inference
  • Listcompute shows managed endpoints

Example prompts

  • “/managed-model-endpoints”

Requirements

  • Python 3
  • Docker
  • A credential in INFER_API_KEY
  • A credential in NVIDIA_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit fb309c3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NVIDIA_API_KEY
    • INFER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Managed Model Endpoints loads about 2.7k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 1,012 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JimLiu/science-skills at commit fb309c3, republished under its Apache-2.0 licence (© JimLiu). 1,012 words, ~2,660 tokens.

Download SKILL.mdSave it as .claude/skills/managed-model-endpoints/SKILL.md (or your agent's skills folder).
name
managed-model-endpoints
description
Register a model service in the managed family — a local model server container the daemon starts/stops on demand, or a remote upstream model API (https). Read the runbook, allocate a port (local only), compose idempotent start/stop scripts (local only), register once. Load when the user wants a model service available for inference, or when list_compute shows managed endpoints.
license
Apache-2.0

Managed model endpoints

A managed model endpoint is a model service the daemon owns: you register it once, then every compute_provider cell against it just works — the daemon swaps the resident model off the device (one model at a time, via the resident's own approved stop), runs your approved start script, waits for the readiness route, then runs your cell, streaming its lifecycle progress into the cell as it goes. You never run the container runtime yourself, never poll readiness in cells, and never see the credential value. Two verbs: register() (asks the user once) and ordinary inference cells.

Container specifics — image, registry login, internal port, cache mount target, readiness route — come from the model's own runbook skill; this skill is the translation contract.

Calling a registered endpoint — inference cells

Calling a registered endpoint — use the using-model-endpoint skill (this skill is the REGISTRATION contract; that one documents the call side in full).

The ONLY dispatch form is the compute_provider tool with the endpoint's registered name (list_compute shows them):

compute_provider(provider="boltz2-service", code="""
import requests
r = requests.post(BASE_URL + "/v1/infer", json=payload)
""")

The daemon brings the model up on demand (a first cold start downloads image + weights — minutes; let it run) and preloads BASE_URL into the cell — both as a Python variable (use it directly, as above) and as os.environ["BASE_URL"] (plus INFER_API_KEY for remote endpoints). Endpoints are not kernel environments: environment="boltz2-service" on a plain python cell fails — plain cells get no BASE_URL.

Enablement — once per machine

The user connects the family under Customize → Compute → Model endpoints (the setup flow saves the family credential first — connect-without-key is not a state) and picks ONE mode: Local (container registrations) or URL/remote (https against the configured host). Until connected, free_port()/register() raise a precise error — relay it; in the wrong mode they refuse with a teaching error naming the setting (existing endpoints of the unarmed leg keep dispatching — only NEW registrations refuse). Disconnecting is a full teardown: every active local service is stopped via its approved stop script and every registration (local AND hosted) is removed; caches stay on disk; a failing stop keeps that one row, FAILED. The "Local machine GPU" toggle never gates registration — it governs cell GPU access only; the approval card is the per-registration gate.

Credential contract (platform rule): every registration passes credential="NVIDIA_API_KEY" — the daemon rejects any other name. Locally the value feeds the start script's registry login and never enters your kernel env; for remote endpoints it authenticates the upstream and is delivered only into the inference cell's env (as INFER_API_KEY), never the repl kernel.

Register (repl kernel)

python
port = host.model_endpoints.free_port()        # local only; random 20000-29999

host.model_endpoints.register(
    name="boltz2-service",            # <model>-service -- descriptive, never the
                                      # bare model name (collides/ambiguous)
    url=f"http://127.0.0.1:{port}",   # LITERAL 127.0.0.1 -- `localhost` rejected
    credential="NVIDIA_API_KEY",      # the family credential NAME, never a value
    skill="<model-runbook-skill>",
    start=START_SCRIPT,               # composed below
    stop="docker stop boltz2-service",# exit 0 ONLY once actually stopped
    live="/v1/health/ready",          # readiness ROUTE (200 = model answers;
                                      # "up but loading" must read not-ready)
)

Name endpoints <model>-service (e.g. diffdock-service) — unambiguous in provider lists; never just the bare model name. Name the CONTAINER after the endpoint too (the template above does): the UI then follows the service's own logs live while it starts.

register() always cards the user (scripts verbatim, port, service dir, credential name). One exception: a byte-identical re-registration is silent — same bytes are approved forever; any byte change re-cards. The registration stays inspectable under Customize → Compute. Re-registering to fix scripts: reuse the existing url — never call free_port() again (the port is the endpoint's stable mutex).

Remote endpoints — upstream APIs (no lifecycle)

Pass url="https://<upstream>" and omit start/stop/live — no port, no scripts, no readiness. Requires URL/remote mode (the setup radio; in Local mode https registrations refuse). The url's HOST must equal the configured upstream host exactly — you pick the path leaf, never the authority. After approval, cells are plain HTTP clients of BASE_URL authenticating with $INFER_API_KEY. list_compute labels every row location: "local" | "remote".

Show full SKILL.md (440 more words)Show less

Composing the start script

The daemon hands scripts three things in their process environment (never argv, never sudo): HOST_PORT (the registered port), SERVICE_DIR (this endpoint's persistent directory — put the model cache here), and the credential value under its own name. Nothing else is inherited — ambient tokens are not visible; the ONLY secret a script sees is its registered credential.

The start script must be idempotent (cold create / warm start / crash re-entry), with the port-mismatch guard — the runtime freezes port mappings at container creation, so a container created under an OLD port must be recreated or readiness can never pass:

bash
mkdir -p "$SERVICE_DIR/cache"
# docker login persists auth in $DOCKER_CONFIG/config.json; scope it to the
# service dir so the credential dies with the service (never ~/.docker).
export DOCKER_CONFIG="$SERVICE_DIR/.docker"
create_service() {
  docker run -d --name boltz2-service \
    --restart unless-stopped \
    -p 127.0.0.1:${HOST_PORT}:8000 --gpus all \
    -e NVIDIA_API_KEY \
    -v "$SERVICE_DIR/cache:<cache target from the runbook>" \
    <image from the runbook>
}
if docker inspect boltz2-service >/dev/null 2>&1 && \
   [ "$(docker inspect -f '{{(index (index .HostConfig.PortBindings "8000/tcp") 0).HostPort}}' boltz2-service)" != "$HOST_PORT" ]; then
  docker rm -f boltz2-service          # stale port mapping -- recreate below
fi
if docker inspect boltz2-service >/dev/null 2>&1; then
  docker start boltz2-service          # warm wake -- no credential, no chown needed
else
  echo "$NVIDIA_API_KEY" | docker login <registry> --username '<user>' --password-stdin
  docker pull <image from the runbook>
  # Cache must be writable by the CONTAINER's user, whose uid the image
  # defines (container uid != host uid). chown needs root the script doesn't
  # have; a throwaway root container does it -- and the chmod, which the
  # host user can no longer do once the dir is chowned away -- without sudo.
  CUID="$(docker inspect --format '{{.Config.User}}' <image from the runbook> 2>/dev/null | cut -d: -f1)"
  case "$CUID" in ''|root) CUID=0;; *[!0-9]*) CUID=1000;; esac   # named user -> default 1000; runbook may override
  if [ "$CUID" != "0" ]; then
    docker run --rm -v "$SERVICE_DIR/cache:/c" alpine sh -c "chown -R $CUID:$CUID /c && chmod 700 /c"
  else
    chmod 700 "$SERVICE_DIR/cache" 2>/dev/null || true
  fi
  create_service
fi
# RUNTIME-binding guard (one retry): after a port-conflict crash the engine
# can start the container yet silently skip port programming -- the CONFIG
# still matches $HOST_PORT (so the guard above cannot catch it) but
# `docker port` prints nothing and the model serves to nobody. Recreate.
if [ -z "$(docker port boltz2-service 2>/dev/null)" ]; then
  docker rm -f boltz2-service
  create_service
fi

Translation rules:

  • Keep scripts ASCII — non-ASCII (em dashes, arrows, curly quotes) triggers the approval card's spoofing warning; use -- and -> in comments.
  • export DOCKER_CONFIG="$SERVICE_DIR/.docker" before any docker login — login persists the credential in config.json, and scoping it to the service dir means Remove honestly reclaims it (never ~/.docker, which outlives stop/Remove/Disable).
  • -p 127.0.0.1:${HOST_PORT}:<internal> — loopback-only publish; the internal port comes from the runbook.
  • If the image reads a different env name, bridge env→env at the top: export OTHER_NAME="$NVIDIA_API_KEY" (never argv, never a file).
  • -e NAME bare (argv is world-readable); the key rides the login stdin pipe only.
  • -d, no --rm — managed containers are stopped, never removed: stop parks them with weights loaded; --rm throws the cache away.
  • Cache under $SERVICE_DIR, owned by the container's uid: the runbook states it when it matters; otherwise derive it post-pull with docker inspect --format '{{.Config.User}}' <image> (empty or root ⇒ runs as root, no chown needed; a NAMED user can't be resolved without running the image — default 1000). Getting it wrong is the cache-empty symptom: the container can't write the mount, weights leak into the writable layer and die on recreate (or the image crash-loops on Permission denied). Never 777 — world-writable cache on a multi-user host. The mount TARGET comes from the runbook.
  • Cells need no auth header against local endpoints — the credential is a pull key that never enters your kernel.

Failures

A failed start/stop flips the endpoint FAILED (transcript on the endpoint panel — never echoed into cell errors; ask the user to read it there) and your cell errors with the daemon's one-line cause. FAILED is sticky: further cells fail fast until the user presses Stop or you re-register (byte-identical re-register also clears it). If a stop is stuck (exit 0 but the port never frees), removal is refused while the port is bound — recover out-of-band; the daemon absorbs the freed port on its next probe. A first-ever cold start downloads image + weights — minutes, once; the cell streams the phase lines live and the endpoint detail view streams the full script output, so let it run.

© JimLiu, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/managed-model-endpoints of JimLiu/science-skills.

Open the folder on GitHubat commit fb309c3

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in JimLiu/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Managed Model Endpoints next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Managed Model Endpoints compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Managed Model Endpoints this skillJimLiu/science-skills2272 repos~2.7kAutomated safety check: PassApache-2.0
Trader Memory Coretradermonty/claude-trading-skills3k3 repos~4.3kAutomated safety check: PassMIT
OpenRig Upgrade Proceduremvschwarz/openrig5.9k1 repos~2.9kAutomated safety check: PassApache-2.0
Author Migrationnrwl/nx29k—~12kAutomated safety check: NotesMIT
Write Notes Like Deepseekczm15053/write-notes-like-deepseek484—~2kAutomated safety check: PassNone
GreptimeDB Release RunbookGreptimeTeam/greptimedb6.7k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Trader Memory Core

    tradermonty/claude-trading-skills

    Track investment theses across their lifecycle — from screening idea to closed position with postmortem.

    3k GitHub starsUsed in 3 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • OpenRig Upgrade Procedure

    mvschwarz/openrig

    Walks an agent through upgrading the OpenRig CLI and daemon one observed step at a time, keeping live seats alive and reconciling managed plugin files.

    5.9k GitHub starsUsed in 1 repo~2.9k tokens
    DevOps & CloudAuto-check passed
  • Author or scope a first-party Nx migration. An agent skill from nrwl/nx.

    29k GitHub stars~12k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Write Notes Like Deepseek

    czm15053/write-notes-like-deepseek

    A skill your agent uses when a change is non-trivial by DSH standards (behavior, architecture, cross-file contracts, process/tooling, testing strategy, or on-disk/wire/config formats), when choosing…

    484 GitHub stars~2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • GreptimeDB Release Runbook

    GreptimeTeam/greptimedb

    Runbook for publishing a GreptimeDB version: pick the release branch, verify the Cargo version, then tag, create the GitHub release and open the docs note PR.

    6.7k GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Statem

    henryqin1997/statem

    A skill your agent uses when a long coding or research task should be managed with statem state-machine runbooks, including creating specs, starting or resuming runs, checking current state…

    1.3k GitHub stars~1.2k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from JimLiu/science-skills

All 27 skills in this repo
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    Auto-check passed
  • Compute Env Setup

    JimLiu/science-skills

    Set up a compute environment on a remote provider so Claude Science jobs can run there.

    227 GitHub starsUsed in 2 repos~4.4k tokens
    Auto-check passed
  • Borzoi

    JimLiu/science-skills

    Predict genome-wide functional tracks (RNA-seq, CAGE, DNase, ChIP) from DNA sequence with Borzoi.

    227 GitHub starsUsed in 4 repos~973 tokens
    Auto-check passed
  • Evo2

    JimLiu/science-skills

    Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check passed
  • Fair Esm2

    JimLiu/science-skills

    Embed proteins with Meta AI's ESM-2 (fair-esm package). An agent skill from JimLiu/science-skills.

    227 GitHub starsUsed in 4 repos~1.2k tokens
    Auto-check passed
  • Openfold3

    JimLiu/science-skills

    Structure prediction using OpenFold3, an open-weights PyTorch reproduction of AlphaFold3 from the AlQuraishi Lab.

    227 GitHub starsUsed in 4 repos~1.8k tokens
    Auto-check passed

Categories

Questions about Managed Model Endpoints

What does Managed Model Endpoints do?

Register a model service in the managed family — a local model server container the daemon starts/stops on demand, or a remote upstream model API (https). Managed Model Endpoints is an agent skill from JimLiu/science-skills. Register a model service in the managed family — a local model server container the daemon starts/stops on demand, or a remote upstream model API (https).

When should I use Managed Model Endpoints?

Managed Model Endpoints fits situations like: wants a model service available for inference; listcompute shows managed endpoints.

How do I install Managed Model Endpoints in Claude Code?

Run `npx skills add JimLiu/science-skills --skill managed-model-endpoints -a claude-code`. Or copy the skill folder (skills/managed-model-endpoints in JimLiu/science-skills) into .claude/skills/managed-model-endpoints in your project. Claude Code loads it when a task matches its description.

How do I install Managed Model Endpoints in Codex?

Run `npx skills add JimLiu/science-skills --skill managed-model-endpoints -a codex`. Or copy the skill folder (skills/managed-model-endpoints in JimLiu/science-skills) into .agents/skills/managed-model-endpoints in your project. Codex loads it when a task matches its description.

Can I use Managed Model Endpoints in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JimLiu/science-skills --skill managed-model-endpoints -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/managed-model-endpoints, .gemini/skills/managed-model-endpoints, .github/skills/managed-model-endpoints and .opencode/skills/managed-model-endpoints in your project.

What does Managed Model Endpoints need to run?

Going by SKILL.md and its folder, Managed Model Endpoints needs the command-line tools its instructions call (docker) and credentials named NVIDIA_API_KEY and INFER_API_KEY. Our summary lists: Python 3; Docker; A credential in INFER_API_KEY; A credential in NVIDIA_API_KEY.

Does Managed Model Endpoints access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Managed Model Endpoints safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Managed Model Endpoints use?

Managed Model Endpoints is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Managed Model Endpoints use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Managed Model Endpoints?

Skills that share tags, products or a category with Managed Model Endpoints: Trader Memory Core (tradermonty/claude-trading-skills, 3k stars), OpenRig Upgrade Procedure (mvschwarz/openrig, 5.9k stars), Author Migration (nrwl/nx, 29k stars) and Write Notes Like Deepseek (czm15053/write-notes-like-deepseek, 484 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Managed Model Endpoints?

JimLiu (a GitHub user) maintains it in JimLiu/science-skills, which has 227 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on July 1, 2026.

Source: JimLiu/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.