Agent skill

Serve Model Vibe Test

by open-thoughts in open-thoughts/OpenThoughts-Agent

Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API.

Apache-2.0Auto-check: warningsAI & LLM Engineering

Install Serve Model Vibe Test

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill serve-model-vibe-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent serve-model-vibe-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/serve-model-vibe-test .claude/skills/serve-model-vibe-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
serve-model-vibe-test
GitHub stars
301
Token cost
~1.8k tokens
SKILL.md length
597 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API.

  • Asked to serve a model for testing
  • SKILL.md covers ⚠️ Security — read first, Prerequisites, Pick an unused Pinggy pair and Launch (the one-liner), plus 3 more sections
  • Calls curl, python and uv
  • Throw up a public endpoint

What it does

Serve Model Vibe Test is an agent skill from open-thoughts/OpenThoughts-Agent. Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API. Use when asked to "serve a model for testing", "throw up a public endpoint", "let people play with model X", or to demo a checkpoint. Combines marin-serve (marin6556) + a Pinggy tunnel from our endpoint bank.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Meeting notes and agendas. It works with OpenAI and vLLM. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Asked to serve a model for testing
  • Throw up a public endpoint
  • Let people play with model X
  • Demo a checkpoint

Example prompts

  • “serve a model for testing”
  • “throw up a public endpoint”
  • “let people play with model X”
  • “/serve-model-vibe-test”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • python
    • uv
    • ssh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, uv and ssh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Serve Model Vibe Test loads about 1.8k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 597 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:29
    _SECRET_ENV` sourced (Pinggy uses your `~/.ssh` identity; SSH to `pro.pinggy.io:443`).

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 597 words, ~1,801 tokens.

Download SKILL.mdSave it as .claude/skills/serve-model-vibe-test/SKILL.md (or your agent's skills folder).
name
serve-model-vibe-test
description
Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API. Use when asked to "serve a model for testing", "throw up a public endpoint", "let people play with model X", or to demo a checkpoint. Combines marin-serve (marin#6556) + a Pinggy tunnel from our endpoint bank.

serve-model-vibe-test

Spin up a public endpoint for a HuggingFace repo or gs:// checkpoint. marin-serve runs vLLM on a single-host Iris TPU slice behind the controller proxy; a Pinggy tunnel makes its proxied port public at https://<id>.a.pinggy.link/proxy/serve.<ep>/, with a dashboard and OpenAI-compatible API.

scripts/inference/serve_public.py wraps marin-serve, hpc.pinggy_utils.PinggyTunnel, and the Pinggy bank.

⚠️ Security — read first

The public URL is unauthenticated. Use it only for throwaway vibe testing; keep --timeout-hours short, use an unused Pinggy pair, and tear it down when done.

Prerequisites

  • marin-serve available (marin#6556). serve_public.py invokes ~/Documents/marin/.venv/bin/python -m marin.inference.quick_serve_cli by default. Override with --marin-serve-bin or MARIN_SERVE_BIN.
  • Run from the marin repo root (~/Documents/marin), not from this repo. marin-serve bundles Path.cwd() as the worker workspace (quick_serve_cli.py:227) and the worker runs uv sync --all-packages --extra tpu --extra vllm against it. Only the marin workspace defines the tpu/vllm extras (on marin-core); any other CWD → Extra 'tpu' is not defined in any project's 'optional-dependencies' table and the job dies before vLLM starts. serve_public.py resolves its own repo root via parents[2] for the hpc import, so invoke it by absolute path from the marin CWD.
  • Iris controller access (the laptop's step-ca/GCP creds — same as iris CLI).
  • $DC_AGENT_SECRET_ENV sourced (Pinggy uses your ~/.ssh identity; SSH to pro.pinggy.io:443).
  • Pinggy bank at ~/Documents/notes/ot-agent/pinggy_bank.md (override --pinggy-bank / PINGGY_BANK).
  • Single-host TPU slices only (v6e-8, v6e-4, v5litepod-8, …); multi-host is rejected.

Pick an unused Pinggy pair

The bank has 10 pairs. Before launching, pick one that isn't already serving — a quick probe (a free pair returns a Pinggy "tunnel not found"/503 or connection error; an in-use one returns your model):

bash
for i in 1 2 3; do u=$(sed -n "/## Pair $i$/,/pinggy.link/p" ~/Documents/notes/ot-agent/pinggy_bank.md | grep pinggy.link); \
  echo "pair $i: $u -> $(curl -s -o /dev/null -w '%{http_code}' --max-time 6 https://$u/ || echo down)"; done

Use a pair that shows down/502/503 (free). Pass it as --pair N.

Launch (the one-liner)

bash
cd ~/Documents/marin                               # MUST be the marin repo — see Prerequisites
set -a; source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"; set +a
python -u ~/Documents/OpenThoughts-Agent/scripts/inference/serve_public.py <MODEL> --tpu <SLICE> [--region <R>] \
    [--chat-template <FILE-or-URL>] --pair <N> --timeout-hours 6

It launches marin-serve, waits for READY — dashboard: http://127.0.0.1:<port>/proxy/serve.<ep>/, opens the Pinggy tunnel, and prints:

  dashboard : https://<id>.a.pinggy.link/proxy/serve.<ep>/
  OpenAI    : https://<id>.a.pinggy.link/proxy/serve.<ep>/v1

Keep the process running because it holds both tunnels. For unattended operation, use tmux, not nohup setsid … on macOS:

bash
tmux new-session -d -s serve-public \
  "source \"${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV first}\" && cd ~/Documents/marin && \
   python -u ~/Documents/OpenThoughts-Agent/scripts/inference/serve_public.py <MODEL> ... 2>&1 | tee /tmp/serve_public.log"
# then poll /tmp/serve_public.log (or `tmux attach -t serve-public`) for the "OpenAI :" line

Worked example — the Delphi 9.7B SFT canary (marin#6545)

This is the reference model: laion/delphi-1e22-p33m67-32p07b-lr0_67-54770ae7-wc386k_lr1e5-sft. It needs its own chat template (the repo ships a plain Llama-3 one); use delphi_v0.jinja2. marin-serve auto-derives the 4k context (malformed RoPE) and TP=2 (30 heads on a 4-chip slice).

bash
# run from ~/Documents/marin (the bundled workspace); script is invoked by absolute path
cd ~/Documents/marin && source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV first}"
python -u ~/Documents/OpenThoughts-Agent/scripts/inference/serve_public.py \
    laion/delphi-1e22-p33m67-32p07b-lr0_67-54770ae7-wc386k_lr1e5-sft \
    --tpu v6e-4 --region europe-west4 \
    --chat-template https://raw.githubusercontent.com/open-thoughts/OpenThoughts-Agent/ed4d6f483151f14d6d78cf732f04cd3c8ff5c606/chat_templates/delphi_v0.jinja2 \
    --pair 1 --timeout-hours 6

Verify once the public URL prints:

bash
BASE=https://<id>.a.pinggy.link/proxy/serve.<ep>
# from the laptop you must force the real Pinggy edge IP (ISP DNS poisons *.a.pinggy.link); see Gotchas
curl "$BASE/v1/models"
curl "$BASE/v1/chat/completions" -H 'content-type: application/json' \
  -d '{"model":"laion/delphi-1e22-p33m67-32p07b-lr0_67-54770ae7-wc386k_lr1e5-sft",
       "messages":[{"role":"user","content":"Give me a fun fact about otters."}]}'

(Base/midtrained checkpoints with no chat template auto-default the dashboard to completion mode — omit --chat-template.)

Show full SKILL.md (229 more words)Show less

Teardown

  • Foreground: Ctrl-C (tears down the Pinggy tunnel and the marin-serve job it launched).
  • Detached/other host: iris job stop <job> --cluster marin (the job name is in the log / iris query), then kill the serve_public.py PID. The slice also self-stops at --timeout-hours.

Gotchas

  • Run from the marin repo or the build fails before vLLM starts. Path.cwd() becomes the worker workspace; only marin defines the needed extras. Invoke serve_public.py by absolute path from ~/Documents/marin.
  • Dead tunnel reported live (ps state T): run kill -CONT <ssh_pid> <loop_pid>, then re-probe. PinggyTunnel.start() prevents this with detached stdin, ssh -n, and setsid; serve_public.py prints LIVE ✓ only after a DNS-aware /v1/models HTTP 200 probe.
  • Cold compile: first boot of a model can take tens of minutes; --ready-timeout (default 2700s) bounds the wait.
  • Don't use marin-serve --no-wait for the public path — the local proxied port only exists while the launching process runs; --no-wait returns immediately and there's nothing to tunnel.
  • Pinggy pair collisions: two jobs on the same pair clobber each other — always pick a free pair.
  • Single-host concurrency: one slice = limited throughput; fine for a few testers, not a viral launch (that's inference-broker territory).
  • The public URL carries the /proxy/serve.<ep>/ base path (it's the controller proxy path), so the OpenAI base is https://<id>.a.pinggy.link/proxy/serve.<ep>/v1, not /v1.
  • Local DNS poisoning: verify through the real edge with dig @1.1.1.1 and curl --resolve, or use 1.1.1.1/8.8.8.8 DNS. See .agents/projects/pinggy/pinggy.md.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/serve-model-vibe-test of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Serve Model Vibe Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Serve Model Vibe Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Serve Model Vibe Test this skillopen-thoughts/OpenThoughts-Agent301—~1.8kAutomated safety check: WarnApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT
Vllm Bench Random Syntheticvllm-project/vllm-skills103—~1.5kAutomated safety check: PassApache-2.0
Vllm Bench Servevllm-project/vllm-skills103—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Bench Random Synthetic

    vllm-project/vllm-skills

    Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

    103 GitHub stars~1.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Bench Serve

    vllm-project/vllm-skills

    Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 11 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed

Works with

Questions about Serve Model Vibe Test

What does Serve Model Vibe Test do?

Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API. Serve Model Vibe Test is an agent skill from open-thoughts/OpenThoughts-Agent. Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API.

When should I use Serve Model Vibe Test?

Serve Model Vibe Test fits situations like: asked to serve a model for testing; throw up a public endpoint; let people play with model X; demo a checkpoint.

How do I install Serve Model Vibe Test in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill serve-model-vibe-test -a claude-code`. Or copy the skill folder (.agents/skills/serve-model-vibe-test in open-thoughts/OpenThoughts-Agent) into .claude/skills/serve-model-vibe-test in your project. Claude Code loads it when a task matches its description.

How do I install Serve Model Vibe Test in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill serve-model-vibe-test -a codex`. Or copy the skill folder (.agents/skills/serve-model-vibe-test in open-thoughts/OpenThoughts-Agent) into .agents/skills/serve-model-vibe-test in your project. Codex loads it when a task matches its description.

Can I use Serve Model Vibe Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill serve-model-vibe-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serve-model-vibe-test, .gemini/skills/serve-model-vibe-test, .github/skills/serve-model-vibe-test and .opencode/skills/serve-model-vibe-test in your project.

What does Serve Model Vibe Test need to run?

Going by SKILL.md and its folder, Serve Model Vibe Test needs the command-line tools its instructions call (curl, python, uv and ssh). Our summary lists: Python 3.

Does Serve Model Vibe Test access the network?

SKILL.md contains no URLs. Its commands use curl, uv and ssh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Serve Model Vibe Test safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Serve Model Vibe Test use?

Serve Model Vibe Test is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Serve Model Vibe Test use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Serve Model Vibe Test?

Skills that share tags, products or a category with Serve Model Vibe Test: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Bench Random Synthetic (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Serve Model Vibe Test?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.