Agent skill

Runpod

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…

MITAuto-check passedAI & LLM Engineering

Install Runpod

skills CLI
$ npx skills add ericrisco/rsc-harness --skill runpod -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness runpod --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/runpod .claude/skills/runpod && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
runpod
GitHub stars
156
Token cost
~2.8k tokens
SKILL.md length
1,265 words
Files
6 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…

  • Works in 2 steps: A volume locks the endpoint/pod to one… → Idle volume cost is double. Pod volume…
  • Running GPU compute on RunPod and deciding between Pods (hourly
  • SKILL.md covers When to use, When NOT to use, Decision: Pod vs Serverless and Serverless worker handler, plus 6 more sections
  • Runs Shell scripts from its folder; calls python and bash; needs RUNPOD_API_KEY

What it does

Runpod is an agent skill from ericrisco/rsc-harness. Use when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless endpoints, handler workers, worker Docker templates, network volumes, timeout and worker-count tuning, cold starts, and runaway bills. NOT Python-native serverless GPU with snapshot autoscaling (that is modal), NOT calling hosted prebuilt model APIs (that is replicate), NOT pulling weights from the Hub (that is huggingface).

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/cost-and-scaling.md`).

It sits in AI & LLM Engineering, covering Serverless, GPU and accelerator computing and Model hubs and datasets. It works with Docker, Hugging Face and Python. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Running GPU compute on RunPod and deciding between Pods (hourly
  • Always-on) and Serverless (per-second
  • Autoscaling) for training
  • Inference — serverless endpoints

Example prompts

  • “/runpod”

Requirements

  • Python 3
  • A Bash shell
  • Docker
  • A credential in RUNPOD_API_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. A volume locks the endpoint/pod to one data center and adds network latency. That
  2. Idle volume cost is double. Pod volume disk bills ~$0.10/GB/mo while running but

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • RUNPOD_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Runpod loads about 2.8k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 129 tokens; SKILL.md has 1,265 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~129
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,265 words, ~2,838 tokens.

Download SKILL.mdSave it as .claude/skills/runpod/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
runpod
description
Use when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless endpoints, handler workers, worker Docker templates, network volumes, timeout and worker-count tuning, cold starts, and runaway bills. NOT Python-native serverless GPU with snapshot autoscaling (that is `modal`), NOT calling hosted prebuilt model APIs (that is `replicate`), NOT pulling weights from the Hub (that is `huggingface`).
tags
runpod, gpu, serverless, inference, training, cost-control, vllm, ai-infra
recommends
modal, replicate, huggingface, docker, llm-pipeline, cost-tracking, ollama
origin
risco

RunPod: GPU compute, two products, one bill

RunPod sells GPU time two ways and they bill on opposite philosophies. Get the choice wrong and you either pay a steep premium for idle work or you pay 24/7 for a box that sits warm doing nothing. Everything below is the RunPod-specific operational playbook: which product a workload belongs on, how to write a worker that does not waste cold-start seconds, and which knobs actually move the number on the invoice.

The two products:

  • Pods — rent a GPU container by the hour. It runs continuously while it is up, billed every hour whether busy or idle. Your dev box, your training job, your Jupyter.
  • Serverless — per-second autoscaling workers. Billed only while a worker is actually running a job, from worker start to full stop, rounded up to the second. Costs roughly 2-3x the equivalent hourly pod rate, but idle gaps cost nothing on flex workers.

RunPod charges zero egress/ingress fees — bandwidth in and out is free, unlike the hyperscalers. That removes one variable from cost math: you only reason about GPU-seconds and storage.

When to use

  • Deploying inference as a Serverless endpoint (vLLM LLM serving, image gen, custom model).
  • Running a training / fine-tuning job or a Jupyter / dev box on a GPU Pod.
  • Writing or debugging a serverless worker handler (handler(job), async, streaming).
  • Building a custom worker Docker template / image, or wiring a network volume.
  • Cost control on RunPod: cold starts, idle bleed, runaway max workers, idle volume charges.
  • Picking a GPU SKU and the pod-vs-serverless tradeoff for a given workload.

When NOT to use

  • Python-native serverless GPU with snapshot autoscaling and no Dockerfile → that is modal.
  • Calling hosted, prebuilt model endpoints you do not host → that is replicate.
  • Pulling/pushing weights, datasets, model cards on the Hub → that is huggingface.
  • Running models locally on your own machine → that is ollama.
  • Provider-agnostic cross-cloud spend dashboards → that is cost-tracking.
  • Generic container packaging → that is docker (here we cover only the RunPod image shape).

Decision: Pod vs Serverless

Workload shapePickWhy
Training / fine-tuning, multi-hour runsPodServerless premium + the 600s default execution timeout kill long jobs. You want the box continuously.
Interactive dev / Jupyter / notebooksPodYou need it now and responsive; per-second autoscaling adds cold-start latency for nothing.
Bursty inference with real idle gapsServerless (flex)Idle costs nothing on flex; you pay only for the seconds a request runs.
24/7 steady high-QPS inferenceCompareActive serverless (40% off flex) vs a dedicated Pod. Past roughly 60% utilization a Pod usually wins.

Rule: if the GPU would sit busy more than ~60% of the time, a Pod is cheaper than serverless even with active-worker discount — model the two before committing.

GPU SKU quick-pick (2026 pod $/hr, lower on Community Cloud)
GPUVRAMPod ~$/hrUse for
L424GB$0.39small models, light inference
A4048GB$0.44mid-size inference, budget training
RTX 409024GB$0.697B-class inference, fast/cheap
L40S48GB$0.8613B inference, image gen
A100 80GB80GB$1.39training, large-batch inference
H100 PCIe80GB$2.89the biggest models / fastest training

Rule: pick the smallest GPU the model fits in VRAM. Defaulting to H100 is up to ~7x the cost for zero speedup when the workload is memory-bound and fits on an L40S or 4090.

Serverless worker handler

A worker is a Python file using the runpod SDK. The minimum: a function that reads job["input"], returns a dict, and is registered with runpod.serverless.start.

python
import runpod

def handler(job):
    job_input = job["input"]
    prompt = job_input["prompt"]
    # ... run the model ...
    return {"output": f"echo: {prompt}"}

runpod.serverless.start({"handler": handler})

Async handler (for awaiting model calls) and a streaming generator both work:

python
import runpod

async def handler(job):
    job_input = job["input"]
    return {"output": await run_model(job_input)}

# Streaming: yield chunks from an async generator instead of returning once.
async def stream_handler(job):
    async for token in generate(job["input"]["prompt"]):
        yield {"token": token}

runpod.serverless.start({"handler": stream_handler, "return_aggregate_stream": True})

Test locally before you push an image. A broken handler still burns build minutes and cold-start seconds when discovered on the platform.

bash
# One-shot: reads ./test_input.json, runs the handler once, prints output.
python worker.py

# HTTP server emulating the real endpoint at http://localhost:8000.
python worker.py --rp_serve_api

Concurrency (concurrency_modifier), job cancel, refresh-worker, and the run / runsync / stream / status / cancel / health HTTP endpoints live in references/serverless-workers.md.

Templates & images

A custom template is a Docker image plus environment variables. Pin the base; an unpinned or :latest base re-pulls on cold start and lengthens it.

dockerfile
# Pin the CUDA base — never bare :latest.
FROM runpod/base:0.6.2-cuda12.4.1

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY handler.py .

# The worker entrypoint runs the handler.
CMD ["python", "-u", "handler.py"]

vLLM shortcut. For OpenAI-compatible LLM serving, use the prebuilt worker-vllm image instead of writing a handler. Its AsyncEngineArgs are set via UPPERCASE env vars:

bash
# Template env vars on the endpoint — these map to vLLM AsyncEngineArgs.
MODEL_NAME=mistralai/Mistral-7B-Instruct-v0.3
MAX_MODEL_LEN=8192

Network volumes

Attach a network volume when data outgrows what you want in the image:

  • Datasets larger than container disk.
  • Weights shared across many workers (download once, mount everywhere).
  • Checkpoints that must survive a worker / pod restart.

Two traps to state up front:

  1. A volume locks the endpoint/pod to one data center and adds network latency. That shrinks the pool of available GPUs in that DC — you can get stuck waiting for capacity.
  2. Idle volume cost is double. Pod volume disk bills ~$0.10/GB/mo while running but ~$0.20/GB/mo while the pod is stopped. A forgotten volume on a stopped pod bleeds money.

Rule: bake small, static weights into the image; reserve volumes for large mutable data (datasets, checkpoints). Full storage tables are in references/cost-and-scaling.md.

Show full SKILL.md (468 more words)Show less

Cost control — the four knobs that move the bill

KnobDefaultWhat it does
Idle Timeout5sHow long a worker stays warm (and billed) after a job. Lower for spiky traffic; raise to dodge repeated cold starts.
Execution Timeout600sMax single-job duration (range 5s–7 days). Set it so a hung job cannot run for days.
Max Workers—Your concurrency cap and cost ceiling. Never leave it sky-high; set ~20% over expected peak.
Active Workers0Always-warm minimum: zero cold start but billed 24/7 (at ~40% off the flex rate). Use only when a latency SLA demands it.

Plus FlashBoot: enable it on flex workers to cut cold start (model load into GPU memory) toward sub-200ms by caching, so flex stops feeling slow.

Worked example — 100k requests/day, 2s each on RTX 4090 serverless ($1.10/hr equiv):

  • Compute is ~55.5 GPU-hours/day regardless of mode → roughly $61/day of actual work.
  • Flex adds idle-timeout tails per cold worker but nothing during true idle — best for bursty business-hours traffic.
  • Active workers remove cold starts but bill 24/7; only worth it if traffic is steady enough that the 40% discount beats paying for idle.
  • A dedicated Pod at ~$0.69/hr = ~$497/mo flat — wins only if utilization stays high.

Full active-vs-flex math and monthly scenarios: references/cost-and-scaling.md.

Driving resources

runpodctl is the open-source CLI. It outputs JSON by default (agent-friendly); add --output table or --output yaml for humans. Pods ship with it pre-installed using a pod-scoped key.

bash
runpodctl serverless list                 # JSON by default
runpodctl serverless get <endpoint-id>
runpodctl serverless update <endpoint-id> --output table
runpodctl get pod --output table

Python SDK for programmatic control — key from env, never in source:

python
import os, runpod

runpod.api_key = os.environ["RUNPOD_API_KEY"]  # never a literal

pod = runpod.create_pod(name="train", image_name="my/img:1.0", gpu_type_id="NVIDIA A100 80GB PCIe")
runpod.stop_pod(pod["id"])      # also resume_pod / terminate_pod

ep = runpod.Endpoint("<endpoint-id>")
job = ep.run({"prompt": "hi"})  # async: job.status(), job.output()
out = ep.run_sync({"prompt": "hi"})  # blocks, ~90s max

Rule: read the API key from RUNPOD_API_KEY. A leaked key is a stranger spending on your GPUs.

Anti-patterns

BadGoodWhy
Max Workers left unbounded / sky-highBound it ~20% over peakOne traffic spike scales to an unbounded bill.
H100 by defaultSmallest GPU that fits VRAM~7x cost for zero speedup when the model fits an L40S/4090.
6-hour training run on ServerlessRun it on a PodServerless premium + 600s execution timeout kills long jobs.
Hardcoded API key in sourceos.environ["RUNPOD_API_KEY"]A leaked key = a stranger's GPU bill on your card.
Weights downloaded at cold startBake into image or mount a volumeEvery cold start re-downloads and pays for the wait.
Active workers "just in case"Flex + FlashBootActive bills 24/7; FlashBoot makes flex cold starts cheap.
Volume left on a stopped podDetach / delete when idleIdle volume bills ~$0.20/GB/mo silently.
Push image, debug on the platformpython worker.py --rp_serve_api firstDebugging on cold-start seconds is slow and costs money.

Verify

Run bash scripts/verify.sh <worker-dir> to statically lint a worker directory: it checks for runpod.serverless.start, a handler with a return/yield, no hardcoded API key, a bounded max_workers plus timeout keys in any config, and a pinned FROM + CMD in a Dockerfile. Pure grep/parse — no network, no RunPod account. It exits 0 on an empty dir.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/runpod of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/cost-and-scaling.md
  • references/serverless-workers.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Runpod next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Runpod compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Runpod this skillericrisco/rsc-harness156—~2.8kAutomated safety check: PassMIT
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0
ModalK-Dense-AI/scientific-agent-skills48k1 repos~4.5kAutomated safety check: NotesApache-2.0
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Generate Openenv Envadithya-s-k/FineEnvs421—~2.4kAutomated safety check: PassApache-2.0
bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs13k3 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Modal

    K-Dense-AI/scientific-agent-skills

    Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs.

    48k GitHub starsUsed in 1 repo~4.5k tokens
    Backend & APIsAuto-check: notes
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Generate Openenv Env

    adithya-s-k/FineEnvs

    Builds an OpenEnv (Hugging Face) variant of an RL environment.

    421 GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • bitsandbytes Model Quantization

    Orchestra-Research/AI-Research-SKILLs

    Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

    13k GitHub starsUsed in 3 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face ZeroGPU

    huggingface/skills

    Official

    Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits.

    11k GitHub starsUsed in 2 repos~4.6k tokens
    AI & LLM EngineeringAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed

Questions about Runpod

What does Runpod do?

A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…. Runpod is an agent skill from ericrisco/rsc-harness. Use when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless endpoints, handler workers, worker Docker templates, network volumes, timeout and worker-count tuning, cold starts, and runaway bills.

When should I use Runpod?

Runpod fits situations like: running GPU compute on RunPod and deciding between Pods (hourly; always-on) and Serverless (per-second; autoscaling) for training; inference — serverless endpoints.

How do I install Runpod in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill runpod -a claude-code`. Or copy the skill folder (skills/runpod in ericrisco/rsc-harness) into .claude/skills/runpod in your project. Claude Code loads it when a task matches its description.

How do I install Runpod in Codex?

Run `npx skills add ericrisco/rsc-harness --skill runpod -a codex`. Or copy the skill folder (skills/runpod in ericrisco/rsc-harness) into .agents/skills/runpod in your project. Codex loads it when a task matches its description.

Can I use Runpod in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill runpod -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/runpod, .gemini/skills/runpod, .github/skills/runpod and .opencode/skills/runpod in your project.

What does Runpod need to run?

Going by SKILL.md and its folder, Runpod needs a shell for the scripts in its folder, the command-line tools its instructions call (python and bash) and credentials named RUNPOD_API_KEY. Our summary lists: Python 3; A Bash shell; Docker; A credential in RUNPOD_API_KEY.

Does Runpod access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Runpod safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Runpod use?

Runpod is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Runpod use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Runpod?

Skills that share tags, products or a category with Runpod: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Modal (K-Dense-AI/scientific-agent-skills, 48k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and Generate Openenv Env (adithya-s-k/FineEnvs, 421 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Runpod?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.