Hugging Face Vision Trainer
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…
$ npx skills add ericrisco/rsc-harness --skill runpod -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness runpod --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/runpod .claude/skills/runpod && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "runpod" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/runpod into .claude/skills/runpod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runpod", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/runpodType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill runpod -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness runpod --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/runpod .agents/skills/runpod && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "runpod" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/runpod into .agents/skills/runpod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runpod", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill runpod -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness runpod --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/runpod .cursor/skills/runpod && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "runpod" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/runpod into .cursor/skills/runpod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runpod", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/runpod--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill runpod -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness runpod --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/runpod .gemini/skills/runpod && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "runpod" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/runpod into .gemini/skills/runpod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runpod", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness runpodInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill runpod -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/runpod .github/skills/runpod && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "runpod" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/runpod into .github/skills/runpod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runpod", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill runpod -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness runpod --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/runpod .opencode/skills/runpod && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "runpod" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/runpod into .opencode/skills/runpod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "runpod", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
runpodA skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…
Runpod is an agent skill from ericrisco/rsc-harness. Use when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless endpoints, handler workers, worker Docker templates, network volumes, timeout and worker-count tuning, cold starts, and runaway bills. NOT Python-native serverless GPU with snapshot autoscaling (that is modal), NOT calling hosted prebuilt model APIs (that is replicate), NOT pulling weights from the Hub (that is huggingface).
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/cost-and-scaling.md`).
It sits in AI & LLM Engineering, covering Serverless, GPU and accelerator computing and Model hubs and datasets. It works with Docker, Hugging Face and Python. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
pythonbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
RUNPOD_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Runpod loads about 2.8k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 129 tokens; SKILL.md has 1,265 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,265 words, ~2,838 tokens.
.claude/skills/runpod/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.RunPod sells GPU time two ways and they bill on opposite philosophies. Get the choice wrong and you either pay a steep premium for idle work or you pay 24/7 for a box that sits warm doing nothing. Everything below is the RunPod-specific operational playbook: which product a workload belongs on, how to write a worker that does not waste cold-start seconds, and which knobs actually move the number on the invoice.
The two products:
RunPod charges zero egress/ingress fees — bandwidth in and out is free, unlike the hyperscalers. That removes one variable from cost math: you only reason about GPU-seconds and storage.
handler(job), async, streaming).modal.replicate.huggingface.ollama.cost-tracking.docker (here we cover only the RunPod image shape).| Workload shape | Pick | Why |
|---|---|---|
| Training / fine-tuning, multi-hour runs | Pod | Serverless premium + the 600s default execution timeout kill long jobs. You want the box continuously. |
| Interactive dev / Jupyter / notebooks | Pod | You need it now and responsive; per-second autoscaling adds cold-start latency for nothing. |
| Bursty inference with real idle gaps | Serverless (flex) | Idle costs nothing on flex; you pay only for the seconds a request runs. |
| 24/7 steady high-QPS inference | Compare | Active serverless (40% off flex) vs a dedicated Pod. Past roughly 60% utilization a Pod usually wins. |
Rule: if the GPU would sit busy more than ~60% of the time, a Pod is cheaper than serverless even with active-worker discount — model the two before committing.
| GPU | VRAM | Pod ~$/hr | Use for |
|---|---|---|---|
| L4 | 24GB | $0.39 | small models, light inference |
| A40 | 48GB | $0.44 | mid-size inference, budget training |
| RTX 4090 | 24GB | $0.69 | 7B-class inference, fast/cheap |
| L40S | 48GB | $0.86 | 13B inference, image gen |
| A100 80GB | 80GB | $1.39 | training, large-batch inference |
| H100 PCIe | 80GB | $2.89 | the biggest models / fastest training |
Rule: pick the smallest GPU the model fits in VRAM. Defaulting to H100 is up to ~7x the cost for zero speedup when the workload is memory-bound and fits on an L40S or 4090.
A worker is a Python file using the runpod SDK. The minimum: a function that reads
job["input"], returns a dict, and is registered with runpod.serverless.start.
import runpod
def handler(job):
job_input = job["input"]
prompt = job_input["prompt"]
# ... run the model ...
return {"output": f"echo: {prompt}"}
runpod.serverless.start({"handler": handler})Async handler (for awaiting model calls) and a streaming generator both work:
import runpod
async def handler(job):
job_input = job["input"]
return {"output": await run_model(job_input)}
# Streaming: yield chunks from an async generator instead of returning once.
async def stream_handler(job):
async for token in generate(job["input"]["prompt"]):
yield {"token": token}
runpod.serverless.start({"handler": stream_handler, "return_aggregate_stream": True})Test locally before you push an image. A broken handler still burns build minutes and cold-start seconds when discovered on the platform.
# One-shot: reads ./test_input.json, runs the handler once, prints output.
python worker.py
# HTTP server emulating the real endpoint at http://localhost:8000.
python worker.py --rp_serve_apiConcurrency (concurrency_modifier), job cancel, refresh-worker, and the run / runsync /
stream / status / cancel / health HTTP endpoints live in
references/serverless-workers.md.
A custom template is a Docker image plus environment variables. Pin the base; an unpinned
or :latest base re-pulls on cold start and lengthens it.
# Pin the CUDA base — never bare :latest.
FROM runpod/base:0.6.2-cuda12.4.1
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY handler.py .
# The worker entrypoint runs the handler.
CMD ["python", "-u", "handler.py"]vLLM shortcut. For OpenAI-compatible LLM serving, use the prebuilt worker-vllm image
instead of writing a handler. Its AsyncEngineArgs are set via UPPERCASE env vars:
# Template env vars on the endpoint — these map to vLLM AsyncEngineArgs.
MODEL_NAME=mistralai/Mistral-7B-Instruct-v0.3
MAX_MODEL_LEN=8192Attach a network volume when data outgrows what you want in the image:
Two traps to state up front:
Rule: bake small, static weights into the image; reserve volumes for large mutable data (datasets, checkpoints). Full storage tables are in references/cost-and-scaling.md.
| Knob | Default | What it does |
|---|---|---|
| Idle Timeout | 5s | How long a worker stays warm (and billed) after a job. Lower for spiky traffic; raise to dodge repeated cold starts. |
| Execution Timeout | 600s | Max single-job duration (range 5s–7 days). Set it so a hung job cannot run for days. |
| Max Workers | — | Your concurrency cap and cost ceiling. Never leave it sky-high; set ~20% over expected peak. |
| Active Workers | 0 | Always-warm minimum: zero cold start but billed 24/7 (at ~40% off the flex rate). Use only when a latency SLA demands it. |
Plus FlashBoot: enable it on flex workers to cut cold start (model load into GPU memory) toward sub-200ms by caching, so flex stops feeling slow.
Worked example — 100k requests/day, 2s each on RTX 4090 serverless ($1.10/hr equiv):
Full active-vs-flex math and monthly scenarios: references/cost-and-scaling.md.
runpodctl is the open-source CLI. It outputs JSON by default (agent-friendly); add
--output table or --output yaml for humans. Pods ship with it pre-installed using a
pod-scoped key.
runpodctl serverless list # JSON by default
runpodctl serverless get <endpoint-id>
runpodctl serverless update <endpoint-id> --output table
runpodctl get pod --output tablePython SDK for programmatic control — key from env, never in source:
import os, runpod
runpod.api_key = os.environ["RUNPOD_API_KEY"] # never a literal
pod = runpod.create_pod(name="train", image_name="my/img:1.0", gpu_type_id="NVIDIA A100 80GB PCIe")
runpod.stop_pod(pod["id"]) # also resume_pod / terminate_pod
ep = runpod.Endpoint("<endpoint-id>")
job = ep.run({"prompt": "hi"}) # async: job.status(), job.output()
out = ep.run_sync({"prompt": "hi"}) # blocks, ~90s maxRule: read the API key from RUNPOD_API_KEY. A leaked key is a stranger spending on your GPUs.
| Bad | Good | Why |
|---|---|---|
| Max Workers left unbounded / sky-high | Bound it ~20% over peak | One traffic spike scales to an unbounded bill. |
| H100 by default | Smallest GPU that fits VRAM | ~7x cost for zero speedup when the model fits an L40S/4090. |
| 6-hour training run on Serverless | Run it on a Pod | Serverless premium + 600s execution timeout kills long jobs. |
| Hardcoded API key in source | os.environ["RUNPOD_API_KEY"] | A leaked key = a stranger's GPU bill on your card. |
| Weights downloaded at cold start | Bake into image or mount a volume | Every cold start re-downloads and pays for the wait. |
| Active workers "just in case" | Flex + FlashBoot | Active bills 24/7; FlashBoot makes flex cold starts cheap. |
| Volume left on a stopped pod | Detach / delete when idle | Idle volume bills ~$0.20/GB/mo silently. |
| Push image, debug on the platform | python worker.py --rp_serve_api first | Debugging on cold-start seconds is slow and costs money. |
Run bash scripts/verify.sh <worker-dir> to statically lint a worker directory: it checks
for runpod.serverless.start, a handler with a return/yield, no hardcoded API key, a
bounded max_workers plus timeout keys in any config, and a pinned FROM + CMD in a
Dockerfile. Pure grep/parse — no network, no RunPod account. It exits 0 on an empty dir.
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/runpod of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
Runpod next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Runpod this skillericrisco/rsc-harness | 156 | — | ~2.8k | Automated safety check: Pass | MIT | |
| Hugging Face Vision Trainerhuggingface/skills | 11k | 1 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| ModalK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~4.5k | Automated safety check: Notes | Apache-2.0 | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 3 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| Generate Openenv Envadithya-s-k/FineEnvs | 421 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~2.5k | Automated safety check: Pass | MIT |
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
K-Dense-AI/scientific-agent-skills
Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs.
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
adithya-s-k/FineEnvs
Builds an OpenEnv (Hugging Face) variant of an RL environment.
Orchestra-Research/AI-Research-SKILLs
Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.
huggingface/skills
Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…. Runpod is an agent skill from ericrisco/rsc-harness. Use when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless endpoints, handler workers, worker Docker templates, network volumes, timeout and worker-count tuning, cold starts, and runaway bills.
Runpod fits situations like: running GPU compute on RunPod and deciding between Pods (hourly; always-on) and Serverless (per-second; autoscaling) for training; inference — serverless endpoints.
Run `npx skills add ericrisco/rsc-harness --skill runpod -a claude-code`. Or copy the skill folder (skills/runpod in ericrisco/rsc-harness) into .claude/skills/runpod in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill runpod -a codex`. Or copy the skill folder (skills/runpod in ericrisco/rsc-harness) into .agents/skills/runpod in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill runpod -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/runpod, .gemini/skills/runpod, .github/skills/runpod and .opencode/skills/runpod in your project.
Going by SKILL.md and its folder, Runpod needs a shell for the scripts in its folder, the command-line tools its instructions call (python and bash) and credentials named RUNPOD_API_KEY. Our summary lists: Python 3; A Bash shell; Docker; A credential in RUNPOD_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Runpod is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Runpod: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Modal (K-Dense-AI/scientific-agent-skills, 48k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and Generate Openenv Env (adithya-s-k/FineEnvs, 421 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.