Agent skill

Huggingface

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when running open models or working on the Hugging Face platform — the Inference Providers router or InferenceClient, Hub repos via the hf CLI, a dedicated Inference Endpoint…

MITAuto-check passedAI & LLM Engineering

Install Huggingface

skills CLI
$ npx skills add ericrisco/rsc-harness --skill huggingface -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness huggingface --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/huggingface .claude/skills/huggingface && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
huggingface
GitHub stars
156
Token cost
~2.6k tokens
SKILL.md length
1,014 words
Files
7 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when running open models or working on the Hugging Face platform — the Inference Providers router or InferenceClient, Hub repos via the hf CLI, a dedicated Inference Endpoint…

  • Works in 3 steps: The Hub — versioned git repos for… → Inference — three ways to actually run a… → The catalog — 1M+ open models you choose…
  • Running open models
  • SKILL.md covers Decision: how should I run…, Auth & install, Inference Providers — the… and Hub ops, plus 6 more sections
  • Runs Shell scripts from its folder; calls hf, pip and huggingface-cli; reaches router.huggingface.co; needs HF_TOKEN

What it does

Huggingface is an agent skill from ericrisco/rsc-harness. Use when running open models or working on the Hugging Face platform — the Inference Providers router or InferenceClient, Hub repos via the hf CLI, a dedicated Inference Endpoint with scale-to-zero, a Gradio Space with ZeroGPU, picking an open model by task/license/size, or loading one locally with transformers. NOT serving locally on your own machine (that is ollama), NOT renting your own GPU box (that is runpod), NOT hosted creative image APIs (that is replicate-images), NOT fine-tuning with trl/peft (that is…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/endpoints-and-spaces.md`).

It sits in AI & LLM Engineering, covering Model hubs and datasets and Fine-tuning. It works with Hugging Face, Ollama and Gradio. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Running open models
  • Working on the Hugging Face platform — the Inference Providers router
  • InferenceClient
  • Hub repos via the hf CLI

Example prompts

  • “/huggingface”

Requirements

  • Python 3
  • A Bash shell
  • Docker

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The Hub — versioned git repos for models, datasets, and Spaces. You search it, you
  2. Inference — three ways to actually run a model: the Inference Providers router
  3. The catalog — 1M+ open models you choose from by task, license, and size.

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • hf
    • pip
    • huggingface-cli

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • router.huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Huggingface loads about 2.6k tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 1,014 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,014 words, ~2,560 tokens.

Download SKILL.mdSave it as .claude/skills/huggingface/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
huggingface
description
Use when running open models or working on the Hugging Face platform — the Inference Providers router or InferenceClient, Hub repos via the hf CLI, a dedicated Inference Endpoint with scale-to-zero, a Gradio Space with ZeroGPU, picking an open model by task/license/size, or loading one locally with transformers. NOT serving locally on your own machine (that is `ollama`), NOT renting your own GPU box (that is `runpod`), NOT hosted creative image APIs (that is `replicate-images`), NOT fine-tuning with trl/peft (that is `finetuning`).
tags
huggingface, inference-providers, transformers, model-hub, inference-endpoints, spaces, ai-infra
recommends
ollama, runpod, modal, replicate, together-fireworks, fal, rag, embeddings-search, prompt-engineering, llm-pipeline, ai-media
origin
risco

Hugging Face: Hub, routed/hosted inference, and transformers

Hugging Face is three surfaces, and you should always know which one you are on:

  1. The Hub — versioned git repos for models, datasets, and Spaces. You search it, you hf download / hf upload, you read and write model cards.
  2. Inference — three ways to actually run a model: the Inference Providers router (serverless, you own nothing), a dedicated Inference Endpoint (you own a deployment that autoscales), or local transformers (you own the machine).
  3. The catalog — 1M+ open models you choose from by task, license, and size.

The whole skill is choosing the right surface for the job and proving it works: a 200 router response, a live endpoint URL, a pushed repo commit. If the model is open and the workflow lives on huggingface.co, you are in the right place. Operating the GPU box yourself is ../ollama/SKILL.md (your machine) or ../runpod/SKILL.md (a rented box); training weights is ../finetuning/SKILL.md.

Decision: how should I run this model?

Pick the row before you write a line of code. The cheapest mistake is standing up infra you did not need.

SituationUseWhy
Try a model now, low/dev volume, own no infraInference Providers router (InferenceClient)Fastest path; monthly credits cover dev.
CPU task: embeddings, text-ranking, text-classification, small BERT/GPT-2provider="hf-inference"That is exactly its remaining niche as of July 2025.
Big LLM (8B, 70B, 405B) through HFrouter with a partner provider (Together/Fireworks/Cerebras/DeepInfra…)hf-inference does not serve big LLMs — it will 404 or stall.
Steady prod traffic, need fixed latency/SLAdedicated Inference Endpoint + scale-to-zeroPredictable, autoscaling, billed per minute.
Interactive demo or shareable GPU appSpace (Gradio + ZeroGPU)Free-ish, public URL, GPU only while a call runs.
One-off GPU job (eval, batch convert)hf jobs runNo standing infra; PRO feature.
Offline, data-private, or already on a GPU boxlocal transformers pipeline()No network, no per-call cost.

Auth & install

bash
pip install "huggingface_hub[inference]"   # 1.17.0; needs Python >=3.10
pip install transformers                    # 5.x line, PyTorch-first, optional/local
hf auth login                               # stores a token; or export HF_TOKEN=...
  • The CLI is hf now, shaped hf <resource> <action> (hf auth login, hf download, hf upload, hf repo create, hf jobs run). huggingface-cli still runs but prints a deprecation warning — do not write it into new scripts.
  • Never hardcode a hf_... token in code — tokens leak the moment the file hits git. Read from the environment instead:
python
import os
from huggingface_hub import InferenceClient
client = InferenceClient(api_key=os.environ["HF_TOKEN"])   # never api_key="hf_xxx"
  • Token scopes: read to pull public/gated repos and run inference, write to push, fine-grained to scope to specific repos/orgs — why: a leaked read token cannot overwrite your models.

Inference Providers — the default path

One router reaches 200+ models across partner providers plus hf-inference; HF passes provider cost through with no markup. Two equivalent entry points:

python
# Native client — task methods, NOT the removed .post()
from huggingface_hub import InferenceClient
client = InferenceClient(api_key=os.environ["HF_TOKEN"])
out = client.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "One sentence on diffusion models."}],
    provider="together",          # name a partner; or omit for auto-routing
)
print(out.choices[0].message.content)
python
# OpenAI-compatible — same router, drop-in for existing OpenAI code
from openai import OpenAI
client = OpenAI(
    base_url="https://router.huggingface.co/v1",   # this exact host, nothing else
    api_key=os.environ["HF_TOKEN"],
)
  • InferenceClient.post() was removed (dropped in hub v0.31.0). Use the task methods: chat.completions.create(), text_generation(), feature_extraction() (embeddings), text_to_image(), automatic_speech_recognition().
  • Credits are real and small: Free $0.10/mo, PRO $2.00/mo, Team/Enterprise $2.00 per seat (shared). Past that you are pay-as-you-go and must buy credits. Budget accordingly — why: a chat loop on a 70B model burns the free tier in minutes.
  • A Custom Provider Key bypasses HF billing entirely (the provider bills you; HF credits do not apply). For org billing, pass bill_to="org-name" (header X-HF-Bill-To).
  • Full recipes (embeddings, image, ASR, streaming, rate-limit handling, the provider list) live in references/inference-providers.md.

Hub ops

bash
hf download meta-llama/Llama-3.1-8B-Instruct --include "*.safetensors"
hf repo create my-org/my-model --repo-type model
hf upload my-org/my-model ./out --commit-message "v1 weights"
python
from huggingface_hub import snapshot_download
path = snapshot_download("BAAI/bge-small-en-v1.5")   # full repo, cached, resumable
  • Gated models (Llama, Gemma, many others) need you to accept terms on the model page first, then a token with read access — otherwise the download 403s.
  • A model card is a README.md with YAML front-matter (license, pipeline_tag, tags, base_model). Ship one on every upload — why: an uncarded repo is unsearchable and unusable by anyone but you. Command map and hf jobs run details in references/hub-and-cli.md.
Show full SKILL.md (427 more words)Show less

Choosing a model

Filter the Hub by task + license + size + recent downloads, then read the card before you commit. Match the model to your constraint; do not grab whatever is trending.

  • Check the license: Apache-2.0/MIT are permissive; Llama/Gemma carry commercial terms and are gated; "non-commercial"/"research-only" cards mean you cannot ship them.
  • Check size vs target: a 70B will not fit a single A10G; an embedding model belongs on CPU.
  • Check context length and intended use in the card — the headline number is not always the usable one.

Dedicated Inference Endpoints — when to graduate

Move off the router when you need fixed latency/SLA, or the router's PAYG cost stops being predictable. An Endpoint is your own autoscaling deployment.

  • Pricing: CPU from ~$0.032/core/hr, GPU from ~$0.50/hr (A10G ~$1.00/hr, H100 ~$6.40–8.00/hr), billed per minute even though shown hourly.
  • Enable scale-to-zero for bursty traffic — it parks at $0 when idle and cold-starts on the next request. A bursty 100–1000 req/day workload typically lands at $20–60/mo.
  • Deploy from the UI or with huggingface_hub (create_inference_endpoint(...)). Config and a cost worksheet are in references/endpoints-and-spaces.md.

Spaces + ZeroGPU

A Space hosts a demo app with a public URL. ZeroGPU grabs an H200 MIG slice (~70GB) only while a decorated function runs, then releases it.

python
import spaces
@spaces.GPU                       # GPU acquired for this call only
def generate(prompt: str) -> str:
    ...
  • ZeroGPU is Gradio-SDK only — Streamlit/Docker/static Spaces cannot use it. PRO ($9/mo) gives 8x daily quota, queue priority, and up to 10 owned ZeroGPU Spaces. Details in references/endpoints-and-spaces.md.

Local transformers

python
from transformers import pipeline
pipe = pipeline("text-generation", model="meta-llama/Llama-3.1-8B-Instruct",
                device_map="auto", torch_dtype="auto")
print(pipe("Hello", max_new_tokens=64)[0]["generated_text"])
  • pipeline("task", model=...) for quick use; AutoModelForCausalLM.from_pretrained(...) when you need control over generation/quantization. Set device_map/torch_dtype explicitly.
  • Use local only when you are offline, data-private, or already on a GPU. Otherwise the router is far less ops than babysitting CUDA and weights.

Anti-patterns

Anti-patternWhy it bitesDo instead
InferenceClient.post(...)Removed in hub v0.31.0; raisesTask methods: chat.completions.create(), feature_extraction()
provider="hf-inference" for a 70B/405B LLMCPU niche; 404s or stallsRoute to a partner provider (Together/Fireworks/Cerebras)
api_key="hf_abc123..." in codeToken leaks in git historyRead os.environ["HF_TOKEN"]
Spin up a dedicated Endpoint just to try a modelBurns money idleUse the router first; graduate only on real traffic
Assuming router calls are free/unlimitedFree tier is $0.10/moBudget credits; expect PAYG
ZeroGPU under Streamlit/Docker SDKUnsupported, silently no GPUUse the Gradio SDK
huggingface-cli ... in new scriptsDeprecated, warnsUse hf ...
OpenAI base URL other than https://router.huggingface.co/v1Won't reach the HF routerUse that exact host

verify.sh

scripts/verify.sh [TARGET] is a static, read-only linter (no network, no token). It flags the hard violations above — .post(, hardcoded hf_ tokens, big-LLM-to-hf-inference, wrong router host — and warns on legacy huggingface-cli. It exits 0 on a clean or empty target.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/huggingface of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/endpoints-and-spaces.md
  • references/hub-and-cli.md
  • references/inference-providers.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Huggingface next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Huggingface compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Huggingface this skillericrisco/rsc-harness156—~2.6kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Huggingface Lora Space Buildersickn33/agentic-awesome-skills47k1 repos~8.3kAutomated safety check: PassApache-2.0
Dataset Transformationawslabs/agent-plugins9122 repos~3.5kAutomated safety check: PassApache-2.0
Space Doctorhuggingface/hf-mcp-server302—~1.8kAutomated safety check: PassMIT
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0

Similar skills

  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface Lora Space Builder

    sickn33/agentic-awesome-skills

    Build and publish a Gradio demo on Hugging Face Spaces for a user-provided LoRA.

    47k GitHub starsUsed in 1 repo~8.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    912 GitHub starsUsed in 2 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Space Doctor

    huggingface/hf-mcp-server

    Official

    Diagnose broken Hugging Face Gradio Spaces from their actual logs and pinned source, then prepare a minimal verified source fix as candidate files.

    302 GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface LLM Trainer

    waybarrios/opencode-power-pack

    Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

    533 GitHub stars~3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Huggingface

What does Huggingface do?

A skill your agent uses when running open models or working on the Hugging Face platform — the Inference Providers router or InferenceClient, Hub repos via the hf CLI, a dedicated Inference Endpoint…. Huggingface is an agent skill from ericrisco/rsc-harness. Use when running open models or working on the Hugging Face platform — the Inference Providers router or InferenceClient, Hub repos via the hf CLI, a dedicated Inference Endpoint with scale-to-zero, a Gradio Space with ZeroGPU, picking an open model by task/license/size, or loading one locally with transformers.

When should I use Huggingface?

Huggingface fits situations like: running open models; working on the Hugging Face platform — the Inference Providers router; inferenceClient; hub repos via the hf CLI.

How do I install Huggingface in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill huggingface -a claude-code`. Or copy the skill folder (skills/huggingface in ericrisco/rsc-harness) into .claude/skills/huggingface in your project. Claude Code loads it when a task matches its description.

How do I install Huggingface in Codex?

Run `npx skills add ericrisco/rsc-harness --skill huggingface -a codex`. Or copy the skill folder (skills/huggingface in ericrisco/rsc-harness) into .agents/skills/huggingface in your project. Codex loads it when a task matches its description.

Can I use Huggingface in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill huggingface -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huggingface, .gemini/skills/huggingface, .github/skills/huggingface and .opencode/skills/huggingface in your project.

What does Huggingface need to run?

Going by SKILL.md and its folder, Huggingface needs a shell for the scripts in its folder, the command-line tools its instructions call (hf, pip and huggingface-cli) and credentials named HF_TOKEN. Our summary lists: Python 3; A Bash shell; Docker.

Does Huggingface access the network?

SKILL.md names 1 domain. In commands or code: router.huggingface.co; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Huggingface safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Huggingface use?

Huggingface is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Huggingface use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Huggingface?

Skills that share tags, products or a category with Huggingface: Hugging Face LLM Trainer (huggingface/skills, 11k stars), Huggingface Lora Space Builder (sickn33/agentic-awesome-skills, 47k stars), Dataset Transformation (awslabs/agent-plugins, 912 stars) and Space Doctor (huggingface/hf-mcp-server, 302 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Huggingface?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.