Agent skill

Replicate

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when running, packaging, deploying or scaling a model on the Replicate platform from code — blocking run versus async predictions, handling FileOutput, deployments with warm…

MITAuto-check passedBackend & APIs

Install Replicate

skills CLI
$ npx skills add ericrisco/rsc-harness --skill replicate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness replicate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/replicate .claude/skills/replicate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
replicate
GitHub stars
156
Token cost
~2.6k tokens
SKILL.md length
1,119 words
Files
7 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when running, packaging, deploying or scaling a model on the Replicate platform from code — blocking run versus async predictions, handling FileOutput, deployments with warm…

  • Scaling a model on the Replicate platform from code — blocking run versus async predictions
  • SKILL.md covers Decision: how should this…, Auth & install, Run a model and Long-running & async, plus 6 more sections
  • Runs Shell scripts from its folder; calls pip and npm; needs REPLICATE_API_TOKEN
  • Handling FileOutput

What it does

Replicate is an agent skill from ericrisco/rsc-harness. Use when running, packaging, deploying or scaling a model on the Replicate platform from code — blocking run versus async predictions, handling FileOutput, deployments with warm private endpoints and autoscaling, packaging with Cog, verifying webhook signatures, and cutting GPU spend. NOT crafting image prompts and parameters or picking image model families (that is replicate-images).

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/cog-packaging.md`).

It sits in Backend & APIs, covering Webhooks and Deployment. It works with Python. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Scaling a model on the Replicate platform from code — blocking run versus async predictions
  • Handling FileOutput
  • Deployments with warm private endpoints and autoscaling
  • Packaging with Cog

Example prompts

  • “/replicate”

Requirements

  • Python 3
  • Node.js
  • A Bash shell
  • Docker
  • A credential in REPLICATE_API_TOKEN

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • REPLICATE_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Replicate loads about 2.6k tokens when it runs, and up to ~5.5k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 1,119 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,119 words, ~2,617 tokens.

Download SKILL.mdSave it as .claude/skills/replicate/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
replicate
description
Use when running, packaging, deploying or scaling a model on the Replicate platform from code — blocking run versus async predictions, handling FileOutput, deployments with warm private endpoints and autoscaling, packaging with Cog, verifying webhook signatures, and cutting GPU spend. NOT crafting image prompts and parameters or picking image model families (that is `replicate-images`).
tags
replicate, cog, deployments, webhooks, model-serving, gpu, ai-infra
recommends
replicate-images, modal, runpod, huggingface, webhooks, cost-tracking, docker
origin
risco

Replicate platform operations

This skill is about how a model runs in production on Replicate — clients, async, deployments, Cog packaging, webhooks, scaling, and spend. It is the platform-engineering counterpart to image prompt craft. If the question is what prompt, aspect ratio, or model family produces a good image, that is replicate-images, not this skill. Here the mental model is: a prediction is a job. You either wait for it, poll it, or get pinged about it — and where it runs (shared cold pool vs a private warm deployment) is a cost-and-latency dial you set deliberately.

Pinned facts (verified 2026-06-02): Python client replicate 1.x (latest 1.0.7), Python 3.8+; JS client replicate on npm; auth via REPLICATE_API_TOKEN. A 2.0.0aN alpha exists on PyPI but is NOT the default — pin replicate>=1,<2 so a fresh install never silently pulls it.

Decision: how should this model run?

Pick the row by latency tolerance and whether your process can block. Do not default to run() for everything — a 10-minute job inside a web request will time out and burn a worker.

PatternLatencyBlocks your process?Cost shapeUse when
replicate.run(...)secondsYes — waits to completionper-prediction, shared poolinteractive/quick calls, scripts, CLIs
predictions.create() + pollminutesYes, but you control the loopper-prediction, shared poollong job, a worker can babysit it
predictions.create(webhook=...)minutes+No — fire and forgetper-prediction, shared poollong job, the request must return now
Deployment (private endpoint)low + steadydepends on call style abovewarm floor + per-predictionsustained traffic, need warm/private/autoscale cap

Rule: if a human or HTTP request is waiting longer than a few seconds, do not block on run() — switch to predictions + webhook. Why: synchronous timeouts kill the request but the GPU job keeps running and billing.

Auth & install

bash
export REPLICATE_API_TOKEN=r8_...           # clients read this env var automatically
pip install 'replicate>=1,<2'               # pin: 2.0.0aN is alpha; unpinned can pull it
npm install replicate                        # Node client

Python 3.8+ is required. Never pip install replicate unpinned in a Dockerfile or requirements file — a rebuild months later can resolve to the 2.0 alpha and break your imports. Why: the alpha is a full Stainless/httpx rewrite with a different surface.

Run a model

python
import replicate

output = replicate.run(
    "black-forest-labs/flux-schnell",
    input={"prompt": "a red bicycle", "num_outputs": 1},
)

Since client 1.0.0, file outputs come back as FileOutput objects, not URL strings. Treating one as a string is the single most common bug.

python
# Bad — output[0] is a FileOutput, not a str; this writes the repr, not the bytes
open("out.png", "w").write(output[0])

# Good — read the bytes, or take the URL explicitly
with open("out.png", "wb") as f:
    f.write(output[0].read())          # bytes
print(output[0].url)                    # hosted URL if you'd rather link

Rule: call .read() for bytes or .url for the link. Why: silently coercing a FileOutput to a string corrupts the file and the error surfaces far from the cause.

Long-running & async

For jobs over a few seconds, create a prediction instead of blocking:

python
client = replicate.Client()
prediction = client.predictions.create(
    model="owner/model",
    input={"prompt": "..."},
)
prediction.reload()                      # refresh status from the API
while prediction.status not in ("succeeded", "failed", "canceled"):
    time.sleep(2)
    prediction.reload()

Set a deadline so a stuck job auto-cancels instead of billing forever, and cancel() on cleanup paths. Why: a hung prediction with no deadline is silent, open-ended GPU spend. Poll with a small backoff, not a tight loop — you are charged for the prediction, not the polling, but a tight loop wastes your own process and rate budget. Full polling loop with backoff and 5xx handling is in references/webhooks-and-async.md.

Webhooks

For fire-and-forget, hand Replicate a URL and filter to the events you care about:

python
client.predictions.create(
    model="owner/model",
    input={"prompt": "..."},
    webhook="https://your.app/hooks/replicate",
    webhook_events_filter=["completed"],   # not every intermediate "logs" event
)

Rule: ALWAYS verify the signature before trusting a webhook body. Why: the URL is public — anyone can POST forged completions to it. Replicate signs each delivery; the secret has a whsec_ prefix. Reconstruct signed content as {webhook-id}.{webhook-timestamp}.{body}, HMAC-SHA256 with the base64-decoded secret, base64-encode, and constant-time compare against the webhook-signature header (a space-separated v1,<sig> list). The clients expose a verification helper; the full Python and Node recipe (including idempotency via webhook-id) is in references/webhooks-and-async.md.

Deployments

A deployment is a private, dedicated API endpoint for one model version that autoscales from zero to hundreds of instances. Reach for it when you need warm instances, a private endpoint, or a hard spend cap — not for one-off runs.

  • min_instances — the warm floor. Set >0 to kill cold starts for latency-sensitive traffic; every warm instance bills whether or not it serves a request.
  • max_instances — the spend cap. The ceiling on concurrent instances; protects you from a traffic spike turning into a surprise bill.
  • Hardware (NVIDIA T4, A100, H100, ...) is switchable without touching code — change the deployment config, not predict.py.
  • Rolling updates ship a new version with no downtime; built-in monitoring covers latency, throughput, error rate, and GPU memory.

You still call a deployment with run() / predictions.create() — it just routes to your private instances. Create/update via HTTP API, the clients, or CLI; fields and the rolling-update flow are in references/deployments-api.md.

Show full SKILL.md (411 more words)Show less

Package a custom model with Cog

Cog packages a model into a production container. You need two files; Replicate builds the API server for you on push.

yaml
# cog.yaml
build:
  gpu: true
  python_version: "3.11"
  python_packages:
    - "torch==2.4.0"
predict: "predict.py:Predictor"
python
# predict.py
from cog import BasePredictor, Input, Path

class Predictor(BasePredictor):
    def setup(self):
        # load weights ONCE here, not per request
        self.model = load_model("weights.pth")

    def predict(self, prompt: str = Input(description="text prompt")) -> Path:
        result = self.model(prompt)
        return Path(result)            # Cog uploads the file
bash
cog predict -i prompt="hello"          # run locally (needs Docker)
cog push r8.im/owner/model             # build + push; Replicate hosts it

Rule: load weights in setup(), never in predict(). Why: setup() runs once per instance; predict() runs every request — loading weights per request makes every call pay the model-load cost. Full cog.yaml (system packages, run steps), typed Input(...), GPU config, version pinning, and common build failures are in references/cog-packaging.md. Building requires Docker.

Cost & scaling levers

  • Scale to zero is the default — idle deployments stop billing. Only raise min_instances above zero when cold-start latency actually hurts users, and treat the warm floor as a line item.
  • Deadlines on predictions cap the worst case so a stuck job cannot bill open-ended.
  • Right-size the GPU: do not run a 1B model on an H100. Hardware is a deployment-config change.
  • Cache weights in setup() so per-request work is just inference.

This skill covers Replicate-specific levers only. Tracking total AI spend across many providers as a discipline is cost-tracking.

Anti-patterns

Anti-patternWhy it bitesDo instead
Treating a FileOutput as a URL stringCorrupts files / writes a repr; error surfaces far away.read() for bytes, .url for the link
Blocking run() for a 10-min job in a web requestRequest times out; the GPU job keeps running and billingpredictions.create(webhook=...), return now
Webhook handler with no signature checkThe URL is public; anyone can forge completionsVerify HMAC-SHA256 against whsec_ secret
High min_instances "just in case"Every warm instance bills 24/7 idleScale to zero; raise floor only when cold starts hurt
Loading weights inside predict()Every request pays the model-load costLoad once in setup()
pip install replicate unpinnedA rebuild can pull the 2.0 alpha and break importsPin replicate>=1,<2
A deployment for a one-off runPays for a private endpoint you call onceUse the shared pool via run()
No deadline on a long predictionA hung job bills open-ended, silentlySet a deadline; cancel() on cleanup

References

  • references/cog-packaging.md — full cog.yaml + predict.py, GPU config, build/push, version pinning, build failures.
  • references/webhooks-and-async.md — signature verification (Python + Node), event filters, idempotency, polling with backoff, 5xx retries.
  • references/deployments-api.md — create/update/get deployment via HTTP + clients, autoscaling fields, rolling updates, monitoring metrics, CLI.

scripts/verify.sh statically checks an emitted artifact dir: a cog.yaml declaring build: and predict:, a predict.py Predictor with setup + predict, and warns on an unpinned replicate dependency or a FileOutput written as a string. Presence + key checks only — it does not run Docker.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/replicate of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/cog-packaging.md
  • references/deployments-api.md
  • references/webhooks-and-async.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Replicate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Replicate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Replicate this skillericrisco/rsc-harness156—~2.6kAutomated safety check: PassMIT
Trigger.dev Configurationpapermark/papermark9.2k—~1.2kAutomated safety check: PassCustom licence
Golivemikehasa/golive-skill1.2k—~13kAutomated safety check: NotesMIT
Channel Debug Corevercel-labs/vercel-openclaw-archived117—~2kAutomated safety check: NotesMIT
Dingtalk Messageagentscope-ai/ReMe3.6k—~1.6kAutomated safety check: PassApache-2.0
Stripe Best Practiceskanchengw/cnllm1753 repos~925Automated safety check: PassApache-2.0

Similar skills

  • Trigger.dev Configuration

    papermark/papermark

    Configures Trigger.dev projects through trigger.config.ts, with build extensions for Prisma, Playwright, Puppeteer, FFmpeg, Python and system packages.

    9.2k GitHub stars~1.2k tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Golive

    mikehasa/golive-skill

    Take an agent-written app from repo to live production on the user's OWN accounts, with providers they choose (hosting, database, auth, payments, email, domain/DNS).

    1.2k GitHub stars~13k tokensUpdated 3 days ago
    Backend & APIsAuto-check: notes
  • Channel Debug Core

    vercel-labs/vercel-openclaw-archived

    Official

    Channel webhook triage for vercel-openclaw Slack/Telegram/Discord/WhatsApp issues: prove deployment state, collect admin readiness endpoints, build evidence-first handoff before fixes.

    117 GitHub stars~2k tokensUpdated 4 mo ago
    Backend & APIsAuto-check: notes
  • Dingtalk Message

    agentscope-ai/ReMe

    钉钉消息发送技能。支持企业内部机器人(批量单聊/群聊)和 Webhook 自定义机器人两种接入方式,支持多机器人管理,支持文本、Markdown、链接、ActionCard、FeedCard等多种消息类型。

    3.6k GitHub stars~1.6k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Stripe Best Practices

    kanchengw/cnllm

    Guides Stripe integration decisions — API selection (Checkout Sessions vs PaymentIntents), Connect platform setup (Accounts v2, controller properties), billing/subscriptions, Treasury financial…

    175 GitHub starsUsed in 3 repos~925 tokens
    Backend & APIsAuto-check passed
  • Tlgr

    tlgrcli/tlgr

    Read and act on a personal Telegram account from the terminal with the tlgr CLI (MTProto user account, not a bot).

    191 GitHub stars~1.5k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Works with

Questions about Replicate

What does Replicate do?

A skill your agent uses when running, packaging, deploying or scaling a model on the Replicate platform from code — blocking run versus async predictions, handling FileOutput, deployments with warm…. Replicate is an agent skill from ericrisco/rsc-harness. Use when running, packaging, deploying or scaling a model on the Replicate platform from code — blocking run versus async predictions, handling FileOutput, deployments with warm private endpoints and autoscaling, packaging with Cog, verifying webhook signatures, and cutting GPU spend.

When should I use Replicate?

Replicate fits situations like: scaling a model on the Replicate platform from code — blocking run versus async predictions; handling FileOutput; deployments with warm private endpoints and autoscaling; packaging with Cog.

How do I install Replicate in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill replicate -a claude-code`. Or copy the skill folder (skills/replicate in ericrisco/rsc-harness) into .claude/skills/replicate in your project. Claude Code loads it when a task matches its description.

How do I install Replicate in Codex?

Run `npx skills add ericrisco/rsc-harness --skill replicate -a codex`. Or copy the skill folder (skills/replicate in ericrisco/rsc-harness) into .agents/skills/replicate in your project. Codex loads it when a task matches its description.

Can I use Replicate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill replicate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/replicate, .gemini/skills/replicate, .github/skills/replicate and .opencode/skills/replicate in your project.

What does Replicate need to run?

Going by SKILL.md and its folder, Replicate needs a shell for the scripts in its folder, the command-line tools its instructions call (pip and npm) and credentials named REPLICATE_API_TOKEN. Our summary lists: Python 3; Node.js; A Bash shell; Docker; A credential in REPLICATE_API_TOKEN.

Does Replicate access the network?

SKILL.md contains no URLs. Its commands use pip and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Replicate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Replicate use?

Replicate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Replicate use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Replicate?

Skills that share tags, products or a category with Replicate: Trigger.dev Configuration (papermark/papermark, 9.2k stars), Golive (mikehasa/golive-skill, 1.2k stars), Channel Debug Core (vercel-labs/vercel-openclaw-archived, 117 stars) and Dingtalk Message (agentscope-ai/ReMe, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Replicate?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.