Agent skill

Scaling

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when traffic is growing or about to spike and the system bends under concurrency — deciding what to add and in what order (cache, connection pool, async queue, read replica…

MITAuto-check passedDatabases

Install Scaling

skills CLI
$ npx skills add ericrisco/rsc-harness --skill scaling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness scaling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scaling .claude/skills/scaling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scaling
GitHub stars
167
Token cost
~2.8k tokens
SKILL.md length
1,376 words
Files
6 (incl. scripts, references)
Skills in repo
227
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when traffic is growing or about to spike and the system bends under concurrency — deciding what to add and in what order (cache, connection pool, async queue, read replica…

  • Traffic is growing
  • SKILL.md covers Start here — order of operations, Step 0 — find the bottleneck…, Lever 1 — Caching (cheapest,… and Lever 2 — Connection pooling +…, plus 5 more sections
  • Runs JavaScript and Shell scripts from its folder
  • About to spike and the system bends under concurrency — deciding what to add and in what order (cache

What it does

Scaling is an agent skill from ericrisco/rsc-harness. Use when traffic is growing or about to spike and the system bends under concurrency — deciding what to add and in what order (cache, connection pool, async queue, read replica, more instances) and proving it with a load test against explicit RPS and p95 targets rather than guessing. NOT making one slow request faster or profiling an N+1 (that is performance), NOT race-free Redis caches, locks and queue semantics (that is redis).

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/load-testing-k6.md`).

It sits in Databases, covering Caching and Load testing. It works with Redis. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Traffic is growing
  • About to spike and the system bends under concurrency — deciding what to add and in what order (cache
  • Connection pool
  • More instances) and proving it with a load test against explicit RPS and p95 targets rather than guessing

Example prompts

  • “/scaling”

Requirements

  • Node.js
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (JavaScript and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scaling loads about 2.8k tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 1,376 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,376 words, ~2,781 tokens.

Download SKILL.mdSave it as .claude/skills/scaling/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
scaling
description
Use when traffic is growing or about to spike and the system bends under concurrency — deciding what to add and in what order (cache, connection pool, async queue, read replica, more instances) and proving it with a load test against explicit RPS and p95 targets rather than guessing. NOT making one slow request faster or profiling an N+1 (that is `performance`), NOT race-free Redis caches, locks and queue semantics (that is `redis`).
tags
scaling, capacity, load-testing, caching, read-replicas, pgbouncer, k6, devops
recommends
performance, redis, postgresdb, monitoring, deployment, backups, fly-io
origin
risco

Scaling: survive many requests at once

Performance makes one request faster. Scaling makes many requests survive at the same time. Different problem, different toolbox — don't reach for this one when a single endpoint is slow for a single user (that is ../performance/SKILL.md).

The whole job in one line: diagnose the bottleneck, apply the cheapest lever that moves it, re-measure under load. Repeat until the next bottleneck appears or you hit your target.

Prime directive: never add infrastructure without a measurement first. A replica, a queue, or a third app instance you bought on a hunch costs money every month and usually moves the wrong tier. Measure, then add.

Start here — order of operations

Levers are ordered by payoff per dollar. Caching is nearly free and wins biggest; replicas and autoscaling cost forever. Climb the ladder in order; stop the moment the symptom clears.

SymptomLikely bottleneck tierFirst leverSibling that wires it
Slow only under load, fine solounknown — measure firstUSE-method triage (Step 0)../monitoring/SKILL.md
Same reads recomputed for everyoneapp/DB doing repeat workLever 1 — cache../redis/SKILL.md
too many clients alreadyDB connection slotsLever 2 — pooler../postgresdb/SKILL.md
Spiky writes time out / dropsynchronous write pathLever 2 — async queue../redis/SKILL.md
Reads dominate, primary CPU hotDB read capacityLever 3 — read replica../postgresdb/SKILL.md
All tiers healthy, just need throughputapp instance countLever 4 — horizontal / autoscale../deployment/SKILL.md

Step 0 — find the bottleneck before you add anything

Use the USE method (Utilization, Saturation, Errors) on every resource — CPU, memory, disk, network, and the DB connection pool. For each one ask: how busy (U), how much is queued/waiting (S), and any errors (E).

  • Saturation predicts collapse earlier than utilization. A CPU at 70% util with a growing run-queue is closer to falling over than one at 90% with no queue. Watch queue depth, run-queue length, and connection-pool wait time — those spike before throughput craters.
  • Read p95/p99, never the mean. The mean hides the tail your users actually feel; if p50 is 80 ms and p95 is 4 s, a meaningful slice of traffic is having a bad time and the average lies about it.
  • The dashboards, alerts, and SLOs that produce these numbers are ../monitoring/SKILL.md / observability's job. Scaling consumes USE signals; it doesn't build the collectors.

Output of Step 0 is one sentence: "the bottleneck is the DB connection pool / app CPU / origin cache-miss rate." Don't proceed without it.

Lever 1 — Caching (cheapest, do first)

Cache layers, outermost to innermost — each one removes work the layer behind it would have done:

LayerRemovesTypical TTL
CDN / edgeorigin round-trip for static + cacheable HTMLminutes–hours
HTTP cache headers (Cache-Control, ETag)re-downloads; enables 304sper-resource
App cache (in-proc / Redis)recomputed views, serialized payloadsseconds–minutes
Query-result cacherepeated identical DB readsseconds
  • Cache the expensive read, not the cheap one. Caching a 2 ms lookup adds a network hop and an invalidation bug for nothing; cache the 400 ms aggregate everyone hits.
  • Target >80% hit ratio for general traffic (static-heavy/CDN routinely hits 95%+). A ratio consistently <60% means a broken strategy — wrong cache keys, TTL too low, or churny data that shouldn't be cached.
  • Beware the cold-cache miss storm. On deploy or cache flush, every request misses at once and stampedes the origin — the cache that was protecting you now amplifies the load. Mitigating that (stampede protection, single-flight, request coalescing) is cache correctness → ../redis/SKILL.md.
  • If the underlying query is just slow, caching only hides it. Fix the query — N+1, missing index — via ../performance/SKILL.md before papering over it with TTL.

Lever 2 — Connection pooling + queues

Pool DB connections. Each Postgres connection is a backend process with real memory cost; apps that open a connection per request exhaust max_connections fast.

text
Bad:  app → opens a fresh DB connection per request → "too many clients already"
Good: app → PgBouncer (transaction mode) → small pool of reused server connections
  • Use transaction pooling for stateless web apps: a server connection is held only for the duration of a transaction and released on COMMIT/ROLLBACK — reported real-world effect is roughly 100× effective connection capacity.
  • Sizing rule: keep (number_of_pools × default_pool_size) < max_connections − ~15 (leave headroom for superuser/admin slots). Set default_pool_size ≈ 1.5–2× vCores for CPU-bound OLTP — more connections than cores just adds context-switch contention, not throughput.
  • Gotcha — transaction pooling breaks session-scoped features: prepared statements (pre-PG14 protocol), SET / session GUCs, advisory session locks, and LISTEN/NOTIFY. Route those to a session-pooling pool or refactor them out. Don't discover this in production.

Shed spiky writes into a queue. Queue-based load leveling puts a queue between a bursty producer and a constrained consumer so the consumer drains at its own steady rate; the queue absorbs the spike instead of the synchronous tier melting.

  • Move anything that doesn't need a synchronous answer — emails, webhooks, thumbnails, exports — off the request path.
  • Queue semantics (race-free rate limits, stalled-job recovery, fencing tokens, durable SKIP LOCKED queues) belong to ../redis/SKILL.md and ../postgresdb/SKILL.md. Scaling decides that you defer work; those decide it's done correctly.
Show full SKILL.md (580 more words)Show less

Lever 3 — Read replicas

Reach for a replica only when Step 0 says reads dominate and the primary is read-saturated — not as a reflex.

  • Replicas serve reads, never the source of truth for read-after-write. Async streaming replication lags; a user who just wrote and immediately reads from a replica may see stale data. Route post-write reads (or that user's whole session for a window) to the primary, or use replica-lag-aware routing.
  • A replica is not a write-scaling story and not, by itself, an HA/backup story. Writes still all hit one primary; durability and recovery are ../backups/SKILL.md.
  • Wiring the replication itself — primary_conninfo, slots, promotion — is ../postgresdb/SKILL.md. Scaling decides add a replica and route reads to it; postgresdb makes it real.

Lever 4 — Horizontal scaling & autoscaling

  • Make the app stateless first. You can't horizontally scale what holds local state — in-memory sessions, sticky uploads on local disk, per-instance caches that must agree. Push session/state to Redis or the DB, then add instances freely.
  • Autoscale on the saturation metric that actually binds, not CPU alone. If the real limit is DB connections, scaling app instances on CPU just opens more connections and topples the DB faster. Scale on the bottleneck Step 0 found.
  • Choosing a host, rolling deploys, and the autoscaling dials are ../deployment/SKILL.md plus the platform skill (../fly-io/SKILL.md, and siblings for railway/render/vercel). Scaling gives the strategy — how many and triggered by what; the platform gives the knobs.

Prove it — load testing with k6

Don't claim the system survives. Measure that it does. k6 (Go core, JS test scripts) reached v1.0.0 on 2025-04-28 under SemVer and is the default OSS load-test tool; the current v1 line is v1.7.x and v2.0.0 shipped in May 2026 (GrafanaCON 2026).

  • Thresholds are the test's pass/fail SLO — codify the target so the run goes red on its own instead of you eyeballing a graph.
  • Test a prod-like target, never localhost. Localhost has no network latency, no real DB, no CDN — it measures your laptop, not your system.
  • Run a ladder, not one shot: smoke → load → stress → spike → soak (find the knee where latency turns vertical).
javascript
import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  stages: [
    { duration: '1m', target: 50 },   // ramp up to 50 virtual users
    { duration: '3m', target: 50 },   // hold (steady-state load test)
    { duration: '1m', target: 0 },    // ramp down
  ],
  thresholds: {
    http_req_duration: ['p(95)<500'], // SLO gate: 95% of requests under 500 ms
    http_req_failed: ['rate<0.01'],   // SLO gate: under 1% errors
  },
};

export default function () {
  const res = http.get(`${__ENV.TARGET_URL}/api/health`);
  check(res, { 'status is 200': (r) => r.status === 200 });
  sleep(1);
}

Full ladder (stage configs for each test type), CI gate snippet, and how to read the summary (p95/p99, http_req_failed, spotting the knee) are in references/load-testing-k6.md.

Anti-patterns

Anti-patternWhy it bitesDo instead
Scaling before measuringYou spend on the wrong tier; symptom persistsStep 0 USE triage, name the bottleneck first
Scaling a stateful app horizontallyInstances disagree; sessions vanish on routingMake it stateless, externalize state, then scale
Caching cheap work / no TTL strategyAdds a hop + invalidation bugs for no gainCache the expensive read; set deliberate TTLs
Load-testing localhostMeasures your laptop, not productionTest a prod-like target over the network
Reporting mean latencyHides the tail users actually feelGate on p95/p99
Read replica to absorb writesWrites still hit one primary; you gain nothingReplica is read-only; queue/shard writes
DB with no connection poolertoo many clients already under any spikePgBouncer transaction mode + the sizing rule
Autoscaling on CPU while DB connections saturateMore instances = more connections = faster DB deathAutoscale on the binding saturation metric

Stop rule & cost

Scale to the next bottleneck, then re-measure — don't pre-buy capacity for traffic you don't have. Every lever has a price: caching is ~free, a pooler is cheap, a read replica and autoscaling cost every month and add operational surface. Climb one rung, re-run the load test, and stop when you clear the target. Survived, proven, no further — that's done.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/scaling of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/load-testing-k6.md
  • scripts/example.load.js
  • scripts/verify.sh

Open the folder on GitHubat commit e3d5b33

Compare with similar skills

Scaling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scaling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scaling this skillericrisco/rsc-harness167—~2.8kAutomated safety check: PassMIT
Configure REST Cachestrapi-community/plugin-rest-cache155—~1.1kAutomated safety check: PassMIT
Commandkit Cacheneplexlabs/commandkit165—~506Automated safety check: PassMIT
Redissickn33/agentic-awesome-skills47k2 repos~2.6kAutomated safety check: NotesMIT
Redis Coreredis/agent-skills1652 repos~759Automated safety check: PassMIT
Redis Connectionsredis/agent-skills1651 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • Configure REST Cache

    strapi-community/plugin-rest-cache

    Choose and write a Strapi REST Cache configuration for a specific use case.

    155 GitHub stars~1.1k tokensUpdated 10 days ago
    DatabasesAuto-check passed
  • Commandkit Cache

    neplexlabs/commandkit

    Implement deterministic caching with @commandkit/cache. An agent skill from neplexlabs/commandkit.

    165 GitHub stars~506 tokensUpdated 11 days ago
    DatabasesAuto-check passed
  • Redis

    sickn33/agentic-awesome-skills

    Configure Redis for caching and data storage. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~2.6k tokens
    DatabasesAuto-check: notes
  • Redis Core

    redis/agent-skills

    Official

    Core Redis modeling guidance — choose the right data structure (String, Hash, List, Set, Sorted Set, JSON, Stream, Vector Set) and use consistent colon-separated key names.

    165 GitHub starsUsed in 2 repos~759 tokens
    DatabasesAuto-check passed
  • Redis Connections

    redis/agent-skills

    Official

    Redis client and connection guidance covering connection pooling, multiplexing, pipelining, client-side caching with RESP3, avoiding slow commands (KEYS, SMEMBERS, HGETALL), and tuning socket…

    165 GitHub starsUsed in 1 repo~1.3k tokens
    DatabasesAuto-check passed
  • Redis Semantic Cache

    redis/agent-skills

    Official

    Redis LangCache guidance for semantic caching of LLM responses on Redis Cloud — calling search/set via the SDK or REST API, tuning the similarity threshold, separating caches per task type, and…

    165 GitHub starsUsed in 1 repo~1k tokens
    DatabasesAuto-check passed

More from ericrisco/rsc-harness

All 227 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    167 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    167 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    167 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    167 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    167 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Works with

Questions about Scaling

What does Scaling do?

A skill your agent uses when traffic is growing or about to spike and the system bends under concurrency — deciding what to add and in what order (cache, connection pool, async queue, read replica…. Scaling is an agent skill from ericrisco/rsc-harness. Use when traffic is growing or about to spike and the system bends under concurrency — deciding what to add and in what order (cache, connection pool, async queue, read replica, more instances) and proving it with a load test against explicit RPS and p95 targets rather than guessing.

When should I use Scaling?

Scaling fits situations like: traffic is growing; about to spike and the system bends under concurrency — deciding what to add and in what order (cache; connection pool; more instances) and proving it with a load test against explicit RPS and p95 targets rather than guessing.

How do I install Scaling in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill scaling -a claude-code`. Or copy the skill folder (skills/scaling in ericrisco/rsc-harness) into .claude/skills/scaling in your project. Claude Code loads it when a task matches its description.

How do I install Scaling in Codex?

Run `npx skills add ericrisco/rsc-harness --skill scaling -a codex`. Or copy the skill folder (skills/scaling in ericrisco/rsc-harness) into .agents/skills/scaling in your project. Codex loads it when a task matches its description.

Can I use Scaling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill scaling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scaling, .gemini/skills/scaling, .github/skills/scaling and .opencode/skills/scaling in your project.

What does Scaling need to run?

Going by SKILL.md and its folder, Scaling needs JavaScript and a shell for the scripts in its folder. Our summary lists: Node.js; A Bash shell.

Does Scaling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scaling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Scaling use?

Scaling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scaling use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Scaling?

Skills that share tags, products or a category with Scaling: Configure REST Cache (strapi-community/plugin-rest-cache, 155 stars), Commandkit Cache (neplexlabs/commandkit, 165 stars), Redis (sickn33/agentic-awesome-skills, 47k stars) and Redis Core (redis/agent-skills, 165 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scaling?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.