Agent skill

Building Agents

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server —…

MITAuto-check passedAI & LLM Engineering

Install Building Agents

skills CLI
$ npx skills add ericrisco/rsc-harness --skill building-agents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness building-agents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/building-agents .claude/skills/building-agents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
building-agents
GitHub stars
167
Token cost
~5k tokens
SKILL.md length
959 words
Files
9 (incl. scripts, references)
Skills in repo
227
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server —…

  • Works in 6 steps: Adapter first — define the LLMProvider… → Smallest loop that works — single-agent… → Tools are typed contracts — schema +… → …
  • Restructuring an LLM agent — provider adapter
  • SKILL.md covers The one rule, Decision rules (read before…, The provider adapter (the… and Good vs Bad, plus 9 more sections
  • Runs Shell scripts from its folder; calls bash

What it does

Building Agents is an agent skill from ericrisco/rsc-harness. Use when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change. NOT vector-store SQL alone (that is postgresdb) or service deployment (that is deployment).

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/agent-loops-and-harness.md`).

It sits in AI & LLM Engineering, covering Structured output and tool calling, Building AI agents and Autonomous loops. It works with Model Context Protocol, OpenAI and SQL. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Restructuring an LLM agent — provider adapter
  • Structured output
  • MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change

Example prompts

  • “/building-agents”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Adapter first — define the LLMProvider Protocol before any provider call.
  2. Smallest loop that works — single-agent before multi-agent; ReAct only when the path is uncertain; plan-execute when steps are knowable…
  3. Tools are typed contracts — schema + validation + idempotency key on every side-effecting tool; no catch-all tools.
  4. Retrieve, don't stuff — RAG when ground truth lives in data; cite or refuse.
  5. Eval before ship — a golden set + regression gate in CI, or it's not production.
  6. Cheapest model that passes the eval — route/cascade up, never default to flagship.

What it can do on your machine

Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Building Agents loads about 5k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 91 tokens; SKILL.md has 959 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 959 words, ~4,959 tokens.

Download SKILL.mdSave it as .claude/skills/building-agents/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
building-agents
description
Use when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change. NOT vector-store SQL alone (that is `postgresdb`) or service deployment (that is `deployment`).
tags
agents, llm, mcp, rag, evals, ai
recommends
secure-coding, deployment
origin
risco

Building production LLM agents (model-agnostic)

A thin provider adapter, a disciplined agent loop, schema-validated tools, provider-neutral RAG, eval gates, OTel tracing, and optionally an MCP server — so swapping OpenAI ↔ Anthropic ↔ Gemini ↔ OSS is a config change, not a rewrite.

The one rule

Program against a capability interface, never a vendor SDK. Vendor specifics (model id, tool-schema shape, JSON mode, caching, token limits) live behind one adapter resolved from config. Model names and prices rot — if one appears in business logic it's a bug, and re-verify the dated tables before quoting a number.

Hand off instead when: a new non-trivial feature has no approved spec + plan under 02-DOCS/wiki/sdd/ → stop and run specify first (method: sdd), which routes back here once the plan is approved; one-line/low-risk changes go straight through. Anthropic-SDK internals (caching, thinking, batch) in a file that only imports anthropic → claude-api if your environment has it, since this skill stays multi-provider. Workspace scaffolding → harness. Choosing which coding agent to use → agent-eval territory. Pure prompt-wording tuning with no architecture change → prompt engineering, not this. A one-shot throwaway prompt, or no retrieval/tools/loop/evals at all → you don't need an agent; call the SDK directly and say so.

Decision rules (read before writing code)

  1. Adapter first — define the LLMProvider Protocol before any provider call.
  2. Smallest loop that works — single-agent before multi-agent; ReAct only when the path is uncertain; plan-execute when steps are knowable. Multi-agent means orchestrator-worker with a semaphore-bounded parallel fan-out, never a free-for-all.
  3. Tools are typed contracts — schema + validation + idempotency key on every side-effecting tool; no catch-all tools.
  4. Retrieve, don't stuff — RAG when ground truth lives in data; cite or refuse.
  5. Eval before ship — a golden set + regression gate in CI, or it's not production.
  6. Cheapest model that passes the eval — route/cascade up, never default to flagship.

The provider adapter (the heart of the skill)

The one payload to internalize. Python 3.12+, Pydantic v2, async so it composes directly with the bounded loop (and orchestrator-worker fan-out) in references/agent-loops-and-harness.md. Structured output is the quirk that differs most per vendor: strict JSON Schema (OpenAI), tool-forcing (Anthropic), response_json_schema (Gemini). Streaming, the Gemini and OSS/litellm adapters, tool-result plumbing, and a route() registry live in references/provider-abstraction.md — this excerpt is the load-bearing core, not the whole interface.

python
from __future__ import annotations

import os
from typing import Literal, Protocol, runtime_checkable

from pydantic import BaseModel, Field


class Message(BaseModel):
    role: Literal["system", "user", "assistant", "tool"]
    content: str


class ToolSpec(BaseModel):
    name: str
    description: str
    parameters: dict  # JSON Schema for the tool's arguments


class Usage(BaseModel):
    input_tokens: int = 0
    output_tokens: int = 0
    cost_usd: float = 0.0


class CompletionRequest(BaseModel):
    model: str  # resolved from config, e.g. "claude-sonnet-4-6" — never literal in logic
    messages: list[Message]
    tools: list[ToolSpec] = Field(default_factory=list)
    response_schema: dict | None = None  # JSON Schema -> structured output
    temperature: float = 0.0
    max_tokens: int = 1024


class CompletionResponse(BaseModel):
    text: str = ""
    tool_calls: list[dict] = Field(default_factory=list)  # [{id, name, arguments}]
    usage: Usage = Field(default_factory=Usage)
    raw: dict | None = None


@runtime_checkable
class LLMProvider(Protocol):
    # Async so it drives the async agent loop directly. The full interface in
    # references/provider-abstraction.md adds stream() and embed().
    async def complete(self, req: CompletionRequest) -> CompletionResponse: ...


class OpenAIAdapter:
    def __init__(self, model: str) -> None:
        from openai import AsyncOpenAI

        self.model, self.client = model, AsyncOpenAI()

    async def complete(self, req: CompletionRequest) -> CompletionResponse:
        # Chat Completions shape (universal, still current); references/provider-abstraction.md
        # gives the preferred Responses-API adapter. system stays a `system` role message here.
        kwargs: dict = {"model": self.model, "messages": [m.model_dump() for m in req.messages],
                        "temperature": req.temperature, "max_tokens": req.max_tokens}
        if req.tools:
            kwargs["tools"] = [{"type": "function", "function": {"name": t.name, "description": t.description, "parameters": t.parameters}} for t in req.tools]
        if req.response_schema:
            kwargs["response_format"] = {"type": "json_schema", "json_schema": {"name": "out", "schema": req.response_schema, "strict": True}}
        r = await self.client.chat.completions.create(**kwargs)
        msg = r.choices[0].message
        calls = [{"id": c.id, "name": c.function.name, "arguments": c.function.arguments} for c in (msg.tool_calls or [])]
        return CompletionResponse(text=msg.content or "", tool_calls=calls, raw=r.model_dump(),
            usage=Usage(input_tokens=r.usage.prompt_tokens, output_tokens=r.usage.completion_tokens))


class AnthropicAdapter:
    def __init__(self, model: str) -> None:
        from anthropic import AsyncAnthropic

        self.model, self.client = model, AsyncAnthropic()

    async def complete(self, req: CompletionRequest) -> CompletionResponse:
        # QUIRKS: system is a top-level param (not a message); tools use input_schema (not function).
        system = "\n".join(m.content for m in req.messages if m.role == "system") or None
        turns = [{"role": m.role, "content": m.content} for m in req.messages if m.role != "system"]
        kwargs: dict = {"model": self.model, "system": system, "messages": turns, "max_tokens": req.max_tokens, "temperature": req.temperature}
        if req.tools:
            kwargs["tools"] = [{"name": t.name, "description": t.description, "input_schema": t.parameters} for t in req.tools]
        if req.response_schema:  # structured output via tool-forcing
            kwargs["tools"] = [{"name": "out", "description": "Emit the result", "input_schema": req.response_schema}]
            kwargs["tool_choice"] = {"type": "tool", "name": "out"}
        r = await self.client.messages.create(**kwargs)
        text = "".join(b.text for b in r.content if b.type == "text")
        calls = [{"id": b.id, "name": b.name, "arguments": b.input} for b in r.content if b.type == "tool_use"]
        return CompletionResponse(text=text, tool_calls=calls, raw=r.model_dump(),
            usage=Usage(input_tokens=r.usage.input_tokens, output_tokens=r.usage.output_tokens))


def get_provider(spec: str | None = None) -> LLMProvider:
    """Parse 'provider:model' (default from env LLM) into a concrete adapter."""
    provider, _, model = (spec or os.environ["LLM"]).partition(":")
    if provider == "openai":
        return OpenAIAdapter(model)
    if provider == "anthropic":
        return AnthropicAdapter(model)
    raise ValueError(f"unknown provider: {provider!r}")
# Gemini + OSS/litellm adapters, streaming, tool-result plumbing, and route() registry
# -> references/provider-abstraction.md

Good vs Bad

Call-sites use the adapter and never name a model: provider = get_provider(settings.llm) (e.g. "anthropic:claude-sonnet-4-6"), then await provider.complete(req). The two failures that survive that discipline:

python
# BAD — parse-and-pray; wrong shape fails silently at 3am.
raw = (await provider.complete(req)).text
try:
    data = json.loads(raw)
except json.JSONDecodeError:
    data = {}  # the bug is now invisible
python
# GOOD — strict structured output + schema validation that fails loudly on drift.
class Answer(BaseModel):
    sentiment: Literal["pos", "neg", "neu"]
    score: float

req.response_schema = Answer.model_json_schema()
ans = Answer.model_validate_json((await provider.complete(req)).text)
python
# BAD — unbounded loop; no cap/timeout/idempotency. Burns budget, repeats side effects, wedges.
while True:
    resp = await provider.complete(req)
    if not resp.tool_calls:
        break
    for call in resp.tool_calls:
        await run_tool(call)
python
# GOOD — bounded loop: step cap + per-tool timeout + idempotency key (safe to retry).
for step in range(max_steps):
    resp = await provider.complete(req)
    if not resp.tool_calls:
        break
    for call in resp.tool_calls:
        async with asyncio.timeout(tool_timeout_s):
            await run_tool(call, idempotency_key=call["id"])
# full loop, budgets, recovery -> references/agent-loops-and-harness.md

Tools & structured output (minimum viable)

python
from typing import Callable, Literal

from pydantic import BaseModel, ConfigDict, Field, ValidationError


class CreateInvoiceArgs(BaseModel):
    model_config = ConfigDict(extra="forbid")  # reject unknown keys from the model
    customer_id: str = Field(min_length=1)
    amount_cents: int = Field(gt=0)
    currency: Literal["EUR", "USD"] = "EUR"


class ToolResult(BaseModel):
    status: Literal["success", "warning", "error"]
    summary: str
    data: dict | None = None
    next_actions: list[str] = Field(default_factory=list)


def _create_invoice(args: CreateInvoiceArgs) -> ToolResult:
    invoice_id = f"inv_{args.customer_id}_{args.amount_cents}"  # real impl: DB insert + idempotency
    return ToolResult(status="success", summary=f"Created {invoice_id}", data={"id": invoice_id})


TOOLS: dict[str, tuple[type[BaseModel], Callable]] = {
    "create_invoice": (CreateInvoiceArgs, _create_invoice),
}


def dispatch(name: str, raw_args: dict) -> ToolResult:
    spec = TOOLS.get(name)
    if spec is None:
        return ToolResult(status="error", summary=f"unknown tool {name!r}", next_actions=["pick a registered tool"])
    args_model, handler = spec
    try:
        args = args_model.model_validate(raw_args)  # validate BEFORE side effects
    except ValidationError as e:
        return ToolResult(status="error", summary="invalid args", data={"errors": e.errors()},
                          next_actions=["fix the arguments and retry"])
    return handler(args)

Schema design, sandboxing, idempotency, DI-scoped DB sessions, plus the RAG internals below — chunking, hybrid RRF, rerank, the citation grader, memory — are in references/tools-and-rag.md.

RAG in 30 lines (provider-agnostic embeddings)

sql
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE IF NOT EXISTS docs (
    id        bigserial PRIMARY KEY,
    content   text NOT NULL,
    embedding vector(1536) NOT NULL,
    meta      jsonb NOT NULL DEFAULT '{}'
);
CREATE INDEX IF NOT EXISTS docs_embedding_hnsw
    ON docs USING hnsw (embedding vector_cosine_ops);
python
async def embed(texts: list[str]) -> list[list[float]]:
    # Same provider interface as completions; impl in references/tools-and-rag.md.
    return await provider.embed(texts)  # returns one 1536-d vector per text


async def retrieve(query: str, k: int = 5, min_sim: float = 0.25) -> list[dict]:
    [q] = await embed([query])
    rows = await db.fetch(  # cosine distance <=>; similarity = 1 - distance
        "SELECT id, content, 1 - (embedding <=> $1) AS sim "
        "FROM docs ORDER BY embedding <=> $1 LIMIT $2",
        q, k,
    )
    return [dict(r) for r in rows if r["sim"] >= min_sim]


async def answer(query: str) -> str:
    chunks = await retrieve(query)
    if not chunks:                                   # refuse rather than hallucinate
        return "I don't have grounded information to answer that."
    context = "\n".join(f"[{c['id']}] {c['content']}" for c in chunks)
    req = CompletionRequest(
        model=settings.model_id,
        messages=[Message(role="system", content="Answer ONLY from context; cite chunk ids like [12]."),
                  Message(role="user", content=f"{context}\n\nQ: {query}")],
    )
    return (await provider.complete(req)).text

Evals & cost gates (the production line)

python
import json
import statistics
import sys
import time


async def run_eval(golden_path: str, graders: list, thresholds: dict[str, float]) -> None:
    cases = [json.loads(line) for line in open(golden_path)]  # {"input","expected","meta"}
    results = []
    for case in cases:
        t0 = time.perf_counter()
        out = await provider.complete(CompletionRequest(model=settings.model_id,
              messages=[Message(role="user", content=case["input"])]))
        scores = {g.name: g.grade(case, out) for g in graders}  # exact / schema / LLM-judge
        results.append({"scores": scores, "cost": out.usage.cost_usd,
                        "ms": (time.perf_counter() - t0) * 1000})
    n = len(results)
    metrics = {
        "accuracy": sum(r["scores"]["exact"] for r in results) / n,
        "faithfulness": sum(r["scores"]["judge"] for r in results) / n,
        "p95_latency_ms": statistics.quantiles([r["ms"] for r in results], n=20)[-1],
        "cost_per_task": sum(r["cost"] for r in results) / n,
    }
    failed = [k for k, lo in thresholds.items() if metrics[k] < lo]
    print(json.dumps(metrics, indent=2))
    sys.exit(1 if failed else 0)  # CI gate: non-zero blocks the merge

Routing cascade in one line: route(task) → cheapest model whose eval passes; escalate only on a failed self-check. Full runner, judge, CI gate, caching, batching and budgets → references/evals-and-observability.md.

Observability (OTel GenAI, vendor-neutral)

python
from opentelemetry import trace

tracer = trace.get_tracer("agent")


async def traced_complete(provider: LLMProvider, req: CompletionRequest) -> CompletionResponse:
    with tracer.start_as_current_span("chat") as span:
        span.set_attribute("gen_ai.system", settings.llm.split(":")[0])
        span.set_attribute("gen_ai.request.model", req.model)
        resp = await provider.complete(req)
        span.set_attributes({"gen_ai.usage.input_tokens": resp.usage.input_tokens,
                             "gen_ai.usage.output_tokens": resp.usage.output_tokens,
                             "gen_ai.usage.cost_usd": resp.usage.cost_usd})
        return resp
# Langfuse / Phoenix / Braintrust are swappable OTLP backends: emit spans, swap the exporter.
# span-per-tool, trace-id propagation, exporters -> references/evals-and-observability.md

MCP: when and the smallest server

Native tools when the agent and tools share a process/repo. MCP when tools must be reused across clients/teams or run out-of-process — accept the MCP cost (schema tokens, transport, ops) in exchange for reuse. TypeScript server, transports, HTTP+auth and testing are in references/mcp-servers.md.

python
from fastmcp import FastMCP  # standalone fastmcp 2.x

mcp = FastMCP("invoices")


@mcp.tool()
def create_invoice(customer_id: str, amount_cents: int, currency: str = "EUR") -> dict:
    """Create an invoice. amount_cents must be > 0."""
    if amount_cents <= 0:
        raise ValueError("amount_cents must be positive")
    return {"id": f"inv_{customer_id}_{amount_cents}", "currency": currency}


@mcp.resource("invoice://{invoice_id}")
def read_invoice(invoice_id: str) -> str:
    """Read-only invoice lookup by id."""
    return f"Invoice {invoice_id}: status=open"


if __name__ == "__main__":
    mcp.run()  # stdio transport
# (MCP spec 2025-11-25; stateless-core RC 2026-07-28; verify before quoting)
Show full SKILL.md (440 more words)Show less

Anti-patterns

Anti-patternReality
"I'll just call the OpenAI SDK directly, we'll never switch"The adapter is ~40 lines; retrofitting it across 30 call-sites later is a rewrite. Adapter first.
"JSON output is usually valid, I'll parse it""Usually" = pages at 3am. Use strict structured output + schema validation.
"The agent loop works, I don't need a step cap"Unbounded loops burn budget and wedge on errors. Cap steps, timeouts, and budget.
"One mega-tool that takes a freeform command is flexible"It's unobservable and unsafe. Narrow typed tools with idempotency keys.
"We can eval by eyeballing outputs"Vibes don't gate CI. Golden set + graders + threshold or it's not production.
"Default everything to the flagship model, it's smartest"5–20× cost for no measured gain. Route to the cheapest model that passes the eval.
"Stuff the whole doc in the prompt instead of RAG"Blows context + cost and still hallucinates. Retrieve + cite + refuse.
"Retry on every exception"Retrying a 400/401 wastes budget. Retry only transient (429/5xx/timeout) with backoff+jitter.
"Hardcode the model name, it's fine"Names rot (Opus 4.7 → 4.8 in weeks). Resolve from config/registry.
"MCP for everything"In-process native tools are simpler and faster when reuse isn't needed. MCP only for cross-client reuse.
"Tool results just return the raw API blob"Give the model status/summary/next_actions; raw blobs waste context and stall recovery.
"Prompt caching is Anthropic-only so skip caching"Each provider has its own caching/dedup; abstract it behind the adapter, don't skip it.

verify.sh

scripts/verify.sh lints example agent code and dry-runs the eval smoke test in the user's project — not in this skill repo. It detects each tool (ruff, mypy, tsc/node, go, the eval entrypoint, markdownlint) and skips any that are missing with a yellow WARN; a missing tool never fails the run. Invoke it with bash scripts/verify.sh from the project root. Exit 0 means clean (or only skips); a non-zero exit means a real lint/typecheck/vet/eval failure.

Project grounding

In a project with a 02-DOCS/ layer (harness), read 02-DOCS/wiki/stack/agents.md first on every use and stay consistent with it. If it is missing or stale, write this project's real choices there — provider(s) and model routing, where the adapter lives, tool/RAG conventions, eval gates, observability backend — as a type: stack article per the harness wiki-article-template.md, index it in 02-DOCS/wiki/index.md, and bump its timestamp in the same change as any convention change. No 02-DOCS/? Skip silently — technical conventions here are recorded, not gated; never block the task on this.

See also

  • Stacks the examples target: fastapi, nextjs, go, postgresdb, flutter. Harden with secure-coding, ship with deployment.
  • External (no sibling here; use if your environment provides them): claude-api for Anthropic-SDK-only tuning, deep-research for the research-harness fan-out / verify pattern.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/building-agents of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/agent-loops-and-harness.md
  • references/evals-and-observability.md
  • references/mcp-servers.md
  • references/provider-abstraction.md
  • references/tools-and-rag.md
  • scripts/verify.sh

Open the folder on GitHubat commit e3d5b33

Compare with similar skills

Building Agents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Building Agents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Building Agents this skillericrisco/rsc-harness167—~5kAutomated safety check: PassMIT
LangchainOrchestra-Research/AI-Research-SKILLs13k2 repos~3.2kAutomated safety check: PassMIT
Agent Harness DesignAnastasiyaW/codex-claude-code-config154—~764Automated safety check: PassMIT
Agent Squad for TypeScript2FastLabs/agent-squad7.8k—~4.3kAutomated safety check: PassApache-2.0
Tool Designagentailor/fullstack-langgraph-nextjs-agent132—~3.2kAutomated safety check: PassMIT
Qianwenai Wikichujianyun/skills740—~717Automated safety check: PassCustom licence

Similar skills

  • Langchain

    Orchestra-Research/AI-Research-SKILLs

    Framework for building LLM-powered applications with agents, chains, and RAG.

    13k GitHub starsUsed in 2 repos~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Harness Design

    AnastasiyaW/codex-claude-code-config

    Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…

    154 GitHub stars~764 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Agent Squad for TypeScript

    2FastLabs/agent-squad

    Guide to building Node.js and TypeScript apps on the agent-squad package: orchestrator, agent types, classifier routing, storage, retrievers and MCP tools.

    7.8k GitHub stars~4.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Tool Design

    agentailor/fullstack-langgraph-nextjs-agent

    Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

    132 GitHub stars~3.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Qianwenai Wiki

    chujianyun/skills

    千问AI平台(Qianwen AI Platform / DashScope)官方文档离线知识库,用于检索并回答模型选择、API Key、OpenAI 兼容接口、DashScope SDK、文本与多模态生成、图像/视频/语音、Realtime API、Embedding、Reranking、Function Calling、MCP、批量调用、计费、Token Plan、API/SDK/CLI…

    740 GitHub stars~717 tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub starsUsed in 1 repo~3.9k tokens
    AI & LLM EngineeringAuto-check passed

More from ericrisco/rsc-harness

All 227 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    167 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    167 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    167 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    167 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    167 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Building Agents

What does Building Agents do?

A skill your agent uses when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server —…. Building Agents is an agent skill from ericrisco/rsc-harness. Use when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change.

When should I use Building Agents?

Building Agents fits situations like: restructuring an LLM agent — provider adapter; structured output; MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change.

How do I install Building Agents in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill building-agents -a claude-code`. Or copy the skill folder (skills/building-agents in ericrisco/rsc-harness) into .claude/skills/building-agents in your project. Claude Code loads it when a task matches its description.

How do I install Building Agents in Codex?

Run `npx skills add ericrisco/rsc-harness --skill building-agents -a codex`. Or copy the skill folder (skills/building-agents in ericrisco/rsc-harness) into .agents/skills/building-agents in your project. Codex loads it when a task matches its description.

Can I use Building Agents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill building-agents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/building-agents, .gemini/skills/building-agents, .github/skills/building-agents and .opencode/skills/building-agents in your project.

What does Building Agents need to run?

Going by SKILL.md and its folder, Building Agents needs a shell for the scripts in its folder and the command-line tools its instructions call (bash). Our summary lists: Python 3; A Bash shell.

Does Building Agents access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Building Agents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Building Agents use?

Building Agents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Building Agents use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.

What are the alternatives to Building Agents?

Skills that share tags, products or a category with Building Agents: Langchain (Orchestra-Research/AI-Research-SKILLs, 13k stars), Agent Harness Design (AnastasiyaW/codex-claude-code-config, 154 stars), Agent Squad for TypeScript (2FastLabs/agent-squad, 7.8k stars) and Tool Design (agentailor/fullstack-langgraph-nextjs-agent, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Building Agents?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.