Langchain
Orchestra-Research/AI-Research-SKILLs
Framework for building LLM-powered applications with agents, chains, and RAG.
A skill your agent uses when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server —…
$ npx skills add ericrisco/rsc-harness --skill building-agents -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness building-agents --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/building-agents .claude/skills/building-agents && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "building-agents" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/building-agents into .claude/skills/building-agents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-agents", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/building-agentsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill building-agents -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness building-agents --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/building-agents .agents/skills/building-agents && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "building-agents" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/building-agents into .agents/skills/building-agents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-agents", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill building-agents -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness building-agents --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/building-agents .cursor/skills/building-agents && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "building-agents" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/building-agents into .cursor/skills/building-agents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-agents", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/building-agents--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill building-agents -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness building-agents --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/building-agents .gemini/skills/building-agents && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "building-agents" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/building-agents into .gemini/skills/building-agents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-agents", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness building-agentsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill building-agents -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/building-agents .github/skills/building-agents && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "building-agents" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/building-agents into .github/skills/building-agents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-agents", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill building-agents -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness building-agents --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/building-agents .opencode/skills/building-agents && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "building-agents" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/building-agents into .opencode/skills/building-agents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-agents", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
building-agentsA skill your agent uses when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server —…
Building Agents is an agent skill from ericrisco/rsc-harness. Use when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change. NOT vector-store SQL alone (that is postgresdb) or service deployment (that is deployment).
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/agent-loops-and-harness.md`).
It sits in AI & LLM Engineering, covering Structured output and tool calling, Building AI agents and Autonomous loops. It works with Model Context Protocol, OpenAI and SQL. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
bashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Building Agents loads about 5k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 91 tokens; SKILL.md has 959 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 959 words, ~4,959 tokens.
.claude/skills/building-agents/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.A thin provider adapter, a disciplined agent loop, schema-validated tools, provider-neutral RAG, eval gates, OTel tracing, and optionally an MCP server — so swapping OpenAI ↔ Anthropic ↔ Gemini ↔ OSS is a config change, not a rewrite.
Program against a capability interface, never a vendor SDK. Vendor specifics (model id, tool-schema shape, JSON mode, caching, token limits) live behind one adapter resolved from config. Model names and prices rot — if one appears in business logic it's a bug, and re-verify the dated tables before quoting a number.
Hand off instead when: a new non-trivial feature has no approved spec + plan under 02-DOCS/wiki/sdd/ → stop and run specify first (method: sdd), which routes back here once the plan is approved; one-line/low-risk changes go straight through. Anthropic-SDK internals (caching, thinking, batch) in a file that only imports anthropic → claude-api if your environment has it, since this skill stays multi-provider. Workspace scaffolding → harness. Choosing which coding agent to use → agent-eval territory. Pure prompt-wording tuning with no architecture change → prompt engineering, not this. A one-shot throwaway prompt, or no retrieval/tools/loop/evals at all → you don't need an agent; call the SDK directly and say so.
LLMProvider Protocol before any provider call.The one payload to internalize. Python 3.12+, Pydantic v2, async so it composes directly with the bounded loop (and orchestrator-worker fan-out) in references/agent-loops-and-harness.md. Structured output is the quirk that differs most per vendor: strict JSON Schema (OpenAI), tool-forcing (Anthropic), response_json_schema (Gemini). Streaming, the Gemini and OSS/litellm adapters, tool-result plumbing, and a route() registry live in references/provider-abstraction.md — this excerpt is the load-bearing core, not the whole interface.
from __future__ import annotations
import os
from typing import Literal, Protocol, runtime_checkable
from pydantic import BaseModel, Field
class Message(BaseModel):
role: Literal["system", "user", "assistant", "tool"]
content: str
class ToolSpec(BaseModel):
name: str
description: str
parameters: dict # JSON Schema for the tool's arguments
class Usage(BaseModel):
input_tokens: int = 0
output_tokens: int = 0
cost_usd: float = 0.0
class CompletionRequest(BaseModel):
model: str # resolved from config, e.g. "claude-sonnet-4-6" — never literal in logic
messages: list[Message]
tools: list[ToolSpec] = Field(default_factory=list)
response_schema: dict | None = None # JSON Schema -> structured output
temperature: float = 0.0
max_tokens: int = 1024
class CompletionResponse(BaseModel):
text: str = ""
tool_calls: list[dict] = Field(default_factory=list) # [{id, name, arguments}]
usage: Usage = Field(default_factory=Usage)
raw: dict | None = None
@runtime_checkable
class LLMProvider(Protocol):
# Async so it drives the async agent loop directly. The full interface in
# references/provider-abstraction.md adds stream() and embed().
async def complete(self, req: CompletionRequest) -> CompletionResponse: ...
class OpenAIAdapter:
def __init__(self, model: str) -> None:
from openai import AsyncOpenAI
self.model, self.client = model, AsyncOpenAI()
async def complete(self, req: CompletionRequest) -> CompletionResponse:
# Chat Completions shape (universal, still current); references/provider-abstraction.md
# gives the preferred Responses-API adapter. system stays a `system` role message here.
kwargs: dict = {"model": self.model, "messages": [m.model_dump() for m in req.messages],
"temperature": req.temperature, "max_tokens": req.max_tokens}
if req.tools:
kwargs["tools"] = [{"type": "function", "function": {"name": t.name, "description": t.description, "parameters": t.parameters}} for t in req.tools]
if req.response_schema:
kwargs["response_format"] = {"type": "json_schema", "json_schema": {"name": "out", "schema": req.response_schema, "strict": True}}
r = await self.client.chat.completions.create(**kwargs)
msg = r.choices[0].message
calls = [{"id": c.id, "name": c.function.name, "arguments": c.function.arguments} for c in (msg.tool_calls or [])]
return CompletionResponse(text=msg.content or "", tool_calls=calls, raw=r.model_dump(),
usage=Usage(input_tokens=r.usage.prompt_tokens, output_tokens=r.usage.completion_tokens))
class AnthropicAdapter:
def __init__(self, model: str) -> None:
from anthropic import AsyncAnthropic
self.model, self.client = model, AsyncAnthropic()
async def complete(self, req: CompletionRequest) -> CompletionResponse:
# QUIRKS: system is a top-level param (not a message); tools use input_schema (not function).
system = "\n".join(m.content for m in req.messages if m.role == "system") or None
turns = [{"role": m.role, "content": m.content} for m in req.messages if m.role != "system"]
kwargs: dict = {"model": self.model, "system": system, "messages": turns, "max_tokens": req.max_tokens, "temperature": req.temperature}
if req.tools:
kwargs["tools"] = [{"name": t.name, "description": t.description, "input_schema": t.parameters} for t in req.tools]
if req.response_schema: # structured output via tool-forcing
kwargs["tools"] = [{"name": "out", "description": "Emit the result", "input_schema": req.response_schema}]
kwargs["tool_choice"] = {"type": "tool", "name": "out"}
r = await self.client.messages.create(**kwargs)
text = "".join(b.text for b in r.content if b.type == "text")
calls = [{"id": b.id, "name": b.name, "arguments": b.input} for b in r.content if b.type == "tool_use"]
return CompletionResponse(text=text, tool_calls=calls, raw=r.model_dump(),
usage=Usage(input_tokens=r.usage.input_tokens, output_tokens=r.usage.output_tokens))
def get_provider(spec: str | None = None) -> LLMProvider:
"""Parse 'provider:model' (default from env LLM) into a concrete adapter."""
provider, _, model = (spec or os.environ["LLM"]).partition(":")
if provider == "openai":
return OpenAIAdapter(model)
if provider == "anthropic":
return AnthropicAdapter(model)
raise ValueError(f"unknown provider: {provider!r}")
# Gemini + OSS/litellm adapters, streaming, tool-result plumbing, and route() registry
# -> references/provider-abstraction.mdCall-sites use the adapter and never name a model: provider = get_provider(settings.llm) (e.g. "anthropic:claude-sonnet-4-6"), then await provider.complete(req). The two failures that survive that discipline:
# BAD — parse-and-pray; wrong shape fails silently at 3am.
raw = (await provider.complete(req)).text
try:
data = json.loads(raw)
except json.JSONDecodeError:
data = {} # the bug is now invisible# GOOD — strict structured output + schema validation that fails loudly on drift.
class Answer(BaseModel):
sentiment: Literal["pos", "neg", "neu"]
score: float
req.response_schema = Answer.model_json_schema()
ans = Answer.model_validate_json((await provider.complete(req)).text)# BAD — unbounded loop; no cap/timeout/idempotency. Burns budget, repeats side effects, wedges.
while True:
resp = await provider.complete(req)
if not resp.tool_calls:
break
for call in resp.tool_calls:
await run_tool(call)# GOOD — bounded loop: step cap + per-tool timeout + idempotency key (safe to retry).
for step in range(max_steps):
resp = await provider.complete(req)
if not resp.tool_calls:
break
for call in resp.tool_calls:
async with asyncio.timeout(tool_timeout_s):
await run_tool(call, idempotency_key=call["id"])
# full loop, budgets, recovery -> references/agent-loops-and-harness.mdfrom typing import Callable, Literal
from pydantic import BaseModel, ConfigDict, Field, ValidationError
class CreateInvoiceArgs(BaseModel):
model_config = ConfigDict(extra="forbid") # reject unknown keys from the model
customer_id: str = Field(min_length=1)
amount_cents: int = Field(gt=0)
currency: Literal["EUR", "USD"] = "EUR"
class ToolResult(BaseModel):
status: Literal["success", "warning", "error"]
summary: str
data: dict | None = None
next_actions: list[str] = Field(default_factory=list)
def _create_invoice(args: CreateInvoiceArgs) -> ToolResult:
invoice_id = f"inv_{args.customer_id}_{args.amount_cents}" # real impl: DB insert + idempotency
return ToolResult(status="success", summary=f"Created {invoice_id}", data={"id": invoice_id})
TOOLS: dict[str, tuple[type[BaseModel], Callable]] = {
"create_invoice": (CreateInvoiceArgs, _create_invoice),
}
def dispatch(name: str, raw_args: dict) -> ToolResult:
spec = TOOLS.get(name)
if spec is None:
return ToolResult(status="error", summary=f"unknown tool {name!r}", next_actions=["pick a registered tool"])
args_model, handler = spec
try:
args = args_model.model_validate(raw_args) # validate BEFORE side effects
except ValidationError as e:
return ToolResult(status="error", summary="invalid args", data={"errors": e.errors()},
next_actions=["fix the arguments and retry"])
return handler(args)Schema design, sandboxing, idempotency, DI-scoped DB sessions, plus the RAG internals below — chunking, hybrid RRF, rerank, the citation grader, memory — are in references/tools-and-rag.md.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE IF NOT EXISTS docs (
id bigserial PRIMARY KEY,
content text NOT NULL,
embedding vector(1536) NOT NULL,
meta jsonb NOT NULL DEFAULT '{}'
);
CREATE INDEX IF NOT EXISTS docs_embedding_hnsw
ON docs USING hnsw (embedding vector_cosine_ops);async def embed(texts: list[str]) -> list[list[float]]:
# Same provider interface as completions; impl in references/tools-and-rag.md.
return await provider.embed(texts) # returns one 1536-d vector per text
async def retrieve(query: str, k: int = 5, min_sim: float = 0.25) -> list[dict]:
[q] = await embed([query])
rows = await db.fetch( # cosine distance <=>; similarity = 1 - distance
"SELECT id, content, 1 - (embedding <=> $1) AS sim "
"FROM docs ORDER BY embedding <=> $1 LIMIT $2",
q, k,
)
return [dict(r) for r in rows if r["sim"] >= min_sim]
async def answer(query: str) -> str:
chunks = await retrieve(query)
if not chunks: # refuse rather than hallucinate
return "I don't have grounded information to answer that."
context = "\n".join(f"[{c['id']}] {c['content']}" for c in chunks)
req = CompletionRequest(
model=settings.model_id,
messages=[Message(role="system", content="Answer ONLY from context; cite chunk ids like [12]."),
Message(role="user", content=f"{context}\n\nQ: {query}")],
)
return (await provider.complete(req)).textimport json
import statistics
import sys
import time
async def run_eval(golden_path: str, graders: list, thresholds: dict[str, float]) -> None:
cases = [json.loads(line) for line in open(golden_path)] # {"input","expected","meta"}
results = []
for case in cases:
t0 = time.perf_counter()
out = await provider.complete(CompletionRequest(model=settings.model_id,
messages=[Message(role="user", content=case["input"])]))
scores = {g.name: g.grade(case, out) for g in graders} # exact / schema / LLM-judge
results.append({"scores": scores, "cost": out.usage.cost_usd,
"ms": (time.perf_counter() - t0) * 1000})
n = len(results)
metrics = {
"accuracy": sum(r["scores"]["exact"] for r in results) / n,
"faithfulness": sum(r["scores"]["judge"] for r in results) / n,
"p95_latency_ms": statistics.quantiles([r["ms"] for r in results], n=20)[-1],
"cost_per_task": sum(r["cost"] for r in results) / n,
}
failed = [k for k, lo in thresholds.items() if metrics[k] < lo]
print(json.dumps(metrics, indent=2))
sys.exit(1 if failed else 0) # CI gate: non-zero blocks the mergeRouting cascade in one line: route(task) → cheapest model whose eval passes; escalate only on a failed self-check.
Full runner, judge, CI gate, caching, batching and budgets →
references/evals-and-observability.md.
from opentelemetry import trace
tracer = trace.get_tracer("agent")
async def traced_complete(provider: LLMProvider, req: CompletionRequest) -> CompletionResponse:
with tracer.start_as_current_span("chat") as span:
span.set_attribute("gen_ai.system", settings.llm.split(":")[0])
span.set_attribute("gen_ai.request.model", req.model)
resp = await provider.complete(req)
span.set_attributes({"gen_ai.usage.input_tokens": resp.usage.input_tokens,
"gen_ai.usage.output_tokens": resp.usage.output_tokens,
"gen_ai.usage.cost_usd": resp.usage.cost_usd})
return resp
# Langfuse / Phoenix / Braintrust are swappable OTLP backends: emit spans, swap the exporter.
# span-per-tool, trace-id propagation, exporters -> references/evals-and-observability.mdNative tools when the agent and tools share a process/repo. MCP when tools must be reused across clients/teams or run out-of-process — accept the MCP cost (schema tokens, transport, ops) in exchange for reuse. TypeScript server, transports, HTTP+auth and testing are in references/mcp-servers.md.
from fastmcp import FastMCP # standalone fastmcp 2.x
mcp = FastMCP("invoices")
@mcp.tool()
def create_invoice(customer_id: str, amount_cents: int, currency: str = "EUR") -> dict:
"""Create an invoice. amount_cents must be > 0."""
if amount_cents <= 0:
raise ValueError("amount_cents must be positive")
return {"id": f"inv_{customer_id}_{amount_cents}", "currency": currency}
@mcp.resource("invoice://{invoice_id}")
def read_invoice(invoice_id: str) -> str:
"""Read-only invoice lookup by id."""
return f"Invoice {invoice_id}: status=open"
if __name__ == "__main__":
mcp.run() # stdio transport
# (MCP spec 2025-11-25; stateless-core RC 2026-07-28; verify before quoting)| Anti-pattern | Reality |
|---|---|
| "I'll just call the OpenAI SDK directly, we'll never switch" | The adapter is ~40 lines; retrofitting it across 30 call-sites later is a rewrite. Adapter first. |
| "JSON output is usually valid, I'll parse it" | "Usually" = pages at 3am. Use strict structured output + schema validation. |
| "The agent loop works, I don't need a step cap" | Unbounded loops burn budget and wedge on errors. Cap steps, timeouts, and budget. |
| "One mega-tool that takes a freeform command is flexible" | It's unobservable and unsafe. Narrow typed tools with idempotency keys. |
| "We can eval by eyeballing outputs" | Vibes don't gate CI. Golden set + graders + threshold or it's not production. |
| "Default everything to the flagship model, it's smartest" | 5–20× cost for no measured gain. Route to the cheapest model that passes the eval. |
| "Stuff the whole doc in the prompt instead of RAG" | Blows context + cost and still hallucinates. Retrieve + cite + refuse. |
| "Retry on every exception" | Retrying a 400/401 wastes budget. Retry only transient (429/5xx/timeout) with backoff+jitter. |
| "Hardcode the model name, it's fine" | Names rot (Opus 4.7 → 4.8 in weeks). Resolve from config/registry. |
| "MCP for everything" | In-process native tools are simpler and faster when reuse isn't needed. MCP only for cross-client reuse. |
| "Tool results just return the raw API blob" | Give the model status/summary/next_actions; raw blobs waste context and stall recovery. |
| "Prompt caching is Anthropic-only so skip caching" | Each provider has its own caching/dedup; abstract it behind the adapter, don't skip it. |
scripts/verify.sh lints example agent code and dry-runs the eval smoke test in the user's project — not in this skill repo. It detects each tool (ruff, mypy, tsc/node, go, the eval entrypoint, markdownlint) and skips any that are missing with a yellow WARN; a missing tool never fails the run. Invoke it with bash scripts/verify.sh from the project root. Exit 0 means clean (or only skips); a non-zero exit means a real lint/typecheck/vet/eval failure.
In a project with a 02-DOCS/ layer (harness), read 02-DOCS/wiki/stack/agents.md first on every use and stay consistent with it. If it is missing or stale, write this project's real choices there — provider(s) and model routing, where the adapter lives, tool/RAG conventions, eval gates, observability backend — as a type: stack article per the harness wiki-article-template.md, index it in 02-DOCS/wiki/index.md, and bump its timestamp in the same change as any convention change. No 02-DOCS/? Skip silently — technical conventions here are recorded, not gated; never block the task on this.
fastapi, nextjs, go, postgresdb, flutter. Harden with secure-coding, ship with deployment.claude-api for Anthropic-SDK-only tuning, deep-research for the research-harness fan-out / verify pattern.© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references) in skills/building-agents of ericrisco/rsc-harness.
Open the folder on GitHubat commit e3d5b33
Building Agents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Building Agents this skillericrisco/rsc-harness | 167 | — | ~5k | Automated safety check: Pass | MIT | |
| LangchainOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Agent Harness DesignAnastasiyaW/codex-claude-code-config | 154 | — | ~764 | Automated safety check: Pass | MIT | |
| Agent Squad for TypeScript2FastLabs/agent-squad | 7.8k | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | |
| Tool Designagentailor/fullstack-langgraph-nextjs-agent | 132 | — | ~3.2k | Automated safety check: Pass | MIT | |
| Qianwenai Wikichujianyun/skills | 740 | — | ~717 | Automated safety check: Pass | Custom licence |
Orchestra-Research/AI-Research-SKILLs
Framework for building LLM-powered applications with agents, chains, and RAG.
AnastasiyaW/codex-claude-code-config
Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…
2FastLabs/agent-squad
Guide to building Node.js and TypeScript apps on the agent-squad package: orchestrator, agent types, classifier routing, storage, retrievers and MCP tools.
agentailor/fullstack-langgraph-nextjs-agent
Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).
chujianyun/skills
千问AI平台(Qianwen AI Platform / DashScope)官方文档离线知识库,用于检索并回答模型选择、API Key、OpenAI 兼容接口、DashScope SDK、文本与多模态生成、图像/视频/语音、Realtime API、Embedding、Reranking、Function Calling、MCP、批量调用、计费、Token Plan、API/SDK/CLI…
langchain-ai/langchain-skills
INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server —…. Building Agents is an agent skill from ericrisco/rsc-harness. Use when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change.
Building Agents fits situations like: restructuring an LLM agent — provider adapter; structured output; MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change.
Run `npx skills add ericrisco/rsc-harness --skill building-agents -a claude-code`. Or copy the skill folder (skills/building-agents in ericrisco/rsc-harness) into .claude/skills/building-agents in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill building-agents -a codex`. Or copy the skill folder (skills/building-agents in ericrisco/rsc-harness) into .agents/skills/building-agents in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill building-agents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/building-agents, .gemini/skills/building-agents, .github/skills/building-agents and .opencode/skills/building-agents in your project.
Going by SKILL.md and its folder, Building Agents needs a shell for the scripts in its folder and the command-line tools its instructions call (bash). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Building Agents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Building Agents: Langchain (Orchestra-Research/AI-Research-SKILLs, 13k stars), Agent Harness Design (AnastasiyaW/codex-claude-code-config, 154 stars), Agent Squad for TypeScript (2FastLabs/agent-squad, 7.8k stars) and Tool Design (agentailor/fullstack-langgraph-nextjs-agent, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.