Deep Agents
langchain-ai/docs
Build batteries-included agents with planning, context management, subagent delegation, and sandboxed execution.
A skill your agent uses when designing, reviewing, or debugging how an agent's context window gets filled, pruned, or shared — choosing what loads at boot versus on demand, sizing an install or an…
$ npx skills add mvschwarz/openrig --skill context-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mvschwarz/openrig context-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mvschwarz/openrig.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/_canonical/process/context-engineering .claude/skills/context-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "context-engineering" agent skill from https://github.com/mvschwarz/openrig/tree/main/skills/_canonical/process/context-engineering into .claude/skills/context-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "context-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mvschwarz/openrig/tree/main/skills/_canonical/process/context-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mvschwarz/openrig --skill context-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mvschwarz/openrig context-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mvschwarz/openrig.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/_canonical/process/context-engineering .agents/skills/context-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "context-engineering" agent skill from https://github.com/mvschwarz/openrig/tree/main/skills/_canonical/process/context-engineering into .agents/skills/context-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "context-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mvschwarz/openrig --skill context-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mvschwarz/openrig context-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mvschwarz/openrig.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/_canonical/process/context-engineering .cursor/skills/context-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "context-engineering" agent skill from https://github.com/mvschwarz/openrig/tree/main/skills/_canonical/process/context-engineering into .cursor/skills/context-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "context-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mvschwarz/openrig.git --path skills/_canonical/process/context-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mvschwarz/openrig --skill context-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mvschwarz/openrig context-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mvschwarz/openrig.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/_canonical/process/context-engineering .gemini/skills/context-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "context-engineering" agent skill from https://github.com/mvschwarz/openrig/tree/main/skills/_canonical/process/context-engineering into .gemini/skills/context-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "context-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mvschwarz/openrig context-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mvschwarz/openrig --skill context-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mvschwarz/openrig.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/_canonical/process/context-engineering .github/skills/context-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "context-engineering" agent skill from https://github.com/mvschwarz/openrig/tree/main/skills/_canonical/process/context-engineering into .github/skills/context-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "context-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mvschwarz/openrig --skill context-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mvschwarz/openrig context-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mvschwarz/openrig.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/_canonical/process/context-engineering .opencode/skills/context-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "context-engineering" agent skill from https://github.com/mvschwarz/openrig/tree/main/skills/_canonical/process/context-engineering into .opencode/skills/context-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "context-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
context-engineeringA skill your agent uses when designing, reviewing, or debugging how an agent's context window gets filled, pruned, or shared — choosing what loads at boot versus on demand, sizing an install or an…
Context Engineering is an agent skill from mvschwarz/openrig. Use when designing, reviewing, or debugging how an agent's context window gets filled, pruned, or shared — choosing what loads at boot versus on demand, sizing an install or an always-loaded file, fixing an agent that drifts, repeats itself, or forgets constraints mid-task, planning compaction or summarization, deciding single-agent versus subagents, engineering handoffs between agents, or picking a tool loadout. NOT for rewording a prompt's tone, choosing which model to pin, or debugging business logic — those…
Its SKILL.md is about 8.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Context engineering and Summarization. It works with LangChain. The repository describes itself as: Build your own network of agents from Claude Code, Codex and Pi: persistent teams with roles, shared context and owned work. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a7fed63. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
anthropic.comdbreunig.comclaude.comcode.claude.complatform.claude.comtrychroma.comcognition.commanus.imlangchain.comphilschmid.degithub.comarxiv.orgopenai.comdevelopers.openai.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context Engineering loads about 8.2k tokens when it runs. Until then it costs about 172 tokens; SKILL.md has 4,150 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mvschwarz/openrig at commit a7fed63, republished under its Apache-2.0 licence (© mvschwarz). 4,150 words, ~8,162 tokens.
.claude/skills/context-engineering/SKILL.md (or your agent's skills folder).This skill is a provisional historical research snapshot of context-engineering practice as published in 2024–2025. It is NOT normative. Do not apply it prescriptively to current frontier models without re-verification. On any conflict, OpenRig current skills, explicit user rulings, and directly measured OpenRig practice PREVAIL over this document. Treat these areas as particularly suspect pending re-verification: compaction/summarization guidance, minimal-upfront versus broad orientation, just-in-time lookup assumptions, fixed context ceilings, and single-agent versus multi-agent advice.
Curated distillation of the best publicly available expertise on context engineering for coding agents, drawn from primary sources at Anthropic, OpenAI, and leading practitioners (Manus, Cognition, Chroma, LangChain, Drew Breunig, and others). Load on the moments the description names; it is deliberately not part of any base walk.
Definition. Context engineering is "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference" — everything that lands in the window: system instructions, tool definitions, retrieved data, message history, and tool outputs, not just the prompt text (Anthropic, Effective context engineering for AI agents). Andrej Karpathy's framing, popularized via LangChain: "the delicate art and science of filling the context window with just the right information for the next step" — the LLM is a CPU and the context window is its RAM, and your job is deciding what gets loaded into RAM at each step (LangChain, Context Engineering for Agents).
Why it superseded prompt engineering. A chatbot answers one question with whatever fits in one turn. An agent runs in a loop, accumulating tool results, file contents, and history across dozens or hundreds of steps. The improvements stop coming from rewording instructions and start coming from rewiring — what the agent retrieves, in what order, and what gets evicted when the window fills (Anthropic, ibid.). Philipp Schmid's formulation of the practical consequence: "Agent failures aren't only model failures; they are context failures." Most of the time when a capable model does something dumb, the context it was given made the dumb thing likely (Schmid, The New Skill in AI is Context Engineering).
The physical constraint: attention is a budget, not a bucket. Three mechanisms make context a scarce resource rather than free storage:
The one-sentence discipline. From Anthropic: "Find the smallest set of high-signal tokens that maximize the likelihood of some desired outcome." Everything else in this pack is a technique in service of that sentence.
The components you are engineering. Schmid's inventory is a useful checklist of what actually occupies the window: (1) system instructions, (2) the user's immediate request, (3) state/history of the current session, (4) long-term memory, (5) retrieved external information, (6) tool definitions, (7) output-format specifications (Schmid, ibid.). Each is a separate dial. When an agent misbehaves, walk this list asking "which of these is missing, stale, bloated, or contradictory?"
P1 — Minimal ≠ short; curate for signal density. The goal is the smallest sufficient set of tokens, not the shortest prompt. A system prompt should fully outline expected behavior; the sin is low-signal filler, not length (Anthropic, Effective context engineering).
P2 — Engineer for the next step, not the whole task. Context is curated per inference step (Karpathy via LangChain). The question is never "what might the agent ever need" but "what does this step need to succeed." This is why loading everything upfront loses to just-in-time retrieval on long tasks.
P3 — Degradation precedes exhaustion. Budget context well below the marketed window. Chroma showed serious degradation mid-window; Breunig collects the operational evidence: a Gemini agent's planning quality collapsed beyond ~100K tokens [perishable snapshot, 2025–2026: model/vendor-specific — teach the mechanism, re-verify the number] into repeating past actions, and Databricks found correctness falling around 32K for Llama 3.1 405B (Breunig, How Long Contexts Fail). [perishable snapshot, 2025–2026: model/vendor-specific — teach the mechanism, re-verify the number] These specific ceilings are model- and time-specific — the durable lesson is that every model has one, and it is lower than the spec sheet.
P4 — Stability is money and latency: design append-only. Manus calls KV-cache hit rate "the single most important metric for a production-stage AI agent": cached input tokens can cost 10x less than uncached (their figure: $0.30 vs $3.00/MTok). [perishable snapshot, 2025–2026: model/vendor-specific — teach the mechanism, re-verify the number] One changed token invalidates the cache for everything after it. Therefore: stable prompt prefixes (never embed a timestamp at the top), append-only context (never rewrite history mid-session), deterministic serialization (Manus, Context Engineering for AI Agents). Anthropic's caching docs confirm the mechanics: exact prefix matching, cache reads at 0.1x base price, and a strict tools → system → messages hierarchy where a change at any level invalidates everything below it (Anthropic, prompt caching docs).
P5 — Attention has a shape; place and refresh accordingly. Because of the U-curve (Liu et al.) and recency effects, put durable instructions at the start, and re-surface the current objective near the end. Manus operationalizes this as recitation: the agent rewrites a todo.md and appends it late in context on every step, "reciting its objectives into the end of the context" to prevent goal drift across ~50-tool-call tasks (Manus, ibid.).
P6 — Failures are context, not garbage. Keep failed actions and stack traces in context; the model updates its implicit beliefs and stops repeating the mistake (Manus, ibid.). The 12-Factor Agents version: "Compact Errors into Context Window" — represent failures efficiently so they inform the next step rather than either vanishing or flooding the window (HumanLayer, 12-Factor Agents, Factor 9).
P7 — Share decisions, not just facts. Cognition's two principles: "Share context, and share full agent traces, not just individual messages" and "Actions carry implicit decisions, and conflicting decisions carry bad results." Two workers given the same task summary but not each other's traces will make incompatible implicit choices (their example: subagents building visually clashing pieces of the same game) (Cognition, Don't Build Multi-Agents). Any handoff or summary that transmits conclusions without the decisions behind them is lossy in the way that breaks systems.
P8 — Calibrate instruction altitude. System prompts fail in two directions: hardcoded brittle if-else logic (fragile, high-maintenance) and vague high-level guidance that "falsely assumes shared context." Aim for "specific enough to guide behavior effectively, yet flexible enough to provide strong heuristics." Start minimal with a capable model, then add instructions driven by observed failure modes — not speculation (Anthropic, Effective context engineering).
P9 — Own the window; treat the agent as a function of its context. Deliberately control what the model receives rather than accepting framework defaults ("Own your context window," Factor 3), and design the agent as "a stateless reducer" — output is a pure function of the context you assembled, which makes context bugs reproducible and testable (HumanLayer, 12-Factor Agents, Factors 3 and 12).
LangChain's taxonomy organizes nearly everything into four moves: write (persist outside the window), select (pull the right things in), compress (shrink what's there), isolate (split across contexts) (LangChain, Context Engineering for Agents). The named techniques:
What: Structure knowledge in layers so the agent loads only what the current task needs. Anthropic's Agent Skills are the canonical design: Level 1 is name + description metadata (always in the system prompt — just enough to know when the skill applies), Level 2 is the SKILL.md body (loaded when relevant), Level 3+ is bundled files and scripts the agent navigates "only as needed." Context becomes "effectively unbounded" because nothing loads until demanded (Anthropic, Equipping agents for the real world with Agent Skills).
Why it works: It converts a token cost into a pointer cost. The analogy Anthropic uses is "an onboarding guide for a new hire" — compartmentalized knowledge absorbed progressively.
When: Any recurring domain knowledge, workflow, or reference material. The rule of thumb from Claude Code's docs: always-loaded files (CLAUDE.md) get only what applies broadly to every session; anything situational belongs in an on-demand skill (Claude Code best practices).
What: The agent maintains lightweight identifiers — file paths, queries, URLs — and loads data at runtime with tools, instead of receiving everything up front. This mirrors human cognition: we don't memorize corpora, we keep organization systems and retrieve on demand. Metadata itself (folder hierarchies, naming conventions, timestamps) is signal (Anthropic, Effective context engineering).
Trade-off: Runtime exploration is slower than pre-computed retrieval (embeddings/RAG). The production answer is usually hybrid: some context up front for speed, plus tools for autonomous exploration — Claude Code's CLAUDE.md-plus-grep pattern (Anthropic, ibid.). Classic RAG — "selectively adding relevant information to help the LLM generate a better response" — remains the right tool when the corpus is large and unindexed by structure (Breunig, How to Fix Your Context).
When: Prefer just-in-time for coding agents in navigable environments (filesystems, git, APIs); prefer indexed retrieval for large unstructured corpora; hybridize when latency matters.
What: When the window approaches its limit, summarize the trajectory, reinitialize with the summary, and continue. Claude Code auto-compacts near the window limit (LangChain reports at ~95%); the model distills decisions, code patterns, and open threads while discarding redundant tool outputs (Anthropic, Effective context engineering; LangChain, ibid.).
How to tune it: "Start by maximizing recall to ensure your compaction prompt captures every relevant piece of information from the trace, then iterate to improve precision by eliminating superfluous content" (Anthropic, ibid.). Good summarization prompts enforce temporal ordering, structured sections (environment, steps tried, current status), and explicit "UNVERIFIED" marking on uncertain facts — because "if a bad fact enters the summary, it can poison future behavior" (OpenAI Cookbook, Session memory).
Trimming vs. summarizing (OpenAI Cookbook, ibid.): keep-last-N-turns trimming is deterministic, zero-latency, easy to reason about — but loses old constraints abruptly. LLM summarization preserves long-range decisions compactly — but adds latency spikes, drift risk, and observability burden. Trimming fits tool-heavy, independent tasks; summarization fits long-horizon work where accumulated decisions matter.
Cognition's caution: a dedicated compressor model that distills "key details, events, and decisions" is their recommended path for long tasks — and they note it "is hard to get right." Getting it right means capturing decisions and their rationale, not just facts (Cognition, ibid.).
What: The lightest-touch compaction: automatically drop stale tool calls and results deep in history while preserving the conversational thread. Anthropic ships this as "context editing"; measured effects: context editing alone gave 29% improvement on internal agentic-search evals, combined with the memory tool 39%; on a 100-turn web-search eval it let agents complete workflows that would otherwise die of context exhaustion while cutting token consumption 84% (Anthropic/Claude, Context management).
Tension to know: this conflicts with P4 (append-only for cache) and P6 (keep errors). The reconciliation: clear bulky, stale, already-acted-upon tool outputs (a 30K-token file read from 40 turns ago), keep decisions and failures. Manus's version keeps information restorable — drop a webpage's content but keep its URL, drop a document's body but keep its path (Manus, ibid.).
What: The agent writes notes to persistent storage outside the window — a NOTES.md, a todo.md, a memory directory — and pulls them back when relevant. Persistent memory with minimal context overhead. Anthropic's example: Claude playing Pokémon "maintains precise tallies across thousands of game steps," builds maps, and remembers strategies across multi-hour sessions (Anthropic, Effective context engineering). Anthropic's memory tool productizes this as file-based CRUD in a client-side memory directory persisting across conversations (Anthropic/Claude, Context management).
The general form — filesystem as ultimate context: Manus treats the filesystem as memory that is "unlimited in size, persistent by nature, and directly operable by the agent itself" (Manus, ibid.). Breunig's name for the family is context offloading; even a simple scratchpad ("think" tool) produced up to 54% improvement on specialized-agent benchmarks (Breunig, How to Fix Your Context).
Memory layers in practice (synthesis of LangChain + Anthropic + OpenAI): (1) in-context working memory — the current window; (2) session-scoped scratchpads/todo files; (3) persistent cross-session memory — files or stores, retrieved by relevance; (4) always-loaded curated core (CLAUDE.md-class files), kept ruthlessly small. Information should flow down this stack as it proves durable, and each layer buys persistence at the price of retrieval reliability.
What: Breunig's "context quarantine": isolate work in dedicated threads, each with its own window (Breunig, How to Fix Your Context). Anthropic's research system is the flagship: an orchestrator spawns parallel subagents, each exploring one facet with a clean window, each returning a condensed summary — typically 1,000–2,000 tokens — to the coordinator. "Subagents facilitate compression by operating in parallel with their own context windows" (Anthropic, Multi-agent research system). The multi-agent system beat single-agent Claude Opus 4 by 90.2% on their internal research eval, at the price of ~15x the tokens of a chat (Anthropic, ibid.).
When it wins: read-heavy, parallelizable, breadth-first work — research, codebase investigation, review — where workers' outputs are reports that the orchestrator integrates. Claude Code's guidance is exactly this: "use subagents to investigate" so exploration burns a disposable context, not your main one; and use a fresh-context subagent for adversarial review, because "a fresh context improves code review since Claude won't be biased toward code it just wrote" (Claude Code best practices).
When it loses: write-heavy, coherence-critical work. Cognition's argument (P7): parallel workers make conflicting implicit decisions and current models can't negotiate them away; "running multiple agents in collaboration only results in fragile systems" for building software. Their prescription is a single agent with full traces plus a compressor (Cognition, ibid.). OpenAI's builder guidance points the same direction: maximize a single agent's capability first and reach for multi-agent orchestration only when single-agent complexity demonstrably fails (OpenAI, A practical guide to building agents). See §6 for the reconciliation.
What: When context must cross an agent boundary, the transfer is an engineered artifact, not a vibe. Anthropic's hard-won spec for what every delegated task must contain: an objective, an output format, guidance on tools and sources, and clear task boundaries. Without it, "agents duplicate work, leave gaps, or fail to find necessary information" (Anthropic, Multi-agent research system).
Effort scaling belongs in the contract: encode explicit rules — "simple fact-finding requires just 1 agent with 3–10 tool calls, direct comparisons might need 2–4 subagents with 10–15 calls each" — or orchestrators over-provision wildly (early versions spawned 50 subagents for simple queries) (Anthropic, ibid.).
Returns are contracts too: the worker's report back is a compression step (1–2K tokens, per §3.6) and inherits every summarization risk in §3.3 — a handoff that reports conclusions without decisions violates P7. The spec-then-fresh-session pattern is the single-player version: interview, write a self-contained SPEC.md (files, interfaces, out-of-scope, end-to-end verification step), then start a clean session to execute it (Claude Code best practices).
Selection: "Every model performs worse when provided with more than one tool" is the provocative headline from Berkeley's function-calling data; a quantized Llama 3.1 8B failed with 46 tools and succeeded with 19 (Breunig, How Long Contexts Fail). Dynamic tool selection — RAG over tool descriptions — improved Llama 3.1 8B performance 44% (Breunig, How to Fix Your Context); LangChain cites ~3x tool-selection accuracy from the same family of techniques (LangChain, ibid.). Anthropic's design rule: tools must have minimal overlap — "if a human engineer can't definitively say which tool should be used in a given situation, an agent can't be expected to do better" (Anthropic, Effective context engineering).
Response design: tool outputs are context injections; engineer them. Paginate, filter, and
truncate with sensible defaults (Claude Code caps tool responses at 25,000 tokens [perishable snapshot, 2025–2026: model/vendor-specific — teach the mechanism, re-verify the number]); prefer
semantically meaningful names over UUIDs; offer a response_format: concise|detailed knob;
consolidate chains of granular calls into one higher-level tool (schedule_event, not
list_users + list_events + create_event); use error messages to steer the agent toward
efficient strategies like "many small targeted searches" (Anthropic, Writing effective tools
for agents). CLI tools are often the most context-efficient integration surface of all
(Claude Code best practices).
Masking over removal: dynamically removing tools mid-session invalidates the KV cache and confuses the model about past references; prefer masking token logits to constrain choice while keeping definitions stable (Manus, ibid.).
What: Treat the window as a budgeted resource with an explicit spending plan: how much for system + tools, how much reserved for the task, at what fill level compaction triggers. Claude Code's docs are blunt: "The context window is the most important resource to manage," and its UX (a /context inspector, status-line usage tracking, /clear, targeted /compact) is budget tooling (Claude Code best practices). Chroma's practical corollary: set working budgets far below the advertised window for high-accuracy work (Chroma, ibid.).
Caching discipline (the mechanics behind P4): structure prompts static-first (tools → system → messages); place cache breakpoints on the last stable block, never on content that changes per request; verify with cache-read/cache-write token counts rather than assuming. A timestamp above the fold means you pay cache-write prices on every request forever (Anthropic, prompt caching docs; Manus, ibid.).
Few-shot rut: uniform repeated action-observation patterns in context cause the model to mimic rhythm over substance — "drift, overgeneralization, or sometimes hallucination." Inject structured variation in serialization and phrasing (Manus, ibid.).
Breunig's four-way taxonomy (How Long Contexts Fail) is the field's shared vocabulary; know these by name:
Add the operational failures from production systems:
They measure token-shaped things. Anthropic found token usage alone explains 80% of performance variance in their research eval (Anthropic, Multi-agent); Manus optimizes KV-cache hit rate as the top-line production metric (Manus). Elite teams instrument context — fill levels, cache hits, tokens per task — before theorizing.
They iterate from observed failures, with the model in the loop. Anthropic's tool and skill guidance is evaluation-first: build representative tasks, watch real transcripts, let the model critique its own failures, refine, re-run against held-out sets (Anthropic, Writing effective tools; Agent Skills). Prompts start minimal and grow only where failures demand (P8). And automated evals aren't sufficient: human testers caught Anthropic's agents "consistently choosing SEO-optimized content farms over authoritative sources" — a context-quality failure no programmatic eval flagged (Anthropic, Multi-agent).
They keep the always-loaded core tiny and push everything else behind disclosure. Concise checked-in CLAUDE.md-class files for what applies every session; skills for everything situational; pruning as routine maintenance ("treat CLAUDE.md like code") (Claude Code best practices).
They design the write path around verification and fresh contexts. Explore → plan → implement → verify, with plan mode separating research from execution; a check the agent can run (tests, build, screenshot diff) so the loop closes without a human; specs executed in fresh sessions; adversarial review in a fresh subagent context (Claude Code best practices).
They engineer for the cache and the filesystem. Stable prefixes, append-only histories, masked (not removed) tools, files as unlimited restorable memory, recitation for goal stability, errors left visible (Manus). These are production-economics practices as much as quality practices.
They match architecture to task shape. Anthropic's own selection heuristic: compaction for long conversational flows; note-taking for iterative work with milestones (coding fits here); multi-agent for parallel-explorable research (Anthropic, Effective context engineering).
| Situation | Reach for | Source |
|---|---|---|
| Recurring domain knowledge, sometimes needed | Skill w/ progressive disclosure | Anthropic Skills |
| Rules needed every session | Tiny curated always-loaded file; prune ruthlessly | Claude Code docs |
| Large explorable environment (repo, filesystem) | Just-in-time retrieval via tools; hybrid if latency-bound | Anthropic CE |
| Large unstructured corpus | Indexed retrieval (RAG) | Breunig |
| Long task nearing window limit | Compaction (recall-first, then precision); or clear stale tool results | Anthropic CE / context mgmt |
| Long task with milestones | Structured notes + todo recitation | Anthropic CE, Manus |
| Cross-session persistence | File-based memory outside the window | Anthropic context mgmt, Manus |
| Breadth-first research / investigation / review | Parallel subagents, condensed returns, explicit delegation contract + effort scaling | Anthropic multi-agent |
| Coherent build/edit on one artifact | Single agent, full traces, compressor for length — not parallel workers | Cognition |
| Many tools available | Loadout selection; consolidate; namespace; no overlap | Breunig, Anthropic tools |
| Cost/latency pressure | Stable prefix + append-only + cache breakpoints; mask don't remove | Manus, Anthropic caching |
| Agent repeating itself / drifting | Suspect distraction: check fill level, compact or clear, recite goals | Breunig, Manus |
| Agent confidently wrong about earlier "facts" | Suspect poisoning: audit summaries and notes, restart from clean spec | Breunig, OpenAI Cookbook |
© mvschwarz, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/_canonical/process/context-engineering of mvschwarz/openrig.
Open the folder on GitHubat commit a7fed63
Context Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Context Engineering this skillmvschwarz/openrig | 5.5k | — | ~8.2k | Automated safety check: Pass | Apache-2.0 | |
| Deep Agentslangchain-ai/docs | 424 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Self Managed Contextguanyang/open-agent-hub | 973 | 1 repos | ~6k | Automated safety check: Pass | MIT | |
| Context Engineeringa5c-ai/babysitter | 1.8k | — | ~1.8k | Automated safety check: Notes | MIT | |
| Context Window Managementaiskillstore/marketplace | 430 | 3 repos | ~2.4k | Automated safety check: Pass | None | |
| Context Crushermohitagw15856/pm-claude-skills | 1.4k | — | ~1.4k | Automated safety check: Pass | MIT |
langchain-ai/docs
Build batteries-included agents with planning, context management, subagent delegation, and sandboxed execution.
guanyang/open-agent-hub
This skill should be used when a model gets read-write control over its own live context window instead of a harness-scheduled compaction policy: the context exposed as an editable file the model…
a5c-ai/babysitter
Context window monitoring and budget management. An agent skill from a5c-ai/babysitter.
aiskillstore/marketplace
Strategies for managing LLM context windows including summarization, trimming, routing, and avoiding context rot
mohitagw15856/pm-claude-skills
Compress tool outputs, logs, and JSON before they enter the context window — structural compression via a deterministic stdlib script (schema + samples + stats instead of 300 raw rows), no API, no…
thiswillbeyourgithub/wdoc
Quick reference for wdoc, a command-line and Python tool that summarizes, searches and answers questions over documents of many file types.
mvschwarz/openrig
Walks an agent through upgrading the OpenRig CLI and daemon one observed step at a time, keeping live seats alive and reconciling managed plugin files.
mvschwarz/openrig
Helps set up a continuing agent software team for a real repository with OpenRig, choosing between manual work, queue handoffs and an explicit Workflow.
mvschwarz/openrig
Separates a stable agent seat's identity from its changing occupant, and records honest, two-part provenance whenever one occupant replaces another.
mvschwarz/openrig
Loads one section of a Markdown file by its path#h2-slug address with a bundled resolver script, for use outside OpenRig's context library.
mvschwarz/openrig
A skill your agent uses when checking the health of this rig's HashiCorp Vault or writing, reading, listing, deleting or explaining its secrets.
mvschwarz/openrig
Re-grounds a long-running agent in the current product outcome by running a path-based trace to the root of its topology and work trees.
Works with
Categories
A skill your agent uses when designing, reviewing, or debugging how an agent's context window gets filled, pruned, or shared — choosing what loads at boot versus on demand, sizing an install or an…. Context Engineering is an agent skill from mvschwarz/openrig. Use when designing, reviewing, or debugging how an agent's context window gets filled, pruned, or shared — choosing what loads at boot versus on demand, sizing an install or an always-loaded file, fixing an agent that drifts, repeats itself, or forgets constraints mid-task, planning compaction or summarization, deciding single-agent versus subagents, engineering handoffs between agents, or picking a tool loadout.
Context Engineering fits situations like: debugging how an agents context window gets filled; shared — choosing what loads at boot versus on demand; sizing an install; an always-loaded file.
Run `npx skills add mvschwarz/openrig --skill context-engineering -a claude-code`. Or copy the skill folder (skills/_canonical/process/context-engineering in mvschwarz/openrig) into .claude/skills/context-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mvschwarz/openrig --skill context-engineering -a codex`. Or copy the skill folder (skills/_canonical/process/context-engineering in mvschwarz/openrig) into .agents/skills/context-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mvschwarz/openrig --skill context-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/context-engineering, .gemini/skills/context-engineering, .github/skills/context-engineering and .opencode/skills/context-engineering in your project.
SKILL.md names no scripts, command-line tools or credentials: Context Engineering is instructions for the agent only.
SKILL.md names 14 domains. As links in the text: anthropic.com, dbreunig.com, claude.com, code.claude.com, platform.claude.com, trychroma.com, cognition.com, manus.im, langchain.com, philschmid.de, github.com, arxiv.org, openai.com and developers.openai.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Context Engineering is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.2k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Context Engineering: Deep Agents (langchain-ai/docs, 424 stars), Self Managed Context (guanyang/open-agent-hub, 973 stars), Context Engineering (a5c-ai/babysitter, 1.8k stars) and Context Window Management (aiskillstore/marketplace, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mvschwarz (a GitHub user) maintains it in mvschwarz/openrig, which has 5,542 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 7, 2026.
Source: mvschwarz/openrig on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.