Agent skill

Compact Memory Implementation

by simbajigege in simbajigege/book2skills

A developer guide to adding compact memory to an agent: when to trigger compaction, how to fork a compactor sub-agent, what the summary holds, and how to restore it.

MITAuto-check passedAI & LLM Engineering

Install Compact Memory Implementation

skills CLI
$ npx skills add simbajigege/book2skills --skill compact-memory-implementation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install simbajigege/book2skills compact-memory-implementation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/simbajigege/book2skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/compact-memory-implementation .claude/skills/compact-memory-implementation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compact-memory-implementation
GitHub stars
184
Token cost
~2.5k tokens
SKILL.md length
439 words
Files
3 (incl. scripts)
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

A developer guide to adding compact memory to an agent: when to trigger compaction, how to fork a compactor sub-agent, what the summary holds, and how to restore it.

  • Works in 7 steps: Understand the setup → When to trigger compact → Fork agent for compaction → …
  • Adding context compression to an agent that runs long sessions
  • SKILL.md covers Step 1 — Understand the setup, Step 2 — When to trigger compact, Step 3 — Fork agent for… and Step 4 — How to compact:…, plus 4 more sections
  • Runs Python scripts from its folder

What it does

This skill walks a developer through building context compaction into an agent built on the Claude Agent SDK or the Anthropic API. It begins by clarifying the SDK and language, the agent architecture, whether sessions are long-running or short, and what must survive compaction: task state, decisions, tool results or conversation history.

It then covers three triggers. A token threshold, recommended at roughly 70 to 80 percent of the model's context limit, checks the previous response's input token usage. A turn count compacts every N turns. A phase boundary compacts between tasks such as research and implementation. Compaction itself runs in a separate forked agent call, which can use a cheaper model, so the main agent waits for a fresh structured summary instead of summarizing its own drifted context.

Later steps define the summary format and prompt and how the next session restores the stored memory. A script, scripts/pre_compact_extract.py, is bundled, and the code examples are in Python.

When your agent uses it

  • Adding context compression to an agent that runs long sessions
  • Choosing between token-threshold, turn-count and phase-boundary triggers
  • Designing the summary format a compactor sub-agent returns
  • Restoring compacted memory when a new session starts

Example prompts

  • “How should I implement compact memory in my Claude Agent SDK agent?”
  • “Add a token-threshold compaction step to my Python agent loop.”
  • “Design a structured summary schema so my agent keeps decisions and tool results after compaction.”

Requirements

  • An agent built with the Claude Agent SDK or the Anthropic API
  • Python, for the bundled extraction script

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Understand the setup
  2. When to trigger compact
  3. Fork agent for compaction
  4. How to compact: format and prompt
  5. How to use after compacting: memory restoration
  6. Full agent loop
  7. Chaining compacts across sessions

What it can do on your machine

Read from SKILL.md and the folder at commit e5ba66c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compact Memory Implementation loads about 2.5k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 439 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from simbajigege/book2skills at commit e5ba66c, republished under its MIT licence (© simbajigege). 439 words, ~2,492 tokens.

Download SKILL.mdSave it as .claude/skills/compact-memory-implementation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
compact-memory-implementation
description
Developer implementation guide for adding compact memory to an Agent — covers fork agent pattern for compaction, trigger strategy, summary format design, and memory restoration in subsequent sessions. Use when a developer asks how to implement compact memory, context compression, or memory persistence in their agent built with Claude Agent SDK or Anthropic API.

compact-memory-implementation

A developer guide for building compact memory into an Agent: detect when to compress, fork a compactor sub-agent, produce a structured summary, and restore it in the next session.

Step 1 — Understand the setup

Before designing anything, clarify:

  • SDK / language: Claude Agent SDK? Direct Anthropic API? Python or TypeScript?
  • Agent architecture: single-agent loop, multi-agent, tool-calling?
  • Session model: one long-running session or multiple short sessions?
  • What must survive compaction: task state, decisions, tool results, conversation history?

This determines which pattern fits.


Step 2 — When to trigger compact

Three strategies, pick based on your session model:

1. Token threshold (recommended) Check usage.input_tokens from the previous response. When it exceeds ~70–80% of your model's context limit, trigger compact.

python
COMPACT_THRESHOLD = 150_000  # adjust per model

if response.usage.input_tokens > COMPACT_THRESHOLD:
    compact = compact_memory(history)
    history = []  # reset — compact moves to system prompt

2. Turn count Compact every N turns. Simpler but less adaptive — misses sessions with a few very long turns.

python
COMPACT_EVERY_N = 30

if turn_count % COMPACT_EVERY_N == 0:
    compact = compact_memory(history)

3. Phase boundary Compact at natural task boundaries (after research, before implementation). Requires the agent to detect phases. Produces summaries that align with meaningful milestones, but harder to implement reliably.

Recommended default: token threshold at 70%, with turn-count fallback at N=40.


Step 3 — Fork agent for compaction

The compactor is a separate agent call whose only job is to read the current state and return a structured summary. Fork it synchronously — the main agent waits for the result before continuing.

python
def compact_memory(history: list[dict]) -> dict:
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",  # cheaper model is fine for compaction
        max_tokens=4096,
        system=COMPACTOR_SYSTEM_PROMPT,
        messages=[
            {
                "role": "user",
                "content": format_history_for_compact(history),
            }
        ],
    )
    return json.loads(response.content[0].text)

Why fork instead of self-compact:

  • The main agent may have drifted in focus; the compactor starts fresh with the full picture
  • Compaction is a different cognitive task — summarizing vs. executing
  • A cheaper, smaller model (Haiku) can do compaction; save the expensive model for main work
  • Clean separation makes the compact output easier to validate and test

Show full SKILL.md (168 more words)Show less

Step 4 — How to compact: format and prompt

Compact output schema
json
{
  "task": "What the agent is working on and why — the goal, not the steps",
  "current_state": "Exact status at compaction point: what is done, what is not, what is in progress",
  "key_decisions": [
    { "decision": "...", "reason": "...", "constraint": "..." }
  ],
  "eliminated_approaches": [
    { "approach": "...", "reason_ruled_out": "..." }
  ],
  "open_questions": ["..."],
  "next_steps": ["..."],
  "relevant_tool_results": {
    "key": "Only results future steps will need — summarized, not raw dumps"
  },
  "compacted_at_turn": 42
}
Compactor system prompt
You are a conversation compactor. Read the provided conversation and produce a JSON summary that captures everything a fresh agent needs to continue the work without asking what happened.

Include:
- Current task and goal (not the steps taken to get here)
- Exact current state — what is done and what is not
- Decisions made and WHY (reasoning, not just the choice)
- Approaches tried and ruled out with reasons (prevents re-exploration)
- Open questions and blockers
- Concrete next steps in priority order
- Tool results that future steps will need (summarize, don't dump raw output)

Omit:
- Intermediate reasoning that led nowhere
- Completed sub-tasks with no future relevance
- Raw tool output that has already been acted on
- Anything derivable by reading the code or running a command

Output valid JSON matching the schema provided. No prose outside the JSON.
Format history for compactor
python
def format_history_for_compact(history: list[dict]) -> str:
    lines = ["Conversation to compact:\n"]
    for msg in history:
        role = msg["role"].upper()
        content = msg["content"] if isinstance(msg["content"], str) else "[tool use]"
        lines.append(f"[{role}]: {content[:2000]}")  # cap very long messages
    return "\n".join(lines)

Step 5 — How to use after compacting: memory restoration

The compact object becomes the "memory" for the next turn or session. Inject it into the system prompt so it's always visible to the agent.

python
MEMORY_BLOCK_TEMPLATE = """
## Restored memory (compacted at turn {turn})

**Task**: {task}

**Current state**: {current_state}

**Key decisions**:
{decisions}

**Ruled out approaches**:
{eliminated}

**Next steps**:
{next_steps}

Begin from current state above. Do not re-explore eliminated approaches.
"""

def build_system_with_memory(base_system: str, compact: dict | None) -> str:
    if compact is None:
        return base_system
    memory = MEMORY_BLOCK_TEMPLATE.format(
        turn=compact["compacted_at_turn"],
        task=compact["task"],
        current_state=compact["current_state"],
        decisions="\n".join(f"- {d['decision']} (because {d['reason']})"
                            for d in compact["key_decisions"]),
        eliminated="\n".join(f"- {e['approach']}: {e['reason_ruled_out']}"
                             for e in compact["eliminated_approaches"]),
        next_steps="\n".join(f"- {s}" for s in compact["next_steps"]),
    )
    return base_system + "\n\n" + memory
Pattern B — First message injection (for stateless API callers)
python
messages = [
    {
        "role": "user",
        "content": f"[Resuming from compacted state — turn {compact['compacted_at_turn']}]\n"
                   f"{json.dumps(compact, indent=2)}\n\n"
                   f"Continue from the next steps listed above.",
    }
]
Persistence across sessions
python
import json, pathlib

MEMORY_DIR = pathlib.Path("memory")
MEMORY_DIR.mkdir(exist_ok=True)

def save_compact(session_id: str, compact: dict) -> None:
    (MEMORY_DIR / f"{session_id}.json").write_text(json.dumps(compact, indent=2))

def load_compact(session_id: str) -> dict | None:
    path = MEMORY_DIR / f"{session_id}.json"
    return json.loads(path.read_text()) if path.exists() else None

Step 6 — Full agent loop

python
def run_agent(session_id: str, user_input: str) -> str:
    compact = load_compact(session_id)
    system = build_system_with_memory(BASE_SYSTEM, compact)
    history = []
    turn = 0

    while True:
        response = client.messages.create(
            model="claude-opus-4-7",
            system=system,
            messages=history + [{"role": "user", "content": user_input}],
            max_tokens=8192,
        )

        # Trigger compact if context is growing too large
        if response.usage.input_tokens > COMPACT_THRESHOLD:
            compact = compact_memory(history)
            save_compact(session_id, compact)
            system = build_system_with_memory(BASE_SYSTEM, compact)
            history = []  # reset history — compact is now in system
            turn = 0
            continue

        if response.stop_reason == "end_turn":
            return response.content[0].text

        history.append({"role": "assistant", "content": response.content})
        user_input = handle_tool_calls(response)  # your tool dispatch
        turn += 1

Step 7 — Chaining compacts across sessions

If a session resumes multiple times, don't stack compacts — re-compact instead:

python
COMPACTOR_WITH_PRIOR = """
You are updating an existing memory compact with new information from a continuation session.

Prior compact:
{prior_compact}

New conversation turns since last compact:
{new_turns}

Produce an updated compact that:
- Merges both sources
- Removes resolved items and completed steps
- Adds new decisions, eliminations, and open questions
- Keeps next_steps current

Output valid JSON. No prose outside the JSON.
"""

def compact_memory_with_prior(history: list[dict], prior: dict) -> dict:
    prompt = COMPACTOR_WITH_PRIOR.format(
        prior_compact=json.dumps(prior, indent=2),
        new_turns=format_history_for_compact(history),
    )
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",
        max_tokens=4096,
        system=prompt,
        messages=[{"role": "user", "content": "Update the compact."}],
    )
    return json.loads(response.content[0].text)

Common pitfalls

PitfallFix
Compact loses tool results needed laterInclude summarized results in relevant_tool_results
Fresh session ignores compactInject into system prompt, not buried in messages
Compactor uses the same expensive modelUse Haiku for compaction, Opus for main work
Compact grows unbounded across sessionsRe-compact using "chaining compacts" pattern above
Compacting too often (every turn)Use token threshold, not turn frequency
Compact JSON fails to parseAdd retry with explicit error feedback to compactor

© simbajigege, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/compact-memory-implementation of simbajigege/book2skills.

  • SKILL.md
  • README.md
  • scripts/pre_compact_extract.py

Open the folder on GitHubat commit e5ba66c

Compare with similar skills

Compact Memory Implementation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compact Memory Implementation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compact Memory Implementation this skillsimbajigege/book2skills184—~2.5kAutomated safety check: PassMIT
Agent BuilderMathews-Tom/armory327—~1.7kAutomated safety check: PassMIT
Pydantic AI Harnesspydantic/pydantic-ai20k—~4.9kAutomated safety check: PassMIT
Agent Squad Python Guide2FastLabs/agent-squad7.8k—~4.7kAutomated safety check: PassApache-2.0
Deep Agents Corelangchain-ai/langchain-skills1.3k1 repos~3.1kAutomated safety check: PassMIT
Omnigent Framework Detectionomnigent-ai/omnigent11k—~610Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    Mathews-Tom/armory

    Build AI agents and automate Claude Code programmatically via the Claude Agent SDK and headless CLI mode.

    327 GitHub stars~1.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Pydantic AI Harness

    pydantic/pydantic-ai

    Official

    Adds optional capabilities to Pydantic AI agents from pydantic-ai-harness, led by Code Mode, which runs many tool calls as one sandboxed Python script.

    20k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Agent Squad Python Guide

    2FastLabs/agent-squad

    Map of the agent-squad Python framework for async multi-agent orchestration: which agent, classifier, storage and tool provider to pick, and the pitfalls to avoid.

    7.8k GitHub stars~4.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Deep Agents Core

    langchain-ai/langchain-skills

    Official

    Explains how to build agents with the Deep Agents framework: create_deep_agent, the built-in middleware, the harness, SKILL.md format and configuration options.

    1.3k GitHub starsUsed in 1 repo~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Omnigent Framework Detection

    omnigent-ai/omnigent

    Scans Python agent code for framework imports and recommends the matching Omnigent executor type, or says when the framework is not natively supported yet.

    11k GitHub stars~610 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Explains how cognee stores session memory by session_id and bridges it into the permanent graph with improve(), including the stages, results and settings.

    32k GitHub stars~3k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from simbajigege/book2skills

All 37 skills in this repo
  • MEMORY.md Restructuring

    simbajigege/book2skills

    Reorganizes an overgrown MEMORY.md into a short pointer index plus separate topic files, and fixes or deletes outdated memories instead of archiving them.

    184 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Semantic Line-Art SVG Diagrams

    simbajigege/book2skills

    Turns text, screenshots, or existing diagrams into minimal, accessible line-art SVGs for teaching material, with an optional Mermaid relationship spec.

    184 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Fail-Closed Agent Tool Builder

    simbajigege/book2skills

    Helps define agent tools with a fail-closed pattern: one class holding name, schema, security flags and a validate, permission and call execution chain.

    184 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • LLM Query Loop Implementation

    simbajigege/book2skills

    Implements a production-style agent loop in your own AI product, with tool calling, tool results fed back, exit conditions and budget guards.

    184 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Analyzing Financial Reports

    simbajigege/book2skills

    Analyzes Chinese public company financial statements (balance sheet, income statement, cash flow) to assess asset quality, profit authenticity, cash flow health, solvency, and overall investment…

    184 GitHub starsUsed in 1 repo~930 tokens
    Auto-check passed
  • Tool Permission System Design

    simbajigege/book2skills

    Guides designing a layered permission pipeline for agent tools that decides which calls are allowed, need confirmation or are denied, with scopes and hooks.

    184 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Compact Memory Implementation

What does Compact Memory Implementation do?

A developer guide to adding compact memory to an agent: when to trigger compaction, how to fork a compactor sub-agent, what the summary holds, and how to restore it. This skill walks a developer through building context compaction into an agent built on the Claude Agent SDK or the Anthropic API. It begins by clarifying the SDK and language, the agent architecture, whether sessions are long-running or short, and what must survive compaction: task state, decisions, tool results or conversation history.

When should I use Compact Memory Implementation?

Compact Memory Implementation fits situations like: adding context compression to an agent that runs long sessions; choosing between token-threshold, turn-count and phase-boundary triggers; designing the summary format a compactor sub-agent returns; restoring compacted memory when a new session starts.

How do I install Compact Memory Implementation in Claude Code?

Run `npx skills add simbajigege/book2skills --skill compact-memory-implementation -a claude-code`. Or copy the skill folder (skills/compact-memory-implementation in simbajigege/book2skills) into .claude/skills/compact-memory-implementation in your project. Claude Code loads it when a task matches its description.

How do I install Compact Memory Implementation in Codex?

Run `npx skills add simbajigege/book2skills --skill compact-memory-implementation -a codex`. Or copy the skill folder (skills/compact-memory-implementation in simbajigege/book2skills) into .agents/skills/compact-memory-implementation in your project. Codex loads it when a task matches its description.

Can I use Compact Memory Implementation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add simbajigege/book2skills --skill compact-memory-implementation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compact-memory-implementation, .gemini/skills/compact-memory-implementation, .github/skills/compact-memory-implementation and .opencode/skills/compact-memory-implementation in your project.

What does Compact Memory Implementation need to run?

Going by SKILL.md and its folder, Compact Memory Implementation needs Python for the scripts in its folder. Our summary lists: An agent built with the Claude Agent SDK or the Anthropic API; Python, for the bundled extraction script.

Does Compact Memory Implementation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Compact Memory Implementation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Compact Memory Implementation use?

Compact Memory Implementation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compact Memory Implementation use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Compact Memory Implementation?

Skills that share tags, products or a category with Compact Memory Implementation: Agent Builder (Mathews-Tom/armory, 327 stars), Pydantic AI Harness (pydantic/pydantic-ai, 20k stars), Agent Squad Python Guide (2FastLabs/agent-squad, 7.8k stars) and Deep Agents Core (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compact Memory Implementation?

simbajigege (a GitHub user) maintains it in simbajigege/book2skills, which has 184 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on August 26, 2026.

Source: simbajigege/book2skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.