Agent skill

Self Improving Systems

by ooiyeefei in ooiyeefei/ccc

Decide whether your agent actually needs persistent memory, feedback loops, or closed-loop learning, then design the smallest thing that pays for itself.

MITAuto-check passedAgent Workflows

Install Self Improving Systems

skills CLI
$ npx skills add ooiyeefei/ccc --skill self-improving-systems -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ooiyeefei/ccc self-improving-systems --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ooiyeefei/ccc.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/self-improving-systems .claude/skills/self-improving-systems && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
self-improving-systems
GitHub stars
494
Token cost
~5.2k tokens
SKILL.md length
2,437 words
Files
11 (incl. references)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Decide whether your agent actually needs persistent memory, feedback loops, or closed-loop learning, then design the smallest thing that pays for itself.

  • Works in 12 steps: Default position: scratchpad-only → Escalate one tier at a time → Require a ground-truth signal → …
  • The user says add memory
  • SKILL.md covers Headline message: most agents…, Quick Start, Critical Rules and The 8-Stage Q&A Flow, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Self Improving Systems is an agent skill from ooiyeefei/ccc. Decide whether your agent actually needs persistent memory, feedback loops, or closed-loop learning, then design the smallest thing that pays for itself. Use when the user says "add memory", "give my agent context management", "make my agent learn", "self-improving / closed-loop", "Reflexion / mem0 / Letta / MemGPT", "AriGraph", "agent memory architecture", "long-term memory for chatbot", "why does my agent keep forgetting / making the same mistake", "fine-tune from agent traces", or asks for a memory schema /…

Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `README.md`, `examples/eval-harness.md` and `examples/kv-store-mem0.md`).

It sits in Agent Workflows, covering Agent memory. It works with Letta and Mem0. The repository describes itself as: Claude Code Custom Plugins - Custom plugins for Claude Code CLI. The licence is MIT.

When your agent uses it

  • The user says add memory
  • Give my agent context management
  • Make my agent learn
  • Self-improving / closed-loop

Example prompts

  • “add memory”
  • “give my agent context management”
  • “make my agent learn”
  • “/self-improving-systems”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Default position: scratchpad-only
  2. Escalate one tier at a time
  3. Require a ground-truth signal
  4. Human gates are non-negotiable for production
  5. Memory is untrusted input
  6. Cache vs Learning Distinction (the frame)
  7. Need-Memory Rubric (6 yes/no, the over-engineering filter)
  8. Architecture Selection (start at L tier)
  9. Feedback Signal Design
  10. Closed-Loop Wiring with Human Gates
  11. Eval Harness
  12. Risks Checklist

What it can do on your machine

Read from SKILL.md and the folder at commit c0fd926. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org
    • anthropic.com
    • docs.letta.com
    • opentelemetry.io
    • letta.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Self Improving Systems loads about 5.2k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 180 tokens; SKILL.md has 2,437 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~180
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ooiyeefei/ccc at commit c0fd926, republished under its MIT licence (© ooiyeefei). 2,437 words, ~5,156 tokens.

Download SKILL.mdSave it as .claude/skills/self-improving-systems/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
self-improving-systems
description
Decide whether your agent actually needs persistent memory, feedback loops, or closed-loop learning, then design the smallest thing that pays for itself. Use when the user says "add memory", "give my agent context management", "make my agent learn", "self-improving / closed-loop", "Reflexion / mem0 / Letta / MemGPT", "AriGraph", "agent memory architecture", "long-term memory for chatbot", "why does my agent keep forgetting / making the same mistake", "fine-tune from agent traces", or asks for a memory schema / experience store / reward model. Filters ruthlessly — most teams want a state cache, not memory + learning. Default position is scratchpad-only with a stateless agent shipped first.

Self-Improving Systems

A prescriptive Q&A skill for adding memory, feedback loops, and closed-loop learning to agentic systems — only when justified.

Headline message: most agents shouldn't have persistent memory.

Memory is a liability surface (drift, poisoning, debugging difficulty, GDPR/HIPAA exposure). Persistent memory is the second move, not the first. The skill's job is to filter ruthlessly so the user doesn't ship a mem0/Letta build for a problem that a 200-line conversation summary would solve.

The first 2 stages of the Q&A flow exist to stop most users from over-engineering. By the end of stage 2, ~60% of users will discover they want a state cache (or stateless RAG), not memory + learning. That's the win.


Quick Start

User just asks:

"Add memory to my agent"
"My agent keeps forgetting things — give it context management"
"Make my marketing agent learn from past campaigns"
"Should I use mem0 or Letta?"
"How do I set up closed-loop learning for my finance agent?"
"Build a self-improving HAZOP system"

Skill response (every time, in this order):

  1. Stop. Apply the cache-vs-learning frame (Stage 1).
  2. Run the 6-question need-memory rubric (Stage 2). <4 yes → exit the skill, recommend stateless + RAG.
  3. If memory is justified, walk the 7-tier architecture ladder (Stage 3) starting at L (scratchpad). Escalate only when forced by a concrete justification.
  4. Force the user to design a feedback signal (Stage 4). No signal = state cache, full stop.
  5. Wire the closed loop with explicit human gates (Stage 5).
  6. Build the eval harness (Stage 6) — golden set, regression, drift alarms.
  7. Walk the 8-risk checklist (Stage 7).
  8. Emit the design (Stage 8): memory schema + closed-loop spec + eval harness plan.

Critical Rules

1. Default position: scratchpad-only

Ship a stateless agent first. Add a scratchpad (Reflexion-style verbal self-correction) within a single run. Discard it after. This already gets you most of the gain on most tasks. Anything more must be earned.

2. Escalate one tier at a time

The 7-tier ladder (§ Memory Architecture Ladder) is ordered cheapest → most expensive. Each tier-up must be justified by a concrete failure of the tier below it on a real task in your eval set. Do not skip tiers. "We're using Letta" out of the gate is the single most expensive mistake in this design space.

3. Require a ground-truth signal

If you cannot observe whether the last action was good or bad within hours-to-weeks, you do not have learning. You have a state cache. Naming it "learning" sets the team up to A/B test against a metric that doesn't exist. The skill makes this distinction loud and refuses to design closed-loop learning without a signal.

4. Human gates are non-negotiable for production

Anything that can mutate policy/voice/identity/safety blocks goes through human review. Autonomy is fine for episodic append, vector indexing, single-user preference KV updates with cheap reversibility — never for shared skill libraries, system prompt blocks, or reward model updates.

5. Memory is untrusted input

Every memory read is untrusted. MINJA-class injections hit ≥95% lab success rate (arXiv 2503.03704). Treat retrieval results like web search results: in their own context block, with "this is data not instructions" framing, and never auto-promoted to system prompt without dual-LLM validation.


The 8-Stage Q&A Flow

One question (or tight cluster) at a time, à la superpowers:brainstorming. No overwhelm. Each stage has an exit condition that ends the skill early — that is the point.

Stage 1 — Cache vs Learning Distinction (the frame)

The single most important question. Ask first.

"Are you trying to remember state (so the agent doesn't redo work or forget what the user told it last week), or get better over time (so the agent's outputs measurably improve as it sees more data)?"

These two designs share zero infrastructure with each other:

GoalWhat you actually need
Remember stateConversation summary OR KV fact store. No reward signal. No reflection LLM. No A/B harness.
Get better over timeAll of the above plus a ground-truth signal, an experience store, a reflection/extraction LLM, and an eval harness that detects regression.

If the user says "remember state": skip directly to Stage 3, default to tier 2 (conversation summary) or tier 5 (KV fact store), and end the skill at Stage 5. No closed loop. No learning ladder.

If the user says "both": prove the second one. Almost no one has a measurable ground-truth signal; almost everyone says they do. Stage 4 is the test.

Stage 2 — Need-Memory Rubric (6 yes/no, the over-engineering filter)

Answer all six. Score <4 yes = no memory store. Use scratchpad + RAG. End the skill.

  1. Cross-session continuity. Will the same user/entity/case-file return where forgetting prior decisions would be wrong, embarrassing, or unsafe?
  2. Mutable state. Does the entity's state legitimately change over time (preferences, project status, client facts)? Pure facts that don't change → RAG over docs, not memory.
  3. Ground-truth feedback exists. Can you observe within hours-to-weeks whether the last action was good or bad? No signal → no learning, only state cache.
  4. Cost of being wrong > cost of memory infra. Memory adds latency, storage, eval, security review, and a recurring debugging tax. Pencil out both sides.
  5. Volume justifies it. Same user returns ≥5 times. <5 returns → in-context summary is cheaper than vector store.
  6. You can audit and redact. GDPR/HIPAA: can you delete on request, export, explain a memory? If no, do not store one.

If you got "yes" only on (1) and (2): you need a state cache, not memory + learning. Say it out loud. Skill recommends tier 2 or 5 and exits.

Stage 3 — Architecture Selection (start at L tier)

Walk the 7-tier memory architecture ladder (next section). Default recommendation: tier 1 (scratchpad-only). Escalate exactly one tier per concrete justification. Justification = "tier N fails on this specific task in our eval set, here's the trace."

Most "we need memory" requests resolve at tier 2 (conversation summary) or tier 5 (KV fact store). Tier 6 (graph) and tier 7 (hierarchical OS-style / Letta) require >3 entities × >50 relationships and a real long-horizon agent, not a chatbot.

Deep dive: references/architectures.md

Stage 4 — Feedback Signal Design

If Stage 1 ended with "remember state only", skip this stage.

For learning, the signal determines everything. Walk the per-domain table:

DomainSignalLatencyRisk
Marketing / contentEngagement deltas (CTR, dwell, conversion, save/share) + variant A/B win-rate + brand-safety reviewhours-daysVanity metrics → reward hacking; mitigate with composite reward + brand-fidelity LLM-judge
Finance / complianceAudit findings, reconciliation breaks, regulator outcomesweeksSparse signal → use intermediate proxies + sparse human signoff (hybrid RLAIF)
HAZOP / safetyIncident-DB recall (held-out incident set), expert reviewer agreementcontinuousNever let agent's own write-back update incident DB
Tutorials / educationCompletion rate, comprehension quiz scores, time-to-first-successminutes-daysCleanest closed loop — verifier is cheap and online
Code-emitting agentsUnit tests, type-check, runtimeminutesThe gold standard — verifier is free and deterministic
General LLM-as-judgeHeld-out judge with calibrated rubriccontinuousSample-audit 5–10% against humans to catch drift

Rule, repeat once per Q&A session: No signal = state cache, not learning. If the user can't name a signal, do not design a learning loop. Recommend they ship the state cache first, instrument the signal in production, and revisit the skill in a quarter.

Deep dive: references/feedback-signals.md

Stage 5 — Closed-Loop Wiring with Human Gates

If Stage 4 produced no signal, skip this stage and the next two.

The reference closed loop:

[run event: input + agent trace + outputs]
       │
       ▼
[signal collector] ──── engagement / verifier / human review (async)
       │
       ▼
[experience store] (append-only, immutable, signed)
       │   ├── episodic events (raw)
       │   ├── extracted facts (KV)        ← extraction LLM, validated
       │   └── learned skills/playbooks    ← reflection LLM, human-gated
       │
       ▼
[retrieval layer] (hybrid: vector + BM25 + entity link)
       │
       ▼
[state mutator]
       │   ├── AUTONOMOUS: low-risk fields (recency, prefs)
       │   └── HUMAN-GATED: anything that changes policy/voice/identity
       │
       ▼
[next run] ─── core memory in prompt + retrieved episodic + skill lookup

Where humans gate (non-negotiable for production):

  • Promotion of any item to "core memory" / system-prompt block
  • Schema changes in graph memory
  • Skill-library additions used by >1 user (Voyager-style accumulation needs review when shared)
  • Reward model updates / fine-tunes from agent feedback

Where it can be autonomous: episodic append, vector indexing, retrieval scoring tweaks, single-user preference KV updates with cheap reversibility, Reflexion-style within-task verbal self-correction (lives in scratchpad, not persistent memory).

Stage 6 — Eval Harness

Six patterns, ship at least the first three before going live:

  1. Golden set — 50–500 hand-curated (input, expected behavior, expected memory side-effect) tuples; include adversarial / poisoning attempts.
  2. Regression on memory side-effects — assert get(user, "allergies") == ["peanut"] after run X.
  3. Drift alarms via OpenTelemetry GenAI semconv — judge-score rolling mean, memory-store size growth rate, retrieval hit-rate distribution, % of runs that mutate core memory.
  4. A/B between agent versions — slice traffic, compare composite reward over fixed window.
  5. LLM-as-judge with human calibration — 5–10% audit; recompute judge–human Cohen's κ weekly.
  6. Held-out human-written tasks — never trained on; detects distribution collapse from self-play.

Deep dive: references/eval-harness.md

Stage 7 — Risks Checklist

Walk all 8 once. Each must have a concrete mitigation in the design doc.

  1. Memory poisoning (MINJA, ≥95% lab injection success)
  2. Prompt injection via memory
  3. Reward hacking
  4. Drift / staleness
  5. Context rot / window blowup (200K models often unreliable past ~130K)
  6. Runaway self-modification
  7. Distribution collapse in self-play
  8. Multi-agent context explosion

Deep dive: references/risks.md

Stage 8 — Output

Produce the design document:

  • Memory schema — chosen tier(s), data model, retention/TTL, redaction hooks
  • Closed-loop spec — signal source, collector, experience store, retrieval, mutator, human-gate list
  • Eval harness plan — golden set sketch, regression assertions, OTel metric list, A/B split, judge-calibration cadence
  • Risk register — 8 risks × 1 mitigation each
  • Build order — what ships in week 1 (state cache only), week 4 (signal collection on production traffic), week 12 (closed loop activated behind feature flag)

Show full SKILL.md (957 more words)Show less

Memory Architecture Ladder (escalate only when justified)

L → L → M → M → M → H → XH
1    2    3    4    5    6    7
#ArchitectureUse caseCostPitfallCitation
1Scratchpad-only (in-run, discarded)Multi-step reasoning within one task; ReAct loops; debate transcriptsLDon't fake durability — make it obvious to LLM and ops nothing persistsReflexion
2Conversation summary (rolling LLM compaction into system prompt)Single-session chat, support tickets, ≤1 day horizonLSummaries lossy-compress unpredictably; pin facts verbatim, summarize narrativeAnthropic context engineering
3Episodic stream (append-only event log, recency × importance × relevance retrieval)Long-running personas, simulations, journal-style apps where order mattersMBespoke scoring; without reflection, bloats fastGenerative Agents (Park et al., 2023)
4Vector RAG over interactionsKnowledge retrieval, FAQ, doc Q&A, low-personalizationMReactive only — won't surface "favorite color" on "birthday" queryLetta — RAG vs Agent Memory
5Key-value fact store (mem0 single-pass ADD)Personalization (name, prefs, history), CRM-like agentsMBad extractors poison store; need write-time validatorsmem0 paper
6Graph memory (mem0g, AriGraph)Multi-hop reasoning over relationshipsHSchema drift kills you; LLM-extended schemas degrade into vector store with extra stepsmem0g
7Hierarchical OS-style (Letta / MemGPT, agent self-edits via tools)High-stakes long-horizon agentsXHSelf-editing memory is prompt-injection bomb on untrusted inputMemGPT, Letta

Default recommendation in the skill: start at #1, escalate one tier at a time. Many "we need memory" requests are actually #2.

Deep dive: references/architectures.md


Anti-Patterns (load-bearing — call out before user picks the wrong path)

Anti-patternTestFix
Memory because it's coolAdding mem0/Letta to a one-shot pipelineSkip memory. Stateless + RAG.
Cache labeled "memory"No feedback signal exists in the user's domainHonest naming: call it a "state cache" not "learning". Design accordingly.
Vector RAG for personalization"What's my favorite color?" returns nothing because the user never asked it; embeddings can't surface unprompted factsKV fact store, not vector RAG
Self-editing memory on untrusted inputLetta with user-pasted content writing into core memoryQuarantined-LLM pattern; never untrusted source → core memory
Reward hacking via vanity metricsEngagement-only signal → clickbait drift; finance "% reviewed" → rubber-stampingComposite rewards: engagement + brand-fidelity judge + sample audit; finance: composite includes materiality threshold + reviewer agreement
Memory as the first moveBuilding memory store before the stateless agent has shippedShip stateless first. Instrument the signal. Decide a quarter later.
Graph memory by defaultModeling 1 brand's 5 competitors as a graphStay in KV+vector until >3 entities × >50 relationships. Graph schemas drift; LLM-extended schemas degrade into vector stores with extra steps.
Self-play with no external verifierAgent training on its own outputs, no held-out signalPin a verifier external to the model. V-STaR / Quiet-STaR loops without external verification narrow capability.
Forgetting context-rotStuffing 130K of memory into context "because the model supports 200K"Compaction + retrieval + sub-agent isolation; 200K models often unreliable past ~130K (Anthropic)

Self-Improvement Playbook Ladder (cheapest first)

Reflexion → Generative Agents → Voyager → mem0 → Letta
   1            2                  3        4      5
TierPatternWhenCitation
1In-loop verbal correction, no persistenceCheapest learning; the first move before ANY memory store. ~91% pass@1 HumanEval at the time of publication. Lives in the scratchpad.Reflexion (Shinn et al., 2023)
2Long-horizon persona / social simsMemory stream + reflection + planning loop. For agents that need to act in character over days/weeks.Generative Agents (Park et al., 2023)
3Skill-library accumulationTool-using agents solving novel-but-related tasks; "what worked for Brand X in vertical Y" patterns.Voyager (Wang et al., 2023)
4Production fact memoryChat-like personalization at scale. 91.6 LoCoMo, ~90% token savings vs full-context.mem0 (arXiv 2504.19413, ECAI 2025)
5Self-editing hierarchical memoryHighest power, highest attack surface. Use only when long-horizon autonomy is the product, not a nice-to-have.MemGPT → Letta

The skill walks the user up this ladder only when justified by a concrete failure of the tier below. Most production systems sit at tier 1 + tier 4. Tier 5 is appropriate for <5% of agentic projects.

Deep dive: references/playbook-ladder.md


Reference Files

FileContents
references/architectures.mdDeep-dive on the 7 memory architectures with cost ratings L→XH
references/feedback-signals.mdPer-domain feedback signal design + the no-signal-no-learning rule
references/eval-harness.mdThe 6 eval patterns: golden set, regression, drift alarms, A/B, judge calibration, held-out tasks
references/risks.mdThe 8 risks with citations and mitigations (MINJA, prompt injection, reward hacking, drift, context rot, runaway self-mod, distribution collapse, multi-agent explosion)
references/playbook-ladder.mdReflexion → Generative Agents → Voyager → mem0 → Letta progression
references/case-studies.mdBrandling Mutation Engine "state cache, not learning" lesson + marketing/finance/HAZOP/tutorial-gen worked examples through the memory/feedback lens

Examples

The examples/ directory will hold:

  • reflexion-loop.md — cheapest first move, scratchpad-only
  • kv-store-mem0.md — production personalization with extraction validation
  • eval-harness.md — golden set runner with regression assertions

Output Contract

A skill run is complete when the user has:

  1. A documented answer to "cache or learning?" (Stage 1).
  2. A scored need-memory rubric (Stage 2).
  3. A chosen architecture tier with justification for not stopping at the previous tier (Stage 3).
  4. (If learning) a named feedback signal with latency, source, and mitigations (Stage 4).
  5. (If learning) a closed-loop spec with human gates marked explicitly (Stage 5).
  6. An eval harness plan with at least patterns 1–3 from §3.8 (Stage 6).
  7. A risk register: 8 rows × 1 mitigation each (Stage 7).
  8. A build order showing what ships when (Stage 8).

If the user wants to skip steps, the skill refuses. The whole point is the filter.


Design Philosophy

Memory is a liability surface. The cheapest memory is the one you didn't add.

Every memory tier you add carries a recurring debugging tax (why did it remember that? why did it forget this?), a security tax (every read is untrusted input), a privacy tax (GDPR/HIPAA delete-on-request), and an eval tax (regression on memory side-effects). Stateless agents fail in ways you can reproduce by re-running the input. Memoryful agents fail in ways you can't.

The skill's stance: earn each tier with a real failure on a real eval set. When in doubt, ship the lower tier and instrument the signal. Decide next quarter.

© ooiyeefei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/self-improving-systems of ooiyeefei/ccc.

  • SKILL.md
  • README.md
  • examples/eval-harness.md
  • examples/kv-store-mem0.md
  • examples/reflexion-loop.md
  • references/architectures.md
  • references/case-studies.md
  • references/eval-harness.md
  • references/feedback-signals.md
  • references/playbook-ladder.md
  • references/risks.md

Open the folder on GitHubat commit c0fd926

Compare with similar skills

Self Improving Systems next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Self Improving Systems compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Self Improving Systems this skillooiyeefei/ccc494—~5.2kAutomated safety check: PassMIT
Mem0 CLI Memory Commandsmem0ai/mem067k—~2kAutomated safety check: NotesApache-2.0
Mem0 Remember Commandmem0ai/mem067k—~560Automated safety check: PassApache-2.0
Mem0 Memory Scopemem0ai/mem067k—~1.1kAutomated safety check: PassApache-2.0
Mem0 Memory Searchmem0ai/mem067k—~502Automated safety check: PassApache-2.0
Mem0 Project Tourmem0ai/mem067k—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Adds, searches, lists, updates and deletes memories on the Mem0 platform from the terminal with the mem0 command, including a JSON mode built for agents.

    67k GitHub stars~2k tokensUpdated yesterday
    Agent WorkflowsAuto-check: notes
  • Saves a fact, decision or preference the user states into mem0 as written, labeled with a memory type such as decision, convention or user_preference.

    67k GitHub stars~560 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Shows or changes the default Mem0 memory scope, project, session or global, which decides where memories are saved and searched.

    67k GitHub stars~1.1k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Looks up stored agent memories by keyword or ID and prints compact one-line results instead of full detail.

    67k GitHub stars~502 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Shows everything mem0 has stored for the current project, grouped by category, with a compact search mode and an all-projects view.

    67k GitHub stars~1.6k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Adds Mem0 memory to an existing repository with a test-first pipeline that detects the language, lets you choose Platform or open source, and leaves a local feature branch.

    67k GitHub stars~3.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from ooiyeefei/ccc

All 22 skills in this repo
  • Turns meeting recordings into notes with a chain of custody from audio to claim, auditing transcripts for gaps and low-confidence numbers and names.

    494 GitHub stars~876 tokensUpdated 2 mo ago
    Auto-check passed
  • Builds marketing and explainer videos in Remotion from rendered scenes, with one real product capture as proof, and cuts them for each platform's formats.

    494 GitHub stars~2.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Records a sharp product demo video by driving the real app with a browser agent, from storyboard to Xvfb capture and a narration script synced to the frames.

    494 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Generates architecture diagrams as .excalidraw files by analyzing a codebase, with optional PNG or SVG export through Playwright.

    494 GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Sets up GA4 on a website and wires one real conversion event end to end, verified in DebugView before any money goes into ads.

    494 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Builds or rewrites SaaS landing pages by researching the real product, positioning it against alternatives and writing buyer-focused copy, then implementing it in the codebase.

    494 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Categories

Questions about Self Improving Systems

What does Self Improving Systems do?

Decide whether your agent actually needs persistent memory, feedback loops, or closed-loop learning, then design the smallest thing that pays for itself. Self Improving Systems is an agent skill from ooiyeefei/ccc. Decide whether your agent actually needs persistent memory, feedback loops, or closed-loop learning, then design the smallest thing that pays for itself.

When should I use Self Improving Systems?

Self Improving Systems fits situations like: the user says add memory; give my agent context management; make my agent learn; self-improving / closed-loop.

How do I install Self Improving Systems in Claude Code?

Run `npx skills add ooiyeefei/ccc --skill self-improving-systems -a claude-code`. Or copy the skill folder (skills/self-improving-systems in ooiyeefei/ccc) into .claude/skills/self-improving-systems in your project. Claude Code loads it when a task matches its description.

How do I install Self Improving Systems in Codex?

Run `npx skills add ooiyeefei/ccc --skill self-improving-systems -a codex`. Or copy the skill folder (skills/self-improving-systems in ooiyeefei/ccc) into .agents/skills/self-improving-systems in your project. Codex loads it when a task matches its description.

Can I use Self Improving Systems in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ooiyeefei/ccc --skill self-improving-systems -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/self-improving-systems, .gemini/skills/self-improving-systems, .github/skills/self-improving-systems and .opencode/skills/self-improving-systems in your project.

What does Self Improving Systems need to run?

SKILL.md names no scripts, command-line tools or credentials: Self Improving Systems is instructions for the agent only.

Does Self Improving Systems access the network?

SKILL.md names 5 domains. As links in the text: arxiv.org, anthropic.com, docs.letta.com, opentelemetry.io and letta.com. This is read from the text; nothing was executed.

Is Self Improving Systems safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Self Improving Systems use?

Self Improving Systems is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Self Improving Systems use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 17k tokens, read only when the agent opens those files.

What are the alternatives to Self Improving Systems?

Skills that share tags, products or a category with Self Improving Systems: Mem0 CLI Memory Commands (mem0ai/mem0, 67k stars), Mem0 Remember Command (mem0ai/mem0, 67k stars), Mem0 Memory Scope (mem0ai/mem0, 67k stars) and Mem0 Memory Search (mem0ai/mem0, 67k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Self Improving Systems?

ooiyeefei (a GitHub user) maintains it in ooiyeefei/ccc, which has 494 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on July 29, 2026.

Source: ooiyeefei/ccc on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.