Orca CLI
stablyai/orca
Operate Orca-managed worktrees, folder contexts, terminals, repos, automations, artifacts, skill sharing, worktree comments, and Orca's embedded browser…
Manage the cross-session tracking document with restore, create, update, or compact (only user triggered, no agent self-invocation).
$ npx skills add gweslab/cerf --skill tracking -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install gweslab/cerf tracking --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/gweslab/cerf.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tracking .claude/skills/tracking && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tracking" agent skill from https://github.com/gweslab/cerf/tree/main/.claude/skills/tracking into .claude/skills/tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tracking", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/gweslab/cerf/tree/main/.claude/skills/trackingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add gweslab/cerf --skill tracking -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install gweslab/cerf tracking --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gweslab/cerf.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/tracking .agents/skills/tracking && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tracking" agent skill from https://github.com/gweslab/cerf/tree/main/.claude/skills/tracking into .agents/skills/tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tracking", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add gweslab/cerf --skill tracking -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install gweslab/cerf tracking --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gweslab/cerf.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/tracking .cursor/skills/tracking && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tracking" agent skill from https://github.com/gweslab/cerf/tree/main/.claude/skills/tracking into .cursor/skills/tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tracking", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/gweslab/cerf.git --path .claude/skills/tracking--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add gweslab/cerf --skill tracking -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install gweslab/cerf tracking --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gweslab/cerf.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/tracking .gemini/skills/tracking && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tracking" agent skill from https://github.com/gweslab/cerf/tree/main/.claude/skills/tracking into .gemini/skills/tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tracking", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install gweslab/cerf trackingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add gweslab/cerf --skill tracking -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/gweslab/cerf.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/tracking .github/skills/tracking && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tracking" agent skill from https://github.com/gweslab/cerf/tree/main/.claude/skills/tracking into .github/skills/tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tracking", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add gweslab/cerf --skill tracking -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install gweslab/cerf tracking --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/gweslab/cerf.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/tracking .opencode/skills/tracking && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tracking" agent skill from https://github.com/gweslab/cerf/tree/main/.claude/skills/tracking into .opencode/skills/tracking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tracking", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
trackingManage the cross-session tracking document with restore, create, update, or compact (only user triggered, no agent self-invocation).
Tracking is an agent skill from gweslab/cerf. Manage the cross-session tracking document with restore, create, update, or compact (only user triggered, no agent self-invocation).
Its SKILL.md is about 22k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Session handoff. The repository describes itself as: Universal Windows CE Emulator - CE Runtime Foundation. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 462ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tracking loads about 22k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 13,922 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from gweslab/cerf at commit 462ed3d, republished under its MIT licence (© gweslab). 13,922 words, ~22,385 tokens.
.claude/skills/tracking/SKILL.md (or your agent's skills folder).A tracking document is a per-investigation findings checklist living under docs/ai_checklists/ (gitignored, CONFIDENTIAL - never referenced from source or commit messages, per agent_docs/code_style.md). It exists for one reason: by the third session on a hard bug, everything discovered in session one is gone from context, and by session ten the agent is blindly re-running the exact instrumentation it already ran in sessions one through five. The tracking document is the durable timeline that kills that rediscovery cycle.
THE GOVERNING LAW: full preservation, never hiding. This document is an EVIDENCE RECORD, not an accomplishment report. Every fact that cost effort, every decision that was contested, every approach the user /bad'd, every conclusion that is uncertain or disputed, and the task itself with the reason it exists - ALL of it is preserved in full, especially when it makes the session or the agent look bad. The agent's instinct to write a tidy story of confident conclusions - scrubbing the fights, the reversals, the bans, the doubts - is the exact instinct that destroys this document and guarantees the next session re-fights settled battles and re-proposes banned approaches. Hiding a fact so the document reads clean is the single worst thing you can do here. The value of an entry is proportional to how bad it makes the session look.
This skill has four subcommands. The user either types one explicitly (/tracking restore <path>, /tracking create, /tracking update, /tracking compact) or types bare /tracking, in which case you DEDUCE which one from the test conditions below and ANNOUNCE the deduction verbatim before doing anything - except COMPACT, which is destructive restructuring and is NEVER deduced from a bare /tracking; it fires only on an explicit /tracking compact.
You never edit a tracking document on your own initiative. Only the user invoking /tracking create, /tracking update, or /tracking compact authorizes a write. A memory line can also pre-authorize one compaction (see § Automatic compaction). That line is the user's instruction, given ahead of time. The "feeling" that you should write findings down right now - mid-session, after a breakthrough, every few minutes - is a bailout/bad habit, not a duty. Only the user knows when a session ends and when a write is warranted. (See § UPDATE and § Anti-patterns.)
You may NEVER bring up the tracking document yourself. Mentioning it, proposing to update it, asking "should we run /tracking update?", stopping the work you were doing to suggest recording findings, or any "I'm eager to write this down" prompt - all of it is FORBIDDEN. Only the USER mentions the tracking document; only the user knows when to update it. The agent has zero standing to raise it. Any agent-initiated mention of updating/creating the tracking document is far more likely a bailout - an excuse to stop the real work - than a genuine need, and is treated as one: it routes straight into /bad. If you catch yourself about to type "want me to update the tracking doc?" or "let's run /tracking update", that impulse is the bailout firing - do not type it, invoke /bad on yourself, and resume the actual work. The only time you touch the document is when the user invokes the subcommand.
Tracking documents come in two kinds with opposite RESTORE semantics. Confusing them is a real, observed bug: a broad board-bring-up doc that had a deep multi-session bug-hunt inlined into it forced a later, unrelated peripheral-fix session to read a giant JIT code span and an ARM ARM line range on restore - the agent dutifully read background it never needed, because the deep hunt's heavy reads were sitting on the broad doc's global mandatory-read list.
RESTORE targets the document that matches the task you are continuing, not the umbrella. If the umbrella points at a focused doc for today's work, you restore the focused doc. If today's work is a slice the umbrella tracks directly (no focused doc), you restore the umbrella - light in its file fan-out, but still read in full, every session block (see § RESTORE: the document itself is always read whole, archetype only governs how many external files follow).
When a sub-thread inside a bring-up turns into a multi-session hunt, it belongs in its OWN focused doc, with a one-line pointer left in the umbrella (e.g. "rendering→JIT bug: see zune_jit_render_findings.md (resolved S4)") - never inlined into the umbrella's mandatory-read. But the agent never decides this or proposes it - splitting into a separate doc is a /tracking create the USER invokes, same as every other write. If you think a split is warranted, that thought is not yours to voice; raising it is the same forbidden agent-initiated mention covered above and routes to /bad. The user owns document structure.
A PostToolUse hook (check_checklist_edit.py) fires on EVERY Write / Edit under docs/ai_checklists/. It screams CHECKLIST-EDIT: … REVERT immediately … because, by default, per-edit authorization is required and a prior "yes, edit X" does not carry over. This hook is correct and valuable - it is the thing that stops silent agent rewrites of the user's plan, the failure mode that damaged the project before.
A /tracking create, /tracking update, or /tracking compact invocation IS the explicit, full authorization for the writes this skill's protocol prescribes - end-to-end, for the duration of the protocol. While you are executing the create/update/compact protocol:
CHECKLIST-EDIT on each one - that is expected noise, NOT a signal to revert. Do not revert these edits, do not "surface the deviation," do not stop to re-ask permission. The user already authorized them by invoking the subcommand. (COMPACT is the one protocol that legitimately rewrites earlier content of the live document - collapsing old session blocks to one-liners - because the full text is preserved verbatim in the pre-compact/<name>.md file it just wrote/appended; this is the sole sanctioned exception to the append-only rule, and it is authorized only inside an explicit /tracking compact.)The moment the skill protocol completes, the authorization is gone. Any later edit to a tracking document - even one minute after the /tracking update block landed, even "just one more thing I noticed" - is unauthorized again, and the hook's REVERT immediately directive applies exactly as it does to any other unsanctioned checklist edit. To write again, the user must invoke /tracking update again. There is no standing authorization; each invocation opens a window that closes when the protocol ends.
/tracking with no subcommandThe three test conditions are simple and almost never ambiguous:
COMPACT is never on this deduction list. It restructures the live document (collapsing old blocks), so it fires ONLY on an explicit /tracking compact - never inferred from a bare /tracking. The one exception is the memory-gated chain after a RESTORE (see § Automatic compaction), where the memory line is the standing authorization. If you ever feel the document is "too big" and want to compact it, that urge is the same forbidden agent-initiated mention as proposing an update: do not raise it, route it to /bad. Only the user decides a document needs compaction.
Pick the one whose condition holds. Then, before acting, say this verbatim (filling the bracketed parts with what you actually detected):
Condition is: <the concrete condition you detected, e.g. "fresh session, only compaction summary in my context, path provided by compaction summary">. Executing /tracking <restore|create|update> <path-if-applicable>.
Example, word-for-word in the shape required:
Condition is: fresh session, only compaction summary in my context, path provided my compaction summary. Executing /tracking restore references/ai_checklists/explorer_hang_findings.md
If the conditions are genuinely ambiguous (e.g. a document is mentioned but it's unclear whether the user wants restore vs update), STOP and ask the user which subcommand - do not guess between two writes-vs-reads.
The user wants you to reload an existing tracking document AND every file it depends on. This is invoked in ~99.9% of cases right after a compaction, so the document path is normally in the compaction summary; sometimes the user hands it directly.
RESTORE depth follows the archetype (see § Two archetypes) - but "light" governs the FILES, NEVER the document's own session blocks. The document itself is ALWAYS read in full, every session block, in both archetypes; the archetype only decides how many external files you pull in afterward. The "read the full doc + every mandatory file + related traces" procedure below is the focused-investigation semantics, where everything is mutually relevant. If the path you were given is an umbrella / progress tracker, RESTORE is light in its file fan-out: read the doc in full (it is a coarse index - read all of it anyway), note current state, and follow its pointer to the focused doc for the task you are continuing - restoring THAT focused doc with the full procedure. Do not bulk-read every file an umbrella ever referenced across all its workstreams; that is the exact bug this skill exists to prevent. "Light" never licenses skipping the umbrella's own session blocks - read every one. Restore depth (of files) matches the doc you are actually continuing work in; reading the document itself in full is unconditional.
TASK & WHY FIRST, and in your sign-off state the task + why it exists + the banned approaches in your own words. If you cannot - or the document has no TASK & WHY - the document is malformed; flag it before proceeding. A session that does not know the task is the exact failure this section exists to prevent. If the document carries a UI / INTERACTION GATES map, also state in your own words how the target state is reached - the screen order, which steps are [USER-GATED], and where the known-good/known-bad boundary sits. You may not run CERF toward the symptom before you can state this; an autonomous run that parks at a user gate and gets read as "nothing happens" is the drift § The UI / INTERACTION GATES rule exists to kill.COMPACTION LOG, record its pre-compact archive path as a lookup target for the rest of the session. Do NOT read the archive in full on RESTORE. The archive is unbounded, and a full read cancels the compaction that made the live doc readable. State the path in the sign-off. From that point you MUST grep the archive each time the session needs grounding that the live doc does not carry (see § The pre-compact archive). A restore that misses the archive sends the session to re-derive what a collapsed session already solved.ls the tracing directory of the device (cerf/tracing/<bundle>/) and judge each file against the CURRENT goal. Do not read the directory by default. Judge from the filename and the hook targets, not from a full read. For each file, state a 0-100% figure for how much it serves the NEXT target the document names. Read a file only when you can name the thing in that target it serves. Zero files read is a good outcome. It needs no apology. A device carries probes from investigations that closed long ago. Those probes buy nothing in an unrelated task, and they cost the whole read. The mandatory list stays the floor. This directory is a candidate pool, never a second floor.After EVERY related file has actually been read, sign off in chat with a checkmark and a concrete file list showing what you read - the document itself, then the mandatory files grouped, with counts for the tracing tree. Shape:
✅ /tracking restore complete.
- Tracking document:
docs/ai_checklists/<name>.md- read in FULL, all<X>session blocks (Session #1 … Session #<X>), every section- Mandatory reads (from the document's section):
<file>,<file>, … (M files, all read)- Tracing dir
cerf/tracing/<bundle>/:<N>files listed,<K>read -<file><%> (serves:<what>), … - or "0 read, nothing served the current target"- Live doc size:
<L>lines. <the compaction line for that count, if any>- Pre-compact archive:
docs/ai_checklists/pre-compact/<name>.md(Sessions #1-#<K>collapsed) - NOT read in full, to grep for any grounding the live doc does not carry - or "no COMPACTION LOG, the document was never compacted"- <any extra mandatory files the document specified>
- Stats:
<W>weeks in,<N>sessions,<B>bans earned,<S>archived.
Count the lines of the live document (wc -l), then add the matching line to the sign-off. Under 2000 lines, add nothing.
| Lines | Line to add |
|---|---|
| >= 2000 | 🟡 You can run /tracking compact to optimize read amount |
| >= 2500 | 🔴 You should run /tracking compact to optimize read amount |
| >= 3000 | 🔴💀 You MUST run /tracking compact to optimize read amount |
Use the highest tier the count reaches. This line sits inside a user-invoked RESTORE, so it is not the forbidden agent-initiated mention of the document. Never run /tracking compact off the back of it, unless § Automatic compaction says you must. The recommendation goes in the sign-off, and the user decides.
User memory carries the switch, and it reads exactly:
/tracking skill AUTOMATIC compact on restore is ENABLED
Read memory for that line on every RESTORE. The line is a standing authorization from the user. It lifts the explicit-only rule for this one path, and for nothing else.
/tracking compact yourself, straight after the RESTORE sign-off. Do not ask, and do not propose. The memory entry is already the answer.Write or remove that memory line only when the user asks to enable or disable the switch. Every other rule about user memory still holds.
Close the tier line with the state of the switch:
| Condition | Suffix |
|---|---|
| no memory line | (autocompact unknown) |
| memory says DISABLED | (autocompact disabled) |
| enabled, 2750 lines or more | (autocompact NOW) |
| enabled, 2500-2749 | (autocompact very soon) |
| enabled, 2250-2499 | (autocompact soon) |
| enabled, under 2250 | (autocompact not soon) |
When one more session block carries the document past 2750, write (autocompact probably next session) in place of the band word. Judge that from the size of the recent blocks.
Close both sign-offs with one stats line. Every number comes from the document and its archive:
- Stats:
<W>weeks in,<N>sessions,<B>bans earned,<S>archived.
Session #1 timestamp to today's real date, taken from a tool.Session #N in the document.TASK & WHY → BANNED APPROACHES.The block count is not decorative: stating "all <X> session blocks (Session #1 … Session #<X>)" forces you to account for every block, and you may write it ONLY if you actually read every one of them. A sign-off claiming all 8 blocks when you read 2 is a fabricated success exactly like a fake-success stub.
Any skipped read is a complete violation and a failed RESTORE. You must NOT sign the checkmark unless every session block, every mandatory-list file, AND every trace file you judged worth reading was actually read in this session. Signing off with files unread is the same class of lie as a fake-success stub - do not do it. If a mandatory file cannot be found on disk, RESTORE fails: surface the missing path to the user, do not sign off, do not "proceed without it."
What to do next? Do not ask user their direction if the next step is clear: just continue.
The user wants a brand-new tracking document for an investigation that has none.
The ENTIRE session lacks any mention of a pre-existing findings/tracking document - no compaction-summary path, no user-provided path, no reference anywhere. That absence IS the signal: the user wants to start the chain.
docs/ai_checklists/ for something close, no "I found a related findings file, shall I extend that?" There is no document - create one.docs/ai_checklists/ describing the investigation (e.g. explorer_hang_findings.md, <device>_<symptom>_tracking.md). Match the naming of neighbors in that directory.TASK & WHY (the task + why it exists, captured from the user's framing; FORBIDDEN CONCLUSIONS / BANNED APPROACHES start empty) and Files mandatory to read - then the first session's block (Session #1, with a real timestamp). A document created without TASK & WHY is malformed from birth./tracking create. That is the only authorization needed; do not also ask "should I create it?" after they already said create.zune_jit_render_findings.md") - and nothing else in the umbrella. Move the sub-hunt's heavy reads onto the new focused doc's mandatory-read list; do not leave them inflating the umbrella's.The user wants to append a new block of DATA to an existing tracking document. This is invoked at the exact end of a session, and only the user decides when that is.
/tracking update, and you do NOT even suggest one - proposing an update, asking "should I record this?", or stopping work to raise the document is itself the bailout and routes to /bad (see the intro and § Anti-patterns). Only the user knows when the findings have settled enough to commit, and only the user raises it.Files mandatory to read every entry that this restore read and that served nothing. Leave an entry in place when the NEXT target still needs it. A file that costs a read and serves nothing is not a floor entry. It is a tax on every restore that follows. State the judgment in the sign-off.Session #X (increment from the highest existing session number in the document) and a real timestamp obtained from a tool (date via Bash / Get-Date via PowerShell) - never a guessed or invented time. Agents chronically skip the timestamp, which turns a 50-block single-day document into an unreadable mess with no timeline. Get the real time./verify verdict from this session (especially CRITICAL PROBLEM FOUND), any un-committable state, any regression the session caused goes there verbatim with file:line. This is the highest-value part of the update and the one agents drop because it contradicts the accomplishment narrative - see § Document structure → The most-critical-data rule. If you ran /verify this session and its verdict is not in this block, the UPDATE is incomplete.DISPUTED / CORRECTED / REVERSED, the /bad, and the WHAT THIS SESSION ACTUALLY DID sub-sections (see § Document structure) - mandatory whenever they apply. These are the records agents scrub to make a session read clean, and their omission is what sends the next session to re-fight settled battles and re-propose banned approaches. Mirror any new ban into the global TASK & WHY → BANNED APPROACHES, and sharpen the WHY if this session clarified it. Then verify the whole UPDATE against § Preservation-not-hiding - the three completion tests - before declaring it done.UI / INTERACTION GATES sub-section (see § Document structure) - current and complete. The boot/UI sequence with verbatim on-screen text samples, every [USER-GATED] step the agent cannot perform itself, and the known-good/known-bad boundary. If a prior block already holds the map and nothing changed, restate it or point at that block explicitly; if this session moved the boundary (a step got cleared, the symptom moved), the map is updated to say so. An UPDATE on a GUI-reachable investigation that leaves the next session unable to state "how does the user reach the symptom, and at which step does the agent have to hand over to the user" is incomplete - it is the exact omission that sends the next agent to re-crack a step already known to pass (see § The UI / INTERACTION GATES rule).TASK & WHY ban/forbidden-conclusion ITSELF as ONE LINE - the rule + ×N count + (S#) - and leave the full rationale in this session's block, never inline in TASK & WHY. This is the bloat-prevention rule: a ban whose back-story is re-narrated as a multi-paragraph essay in the global section makes TASK & WHY grow without bound (it is append-only-clarify, so the essay never leaves on its own). The session block is where the evidence belongs; the global ban is a pointer to it. Shape: ❌ <one-line rule> - /bad'd ×2 (S3, S7), see S7. The full "why it was banned, what we tried, how it failed" stays in the S7 block.✅ /tracking update complete. Session #
<X>appended.
- Restore value: earned its read -
<file>,<file>. Served nothing -<file>,<file>.- Dropped from mandatory-read:
<file>, … - or "nothing dropped"- Stats:
<W>weeks in,<N>sessions,<B>bans earned,<S>archived. (see § Session stats)
The served-nothing half is the point of the line. A file named there is a file the next RESTORE does not read.
The user wants to shrink a tracking document that has grown too big to read in one call - the classic symptom is an umbrella tracker whose per-session blocks have piled up (or a long focused hunt) until a single RESTORE read no longer fits comfortably and leaves no window for several more sessions of work. COMPACT trades the narrative bulk of old sessions for one-liners + a pointer, while preserving every piece of critical data in full inside the live document.
The user explicitly typed /tracking compact (optionally with a path), or a RESTORE chained it under § Automatic compaction. COMPACT is NEVER deduced from a bare /tracking and the agent NEVER proposes it (see § Deduction and § Anti-patterns). If no path is given and exactly one tracking document is live in this session, that is the target; if it's ambiguous, STOP and ask which document.
Every tracking document has at most ONE archive: docs/ai_checklists/pre-compact/<original_file_name>.md (same base name, same .md, NO numbering). It is the complete, never-lossy history - a reader who needs the full uncompacted detail opens this one file, always.
So the pre-compact file accumulates the full text of every session ever written, in order, forever; the live document is the shrunk working copy that points at it. The compacted live document ALWAYS carries a pointer to this file (see the COMPACTION LOG step).
This file is the searchable long-term memory of the investigation. Every later session MUST grep it for grounding that the live doc no longer carries. § The pre-compact archive governs that use. Compaction is safe only because that obligation holds.
Compaction is a length operation, not a forgetting operation. The pre-compact file guarantees nothing is ever truly lost. No section is exempt from compaction: a section left in full-prose essay form (TASK & WHY bans, CARRIED-FORWARD, the log) accretes forever and becomes the bulk of the document, which collapsing session blocks alone cannot bound. The rule:
Nothing critical is ever DROPPED or SOFTENED from the live document - but nothing is kept in essay form either. Every critical item is reduced to its bounded actionable form (the rule / the map / the fact) + a see S# pointer; the full prose lives in the archive. This is the same move the session-digest makes (collapse the narration, keep a terse actionable line), applied to EVERY section so the whole document is bounded, not just the session tail.
The distinction that makes this safe - applied per item type:
×N count, and which session earned it - is immutable: never dropped, never softened, never de-counted, stays exactly as loud. The multi-paragraph essay re-narrating the investigation that produced it compresses to a one-line rule + see S#; the essay is in that session's block (digested) and verbatim in the archive. Compressing the essay is NOT softening the ban - the ban is just as binding at one line as at thirteen.[USER-GATED] markers, known-good boundary all stay in full while the investigation is live. A map is already the bounded actionable form (it has no "essay" to shed); collapsing it to a label re-creates the rediscovery cost it exists to prevent, so a live map stays in full. But a map for an axis that is superseded / refuted / resolved (the block itself says "DEAD" / "lives only in the archive") is no longer live - it DROPS from the live doc entirely (it is in the archive).When in doubt whether something is "live," keep it - a slightly-too-long doc that kept a live map beats a tidy one that dropped it. But "live" is the test, not "exists": a refuted map, a resolved defect, and a ban's back-story are all kept in the archive, not the live doc.
Two numbers govern the finished size: ~700 lines is the TARGET, ~1500 lines is the HARD limit. Aim for 700. That size reads in one call and leaves headroom for several more sessions. Never finish with more than 1500 lines, because a document past that costs a whole read before any work starts. Collapsed session blocks alone do not make a compaction "done". It is done when the WHOLE document is inside the limit. On a long hunt that means you compress the bans, the carried-forward maps, and the log, not just the sessions.
Decide how far to go yourself. If the ordinary passes leave the document at more than 1500 lines, continue down the ladder in step 9. Collapse a wider window, merge digest ranges, and drop what has become least important. Do not stop to ask which items can go. You have just read the whole document, so you know which threads are dead and which are load-bearing. The archive holds every word either way. A question here hands off a judgment you are better placed to make. The document also stays oversized until the answer arrives.
So the procedure for each old block is: hoist first, then collapse. Pull any still-live critical item out of the block (into the carried-forward section below, tagged with its origin session), THEN reduce what remains to one line - and separately run the global-section passes (steps 7-8 below, then enforce the step-9 ceiling) over the bans, carried-forward maps, and the log.
Resolve the path (explicit, or the single live document; ambiguous → ask). Read the document in full FIRST if it is not already fully in your context this session - you cannot safely collapse blocks you have not read, and you must know which blocks carry live critical data.
Update the single pre-compact archive at docs/ai_checklists/pre-compact/<original_file_name>.md (same base name as the live doc, no numbering; docs/ai_checklists/ and therefore pre-compact/ is gitignored and confidential - same as every tracking doc):
Session #X heading with a tool (do not eyeball it), then append - verbatim, to the bottom, in order - every live-document session block whose number is greater than that X. Those are exactly the blocks added since the last compaction. Append nothing already present; never rewrite or re-copy existing content; never edit blocks already in the file.Add / refresh the COMPACTION LOG global section in the live document (place it directly under TASK & WHY, above Files mandatory to read). It is pure meta - the pointer to the full reference plus the current state - and is kept SMALL and bounded, never a growing per-compaction narration:
## COMPACTION LOG
This is a COMPACTED document. FULL uncompacted history (every session, verbatim):
docs/ai_checklists/pre-compact/<original_file_name>.md
GREP that file before you re-derive ANY fact this document does not carry. A digest
one-liner or a `see S#` pointer means the full record is in that file, not that it is lost.
State: Sessions #1-#<K> are one-lined in COMPACTED SESSION DIGEST; #<K+1>-#<N> kept full. Last compacted <timestamp>.Overwrite the State: line each compaction - do NOT append a new paragraph per compaction. The blow-by-blow of what each past compaction collapsed/hoisted carries no future value (it is recoverable from the digest + archive); the log must stay at exactly this fixed shape - pointer + grep directive + one State: line - no matter how many compactions have run. If this section currently holds per-compaction narration paragraphs, delete them down to the pointer + State: line. Timestamp comes from a tool (date / Get-Date), never guessed.
Hoist live critical data out of the blocks you are about to collapse into a CARRIED-FORWARD CRITICAL DATA global section (place it directly under Files mandatory to read). Each hoisted item keeps its full content and is tagged with its origin, e.g. (orig Session #3). This section holds ONLY items pulled from collapsed blocks - do-not-rediscover maps, UI / interaction gate maps, still-open defects, unresolved disputes, active bans not already in TASK & WHY. Recent full blocks keep their own critical data in place; do not duplicate it here. (Carried-forward items accumulate across compactions; never drop one because its origin session got collapsed - drop it only when the work it records is actually resolved.)
Decide the keep-window: the most recent 5-10 session blocks stay FULL, untouched. Pick the number that lands the live document inside a single comfortable read with headroom. If even 5 full recent blocks + the preserved globals still overflow, keep 5 and accept a longer file rather than collapsing a recent block - recency is where active work lives.
Collapse every older block to one line under a COMPACTED SESSION DIGEST section (place it just above the first surviving full block), chronological, one line per session. (On a re-compaction, extend the existing digest with the newly-collapsed sessions; keep prior digest lines.)
## COMPACTED SESSION DIGEST
Full text of every session below is in the pre-compact file (see COMPACTION LOG).
- Session #1 - <timestamp> - <terse what-it-established, ≤1 line>
- Session #2 - <timestamp> - <terse>Keep each one-liner factual and specific - "established X, disproved Y" - not "did some work." No per-line path is needed; the single pre-compact file named in COMPACTION LOG holds them all.
Compress the TASK & WHY bans & forbidden-conclusions to one line each. For every BANNED APPROACH and FORBIDDEN CONCLUSION, keep the RULE verbatim - what is forbidden + the running ×N count + the (S#) that earned it - and collapse any multi-paragraph justification essay to a see S# pointer. Shape: ❌ <one-line rule> - /bad'd ×3 (S1, S2, S5), see S5. Drop NO ban, change NO count, soften NO wording of the rule itself; only the back-story prose leaves (it is in that session's block + the archive). A 13-line ban becomes one line; a ban already one line stays. TASK & WHY's task + WHY paragraphs are untouched - only the ban/forbidden lists compress.
Prune CARRIED-FORWARD CRITICAL DATA and Files mandatory to read of dead weight. From CARRIED-FORWARD, DROP every item whose work is resolved/committed or whose axis is superseded/refuted (the entry or its session says "DEAD" / "REFUTED" / "lives only in the archive" / "in-code + archived") - those are history, preserved in the archive, not a standing obligation. KEEP every genuinely-live reusable map in full detail (the DO NOT REDISCOVER map rule governs - never collapse a live map to a label). From Files mandatory to read, remove must-reads for any path fully closed this hunt.
Enforce the ceiling: count with a tool, then descend the ladder until it fits. Count the lines of the live document (wc -l). At 700 lines or fewer, you are done. At more than 1500 lines, you are not finished, however much you already collapsed. Between the two, continue down the ladder while anything on it is spent.
Work the ladder top-down. Stop at the first rung that brings the count inside the limit. Each rung sheds something less important than the rung below it. Everything you shed stays verbatim in the pre-compact archive.
Sessions #1-#40 - board bring-up: INTC/LCD/timer/DMA landed + committed; all resolved. This is the largest and cheapest win on a long hunt. Dozens of lines become one, and no live fact is touched.<subsystem> master-map + call chain - see S<n> in the pre-compact archive. A stub with a real destination is not the forbidden tombstone. The map survives whole in the archive, and the stub says exactly where to find it.Re-count after each rung. Rungs 1 and 2 carry almost every document on their own. Treat rung 5 as a last resort, not a routine move. Four things never go at any rung, because a document without them cannot start a session:
Verify, then sign off. Re-read the live document top to bottom and confirm it passes the three completion tests (§ Preservation-not-hiding) PLUS the compaction-specific checks below. Then sign off (shape in § Sign-off).
Session #X equals the live doc's highest. A gap means the diff append missed a block.CARRIED-FORWARD CRITICAL DATA. A map that is now only a digest one-liner is a FAILED compaction.TASK & WHY with its ×N count and (S#) intact. Compaction may compress a ban's justification essay to a see S# pointer (step 7) but NEVER drops a ban, changes a count, or softens the rule wording. A ban that vanished or lost its count is a FAILED compaction./verify verdict is still stated in full (surviving block or carried-forward), not reduced to a digest line.[USER-GATED] markers, and known-good boundary. Only superseded/refuted/resolved maps were dropped (step 8), and those are in the archive. A live map reduced to a label is a FAILED compaction.✅ /tracking compact complete.
- Pre-compact archive:
docs/ai_checklists/pre-compact/<name>.md- <created as full copy | appended Sessions #<A>-#<B>>; now holds Sessions #1-#<N>in full.- Compaction-log refreshed to pointer + State line (meta-history dropped).
- Collapsed Sessions #1-#
<K>to digest one-liners; kept Sessions #<K+1>-#<N>in full.- Bans compressed to rule + count +
see S#(<B>bans, all counts intact). CARRIED-FORWARD pruned of<D>superseded/resolved items. Mandatory-read pruned of<R>closed paths.- Carried forward in full:
<M live do-not-rediscover maps>, <P open defects/disputes>, all bans - nothing live dropped.- Live doc length: ~
<lines>(was ~<lines>) - target ~700, hard limit ~1500. Ladder rungs applied: <none needed | 1-5>.
If any compaction-specific check fails, you must NOT sign off. Recover the dropped item from the pre-compact file into the live doc, then re-verify.
docs/ai_checklists/pre-compact/<name>.md is the long-term memory of the investigation, and the richest source of grounding you have. The live document is a lossy summary. Compaction collapsed whole sessions to one line each. It compressed the rationale of every ban to see S#. It dropped every map whose axis was superseded. All of that text went into the archive, verbatim and in full. The archive holds the exact register dumps, the exact hook fires, the exact addresses, and the exact reasons an approach failed. This is the concept of the pre-compact file: a dump of the prior work, kept so that a recurring question gets an ANSWER instead of a new investigation.
The observed failure: an agent restores a compacted document and hits a question the live doc does not answer. It reads the digest one-liner "Session #4 - mapped the mmc polling path" as all that survives. It then spends the session on a chain that the archive already records hop by hop, three hundred lines in. The agent never opened that file. The digest line was never the record. It was a POINTER at the record, and the agent read it as a tombstone.
If the current task needs grounding that the live document does not carry, you MUST grep the pre-compact archive BEFORE you derive that grounding yourself. This is not optional and not a judgment call. It is a mandatory prior step, of the same rank as a datasheet read before you write a register handler. When you derive from scratch what the archive already holds, you spend the money of the user a second time.
Grep the archive on ANY of these signals. The list gives examples and is not complete:
see S4, (orig Session #7), a digest one-liner), and you need what that session found.see S#. You must know what was tried and how it failed before you plan an alternative. An alternative that is the banned approach under a new label is still banned.Do NOT read the archive from top to bottom. It is unbounded by design and only grows, and that full read is the cost that compaction removed. Grep it. Then read the full session block around the hit.
Session #4). Read the whole block, not the matched line.<tokens>, no hits"). Do not continue in silence, as if you never looked. An empty search is a real result, and it is what permits you to derive the fact yourself.The archive is read-only to you, always. Every rule in this skill against edits to a tracking document applies to the archive. Step 2 of an explicit /tracking compact is the only writer. Do not append to it. Do not correct it. Do not "tidy" a stale entry you find during a grep. Do not move a finding out of it into the live doc on your own initiative. If a grep finds something you think the live document must carry, that belief is not yours to act on and not yours to speak. To raise it is the same forbidden agent-initiated mention as a proposal to update, and it routes to /bad. Use what you found, continue the work, and let the user decide.
Free-form in detail, but structured by intent - the next agent must be able to read it as a timeline and instantly find "what's already known, what not to repeat, what to read first." Concretely:
Two mandatory global (non-per-session) sections sit at the TOP, in this order. They are the only global blocks; everything else is per-session.
Global 1 - TASK & WHY (immutable). Every tracking document MUST open with this, and RESTORE reads it FIRST. It carries:
/bad'd, ONE LINE each: the rule + RUNNING COUNT + (S#) (e.g. "section-flag injector edit - /bad'd ×2 (S4, S9), see S9"). The full rationale (what was tried, how it failed) lives in that session's block, NOT inline here - re-narrating it inline is the bloat that makes TASK & WHY the largest section of a long document. A banned approach may NOT be re-proposed - not even relabeled "grounded / evidenced / root-caused / it's different this time" - until the USER explicitly re-authorizes it. Re-proposing a banned approach is an automatic violation.
TASK & WHY is append-only-clarify: a session may sharpen the WHY, add a forbidden conclusion, or add a banned approach. It may NEVER narrow, soften, delete, or de-count the task, the WHY, or any ban. (/tracking compact MAY compress a ban's justification prose to its one-line rule + see S# - the rule, count, and (S#) are immutable; the back-story moves to the archive. That is the sole edit compaction makes here, and it is not "softening" - see § COMPACT.)Global 2 - Files mandatory to read. EVERY document is REQUIRED to carry it. It explicitly lists the files a RESTORE must read - most of which are trace files under cerf/tracing/<bundle>/. List them by path, do not hand-wave "the tracing tree"; a device with a gigantic investigation has far more probes than belong on a curated mandatory list, so the section names the specific files that matter. RESTORE reads this section as the floor and additionally reads related-looking trace files in the device dir that the section may not yet name (see § RESTORE). Grow it on UPDATE as new must-read files appear.
Everything else is per-session blocks, appended in order so the timeline is visible. Apart from the two governing global sections above (TASK & WHY, Files mandatory to read), there are NO other global/merged blocks - do not hoist findings into a single global "Findings" or global "Do not repeat" list, because that destroys the timeline and is exactly how the document drifts into the schizo-list failure. The ONLY exception is a document that has been compacted: /tracking compact may add three further global sections - COMPACTION LOG, CARRIED-FORWARD CRITICAL DATA, and COMPACTED SESSION DIGEST (see § COMPACT) - and these are legitimate global blocks created ONLY by an explicit user compaction, never by CREATE/UPDATE and never by the agent on its own. Outside of a compaction they must not appear. (The distinction: TASK & WHY and mandatory-read are cross-session INVARIANTS, not timeline data; findings, disputes, and maps are timeline data and stay per-session.) Each block is one session's self-contained summary.
Each per-session block is headed Session #X - <full timestamp> and contains, as sub-sections. The first one is mandatory and comes FIRST for a reason (see § The most-critical-data rule below):
/verify verdict run this session goes here VERBATIM - especially a CRITICAL PROBLEM FOUND - with each finding's file:line and the one-line fix. If a /verify round is still open, record the path of its prompt file (tmp/verify/<slug>.md) with the verdicts. Any known-broken behavior, any regression the session's own changes caused, any "works at runtime but architecturally wrong" smell, any reason a commit must wait - all here, stated plainly. If code was written this session and you did NOT verify it, say that too ("uncommitted, UNVERIFIED"). This sub-section is the single highest-value thing in the whole document; omitting a known critical defect in your own just-written code is the worst failure an UPDATE can commit.DISPUTED (unresolved) / AGENT WAS WRONG / USER WAS RIGHT.DISPUTED, with the user's position recorded at equal or greater prominence than the agent's. Scrubbing the user's position to present a clean agent conclusion is the cardinal sin of this document./bad THIS SESSION (mandatory whenever one fired). Every /bad the user issued: what the agent was doing that triggered it, and what approach/conclusion is now BANNED. Mirror each ban into the global TASK & WHY → BANNED APPROACHES list with its running count. Omitting a /bad you received is hiding the strongest steering signal the user gave you.[USER-GATED] with what the user does and how completion is observable (the next screen's text, a log line).<branch>", "delete <file>". A NEXT block that dictates any of these is reverted on sight.A DO NOT REDISCOVER entry of the "known path / known facts" kind is worthless unless it preserves the actual content. The recurring failure: an agent spends a session decompiling a path across several binaries - A.exe!sub_X at 0x… calls into B.dll!Foo at 0x… under condition Z, which calls C.dll!Bar at 0x… - and then records it as a single bare line: "DO NOT REDISCOVER: the explorer launch path, already mapped." That line carries no address, no module, no call condition. After the next compaction the document says "don't rediscover it" while giving the next agent nothing to work from - so the next agent re-decompiles the entire chain anyway, which is the exact rediscovery the section was supposed to prevent. The label became a gravestone, not a map.
So:
B.dll!Foo to re-read its body. That is cheap and legitimate. What the map saves is the EXPENSIVE part: discovering which functions, in which binaries, in what order, under what condition form the path. Point the next agent precisely ("decompile B.dll!Foo at 0x… to re-read the field check") instead of making them rediscover that Foo is in the chain at all.The point of per-session blocks (rather than merged global lists) is the timeline: a reader sees session 1 → 2 → 3 and how the understanding evolved, including retractions recorded forward in later blocks rather than rewrites of earlier ones.
The observed failure: an agent spends sessions on a hang inside a guest welcome wizard. The path to the hang is known in-session: the wizard opens with a touch calibration step the agent cannot pass itself (precise taps - the USER passes it), then a step 2 where the user manually enters data, and only then step 3, where the wizard hangs on mmc polling - THAT is the bug. Then /tracking update, /compact, /tracking restore - and the interaction path was never written down. The fresh agent runs CERF autonomously, sees a screen where "nothing happens" (the calibration waiting for taps it cannot perform), has no record that calibration even exists, and drifts: it burns money re-investigating and concludes "the hang is: calibration never finishes" - a step the previous session KNEW the user passes successfully - then spends more time rediscovering that the user must pass it. The solved ground got re-cracked because the UI reality was never preserved.
So:
[USER-GATED], with what the user does and how completion is observable.The observed failure: an agent writes a long, polished UPDATE - full sweep tables, instrumentation, multi-binary maps - and OMITS the single most important fact, "the code I wrote this session has a CRITICAL PROBLEM FOUND verify verdict and is not safe to commit." That omission is not random. It happens for three structural reasons, and you must counter all three:
Prioritize the UPDATE by cost-of-omission, not by how good it makes the session look. Rank every candidate fact by: what damage does the next session take if this is missing? The top of that ranking is always - known defects in just-written code, an un-committable state, a /verify verdict, a regression this session caused. Those are recorded FIRST and are non-negotiable. The "look how much I accomplished" summary is the LEAST important part of an UPDATE; if anything gets shortchanged for space or time, it is the success story, never the defect record. A next session that commits the broken code, or burns a fresh /verify re-discovering defects the last session already knew, is the exact money-waste this document exists to prevent.
A confident WRONG conclusion is worse than an honestly uncertain one: the next session trusts it, builds on it, or has to burn a session disproving it. The words confirmed, verified, proven, carefully confirmed, reliable, definitely, 100%, safe to use are BANNED on any conclusion that is EITHER (a) not runtime-verified (an actual log line / hardware behavior / test result observed THIS session), OR (b) contradicts something the user has asserted. Such conclusions are written HYPOTHESIS or DISPUTED, never as settled fact. Observed failure: a session wrote "DISCRIMINATOR confirmed reliable … both need CE5 layout" and "safe to use GetVersionEx, confirmed very carefully" - the conclusion was never runtime-verified AND directly contradicted the user's repeated assertion that WM5 runs the CE6 model. The next session trusted the word "confirmed," then spent a full session disproving it and re-extracting the truth the user had stated all along. Confidence is earned by runtime evidence, not by the strength of the agent's belief. If you did not see it work, you did not confirm it.
Before any UPDATE is done, it must pass all three. Failing any one means the UPDATE is incomplete regardless of how polished the rest is:
[USER-GATED] steps, known-good boundary)? If no → TASK & WHY / BANNED APPROACHES / UI / INTERACTION GATES is incomplete./bad'd, reversed, disputed, or left uncertain that is NOT written down? If yes → write it before you stop.
The instinct to produce a tidy, confident, accomplishment-shaped document is the instinct that fails all three. Rank every candidate entry by cost-of-omission - how badly the next session is hurt if it's missing - and the entries that make the session look worst sort to the top./tracking create|update. The single worst pattern. The "I should record this now" urge is a bad habit; only the user authorizes a write.check_checklist_edit.py hook screamed. During a create/update protocol the hook's REVERT warning is expected noise on prescribed writes - the invocation already authorized them. Do not revert, do not re-ask. (See § The checklist-edit hook.)/tracking update./tracking update", "I should record this before we forget", or halting work to raise the document - all forbidden, all far more likely a bailout than a real need, all route straight to /bad. Only the user mentions the document; only the user knows when to update. The agent has no standing to raise it.cerf/tracing/<bundle>/ and read related-looking trace files the list hasn't caught up to. Equally, do NOT blindly read the entire tree on a device with a huge probe set - read the mandatory list plus what looks related, and report what you judged unrelated./bad. The user owns document structure and invokes /tracking create when a split is wanted.Session #X from the document's existing max./tracking compact, which one-lines old blocks after preserving their full text in a verbatim archive and hoisting their critical data - and only under an explicit user compaction. Outside COMPACT, the past is never edited.CRITICAL PROBLEM FOUND verdict on your own just-written code, an un-committable state, or a regression the session caused MUST go in the mandatory CODE STATE sub-section, first, verbatim with file:line. It is the highest-value carry-forward in the document; leaving it out so the session reads as a clean "✅ complete" sets the next session up to commit broken code or re-discover defects you already knew. (See § The most-critical-data rule.)[USER-GATED] steps the agent cannot perform itself, the known-good boundary - sets the next session up for the observed drift: it runs CERF autonomously, parks at a user gate (a calibration waiting for taps), reads it as "nothing happens", and burns money re-concluding a hang at a step the previous session KNEW the user passes. (See § The UI / INTERACTION GATES rule.)agent_docs/code_style.md); a code comment or commit that names one leaks it.DISPUTED / CORRECTED / REVERSED, with the user's position at equal-or-greater prominence than the agent's.confirmed / verified / proven / reliable / safe to use on a conclusion that is unverified at runtime or that contradicts a user assertion. See § The confidence-honesty rule. Such conclusions are HYPOTHESIS or DISPUTED./bad'd - even relabeled "grounded / evidenced / root-caused / different this time." It stays banned until the USER explicitly re-authorizes. Re-proposing it is an automatic violation; record it in BANNED APPROACHES instead./bad the user issued. Every steering correction the user gave this session is recorded - what triggered it and what is now banned. Omitting it discards the strongest signal in the session.TASK & WHY is append-only-clarify; the task and its rationale only get sharper, never weaker. A session that cannot state the task from the document must reconstruct TASK & WHY before doing anything else, not proceed./tracking. COMPACT is destructive restructuring; it fires ONLY on an explicit /tracking compact, or on the memory-gated chain in § Automatic compaction. The "it's too long" urge is the same forbidden agent-initiated mention as proposing an update and routes to /bad.docs/ai_checklists/pre-compact/<name>.md destroys it irrecoverably. The pre-compact file is written (first compaction) or appended (later) BEFORE any block is one-lined.Session #X and append only blocks numbered above it. Re-copying duplicates sessions; editing existing content corrupts the single full reference.__001, __NNN) or making more than one per doc. There is exactly ONE pre-compact file per tracking doc, named for the live doc with no index. It is the single growing reference; never a chain of snapshots.CARRIED-FORWARD CRITICAL DATA (tagged with origin session) or kept in a surviving full block - NEVER reduced to a label. A live map that survives only as "Session #3 - mapped the launch path" is a failed compaction: the next agent re-decompiles the chain, the exact rediscovery the map prevented. (A superseded/refuted/resolved map is the opposite case - it is correctly DROPPED from the live doc, since it is dead and preserved in the archive; see step 8.) Critical-data preservation beats length, always.TASK & WHY's bans as multi-paragraph essays, dead maps in CARRIED-FORWARD, and a growing COMPACTION LOG leaves the doc far over the ceiling. Compaction is done when the WHOLE live doc is inside the limit (steps 7-9), not when the digest grew. A "compacted" doc still over the limit because the globals were untouched is a FAILED compaction.see S# only; the rule text, the ×N count, and the (S#) are immutable. A ban that vanished, lost its count, or got reworded weaker is not compression - it is the forbidden softening of TASK & WHY.COMPACTION LOG accumulate a per-compaction narration paragraph. The log is pure meta and stays at its fixed shape (pointer + grep directive + one State: line) forever; it is overwritten each compaction, never appended-to. A log that grew one fat paragraph per compaction is itself a bloat source and must be collapsed (step 3).TASK & WHY on UPDATE. The bloat-prevention failure: writing the new ban as a multi-paragraph essay in the global section instead of a one-line rule + count + see S# with the rationale left in the session block. This is what forces a step-7 compression pass later; write the one-liner from the start.see S# ban pointer as "all that survives". Both "Session #4 - mapped the mmc polling path" and "❌ <rule> - /bad'd ×2 (S3, S7), see S7" are instructions to open the archive at that session. If you read them as the complete record, you map the path a second time. Or you plan an "alternative" that is the banned approach under a new label, because you never read why it was banned. That is the exact loss compaction must make impossible./tracking compact. An append, a correction, a "tidy" of a stale entry you find during a grep, or a move of a finding into the live doc on your own initiative. The archive is read-only to you. Step 2 of an explicit compaction is the only writer. A proposal to move a finding is a forbidden agent-initiated mention, and it routes to /bad.© gweslab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/tracking of gweslab/cerf.
Open the folder on GitHubat commit 462ed3d
Tracking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tracking this skillgweslab/cerf | 103 | — | ~22k | Automated safety check: Pass | MIT | |
| Orca CLIstablyai/orca | 87k | 2 repos | ~593 | Automated safety check: Pass | MIT | |
| Beads Task Memorygastownhall/beads | 28k | — | ~1.2k | Automated safety check: Pass | MIT | |
| Session History Searchslopus/happy | 24k | — | ~3.1k | Automated safety check: Pass | MIT | |
| Paseo Agent Handoffgetpaseo/paseo | 20k | 1 repos | ~606 | Automated safety check: Pass | Custom licence | |
| Memori Long-Term MemoryMemoriLabs/Memori | 17k | — | ~2k | Automated safety check: Notes | Custom licence |
stablyai/orca
Operate Orca-managed worktrees, folder contexts, terminals, repos, automations, artifacts, skill sharing, worktree comments, and Orca's embedded browser…
gastownhall/beads
Tracks multi-session work with dependencies in the bd issue tracker so the agent can find ready tasks and recover its context after conversation compaction.
slopus/happy
Searches past Claude Code, Codex and Cursor sessions and summarizes what was worked on, tried or decided, using extraction scripts instead of reading raw logs.
getpaseo/paseo
Hands off the current task, including context, decisions and failed attempts, to a fresh agent through Paseo by writing a self-contained briefing prompt and launching that agent.
MemoriLabs/Memori
Connects Claude Code to Memori Cloud for long-term memory, recalling stored context before substantive replies and saving new context afterward.
liwp/again
A skill your agent uses when working in a repository that uses bd or Beads for durable project task tracking, issue dependencies, blocker management, multi-session handoff, or shared work memory.
gweslab/cerf
List the project skills and offer the environment doctor. An agent skill from gweslab/cerf.
gweslab/cerf
Add a changelog entry for a change that was just made (only user triggered, no agent self-invocation).
gweslab/cerf
Create a git commit with a short message that describes the diff.
gweslab/cerf
Start the bring-up of a new board or ROM in CERF. An agent skill from gweslab/cerf.
gweslab/cerf
Spawn a hostile reviewer that checks a claim or a diff against the project rules.
gweslab/cerf
Audit a list of options for bailouts and rule violations before the user picks one.
Categories
Manage the cross-session tracking document with restore, create, update, or compact (only user triggered, no agent self-invocation). Tracking is an agent skill from gweslab/cerf. Manage the cross-session tracking document with restore, create, update, or compact (only user triggered, no agent self-invocation).
Tracking fits situations like: tasks that involve Session handoff.
Run `npx skills add gweslab/cerf --skill tracking -a claude-code`. Or copy the skill folder (.claude/skills/tracking in gweslab/cerf) into .claude/skills/tracking in your project. Claude Code loads it when a task matches its description.
Run `npx skills add gweslab/cerf --skill tracking -a codex`. Or copy the skill folder (.claude/skills/tracking in gweslab/cerf) into .agents/skills/tracking in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gweslab/cerf --skill tracking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tracking, .gemini/skills/tracking, .github/skills/tracking and .opencode/skills/tracking in your project.
SKILL.md names no scripts, command-line tools or credentials: Tracking is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tracking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 22k tokens (SKILL.md is roughly 90k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Tracking: Orca CLI (stablyai/orca, 87k stars), Beads Task Memory (gastownhall/beads, 28k stars), Session History Search (slopus/happy, 24k stars) and Paseo Agent Handoff (getpaseo/paseo, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
gweslab (a GitHub organization) maintains it in gweslab/cerf, which has 103 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 7, 2026.
Source: gweslab/cerf on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.