Vercel Composition Patterns
supabase/supabase
React composition patterns that scale. An agent skill from supabase/supabase.
Verify agent session format compatibility for Agent Sessions.
$ npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jazzyalex/agent-sessions agent-session-format-check --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jazzyalex/agent-sessions.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-session-format-check .claude/skills/agent-session-format-check && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-session-format-check" agent skill from https://github.com/jazzyalex/agent-sessions/tree/main/skills/agent-session-format-check into .claude/skills/agent-session-format-check/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-session-format-check", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jazzyalex/agent-sessions/tree/main/skills/agent-session-format-checkType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jazzyalex/agent-sessions agent-session-format-check --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jazzyalex/agent-sessions.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agent-session-format-check .agents/skills/agent-session-format-check && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-session-format-check" agent skill from https://github.com/jazzyalex/agent-sessions/tree/main/skills/agent-session-format-check into .agents/skills/agent-session-format-check/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-session-format-check", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jazzyalex/agent-sessions agent-session-format-check --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jazzyalex/agent-sessions.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agent-session-format-check .cursor/skills/agent-session-format-check && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-session-format-check" agent skill from https://github.com/jazzyalex/agent-sessions/tree/main/skills/agent-session-format-check into .cursor/skills/agent-session-format-check/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-session-format-check", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jazzyalex/agent-sessions.git --path skills/agent-session-format-check--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jazzyalex/agent-sessions agent-session-format-check --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jazzyalex/agent-sessions.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agent-session-format-check .gemini/skills/agent-session-format-check && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-session-format-check" agent skill from https://github.com/jazzyalex/agent-sessions/tree/main/skills/agent-session-format-check into .gemini/skills/agent-session-format-check/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-session-format-check", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jazzyalex/agent-sessions agent-session-format-checkInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jazzyalex/agent-sessions.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agent-session-format-check .github/skills/agent-session-format-check && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-session-format-check" agent skill from https://github.com/jazzyalex/agent-sessions/tree/main/skills/agent-session-format-check into .github/skills/agent-session-format-check/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-session-format-check", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jazzyalex/agent-sessions agent-session-format-check --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jazzyalex/agent-sessions.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agent-session-format-check .opencode/skills/agent-session-format-check && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-session-format-check" agent skill from https://github.com/jazzyalex/agent-sessions/tree/main/skills/agent-session-format-check into .opencode/skills/agent-session-format-check/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-session-format-check", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-session-format-checkVerify agent session format compatibility for Agent Sessions.
Agent Session Format Check is an agent skill from jazzyalex/agent-sessions. Verify agent session format compatibility for Agent Sessions. Use when any agent CLI updates, when monitoring flags drift, or when bumping max verified versions (fixtures + docs + tests). Covers session schema, usage/limits tracking, storage backends, and discovery path contracts for all supported agents.
Its SKILL.md is about 13k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development. The repository describes itself as: Local-first macOS app to browse, search, analyze, and resume supported AI coding-agent session history across Codex, Claude Code, OpenCode, Cursor Agent, Antigravity, Hermes… The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit b789596. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
claudexcodebuildcursorFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
cursor.comAlso links to:
platform.claude.comdevelopers.openai.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYANTHROPIC_API_KEYFACTORY_API_KEYGITHUB_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Session Format Check loads about 13k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 6,826 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jazzyalex/agent-sessions at commit b789596, republished under its MIT licence (© jazzyalex). 6,826 words, ~13,183 tokens.
.claude/skills/agent-session-format-check/SKILL.md (or your agent's skills folder).Answer one precise question for every supported agent:
Can current Agent Sessions code support the latest available session/storage/usage format from the latest available agent build?
The answer must be layered. Do not collapse version checks, stale samples, schema fingerprints, discovery contracts, and usage probes into a vague severity label.
Evidence-first:
Related skill: agent-support-matrix — maintains the matrix YAML, ledger, and
update-checklist workflow. This skill focuses on detection and evidence collection;
agent-support-matrix focuses on recording and gating version bumps.
Process doc: docs/agent-support/monitoring.md — defines the compatibility verdicts,
legacy severity model, cadence, and escalation workflow that feed into this skill.
Run weekly monitoring:
./scripts/agent_watch.py --mode weeklyReport path prints to stdout and is written under
scripts/probe_scan_output/agent_watch/*/report.json.
In report.json, check each agent under results.<agent>:
compatibility.verdict, compatibility.scope, compatibility.blockers,
and compatibility.next_actionverified_version, installed.parsed_version, upstream.parsed_versioncompatibility.latest_status to distinguish current_fetch_known,
cached_latest, and unknown latest-source statesweekly.local_schema (newest local session used for fingerprinting)weekly.schema_diff and evidence.schema_matches_baselineevidence.sample_freshness and evidence.fresh_evidence_sourcecompatibility.latest_real_session_failure when a prebump attempt failedseverity and recommendation only as legacy escalation fieldsSanity-check version identity before acting on it. If a monotonic CLI reports an
installed version newer than upstream.parsed_version, or the fetched tag belongs to a
different release family (for example an SDK release instead of the CLI), treat the
latest source as broken/unknown. Do not follow a generated prebump recommendation for
an agent with no driver, and do not lower a verified ceiling to match a bad source.
Usage / limits reading (Codex + Claude) — always verify every weekly run.
These drift independently of session schema (see §2), so a clean schema does
not imply healthy usage reading. Each agent's
results.<agent>.weekly.probes is a list; for every relevant entry confirm
ok == true and exit_code == 0:
label == "codex_status_probe" (parse codex_status_json): the
active CLI status channel (five_hour, weekly percent-left). The passive
channel is the session JSONL token_count / rate_limits events, covered by
the schema fingerprint above.label == "claude_usage_probe" (parse claude_usage_json): the
authenticated /usage reading (session_5h, week_all_models,
week_opus). Also label == "claude_status" (parse claude_status_json):
status.claude.com indicator/incidents.
A failed or unparsed usage probe is a usage-format or auth regression even when
versions match and the session schema is clean — never skip it, and report each
probe's ok explicitly rather than collapsing it into the compatibility verdict.Classify every new field for value, not just for safety (§1e). A field that parses
cleanly is safe, which is not the same as handled. Do not close a drift finding
until each new key or type has been called handled, noise, watch, or feature-candidate
— and candidates filed in docs/backlog.md, not left as a remark in the ledger.
Interpretation:
supports_latest: latest known build is covered by
evidence.fresh_evidence_source == "latest_prebump_report" and
compatibility.latest_real_session_evidence == true with
compatibility.latest_status == "current_fetch_known". This verdict is not a clean
bill by itself: also require the current weekly.schema_diff.unknown_types == [],
weekly.schema_diff.unknown_keys == {}, healthy probes, and a passing discovery
contract. The compatibility layer can reuse an older clean prebump report while a newer
weekly sample has already drifted; in that case report and investigate the current drift
rather than repeating supports_latest.supports_installed_only: installed build is covered by non-stale real local
evidence, but latest is newer, cached from a prior report, unknown, or lacks
fresh real-session proof.latest_unknown: no configured/reachable latest source or no real-session
driver exists; do not claim latest support. Check STEWARDS.md before calling this
a finding — for a steward-verified agent not installed here, this verdict is the
expected output and the agent belongs to its steward, not to you (§1g).blocked_stale_sample: evidence predates the installed CLI; run prebump before claiming support.blocked_no_fresh_evidence: a version changed but no fresh matching sample proves support.format_drift_detected: unknown schema/storage/usage fields appeared; update fixtures/parsers.blocked_thin_sample: the sample was both narrow and tiny, so it evidenced nothing either
way (§5). Generate a session that actually uses tools — not a one-line prompt.monitoring_broken: latest source, usage probe, or discovery contract failed.real_session_auth_failed in blockers: the real-session driver ran but the
sandboxed agent was not authenticated; re-auth or provide the configured env
token, then rerun prebump.Weekly scanning samples the newest on-disk session, which can predate a CLI
upgrade and give a false "safe to bump" call (the codex 0.120.0 trap and the
copilot session.shutdown trap). When weekly reports
recommendation == run_prebump_validator — or before you stage any
max_verified_version bump or latest-support claim — run the prebump path for
every active agent being claimed. The driver exercises the currently installed
CLI once inside a sandbox and diffs its output against the fixture baseline:
./scripts/agent_watch.py --mode prebump --agent codex --agent claudeExit-code contract:
0 — every requested agent produced a fresh session and the schema
matches baseline. Safe to bump.2 — at least one fresh session's schema does not match baseline.
Do not bump; investigate the schema diff in
scripts/probe_scan_output/agent_watch/<slug>-prebump/report.json.3 — at least one driver failed (timeout, auth, CLI not found, or
discovery contract violation — wrong session root, wrong glob, or
missing required event types).4 — config/invariant error: unknown --agent (or one with no
prebump block), missing/invalid discover_session contract,
credential hygiene failure (oversize / mode), or sandbox breach
(copilot hermeticity gate). Re-run with
--allow-real-home only if you understand your real config dir will
be mutated for that one invocation.Keep the direct prebump report.json for every non-zero run and report its schema diff or
driver error from that artifact. A later weekly scan may select older clean evidence or
fall back to a stale local sample and omit the failed fresh diff; it does not erase what the
just-completed driver observed.
Flags:
--agent <name> (repeatable) — restrict to specific agents. An
unknown agent or one without a prebump config block exits 4.--keep-sandbox — preserve the temp $HOME for debugging.--timeout-seconds N — per-driver timeout. CLI flag overrides
per-agent config; falls back to config, then global default (120s).--force-fresh — suppress staleness evaluation for this run only (records
stale_reason=forced_fresh in the report).--allow-real-home — copilot/real-HOME opt-in after a sandbox-breach
diagnostic; never persistent.A thin prebump can no longer downgrade a rich weekly union. A passing prebump
replaces the weekly schema_diff with its own one-prompt session, and the thin-sample
gate used to score only that: on 2026-08-13 codex reported blocked_thin_sample off a
20-event prebump while its weekly union carried 2808 clean events — adding evidence made
the verdict worse. _sample_is_thin() is now applied to every available sample and only
blocks when all of them are thin (pinned by test_thin_prebump_does_not_override_rich_weekly_union).
Give real_home_session agents a tool-using prompt. Their prebump session lands in the
real store and enters the newest-_LOCAL_SCHEMA_SAMPLE_COUNT window, so a "Say hello"
prompt actively degrades the next weekly sample — two such runs pushed antigravity's union
down to 24 events and a genuine blocked_thin_sample. Claude's "Say hi, then use the Bash tool to run pwd." is the pattern; verify any prompt change with a real run, since a
tool-using prompt can hang or return nothing on CLIs whose one-shot mode does not complete
a tool turn.
A passing prebump is a floor, not a ceiling. Drivers use one-line prompts, so a fresh
session may contain only the four most basic event types and still report
fresh_matches_baseline=True — it proves the CLI still writes parseable output, not that
rich event families are unchanged. For real_home_session: true agents that session also
lands in the real store and becomes the newest sample; §5's multi-session union is what
stops it from masking drift. Check the fresh session's type count before treating a pass as
broad evidence.
Configured real-session drivers today are codex, claude, antigravity,
copilot, opencode, hermes, openclaw, cursor, pi, kimi, and grok.
Droid is legacy-only and excluded from active checks. Qwen, Devin, and fx have no
drivers. Judge them on current weekly evidence plus their stewardship record (§1g), and
never claim latest-build support without a session written by that build. Qwen cannot
produce a newer session on this machine because its OAuth free tier was discontinued
2026-04-15; that remains blocked until a paid plan or alternate provider is configured
(see §1c).
Staleness short-circuits the schema verdict. blocked_stale_sample is reported
instead of drift, so a stale agent can be hiding real drift behind it. Kimi sat at
blocked_stale_sample while every one of its sessions carried two unmodelled event
types; building its driver surfaced them immediately. Treat a stale verdict as
"unknown", never as "clean".
Prebump uses the hybrid env-var-first auth policy: if the relevant API-key
env var (OPENAI_API_KEY, ANTHROPIC_API_KEY,
FACTORY_API_KEY, GITHUB_TOKEN) is set it is forwarded into the sandbox
and real HOME is never read. Otherwise the driver copies the declared
credential file from real HOME into the sandbox after running three hygiene
gates (64 KiB max, mode 0600, ≤90-day mtime warning). v1 drivers:
codex_exec, claude_print, antigravity_print, copilot_prompt,
opencode_run, hermes_oneshot, openclaw_local_agent,
cursor_agent_print, pi_prompt, and kimi_prompt. Some OAuth/keychain-backed CLIs use
real_home_session: true; run them with --allow-real-home so the session
lands in the real agent store instead of copying single-use auth state into a
sandbox.
A version bump is never just a YAML edit. max_verified_version claims that the
app parses sessions written by that build, and the only thing that can prove it is a
session that build actually wrote. So when upstream.parsed_version is newer than
installed.parsed_version, the sequence is fixed:
1. Install the newer CLI (its own updater, brew, or npm — whatever owns it).
2. ./scripts/agent_watch.py --mode prebump --agent <name> # writes a FRESH session
3. Confirm exit 0 and fresh_matches_baseline == true.
4. Only now bump max_verified_version + ledger + tracking log (§6).Step 2 is not optional, and skipping it makes the report actively worse. Installing
an update rewrites nothing on disk: the newest session is still the one the old build
wrote, so it is now older than the CLI binary and the agent flips from
supports_installed_only (a real claim, backed by real evidence) to
blocked_stale_sample (which per §1a means unknown, never clean). That is why the
2026-08-13 pass deliberately did not install Hermes 0.19.0: without a working
one-shot run to generate a fresh session, updating would have destroyed a verdict it
could not replace. Install only if you can complete step 2.
An agent with no prebump driver cannot do this loop at all. It is permanently capped
at supports_installed_only, however clean its weekly looks, because nothing can produce
a session from a build on demand. Build the driver first — grok_single (added
2026-08-17) is the worked example, and see §1d for what building one involves.
Re-read the installed version after the run. Several CLIs self-update because the prebump invoked them — copilot, cursor and antigravity all did on 2026-08-13 — so the build you validated may not be the build you started with. Bump to what the post-run report says is installed, not to what you intended to install.
Some agents are blocked for reasons no driver can fix, and the distinction matters because it decides whether there is engineering work to do:
installed=unknown with an
Exec format error that read like a driver bug; /opt/homebrew/bin/opencode was
actually the npm package's shell stub, because its postinstall never ran. Run
file $(which <agent>) before suspecting anything schema-shaped.qwen -p returns "Run /auth to switch to Coding Plan,
OpenRouter, Fireworks AI, or another provider," and qwen auth is itself "(removed)"
in 0.21.x. No driver, re-login, or fixture can produce a session. Record the verdict
honestly as blocked_stale_sample, say why in the matrix, and stop — it is a billing
decision for the maintainer, not a task.Do not let category three masquerade as category one. A missing driver looks like engineering work right up until you try to authenticate.
Drivers live in scripts/agent_watch_prebump_drivers.py and register into DRIVERS;
their config block is agents.<name>.prebump in
docs/agent-support/agent-watch-config.json. Copy the nearest existing driver rather
than starting from the Protocol.
Non-obvious requirements, each learned from a real failure:
discover_session with the modern roots/globs/required_types keys.
The config gate rejects anything else with exit 4, and the runtime validator then
proves the driver returned the artifact it claimed rather than some other file.MATRIX_KEY_FOR_AGENT (scripts/agent_watch.py). A missing
entry does not error: baseline_paths comes back empty, the diff takes its
"no baseline → nothing diffs" branch, and the prebump reports
fresh_matches_baseline=true having compared nothing. grok was missing from the
prebump copy of that map on 2026-08-17 while present in the weekly copy, so its first
driver would have passed vacuously. The map is now single-sourced and pinned by
test_every_monitored_agent_is_registered_in_the_rebuild_tool.prepare_auth is the
only auth path — never build env from os.environ inside a driver. Some agents have no
key env var at all (grok's only documented one is GROK_SANDBOX), so the credential
copy is their sole route and the 0600 hygiene gate always applies.real_home_session: true agents especially —
their session lands in the real store and enters the newest-5 weekly window, so a
"Say hello" prompt degrades next week's sample (§1a).summary.json is a discovery precondition; returning a transcript whose sidecar never
landed hands the fingerprinter half a session and reports the gap as schema drift.Then verify the driver the way you would verify a finding: run it, confirm
baseline_type_count in the report is non-zero (proving a baseline was consulted), and
confirm the fresh session actually contains a tool call rather than four trivial events.
This check has a structural blind spot, and it is worth stating plainly: it asks only
"does it still parse?" Because the parsers read JSON as dictionaries, the answer is
almost always yes — an unknown key is never read, an unknown type falls to a default:
branch and becomes a .meta event. So every finding exits the funnel as "no parser
change needed," which is true and also the end of the thought. Nothing ever asks
whether upstream just started telling us something a user would want to see.
The cost is not hypothetical. Kimi's turn.ended.durationMs was noted in the
2026-08-13 ledger as "the natural source of a per-turn duration UI" and then sat there,
because a ledger note is a remark, not work. And Qwen shipped with
"usage/rate-limit tracking" listed as an unsupported surface while every one of its
transcripts carried full per-call token accounting — 53 records in one ordinary session.
The data was never missing. Nobody opened it.
So classify every new key or type into exactly one of four buckets, and record which:
| Bucket | Meaning | Action |
|---|---|---|
| handled | current code already reads the field for a user-visible or integrity path | cite the code and a behavior test; refresh the normal fixture, but do not file duplicate work |
| noise | internal plumbing, ids, or telemetry nobody would look at | fixture only; say so once so it is not re-litigated |
| watch | meaningful but not actionable yet — a field that will matter if it starts appearing widely, or that only one source emits | fixture + a line in the ledger note |
| candidate | carries information a user would want on screen | fixture + an entry in docs/backlog.md |
Cheap heuristics for spotting a candidate: it is a number a user would ask about (tokens, cost, duration, context size), it names something currently anonymous (which connector ran this tool, which model, which provenance), or it records a state transition the UI hides (mode changes, rewinds, compaction).
Two rules that keep this honest:
handled, do not infer it from a similar field name. Cite the read path and a
test whose behavior depends on the new shape. A parser's tolerant fall-through or raw
JSON retention means safe, not handled.unsupported_surfaces describes what
the app does, never what the agent emits. Qwen's entry made a present surface look
absent for a full release. When a value pass contradicts a matrix line, the matrix line
is the thing that is wrong.budget_usd carries real dollars and looked like
a free replacement for the estimated-cost path — until counting showed it in 2 of 410
recent sessions, because it only appears when a budget is set. Confirm how often a field
actually occurs, and on which record types, before filing it as a candidate.Candidates go in docs/backlog.md under the matching area section, using that file's own
entry format, with verified stamped to the date of the sweep that found them. A format
check should end with two outputs: a clean bill of health, and a short list of things
upstream started telling us that we are not yet using.
An agent may have a steward: a contributor who uses that agent regularly and re-verifies its format a few times a year, or after a big vendor release. Maintained agents are checked by the maintainer; best-effort agents have no steward yet. A steward is not expected to know any of the above. They run one command:
./scripts/steward_check.py <agent> # ./scripts/steward_check.py --list-agentsIt runs the ordinary weekly scan (§1) restricted to that one agent, against the steward's own local sessions, and answers in plain sentences with one of three exits:
all good: <agent> format matches the baseline (N sessions sampled), plus the matrix's verified version. If their CLI is newer than the
verified version and the schema still matches, it says the matrix entry can be
bumped — that is the §1b evidence a maintainer needs.scripts/probe_scan_output/steward_check/<agent>/redacted-sample/, and a
ready-to-paste GitHub issue body (also saved as issue.md).What it deliberately does not do: write or rebuild any baseline fixture.
Deciding that drift is real and rebuilding a baseline stays a maintainer job
(scripts/rebuild_stage0_baseline.py --agent <agent> --emit, §5/§7).
The sample reuses rebuild_stage0_baseline._redact — the same trimming that
produces committed fixtures — and then re-scans the result for home directories,
emails, key-shaped strings, IPs and long opaque ids. If anything survives, the
sample is withheld entirely and the issue body says so; a steward is never
handed an almost-clean file to paste into a public issue.
Maintainer side of the same command: when a steward's issue arrives, its sample
is already in fixture shape, so it drops straight into
Resources/Fixtures/stage0/agents/<agent>/ for a baseline rebuild. Verify the
redaction yourself anyway before committing (§7).
STEWARDS.md assigns each agent a tier: Maintained (the maintainer runs it daily),
Steward-verified (a named contributor re-checks it on their own machine), or
Best-effort (nobody yet). A sweep must read that tier before deciding an agent is a
problem, because the scan cannot distinguish "this agent is broken" from "this agent is
not installed here." Both surface as installed=unknown, latest_unknown,
cli_binary_unresolved, or a discovery contract failing no_local_file.
For a steward-verified agent that is absent from this machine, all of those are the expected, correct output. They are not findings. Do not:
max_verified_version from a machine with no sessions from that build.
The steward's dated verification is the evidence, and overwriting it with a local
guess destroys the only real record.What a sweep does owe a steward-verified agent, in order:
latest_unknown — steward-verified,
no local CLI" is a complete, honest line. "fx: latest_unknown" alone reads as a hole../scripts/steward_check.py <agent>. That is
a ten-minute job for them and an impossible one here.Read the recommendation through the tier. The verdict vocabulary assumes the maintainer
can run the agent, so for a steward-owned one it gives advice you cannot take. Translate:
blocked_no_fresh_evidence / rec=run_prebump_validator on a steward agent does not
mean "run prebump" — no driver here can produce a session for a CLI that is not installed.
It means upstream moved past what the steward last verified, and the action is the ping in
step 3. Configure the agent's upstream source even when nothing else about it is
checkable locally: it is the only channel that can raise this, and without it the rule
above is unenforceable. fx sat at upstream=unknown for exactly this reason while v0.0.6
and v0.0.7 shipped past a record reading 0.0.5, and the sweep reported nothing.
A steward-verified agent only becomes maintainer work when its steward goes quiet. At
that point the honest move is to drop it to Best-effort in STEWARDS.md and say so,
not to quietly keep claiming a verification nobody performed.
Ordering trap. STEWARDS.md is the human table and the matrix is the machine record;
the file itself says the matrix wins on disagreement. But the tier lives only in
STEWARDS.md — the matrix has no steward: field — so a purely matrix-driven sweep is
structurally blind to stewardship and will keep re-reporting the same non-findings every
week. Read both, and if you find yourself about to file a steward-verified agent as a
gap, that is the bug.
Usage and limits tracking can drift independently of session schema. Monitor both.
token_count / rate_limits event structure. This is
covered by the schema fingerprint only because codex is fingerprinted nested (§5) —
these events live under event_msg.payload, and the flat fingerprint that ran until
2026-08-03 stopped at {payload,timestamp,type} and could never see them. If codex is
ever moved back to the flat fingerprint, this channel goes unwatched again.codex_status_capture.sh output schema — parsed as
codex_status_json by agent_watch.py.results.codex.weekly.probes (a list), the entry with
label == "codex_status_probe" returns ok == true and exit_code == 0.
If not, investigate whether Codex changed its status output format.claude_usage_capture.sh output schema — parsed as
claude_usage_json../scripts/claude-status --json records status.claude.com
indicator/incidents (parsed as claude_status_json).results.claude.weekly.probes (a list), the entry with
label == "claude_usage_probe" returns ok == true — if false, the usage API
response format may have changed, or authentication may be required. The
label == "claude_status" entry (parse claude_status_json) reports
status.claude.com indicator/incidents.parsing_failed, auth required, etc.), treat as high severity
because the UI can break.claude_usage_capture.sh can fail with exit 16 when Claude auth tokens are exhausted.
The /usage TUI command itself stops working in this state. This is not a format
change — resolution requires claude auth login. The failure is intermittent and clears
after re-authentication. The Swift side silently retains the last known good snapshot.
$ burn)The runway's $ presentation prices per-type token rates against a model table, so it
drifts whenever a provider changes prices or ships a model slug we don't know — with no
schema change and no failing probe. Nothing else in this scan catches it. Unlike a
broken probe, stale prices fail silently: the number still renders, just wrong.
Sources of truth (fetch these, don't recall them):
The table lives in two places that MUST stay identical:
docs/prices.json — served to clients from GitHub Pages (corrects shipped apps with
no release)RunwayPriceTable.bundledJSON in AgentSessions/CodexStatus/RunwayPriceTable.swift
— the compiled-in default (offline / pre-first-fetch)Prices unchanged? Compare every key's inputPerMTok / cachedInputPerMTok /
outputPerMTok / cacheWritePerMTok against the official pages. Anthropic cache
columns are derived: read = 0.1x input, 5m write = 1.25x input.
New model slugs? Any tier we don't have a key for is dropped from $ (it
still shows in tk/h), so a new model silently disappears from the cost view.
Check what the local CLIs actually emit rather than guessing:
# Codex: the model lives on turn_context lines
grep -ho '"model":"[^"]*"' ~/.codex/sessions/$(date +%Y/%m)/*/*.jsonl | sort -u
# Claude: message.model on assistant lines
find ~/.claude/projects -name '*.jsonl' -mtime -7 -print0 \
| xargs -0 -n1 jq -r 'select(.type=="assistant") | .message.model' 2>/dev/null | sort -uEvery slug returned must resolve via longest-prefix against a table key.
Known slugs that look alarming but are fine (verified 2026-07-14):
<synthetic> (Claude) — not a model. It appears on assistant lines and does
carry a usage object, but every field is 0, so it forms a zero-rate
component and dollarsPerHour skips it. Do NOT "fix" this by adding a price
key; the zero-rate exemption is what keeps it from dropping the whole session.
If Claude ever gives <synthetic> real tokens, that exemption stops applying
and every Claude session would vanish from $ — re-check this if it changes.codex-auto-review (Codex) — Codex's internal auto-review label, on turn_context.
It bills real tokens, and because an unpriced contributing slice makes
dollarsPerHour return nil for the whole session, a missing key here didn't
understate the cost — it deleted the session from $ entirely. Priced at the
gpt-5.6/sol default since 2026-08-03. This is the failure mode to look for whenever
a session is missing from $: check for a slug with no key before anything else.gpt-5.6-codex (Codex) — no key of its own; resolves to the gpt-5.6 fallback
(sol pricing). OpenAI publishes no separate -codex rate, so that is the best
available assumption. The bare gpt-5.6 key exists for exactly this.Prefix collisions? Keys match by longest prefix, so a shorter key must never
shadow a longer one, and a legacy key must never capture a current slug (e.g.
claude-opus-4-1 must NOT match claude-opus-4-8). Adding a bare claude-opus-4
would break exactly this. testPriceTableLegacyKeysPriceWithoutShadowingCurrent
and testPriceTableBundledAndPrefixMatch pin it — run them after any table edit.
Temporary pricing expired? Introductory/promo rates have end dates — but check
whether the end date still stands before "correcting" anything. Claude Sonnet 5's
$2/$10 was announced as introductory through 2026-08-31; Anthropic has since made
it the standard price and cancelled the scheduled rise to $3/$15, so the explicit
claude-sonnet-5 key at $2/$10 is correct and must not be reverted to the
claude-sonnet $3/$15 fallback (verified 2026-08-30 against the pricing page's
claude-sonnet-5-introductory-pricing note). A promo whose date has simply passed is
the easy case; this is the one that costs you, because the stale plan reads as
authoritative.
Edit docs/prices.json, mirror the identical change into bundledJSON, and always
advance updated. Clients only accept a manifest whose updated is >= their
bundled table's, so a forgotten bump means the correction is ignored — that date is the
only thing preventing a stale cache from shadowing corrected prices. Pushing
docs/prices.json corrects already-shipped apps within a day, with no release.
Verify: xcodebuild test -scheme AgentSessions -only-testing:AgentSessionsTests/CodexUsageParserTests
Monthly is enough — provider prices move rarely, but when they move they move a lot
(Opus went $15/$75 → $5/$25, a 3x overstatement that ran undetected). Also check on any
new model launch, since an unknown slug drops that session from $ entirely.
OpenCode's current local backend is SQLite at ~/.local/share/opencode/opencode.db.
Legacy installs may still have a multi-file JSON tree (storage/session/,
storage/message/, storage/part/). Monitoring is SQLite-first and falls back to
the legacy JSON tree when no database is present.
~/.local/share/opencode/opencode.db
# legacy fallback
~/.local/share/opencode/storage/session/<project>/ses_*.json
~/.local/share/opencode/storage/message/<sessionId>/msg_*.json
~/.local/share/opencode/storage/part/<messageId>/*.jsonsession/message/part would
indicate a storage expansion.version, migration, schema_version fields in
session records or new migration files in the OpenCode repo.opencode_storage_latest_session checks db_roots first and fingerprints
session, message, and part rows from opencode.db._opencode_storage_session_tree_schema_fingerprint()
walks the legacy JSON tree for a session and reports keys per record kind.docs/agent-support/agent-watch-config.json still flag release notes mentioning
storage migrations such as SQLite, BoltDB/bbolt, Badger, or database changes.Cursor uses two storage backends:
~/.cursor/projects/<workspace>/agent-transcripts/<uuid>/<uuid>.jsonl) — primary session data, parsed by Agent Sessions. Subagent transcripts live in a subagents/ subdirectory.~/.cursor/chats/<workspace-hash>/<uuid>/store.db) — supplementary metadata (session name, model, timestamps). Key "0" in the meta table contains hex-encoded JSON.The weekly scan fingerprints JSONL transcripts only. The SQLite probe (cursor_sqlite_probe.py) verifies the meta table is readable — it does not deep-fingerprint the database schema.
What to watch for:
role: user/assistant lines.text, tool_use, tool_result, thinking.agent-transcripts/ directory being renamed or moved.meta table schema changes.Note: Some machines may have a stale PATH shim for cursor even when Cursor.app is installed. The weekly monitor tries the PATH command first, then falls back to the embedded app CLI at /Applications/Cursor.app/Contents/Resources/app/bin/cursor --version.
Cursor CLI latest-source truth comes from the official installer script at
https://cursor.com/install, which embeds downloads.cursor.com/lab/<build>/...
agent CLI package URLs. The Homebrew cursor-cli cask page is a fallback. Do
not use the unrelated npm package named cursor-agent.
Cursor Desktop agent windows use the same local surfaces as Cursor CLI:
~/.cursor/projects/*/agent-transcripts/**/*.jsonl for transcript content and
~/.cursor/chats/*/*/store.db for chat metadata. The weekly
cursor_sqlite_probe must keep reporting the newest Desktop chat DB's
agentId, createdAt, mode/model fields, mtime, and meta-key schema so fresh
Desktop-only windows are visible even when their JSONL transcript is absent or
older.
Each agent has a discovery_path_contract in
docs/agent-support/agent-watch-config.json defining the
expected file layout Agent Sessions uses to discover sessions. If an upstream agent moves
or renames its storage, discovery breaks even if the parser still works.
Weekly monitoring checks these contracts. When a contract fails:
severity escalates to high.Key contracts (simplified from regexes in
docs/agent-support/agent-watch-config.json):
| Agent | Expected pattern |
|---|---|
| Codex | */sessions/YYYY/MM/DD/rollout-*.jsonl |
| Claude | ~/.claude/projects/**/*.{jsonl,ndjson} |
| OpenCode | */opencode/opencode.db or legacy */opencode/storage/session/*/ses_*.json |
| Hermes | ~/.hermes/state.db or legacy ~/.hermes/sessions/session_*.json |
| Antigravity | ~/.gemini/antigravity-cli/brain/<conversation-id>/.system_generated/logs/transcript.jsonl |
| Copilot | ~/.copilot/session-state/*.jsonl or */events.jsonl |
| OpenClaw | */agents/<id>/sessions/*.jsonl |
| Cursor | ~/.cursor/projects/*/agent-transcripts/**/*.jsonl |
| Pi | ~/.pi/agent/sessions/**/*.jsonl |
| Kimi | ~/.kimi-code/sessions/wd_*/session_*/agents/main/wire.jsonl |
| Grok | ~/.grok/sessions/<enc-workdir>/<sessionId>/chat_history.jsonl |
| Qwen | */projects/<project>/chats/<id>.jsonl or chats/archive/<id>.jsonl |
| Devin | ~/.local/share/devin/cli/sessions.db |
| fx | ~/.fx/sessions/<id>/checkpoint.json |
required_companion_files — sidecars discovery refuses to work withoutA contract may also declare required_companion_files, a list of paths resolved
relative to the sampled transcript's own directory. Each entry is either a bare
string (existence is enough) or {path, must_parse, note}, where
must_parse: "json_object" additionally requires the file to load as a JSON object.
A breach fails the whole contract, so it lands as severity: high,
verdict: monitoring_broken, and a probe_or_discovery_failed blocker — the same
escalation as a moved store.
Declare one whenever the app's discovery guard chain refuses a session over a file
that is not the transcript. The schema fingerprint structurally cannot cover this: a
sidecar that fails to load simply contributes no keys, and _schema_diff ignores
missing_keys/missing_types on purpose (a thin sample legitimately lacks baseline
types — that is what coverage_ratio is for). Sibling-union sampling then erases even
that trace, because the other four sampled sessions refill the bucket. Verified on a
copy of the real Grok store with one sidecar hidden: unknown_types: [],
unknown_keys: {}, unknown_only_is_empty: true — a total discovery outage reporting
as no drift. Never expect the schema channel to catch a missing companion file.
A Grok session is a directory, not a file: chat_history.jsonl holds the transcript and
the sibling summary.json holds everything discovery depends on (info.id, info.cwd,
current_model_id, chat_format_version). Fingerprinting only the transcript would leave
that half unwatched, so _grok_session_schema_fingerprint() merges the summary in as a
summary bucket — schema diffs on summary/summary.info are summary.json drift, not
transcript drift.
The sidecar is a discovery precondition, not a schema nicety.
GrokSessionDiscovery.discoverSessionFiles() skips any session directory whose
summary.json is absent, so losing that one file removes every Grok session from the
app. Merging it into the fingerprint does not protect it — see §4
required_companion_files, which is what actually flips the verdict. Grok declares it
with must_parse: "json_object", deliberately stricter than the app's own fileExists
guard: a present-but-corrupt sidecar still lists the session but strips its id, cwd,
title and both timestamps. _grok_session_schema_fingerprint() also records a
summary_error (missing / unreadable / invalid_json / not_json_object) so the
fingerprint stops pretending it read a sidecar it could not.
Grok's parser is tolerant of additive fields. GrokSessionParser used to
decode summary.json through a Codable struct, so one field arriving with a new type
threw and try? dropped the whole sidecar — id, cwd, title, model and both timestamps
at once. It reads JSONSerialization dictionaries field by field now, so vendor
additions are ignored keys and a changed field costs exactly that field. Drift is
therefore an alert here (this scan), never a crash — pinned by
GrokSessionParserTests.testUnknownNewFieldsAreIgnored and
testSidecarFieldOfTheWrongTypeCostsOnlyThatField.
Grok is fingerprinted nested (§5). Flat would stop at {type, content, ...} and hide
the content-part types (user.content:text, user.content:image) and
backend_tool_call.kind.action, which is where its format actually moves. arguments is
opaque — it is a tool's parameter object, not Grok format.
The baseline is two session directories: grok/ (top-level) and grok/subagent/.
That split is deliberate — session_kind appears only on subagent sessions and is
absent from top-level ones, so stamping it on the top-level fixture would have taught the
inverted semantics to anything that later classifies on it. Keys that belong to one kind
go in that kind's fixture; the baseline unions both.
Latest version comes from two sources that disagree. Grok ships a native x.ai binary
through the grok-build Homebrew cask, and the cask lags x.ai releases. The
grok_update_check probe declares latest_version_key: "latestVersion", and
_reconcile_latest_version_from_probes() takes the higher of cask and CLI — max, not
replace, because the CLI answers only for its own pinned channel, so a lower CLI answer
must never hide a newer published release. The chosen number's origin is in
upstream.parsed_version_provenance (cli_probe / upstream_source / both_agree /
cached_prior_report), the full comparison in upstream.reconciliation, and any
disagreement prints on the weekly summary line as
latest_disagree=probe:<v>/source:<v>/used:<v>. Never read upstream.parsed_version
without its provenance.
Still unwatched: subagents/<childId>/meta.json — the parent's sidecar, and the only
place parent_session_id and subagent_type appear, so the hierarchy feature depends on
it (the grok/subagent/ fixture is a subagent session, not that sidecar) — and the
compaction/ subtree. No fixture covers either yet. Also unwatched: the app's 50 MB
defaultFullParseMaxBytes ceiling, above which GrokSessionParser silently declines to
parse a transcript, and a wholesale move of the ~/.grok/sessions root, which surfaces
only as local_schema.error: no_files_found plus a stale sample rather than as a named
contract failure (discovery_contract_failed is gated on a file having been found).
Read this before trusting a clean unknown_types=[].
Nested vs flat. _schema_fingerprint_for_agent() in scripts/agent_watch.py is the one
place that decides depth. Codex, Copilot, Claude, Grok, and Qwen use nested fingerprints
(depth 3); Kimi and several storage-backed agents use bespoke fingerprint functions;
remaining JSONL agents are flat. Baseline and observed sample always go through this same
function — fingerprinting one side flat and the other nested diffs two different alphabets.
Lists are unioned across every element (capped at _NESTED_LIST_SAMPLE_LIMIT), not
sampled by their first item. Claude's message.content mixes text/thinking/tool_use blocks,
so first-item-only made later block types invisible — the exact drift the nesting exists to
catch.
Two rules keep nesting from manufacturing drift:
_NESTED_OPAQUE_KEYS — keys whose values are open-ended maps. Codex's
patch_apply_end.changes is keyed by absolute file path; descending into it invents a
bucket per edited file and writes real user paths into report artifacts. Copilot keys
modelMetrics by model id. Add a key here rather than accepting the noise.:type discriminator applies only at the payload wrapper (depth 0→1), where the
real event union lives (event_msg.payload:token_count). Deeper, type tags enum-like
config variants (sandbox_policy:read-only), and splitting on those makes an ordinary
settings change look like schema drift.Claude's opaque keys include input and toolUseResult — both are TOOL-defined payloads,
not Claude format, so walking them would make every new tool read as schema drift. Qwen's
function_args is opaque for the same reason. Read _NESTED_OPAQUE_KEYS for the complete
current list rather than copying it into this skill.
Still flat or bespoke: OpenClaw, Pi, Droid, Hermes, OpenCode, Cursor, Kimi, Devin, and fx do not use the generic nested walker. Their format-specific fingerprint functions may intentionally inspect deeper storage records, but any payload interior they do not expose remains a blind spot. Verify the adapter before assuming either flat or nested coverage.
Multi-session sampling. Weekly fingerprints the newest _LOCAL_SCHEMA_SAMPLE_COUNT (5)
sessions and unions them. This exists because sampling one session let whichever session was
newest decide the verdict — a 4-line "Say hello" prebump session, left in the real store by
--allow-real-home, once flipped antigravity from format_drift_detected to clean with the
drift still sitting in a 92KB session two files back. When required_types is configured,
sibling sampling must honour it (_newest_files_with_types): OpenClaw's **/*.jsonl
glob otherwise sweeps in audit logs and an embedded codex-home and reports codex's event
types as OpenClaw drift.
blocked_thin_sample fires only when a sample is both narrow (<50% of baseline types)
and tiny (<25 events). Coverage alone is not enough: baselines deliberately contain rare
interactive-only families (ai-title, pr-link, permission-mode) that a perfectly healthy
1000-event session will never contain.
Never build a fixture from a handful of recent sessions. Rare families are rare, so a
5-session sample misses them by construction and the weekly then reports "drift" every time
one surfaces — alerts that mean our baseline was incomplete, not upstream changed. On
2026-08-04 that gap was 13 of 24 Claude attachment subtypes and 9 of 18 Codex event_msg
families. Rebuild from every session on disk instead:
./scripts/rebuild_stage0_baseline.py --agent claude # report the gap
./scripts/rebuild_stage0_baseline.py --agent claude --emit # append redacted coverageIt sweeps all discoverable sessions, greedily harvests the fewest real records that close
the gap, and redacts every scalar (only type/role/subtype/model survive, because
those are the schema). Read its report before --emit: a bucket keyed by a UUID, path
or header name is a free-form map that belongs in _NESTED_OPAQUE_KEYS, not in the fixture.
That is how collab_waiting_end.statuses (keyed by thread id) and system.error.headers
(keyed by HTTP header, carrying set-cookie) were caught.
Baseline semantics. _baseline_type_keys_for_agent() excludes *schema_drift* fixtures.
So once a drifted type is verified and handled, it belongs in the normal baseline fixture
— otherwise it re-reports as drift every week forever (Copilot's session.auto_mode_resolved
and session.usage_checkpoint did exactly that from 2026-07-22 until 2026-08-03).
schema_drift.jsonl is for adversarial/speculative shapes only.
Alongside the report fields below, record the value-pass ruling (§1e) for every new
key or type: handled, noise, watch, or candidate. A finding is not closed until it has
one, and candidates are not closed until they are in docs/backlog.md.
From the weekly report (all agents):
results.<agent>.weekly.local_schema.file (newest session).unknown_types, unknown_keys (additive drift),
missing_types, missing_keys (may mean "not observed in this sample").ok status for usage probes.Optional (recommended when a bump is needed):
scripts/agent_captures/<timestamp>/<agent>/.Resources/Fixtures/stage0/agents/<agent>/.
Put verified-and-handled types in the normal fixture, not schema_drift.jsonl — see §5.
Make the fixture a superset of the old key sets; silently dropping keys shrinks the baseline.session_meta payload keys.token_count, rate_limits, billing) when emitted.blocked_stale_sample.docs/agent-support/agent-support-matrix.yml (agents.<key>.max_verified_version)docs/agent-support/agent-support-ledger.ymldocs/agent-json-tracking.md under "Upstream Version Check Log"./scripts/xcode_test_stable.sh./scripts/xcode_test_stable.sh -only-testing:AgentSessionsTests/SessionParserTestsWhen turning a real session into a committed fixture:
© jazzyalex, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/agent-session-format-check of jazzyalex/agent-sessions.
Open the folder on GitHubat commit b789596
Agent Session Format Check next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Session Format Check this skilljazzyalex/agent-sessions | 894 | — | ~13k | Automated safety check: Pass | MIT | |
| Vercel Composition Patternssupabase/supabase | 111k | 58 repos | ~726 | Automated safety check: Pass | MIT | |
| Finishing a Development Branchobra/superpowers | 297k | 5 repos | ~1.9k | Automated safety check: Pass | MIT | |
| Typescript Advanced Typesrolling-scopes/rsschool-app | 10k | 25 repos | ~4.2k | Automated safety check: Pass | MPL-2.0 | |
| PR Babysitteropeninterpreter/openinterpreter | 69k | 3 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Code Review ChecklistshareAI-lab/learn-claude-code | 78k | 5 repos | ~1.1k | Automated safety check: Pass | MIT |
supabase/supabase
React composition patterns that scale. An agent skill from supabase/supabase.
obra/superpowers
Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.
rolling-scopes/rsschool-app
Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.
openinterpreter/openinterpreter
Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.
shareAI-lab/learn-claude-code
Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.
onyx-dot-app/onyx
Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.
jazzyalex/agent-sessions
Capture deterministic macOS screenshots for testing, docs, release notes, and marketing assets.
jazzyalex/agent-sessions
Create and ship AgentSessions support for a new or changed local AI agent/provider.
jazzyalex/agent-sessions
A skill your agent uses when shipping a release of Agent Sessions — bumping version, updating CHANGELOG, building, signing, notarizing, publishing appcast, and creating a GitHub release.
jazzyalex/agent-sessions
Release/deploy workflow for Agent Sessions (Sparkle appcast + GitHub release).
jazzyalex/agent-sessions
A skill your agent uses when writing or curating the user-facing release copy for an Agent Sessions release — README "What's New", GitHub release notes, Sparkle release notes, or website/launch copy.
jazzyalex/agent-sessions
Maintain Agent Sessions agent support matrix and JSON/JSONL parsing compatibility.
Categories
Verify agent session format compatibility for Agent Sessions. Agent Session Format Check is an agent skill from jazzyalex/agent-sessions. Verify agent session format compatibility for Agent Sessions.
Agent Session Format Check fits situations like: any agent CLI updates; monitoring flags drift; bumping max verified versions (fixtures + docs + tests).
Run `npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a claude-code`. Or copy the skill folder (skills/agent-session-format-check in jazzyalex/agent-sessions) into .claude/skills/agent-session-format-check in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a codex`. Or copy the skill folder (skills/agent-session-format-check in jazzyalex/agent-sessions) into .agents/skills/agent-session-format-check in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jazzyalex/agent-sessions --skill agent-session-format-check -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-session-format-check, .gemini/skills/agent-session-format-check, .github/skills/agent-session-format-check and .opencode/skills/agent-session-format-check in your project.
Going by SKILL.md and its folder, Agent Session Format Check needs the command-line tools its instructions call (claude, xcodebuild and cursor) and credentials named OPENAI_API_KEY, ANTHROPIC_API_KEY, FACTORY_API_KEY and GITHUB_TOKEN. Our summary lists: A credential in OPENAI_API_KEY; A credential in ANTHROPIC_API_KEY.
SKILL.md names 3 domains. In commands or code: cursor.com; the agent is likely to contact it when it follows the instructions. As links in the text: platform.claude.com and developers.openai.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agent Session Format Check is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 13k tokens (SKILL.md is roughly 53k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agent Session Format Check: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 297k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jazzyalex (a GitHub user) maintains it in jazzyalex/agent-sessions, which has 894 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 9, 2026.
Source: jazzyalex/agent-sessions on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.