Memori Long-Term Memory
MemoriLabs/Memori
Adds structured long-term memory to OpenClaw agents, built automatically from sessions, with tools the agent calls to recall facts, summaries and decisions.
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-dataset-token-length --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-dataset-token-length .claude/skills/analyze-dataset-token-length && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "analyze-dataset-token-length" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-dataset-token-length into .claude/skills/analyze-dataset-token-length/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-dataset-token-length", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-dataset-token-lengthType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-dataset-token-length --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/analyze-dataset-token-length .agents/skills/analyze-dataset-token-length && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "analyze-dataset-token-length" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-dataset-token-length into .agents/skills/analyze-dataset-token-length/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-dataset-token-length", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-dataset-token-length --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/analyze-dataset-token-length .cursor/skills/analyze-dataset-token-length && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "analyze-dataset-token-length" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-dataset-token-length into .cursor/skills/analyze-dataset-token-length/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-dataset-token-length", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/analyze-dataset-token-length--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-dataset-token-length --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/analyze-dataset-token-length .gemini/skills/analyze-dataset-token-length && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "analyze-dataset-token-length" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-dataset-token-length into .gemini/skills/analyze-dataset-token-length/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-dataset-token-length", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-dataset-token-lengthInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/analyze-dataset-token-length .github/skills/analyze-dataset-token-length && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "analyze-dataset-token-length" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-dataset-token-length into .github/skills/analyze-dataset-token-length/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-dataset-token-length", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-dataset-token-length --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/analyze-dataset-token-length .opencode/skills/analyze-dataset-token-length && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "analyze-dataset-token-length" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-dataset-token-length into .opencode/skills/analyze-dataset-token-length/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-dataset-token-length", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
analyze-dataset-token-lengthAnalyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
Analyze Dataset Token Length is an agent skill from open-thoughts/OpenThoughts-Agent. Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. "taskcomplete AND < 32768 tokens"). Use when asked how long traces are, how many fit a context window (32k/131k), or to filter a trace dataset by length + a field. Uses the OT-Agent analysis tools + the Qwen3-8B tokenizer. Runs LOCALLY on the Mac (no GPU); full-dataset tokenization of ~10k multi-turn traces takes a few…
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Natural language processing and Context engineering. It works with Qwen. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Analyze Dataset Token Length loads about 1.5k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 560 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 560 words, ~1,467 tokens.
.claude/skills/analyze-dataset-token-length/SKILL.md (or your agent's skills folder).OT-Agent trace datasets are conversation-format (ShareGPT-style): each row is
{"conversations": [{"role","content"}, …], + metadata} (some use "messages"; metadata
fields are e.g. task, result, run_id, trial_name, model, agent). "Token length
of a trace" = the tokenized length of the whole conversation.
scripts/analysis/utils.py — canonical pure conversation/token helpers:
extract_conversation_text(record), render_token_representation(...), and
count_conversation_tokens(...). Every count must explicitly select
serialized, conversation_text, or chat_template; these are different
measurements and must never be silently substituted for one another.scripts/analysis/context_length_compare.py — cross-dataset context-length comparison.Always Qwen/Qwen3-8B (AutoTokenizer.from_pretrained("Qwen/Qwen3-8B", trust_remote_code=True)).
Our trace datasets are Qwen3-8B-tokenized even when named for GLM/Kimi/etc. — those "GLM-4.7-…"
models are Qwen3-8B SFTs (see memory reference_glm47_swesmith_is_qwen3_8b); the served model
name in a row's model field (e.g. hosted_vllm/<numeric-id>) is NOT a usable tokenizer name.
tokenizer(extract_conversation_text(row), add_special_tokens=False) — fast;
slightly under-counts vs training (no chat-template tokens).
Right for distribution/relative comparisons.len(tokenizer.apply_chat_template(conv, tokenize=True, add_generation_prompt=False))
— what an SFT trainer actually tokenizes; use when the question is "does it fit a 32k/131k
training window." If a trace's role shape makes the template raise, report it as
uncountable for this representation and tally it separately; do not substitute a
plain-text count.<think> blocks from earlier assistant turns, so on
thinking-mode traces apply_chat_template can count fewer tokens than plain-concat (which keeps
all thinking) — i.e. more traces "fit" under the template. So the "right" count for a < N filter
depends on whether your SFT template preserves thinking: default Qwen3 (strips) → optimistic
count; a thinking-preserving template (qwen3_thinking_acc.jinja2) → conservative count ≈ plain.
Report BOTH and pick by the training template; for a safe "fits 32k" answer use the larger (plain /
thinking-preserving) count.conversations value, including role
fields and JSON punctuation. This is the legacy datagen-counter measurement;
it is useful for continuity but is neither plain text nor training-faithful.Recipe: load_dataset (non-streaming) → per row compute (a) the token count and (b) a metadata
predicate → count the intersection; report each leg separately so it's auditable.
Instruction text leaks into the trace. Fields like task_complete appear verbatim in the user
instruction of EVERY trace (…include "task_complete": true in your response…), so a naive
'"task_complete": true' in full_text matches all rows (false 100%). Scope the predicate to the
agent's actual emission — i.e. an assistant-role message containing the field, not the prompt:
def agent_complete(conv):
return any(m.get("role") == "assistant" and '"task_complete": true' in (m.get("content") or "")
for m in conv)Always sanity-check the predicate VARIES (not all-true / all-false) before trusting a count — print the per-leg breakdown and an early per-1000-row progress line. (Same caution for any tool-name / status substring: confirm you're matching the agent's output, not the system/user scaffolding.)
Local, otagent python, HF token sourced; full-dataset tokenization of ~10k multi-turn traces is a few minutes → background it:
source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python scripts/analysis/<script>.py # run_in_backgroundFirst run downloads + caches the parquet (~hundreds of MB). A benign 'NoneType' has no attribute 'ArrowInvalid' on streaming-generator teardown can be ignored (use non-streaming load_dataset anyway).
DCAgent2/GLM-4.7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k (9437 traces):
detect completion via the assistant-scoped "task_complete": true (not the instruction prose),
count tokens via apply_chat_template (Qwen3-8B), filter complete AND ct < 32768. Reports three
legs — #complete, #<32k, and the intersection — so the filter is auditable. (Early progress
1000/9437: complete=924 / fit32k=886 / both=854 confirmed the predicate varies ≈92%, i.e. the
confound was correctly excluded.)
© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/analyze-dataset-token-length of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Analyze Dataset Token Length next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Analyze Dataset Token Length this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Memori Long-Term MemoryMemoriLabs/Memori | 17k | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Project Developmentguanyang/open-agent-hub | 977 | 2 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Context DoctorjzOcb/context-doctor | 119 | — | ~642 | Automated safety check: Pass | MIT | |
| Cognee Session Memory and Improvetopoteretes/cognee | 32k | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Caveman Learn Token FixesJuliusBrussee/caveman | 111k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 |
MemoriLabs/Memori
Adds structured long-term memory to OpenClaw agents, built automatically from sessions, with tools the agent calls to recall facts, summaries and decisions.
guanyang/open-agent-hub
This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand, the shape of a multi-stage batch or agent pipeline, token…
jzOcb/context-doctor
Visualize and diagnose OpenClaw context window usage. An agent skill from jzOcb/context-doctor.
topoteretes/cognee
Explains how cognee stores session memory by session_id and bridges it into the permanent graph with improve(), including the stages, results and settings.
JuliusBrussee/caveman
Acts on a Caveman learn report: reviews ranked token sinks, applies cost-lowering edits one at a time with your consent, and reports what each fix returned.
MadAppGang/claudish
CRITICAL - Guide for using Claudish CLI ONLY through sub-agents to run Claude Code with any AI model (OpenRouter, Gemini, OpenAI, local models).
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
open-thoughts/OpenThoughts-Agent
Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.
Works with
Categories
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. Analyze Dataset Token Length is an agent skill from open-thoughts/OpenThoughts-Agent.g.
Analyze Dataset Token Length fits situations like: asked how long traces are; how many fit a context window (32k/131k); filter a trace dataset by length + a field.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a claude-code`. Or copy the skill folder (.agents/skills/analyze-dataset-token-length in open-thoughts/OpenThoughts-Agent) into .claude/skills/analyze-dataset-token-length in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a codex`. Or copy the skill folder (.agents/skills/analyze-dataset-token-length in open-thoughts/OpenThoughts-Agent) into .agents/skills/analyze-dataset-token-length in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-dataset-token-length, .gemini/skills/analyze-dataset-token-length, .github/skills/analyze-dataset-token-length and .opencode/skills/analyze-dataset-token-length in your project.
Going by SKILL.md and its folder, Analyze Dataset Token Length needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Analyze Dataset Token Length is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Analyze Dataset Token Length: Memori Long-Term Memory (MemoriLabs/Memori, 17k stars), Project Development (guanyang/open-agent-hub, 977 stars), Context Doctor (jzOcb/context-doctor, 119 stars) and Cognee Session Memory and Improve (topoteretes/cognee, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.