Agent skill

Analyze Dataset Token Length

by open-thoughts in open-thoughts/OpenThoughts-Agent

Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

Apache-2.0Auto-check passedAgent Workflows

Install Analyze Dataset Token Length

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-dataset-token-length --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-dataset-token-length .claude/skills/analyze-dataset-token-length && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-dataset-token-length
GitHub stars
301
Token cost
~1.5k tokens
SKILL.md length
560 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

  • Asked how long traces are
  • SKILL.md covers The canonical OT-Agent tools…, Tokenizer convention, Three token-count… and Threshold + metadata…, plus 2 more sections
  • Calls python
  • How many fit a context window (32k/131k)

What it does

Analyze Dataset Token Length is an agent skill from open-thoughts/OpenThoughts-Agent. Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. "taskcomplete AND < 32768 tokens"). Use when asked how long traces are, how many fit a context window (32k/131k), or to filter a trace dataset by length + a field. Uses the OT-Agent analysis tools + the Qwen3-8B tokenizer. Runs LOCALLY on the Mac (no GPU); full-dataset tokenization of ~10k multi-turn traces takes a few…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Natural language processing and Context engineering. It works with Qwen. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Asked how long traces are
  • How many fit a context window (32k/131k)
  • Filter a trace dataset by length + a field

Example prompts

  • “taskcomplete AND < 32768 tokens”
  • “/analyze-dataset-token-length”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Dataset Token Length loads about 1.5k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 560 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 560 words, ~1,467 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-dataset-token-length/SKILL.md (or your agent's skills folder).
name
analyze-dataset-token-length
description
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. "task_complete AND < 32768 tokens"). Use when asked how long traces are, how many fit a context window (32k/131k), or to filter a trace dataset by length + a field. Uses the OT-Agent analysis tools + the Qwen3-8B tokenizer. Runs LOCALLY on the Mac (no GPU); full-dataset tokenization of ~10k multi-turn traces takes a few minutes → run it in the background.

analyze-dataset-token-length

OT-Agent trace datasets are conversation-format (ShareGPT-style): each row is {"conversations": [{"role","content"}, …], + metadata} (some use "messages"; metadata fields are e.g. task, result, run_id, trial_name, model, agent). "Token length of a trace" = the tokenized length of the whole conversation.

The canonical OT-Agent tools (don't reinvent)

  • scripts/analysis/utils.py — canonical pure conversation/token helpers: extract_conversation_text(record), render_token_representation(...), and count_conversation_tokens(...). Every count must explicitly select serialized, conversation_text, or chat_template; these are different measurements and must never be silently substituted for one another.
  • scripts/analysis/context_length_compare.py — cross-dataset context-length comparison.

Tokenizer convention

Always Qwen/Qwen3-8B (AutoTokenizer.from_pretrained("Qwen/Qwen3-8B", trust_remote_code=True)). Our trace datasets are Qwen3-8B-tokenized even when named for GLM/Kimi/etc. — those "GLM-4.7-…" models are Qwen3-8B SFTs (see memory reference_glm47_swesmith_is_qwen3_8b); the served model name in a row's model field (e.g. hosted_vllm/<numeric-id>) is NOT a usable tokenizer name.

Three token-count representations — pick by the question

  • conversation_text = tokenizer(extract_conversation_text(row), add_special_tokens=False) — fast; slightly under-counts vs training (no chat-template tokens). Right for distribution/relative comparisons.
  • chat_template = len(tokenizer.apply_chat_template(conv, tokenize=True, add_generation_prompt=False)) — what an SFT trainer actually tokenizes; use when the question is "does it fit a 32k/131k training window." If a trace's role shape makes the template raise, report it as uncountable for this representation and tally it separately; do not substitute a plain-text count.
  • ⚠️ The two can differ by MORE than the wrapper tokens — and in the surprising direction. Qwen3's chat template strips historical <think> blocks from earlier assistant turns, so on thinking-mode traces apply_chat_template can count fewer tokens than plain-concat (which keeps all thinking) — i.e. more traces "fit" under the template. So the "right" count for a < N filter depends on whether your SFT template preserves thinking: default Qwen3 (strips) → optimistic count; a thinking-preserving template (qwen3_thinking_acc.jinja2) → conservative count ≈ plain. Report BOTH and pick by the training template; for a safe "fits 32k" answer use the larger (plain / thinking-preserving) count.
  • serialized = compact JSON of the raw conversations value, including role fields and JSON punctuation. This is the legacy datagen-counter measurement; it is useful for continuity but is neither plain text nor training-faithful.

Threshold + metadata filter-count (the common ask)

Recipe: load_dataset (non-streaming) → per row compute (a) the token count and (b) a metadata predicate → count the intersection; report each leg separately so it's auditable.

Show full SKILL.md (210 more words)Show less
⚠️ The metadata-confound trap (read this before any field predicate)

Instruction text leaks into the trace. Fields like task_complete appear verbatim in the user instruction of EVERY trace (…include "task_complete": true in your response…), so a naive '"task_complete": true' in full_text matches all rows (false 100%). Scope the predicate to the agent's actual emission — i.e. an assistant-role message containing the field, not the prompt:

python
def agent_complete(conv):
    return any(m.get("role") == "assistant" and '"task_complete": true' in (m.get("content") or "")
               for m in conv)

Always sanity-check the predicate VARIES (not all-true / all-false) before trusting a count — print the per-leg breakdown and an early per-1000-row progress line. (Same caution for any tool-name / status substring: confirm you're matching the agent's output, not the system/user scaffolding.)

How to run

Local, otagent python, HF token sourced; full-dataset tokenization of ~10k multi-turn traces is a few minutes → background it:

bash
source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python scripts/analysis/<script>.py   # run_in_background

First run downloads + caches the parquet (~hundreds of MB). A benign 'NoneType' has no attribute 'ArrowInvalid' on streaming-generator teardown can be ignored (use non-streaming load_dataset anyway).

Worked example — swesmith, "task_complete AND < 32768 tok"

DCAgent2/GLM-4.7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k (9437 traces): detect completion via the assistant-scoped "task_complete": true (not the instruction prose), count tokens via apply_chat_template (Qwen3-8B), filter complete AND ct < 32768. Reports three legs — #complete, #<32k, and the intersection — so the filter is auditable. (Early progress 1000/9437: complete=924 / fit32k=886 / both=854 confirmed the predicate varies ≈92%, i.e. the confound was correctly excluded.)

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/analyze-dataset-token-length of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Analyze Dataset Token Length next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Dataset Token Length compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Dataset Token Length this skillopen-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.0
Memori Long-Term MemoryMemoriLabs/Memori17k—~2kAutomated safety check: PassApache-2.0
Project Developmentguanyang/open-agent-hub9772 repos~4.7kAutomated safety check: PassMIT
Context DoctorjzOcb/context-doctor119—~642Automated safety check: PassMIT
Cognee Session Memory and Improvetopoteretes/cognee32k—~3.2kAutomated safety check: PassApache-2.0
Caveman Learn Token FixesJuliusBrussee/caveman111k—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • Memori Long-Term Memory

    MemoriLabs/Memori

    Adds structured long-term memory to OpenClaw agents, built automatically from sessions, with tools the agent calls to recall facts, summaries and decisions.

    17k GitHub stars~2k tokensUpdated 6 days ago
    Agent WorkflowsAuto-check passed
  • Project Development

    guanyang/open-agent-hub

    This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand, the shape of a multi-stage batch or agent pipeline, token…

    977 GitHub starsUsed in 2 repos~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Context Doctor

    jzOcb/context-doctor

    Visualize and diagnose OpenClaw context window usage. An agent skill from jzOcb/context-doctor.

    119 GitHub stars~642 tokensUpdated 6 mo ago
    Agent WorkflowsAuto-check passed
  • Explains how cognee stores session memory by session_id and bridges it into the permanent graph with improve(), including the stages, results and settings.

    32k GitHub stars~3.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Caveman Learn Token Fixes

    JuliusBrussee/caveman

    Acts on a Caveman learn report: reviews ranked token sinks, applies cost-lowering edits one at a time with your consent, and reports what each fix returned.

    111k GitHub stars~2.8k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Claudish Usage

    MadAppGang/claudish

    CRITICAL - Guide for using Claudish CLI ONLY through sub-agents to run Claude Code with any AI model (OpenRouter, Gemini, OpenAI, local models).

    1k GitHub stars~9k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 10 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed
  • Commit

    open-thoughts/OpenThoughts-Agent

    Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.

    301 GitHub stars~2.2k tokensUpdated 10 days ago
    Auto-check: notes

Works with

Questions about Analyze Dataset Token Length

What does Analyze Dataset Token Length do?

Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. Analyze Dataset Token Length is an agent skill from open-thoughts/OpenThoughts-Agent.g.

When should I use Analyze Dataset Token Length?

Analyze Dataset Token Length fits situations like: asked how long traces are; how many fit a context window (32k/131k); filter a trace dataset by length + a field.

How do I install Analyze Dataset Token Length in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a claude-code`. Or copy the skill folder (.agents/skills/analyze-dataset-token-length in open-thoughts/OpenThoughts-Agent) into .claude/skills/analyze-dataset-token-length in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Dataset Token Length in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a codex`. Or copy the skill folder (.agents/skills/analyze-dataset-token-length in open-thoughts/OpenThoughts-Agent) into .agents/skills/analyze-dataset-token-length in your project. Codex loads it when a task matches its description.

Can I use Analyze Dataset Token Length in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-dataset-token-length -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-dataset-token-length, .gemini/skills/analyze-dataset-token-length, .github/skills/analyze-dataset-token-length and .opencode/skills/analyze-dataset-token-length in your project.

What does Analyze Dataset Token Length need to run?

Going by SKILL.md and its folder, Analyze Dataset Token Length needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Analyze Dataset Token Length access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Analyze Dataset Token Length safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyze Dataset Token Length use?

Analyze Dataset Token Length is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Dataset Token Length use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyze Dataset Token Length?

Skills that share tags, products or a category with Analyze Dataset Token Length: Memori Long-Term Memory (MemoriLabs/Memori, 17k stars), Project Development (guanyang/open-agent-hub, 977 stars), Context Doctor (jzOcb/context-doctor, 119 stars) and Cognee Session Memory and Improve (topoteretes/cognee, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Dataset Token Length?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.