SwanLab Experiment Tracking
Orchestra-Research/AI-Research-SKILLs
Shows how to log ML runs, configs, metrics and media with SwanLab and view them in cloud, local or self-hosted dashboards.
A skill your agent uses when planning, launching, tracking, or preserving experiments across projects, especially to separate deterministic operating rules from fuzzy experiment-defining variables…
$ npx skills add tokenbender/agent-guides --skill tokenbending -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install tokenbender/agent-guides tokenbending --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude-skills/tokenbending .claude/skills/tokenbending && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tokenbending" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/tokenbending into .claude/skills/tokenbending/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tokenbending", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/tokenbender/agent-guides/tree/main/claude-skills/tokenbendingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add tokenbender/agent-guides --skill tokenbending -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install tokenbender/agent-guides tokenbending --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .agents/skills && cp -r skills-src/claude-skills/tokenbending .agents/skills/tokenbending && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tokenbending" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/tokenbending into .agents/skills/tokenbending/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tokenbending", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tokenbender/agent-guides --skill tokenbending -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install tokenbender/agent-guides tokenbending --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/claude-skills/tokenbending .cursor/skills/tokenbending && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tokenbending" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/tokenbending into .cursor/skills/tokenbending/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tokenbending", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/tokenbender/agent-guides.git --path claude-skills/tokenbending--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add tokenbender/agent-guides --skill tokenbending -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install tokenbender/agent-guides tokenbending --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/claude-skills/tokenbending .gemini/skills/tokenbending && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tokenbending" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/tokenbending into .gemini/skills/tokenbending/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tokenbending", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install tokenbender/agent-guides tokenbendingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add tokenbender/agent-guides --skill tokenbending -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .github/skills && cp -r skills-src/claude-skills/tokenbending .github/skills/tokenbending && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tokenbending" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/tokenbending into .github/skills/tokenbending/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tokenbending", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tokenbender/agent-guides --skill tokenbending -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install tokenbender/agent-guides tokenbending --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/claude-skills/tokenbending .opencode/skills/tokenbending && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tokenbending" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/tokenbending into .opencode/skills/tokenbending/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tokenbending", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tokenbendingA skill your agent uses when planning, launching, tracking, or preserving experiments across projects, especially to separate deterministic operating rules from fuzzy experiment-defining variables…
Tokenbending is an agent skill from tokenbender/agent-guides. Use when planning, launching, tracking, or preserving experiments across projects, especially to separate deterministic operating rules from fuzzy experiment-defining variables that require clarification before costly or irreversible work, and to ensure any residual compatibility, migration, cleanup, or risk is tracked rather than left hanging.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: one page guides that i let my subscribed/customised agents consume to perform actions. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit a74dd9d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tokenbending loads about 1.7k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 959 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from tokenbender/agent-guides at commit a74dd9d, republished under its Apache-2.0 licence (© tokenbender). 959 words, ~1,713 tokens.
.claude/skills/tokenbending/SKILL.md (or your agent's skills folder).Yes. Examples would make it much more usable because they turn the rule from taste into a decision boundary.
Deterministic Rules With Examples
Use the project's canonical tracking surface as the experiment log.
Example: If the project uses GitHub issues, update the issue when a run starts, when a stage completes, when a failure occurs, and when artifacts are uploaded.
Example: If the project uses a lab notebook or runs/README.md, append the command, config, result summary, and artifact path there.
Keep changes atomic and reviewable. Example: Commit "add experiment spec" separately from "fix distributed launcher." Example: Do not mix README cleanup, runner code, and generated data into one commit.
Verify before reporting completion. Example: Do not say "uploaded" until the destination lists the expected files. Example: Do not say "training is distributed" until process/GPU state or logs confirm multiple workers.
Do not leave hanging work implicit. Example: If a temporary compatibility endpoint, migration alias, fallback script, manual workaround, unresolved verification, or cleanup step remains, create or link a follow-up issue before calling the work done. Example: Do not write "follow-up: none" when any known duct tape, residual risk, or future cleanup still exists; make the remaining work durable in the project's tracking surface.
Preserve reproducible state before deleting compute resources. Example: Upload configs, logs, checkpoints, summaries, and command history before terminating a pod. Example: If upload fails, keep the machine alive or create a smaller fallback bundle before cleanup.
Do not upload secrets, tokens, caches, or third-party base assets unless explicitly authorized. Example: Upload adapter weights and run logs, but exclude API keys and model cache directories. Example: Preserve a manifest saying which external base model must be re-downloaded instead of copying the full base model.
Clean up costly resources after preservation is verified. Example: Delete an experiment VM after artifact upload is listed and checksummed. Example: Leave unrelated shared infrastructure alone unless the user asked to clean it too.
If the user says stop at first failure, stop at the first substantive failure. Example: A syntax check failure means stop, report it, and do not launch the long run. Example: If stage 2 crashes after stage 1 succeeds, preserve stage 1 outputs and do not proceed to stage 3.
Honor exact named constraints. Example: If the user specifies a particular GPU type, verify the actual GPU model before running. Example: If the user specifies a held-out dataset, do not silently swap in a convenient alternative.
Prefer resumability. Example: Save the exact command line, config, environment notes, logs, and output manifest. Example: Make the artifact bundle sufficient for a new machine to continue from the last valid checkpoint.
Gate paid or long accelerator runs with a full-load profile smoke. Example: Before launching full training on paid GPU/TPU compute, sweep the viable batch, gradient accumulation, packing, compile, checkpointing, and logging settings on a realistic smoke run; record step time, tokens/sec, memory, estimated MFU, and the chosen profile. Example: Do not spend the full run on a low-utilization profile unless the user explicitly accepts the efficiency tradeoff or the controlled experiment forbids changing the profile.
Fuzzy Clarification Rules With Examples
Clarify the exact research question when ambiguous. Example: "Are we testing whether the smaller subsystem alone performs the task, or whether the full system works after replacing that subsystem?" Example: "Is the goal best absolute score, smallest viable subsystem, or a clean comparison to a prior method?"
Clarify fixed variables versus allowed variables. Example: "Should dataset, model, metric, and evaluation mode stay fixed while only the training method changes?" Example: "Can I change batch size and launcher details for stability, or are those part of the controlled experiment?"
Clarify the comparison baseline. Example: "Which prior result is the baseline: best known score, most recent run, or same-budget run?" Example: "Should this compare against the unmodified base, an earlier trained variant, or a manually selected region?"
Clarify the success metric. Example: "Is success judged by accuracy, judge score, recovery percentage, latency, cost, or Pareto frontier?" Example: "If two candidates tie on score, should smaller size or lower training cost win?"
Clarify the evaluation mode. Example: "Should we evaluate the modified component inside the full system or isolated from the rest?" Example: "Should non-selected components be zeroed, mean-ablated, patched from a baseline, or left untouched?"
Clarify data roles. Example: "Which split is training, which is calibration/attribution, and which is held out?" Example: "Can seed examples overlap with training data, or must they be separated?"
Clarify mandatory versus discretionary methods. Example: "Is this method required, or can I choose a stronger equivalent?" Example: "Should I reproduce the previous method exactly before adding the new variant?"
Clarify failure policy. Example: "If training crashes, should I stop and preserve state, retry with the same spec, or make an engineering fix and continue?" Example: "If the judge service is down, should I pause or fall back to a secondary metric?"
Clarify run grade. Example: "Is this a quick probe where partial data is acceptable, or a benchmark-grade run requiring full provenance?" Example: "Should I optimize for fastest signal or publishable reproducibility?"
Clarify artifact policy. Example: "Should checkpoints be uploaded, or only logs and summaries?" Example: "Should failed-run artifacts be preserved, deleted, or marked as partial?"
Clarify residual-work policy. Example: "If we keep a compatibility shim or temporary fallback after this change, should I open a follow-up issue now or complete the cleanup before closeout?" Example: "Can this risk remain as documented debt, or is the acceptance bar that all migration/cleanup work is complete before the issue closes?"
Clarify resource authority. Example: "May I create new cloud compute, or must I reuse existing machines?" Example: "May I terminate only resources created for this run, or also older matching resources?"
That would make the guidance both general and enforceable.
© tokenbender, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in claude-skills/tokenbending of tokenbender/agent-guides.
Open the folder on GitHubat commit a74dd9d
Tokenbending next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tokenbending this skilltokenbender/agent-guides | 367 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| SwanLab Experiment TrackingOrchestra-Research/AI-Research-SKILLs | 13k | 1 repos | ~2.4k | Automated safety check: Pass | MIT | |
| Setting Up Experiment Trackingjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~954 | Automated safety check: Pass | MIT | |
| Tracking Token Launchesjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Shipping and Launch Checklistaddyosmani/agent-skills | 103k | 1 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Experiment Tracking Setuprevfactory/harness-100 | 1.3k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Shows how to log ML runs, configs, metrics and media with SwanLab and view them in cloud, local or self-hosted dashboards.
jeremylongshore/tons-of-skills-marketplace
Implement machine learning experiment tracking using MLflow or Weights & Biases.
jeremylongshore/tons-of-skills-marketplace
Track new token launches across DEXes with risk analysis and contract verification.
addyosmani/agent-skills
Prepares a production launch with a pre-launch checklist, monitoring, a staged rollout and a rollback plan so every release is reversible and observable.
revfactory/harness-100
Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology.
affaan-m/ECC
Track and report Claude Code token usage, spending, and budgets from the local ECC cost-tracker metrics log.
tokenbender/agent-guides
Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.
tokenbender/agent-guides
A skill your agent uses for planning, researching, drafting, revising, or auditing technical write-ups, textbooks, papers, reports, READMEs, research notes, PR narratives, and public technical prose.
tokenbender/agent-guides
Filter, compare, and rank papers, posts, captures, threads, bookmarks, product claims, or research ideas for high-entropy mechanistic insight and underpriced leverage.
tokenbender/agent-guides
Trigger when: (1) the user asks for Manim, Manim Community, or ManimCE, (2) code contains from manim import , or (3) the task is to build a mathematical explainer animation.
tokenbender/agent-guides
Issue-led atomic work logging. An agent skill from tokenbender/agent-guides.
tokenbender/agent-guides
A skill your agent uses when the user provides an X/Twitter status URL and needs the full thread, context beyond the first post, comparison, summary, intent analysis, title extraction, or reliable…
A skill your agent uses when planning, launching, tracking, or preserving experiments across projects, especially to separate deterministic operating rules from fuzzy experiment-defining variables…. Tokenbending is an agent skill from tokenbender/agent-guides. Use when planning, launching, tracking, or preserving experiments across projects, especially to separate deterministic operating rules from fuzzy experiment-defining variables that require clarification before costly or irreversible work, and to ensure any residual compatibility, migration, cleanup, or risk is tracked rather than left hanging.
Tokenbending fits situations like: preserving experiments across projects; especially to separate deterministic operating rules from fuzzy experiment-defining variables that require clarification before costly; irreversible work; to ensure any residual compatibility.
Run `npx skills add tokenbender/agent-guides --skill tokenbending -a claude-code`. Or copy the skill folder (claude-skills/tokenbending in tokenbender/agent-guides) into .claude/skills/tokenbending in your project. Claude Code loads it when a task matches its description.
Run `npx skills add tokenbender/agent-guides --skill tokenbending -a codex`. Or copy the skill folder (claude-skills/tokenbending in tokenbender/agent-guides) into .agents/skills/tokenbending in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tokenbender/agent-guides --skill tokenbending -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tokenbending, .gemini/skills/tokenbending, .github/skills/tokenbending and .opencode/skills/tokenbending in your project.
SKILL.md names no scripts, command-line tools or credentials: Tokenbending is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tokenbending is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Tokenbending: SwanLab Experiment Tracking (Orchestra-Research/AI-Research-SKILLs, 13k stars), Setting Up Experiment Tracking (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Tracking Token Launches (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Shipping and Launch Checklist (addyosmani/agent-skills, 103k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
tokenbender (a GitHub user) maintains it in tokenbender/agent-guides, which has 367 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on July 23, 2026.
Source: tokenbender/agent-guides on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.