Agent skill

Deep Trajectory Analysis

by swyxio in swyxio/skills

Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

MITAuto-check passedBusiness, Finance & HR

Install Deep Trajectory Analysis

skills CLI
$ npx skills add swyxio/skills --skill deep-trajectory-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills deep-trajectory-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/deep-trajectory-analysis .claude/skills/deep-trajectory-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deep-trajectory-analysis
GitHub stars
175
Token cost
~1.8k tokens
SKILL.md length
874 words
Files
4 (incl. scripts, references)
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

  • Works in 6 steps: Freeze engine/model, rules,… → Pair baseline and candidate on the exact… → Reproduce sampled trajectories and… → …
  • Move-history audits
  • SKILL.md covers Establish the causal question, Preflight integrity, Audit telemetry semantics and Reconstruct paired trajectories, plus 5 more sections
  • Runs JavaScript scripts from its folder

What it does

Deep Trajectory Analysis is an agent skill from swyxio/skills. Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects. Use for move-history audits, replay/trace analysis, policy regressions, behavior calibration, causal first-divergence studies, reply-survival analysis, agent personality validation, promotion decisions, and reports that must connect aggregate outcomes to exact state-action sequences.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `agents/openai.yaml` and `references/evidence-contract.md`).

It sits in Business, Finance & HR, covering Performance reviews. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • Move-history audits
  • Replay/trace analysis
  • Policy regressions
  • Behavior calibration

Example prompts

  • “/deep-trajectory-analysis”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Freeze engine/model, rules, maps/environments, search budget, opponent/control policies, seeds,
  2. Pair baseline and candidate on the exact physical/random inputs. Exclude unmatched cells.
  3. Reproduce sampled trajectories and verify replay, state, or trace fingerprints.
  4. Separate natural finishes, adjudications, horizon caps, failures, and unexplained outcomes.
  5. Check symmetry, label invariance, hidden-information access, legal actions, and compute parity.
  6. Archive incompatible epochs rather than merging them into one trend.

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deep Trajectory Analysis loads about 1.8k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 114 tokens; SKILL.md has 874 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 874 words, ~1,816 tokens.

Download SKILL.mdSave it as .claude/skills/deep-trajectory-analysis/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
deep-trajectory-analysis
description
Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects. Use for move-history audits, replay/trace analysis, policy regressions, behavior calibration, causal first-divergence studies, reply-survival analysis, agent personality validation, promotion decisions, and reports that must connect aggregate outcomes to exact state-action sequences.

Deep Trajectory Analysis

Determine not merely whether a candidate scored better, but what it changed, whether the named mechanism fired, how the environment replied, and whether the benefit survived to the final outcome.

Read references/evidence-contract.md before defining metrics or building the report. Run scripts/validate-report.mjs on a machine-readable result that follows that contract.

Establish the causal question

Write one sentence for each intended effect before inspecting favorable examples:

When opportunity O is publicly observable, policy P should change action A, immediately alter mechanism M, survive response R, and improve horizon outcome H without violating safety S.

Operationalize every noun. Prefer board-, state-, or trace-derived quantities over labels such as “aggressive,” “safe,” “precise,” or “wild.”

Separate:

  • strength: final win, reward, score margin, rank, or task success;
  • identity: the behavior the policy is meant to exhibit;
  • safety: compute, fairness, legality, information access, termination, and worst-case floors.

Do not let one metric stand in for all three.

Preflight integrity

Stop promotion analysis until these checks pass:

  1. Freeze engine/model, rules, maps/environments, search budget, opponent/control policies, seeds, seats/roles, starter/order, horizon, and terminal scoring.
  2. Pair baseline and candidate on the exact physical/random inputs. Exclude unmatched cells.
  3. Reproduce sampled trajectories and verify replay, state, or trace fingerprints.
  4. Separate natural finishes, adjudications, horizon caps, failures, and unexplained outcomes.
  5. Check symmetry, label invariance, hidden-information access, legal actions, and compute parity.
  6. Archive incompatible epochs rather than merging them into one trend.

If compact results omit actions or states, rerun the exact manifests with decision evidence enabled. Do not infer a causal move story from final aggregates.

Audit telemetry semantics

Trace each metric to the code path that emits it.

  • Distinguish policy exposure, selector activation, and action divergence.
  • A replacement evaluator may be active on every decision while reporting zero selector changes.
  • A selector may activate without producing the named mechanism.
  • A signal may double-count overlapping concepts, especially in two-player or single-opponent cases.
  • A physically retained object or cell may have negative strategic value after opportunity cost.

Rename misleading report labels immediately. Preserve the original field for compatibility if needed, but never use an inapplicable field as the research conclusion.

Reconstruct paired trajectories

For every exact pair:

  1. Align baseline and candidate until the first focal decision with identical public pre-state.
  2. Find the first different action. This is the strongest causal comparison.
  3. Record the shared state, hand/input, legal choices, chosen action, predicted receipt, actual transition, opponent/environment reply, next focal action, and final outcome.
  4. Treat later branch differences as consequences. Do not describe them as independent causal policy choices unless their pre-states are again identical.
  5. Measure both the action’s targeted region and the opponent’s best alternative payoff elsewhere.

Store exact pair keys, turn/step numbers, fingerprints, placements/actions, changed entities, score/reward checkpoints, and terminal evidence.

Build the effect funnel

Count distinct stages rather than compressing them into “activation”:

  1. opportunities;
  2. policy exposures;
  3. selector activations, when applicable;
  4. actual action divergences;
  5. intended immediate effect;
  6. physical survival after the reply;
  7. positive full response-cycle exchange;
  8. medium-horizon conversion;
  9. final outcome conversion.

Report the denominator at every stage. Measure both:

  • physical survival: targeted cells, objects, resources, or state remain;
  • value survival: backed-up score/reward after the best observed or searched reply.

Use response-cycle value for promotion. Treat physical survival and immediate gain as diagnostic mechanism metrics only.

Show full SKILL.md (321 more words)Show less

Combine population and case evidence

Use both layers:

  • Population layer: all paired trajectories, clustered uncertainty, map/environment and opponent/task slices, distribution tails, compute, termination, and funnel rates.
  • Case layer: at least one intended-effect success, one counterexample, and one held-out regression. Select cases by predeclared rules such as largest paired outcome changes—not by visual appeal.

Case studies explain mechanisms; they do not estimate prevalence. Aggregate statistics estimate prevalence; they do not explain the move.

Make the report visual

Always produce work-in-progress charts once real data exists. Prefer:

  1. a population effect chart with experimental-unit dots and uncertainty;
  2. a stage funnel from exposure through final conversion;
  3. paired baseline/candidate score or reward trajectories with the first divergence annotated;
  4. a four-frame state filmstrip: shared state, baseline action, candidate action, reply;
  5. small multiples for map/environment, opponent/task, role/seat, and held-out splits.

Render real state and actions where possible. Use schematic illustrations only when clearly marked. Keep local proxies and final success visually separate; a locally successful move that loses must look like a failed final conversion.

Draw conclusions

Classify each intended effect:

  • fires and converts;
  • fires but is erased by reply;
  • fires but carries excessive opportunity cost;
  • proxy fires, intended mechanism does not;
  • policy is exposed but does not change actions;
  • improves discovery but fails held-out transfer;
  • unmeasurable because instrumentation is semantically wrong.

Recommend the smallest next mechanism that directly represents the failed stage. Examples include reply backup, route-hinge detection, multi-purpose payoff, plan-level variety, calibrated horizon value, or corrected telemetry. Do not solve a missing strategic model by merely increasing a generic weight.

Deliverables

Produce:

  • a concise verdict with quantified findings;
  • a machine-readable report satisfying the evidence contract;
  • a reproducible reconstruction/analyzer command;
  • visual population and trajectory analysis;
  • links to exact replay/trace artifacts;
  • verification results and explicit remaining limitations.

Do not promote or modify a live policy unless the user asked for implementation and the predeclared held-out strength, identity, safety, and compute gates all pass.

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in deep-trajectory-analysis of swyxio/skills.

  • SKILL.md
  • agents/openai.yaml
  • references/evidence-contract.md
  • scripts/validate-report.mjs

Open the folder on GitHubat commit 038ef34

Compare with similar skills

Deep Trajectory Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deep Trajectory Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deep Trajectory Analysis this skillswyxio/skills175—~1.8kAutomated safety check: PassMIT
Wp Performance Reviewelvismdev/claude-wordpress-skills2351 repos~4.5kAutomated safety check: PassMIT
Align Humanagentscope-ai/OpenJudge868—~3.1kAutomated safety check: PassApache-2.0
Performance ReportAffitor/affiliate-skills6991 repos~2.5kAutomated safety check: PassMIT
Run Mv Hoi Reconstructionnvidia-isaac/video_to_data850—~1.5kAutomated safety check: PassCustom licence
Company Analysiszhu1090093659/dsh-trading231—~4.2kAutomated safety check: PassCustom licence

Similar skills

  • Wp Performance Review

    elvismdev/claude-wordpress-skills

    WordPress performance code review and optimization analysis.

    235 GitHub starsUsed in 1 repo~4.5k tokens
    Business, Finance & HRAuto-check passed
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    868 GitHub stars~3.1k tokensUpdated 27 days ago
    Business, Finance & HRAuto-check passed
  • Performance Report

    Affitor/affiliate-skills

    Generate affiliate performance reports with KPIs and recommendations.

    699 GitHub starsUsed in 1 repo~2.5k tokens
    Business, Finance & HRAuto-check passed
  • Run Mv Hoi Reconstruction

    nvidia-isaac/video_to_data

    Run and validate the repository-local multi-view camera calibration and human-object reconstruction pipelines.

    850 GitHub stars~1.5k tokensUpdated yesterday
    Business, Finance & HRAuto-check passed
  • Company Analysis

    zhu1090093659/dsh-trading

    A skill your agent uses when the user wants to analyze a listed company, stock, business, or investment target; challenge or revise an existing company report; compare A/H or primary-listing/ADR…

    231 GitHub stars~4.2k tokensUpdated 5 days ago
    Business, Finance & HRAuto-check passed
  • Windbg Diagnostic Method

    microsoft/win-dev-skills

    Official

    Use with every WinDbg plugin investigation to apply evidence-first reasoning, confidence calibration, contrarian review, structured reporting, and deterministic validation.

    462 GitHub stars~1.9k tokensUpdated yesterday
    Business, Finance & HRAuto-check passed

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    175 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    175 GitHub stars~1.1k tokensUpdated 4 days ago
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    175 GitHub stars~4.3k tokensUpdated 4 days ago
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    175 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    175 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check: warnings
  • Forge

    swyxio/skills

    Operate or diagnose SmolForge repositories and Forge Deploy/Sites when the task requires Forge-specific CLI, authentication, manifest, or release behavior on forge.smol.ai or .sites.smol.ai.

    175 GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed

Questions about Deep Trajectory Analysis

What does Deep Trajectory Analysis do?

Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects. Deep Trajectory Analysis is an agent skill from swyxio/skills. Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

When should I use Deep Trajectory Analysis?

Deep Trajectory Analysis fits situations like: move-history audits; replay/trace analysis; policy regressions; behavior calibration.

How do I install Deep Trajectory Analysis in Claude Code?

Run `npx skills add swyxio/skills --skill deep-trajectory-analysis -a claude-code`. Or copy the skill folder (deep-trajectory-analysis in swyxio/skills) into .claude/skills/deep-trajectory-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Deep Trajectory Analysis in Codex?

Run `npx skills add swyxio/skills --skill deep-trajectory-analysis -a codex`. Or copy the skill folder (deep-trajectory-analysis in swyxio/skills) into .agents/skills/deep-trajectory-analysis in your project. Codex loads it when a task matches its description.

Can I use Deep Trajectory Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill deep-trajectory-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deep-trajectory-analysis, .gemini/skills/deep-trajectory-analysis, .github/skills/deep-trajectory-analysis and .opencode/skills/deep-trajectory-analysis in your project.

What does Deep Trajectory Analysis need to run?

Going by SKILL.md and its folder, Deep Trajectory Analysis needs JavaScript for the scripts in its folder. Our summary lists: Node.js.

Does Deep Trajectory Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Deep Trajectory Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Deep Trajectory Analysis use?

Deep Trajectory Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deep Trajectory Analysis use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Deep Trajectory Analysis?

Skills that share tags, products or a category with Deep Trajectory Analysis: Wp Performance Review (elvismdev/claude-wordpress-skills, 235 stars), Align Human (agentscope-ai/OpenJudge, 868 stars), Performance Report (Affitor/affiliate-skills, 699 stars) and Run Mv Hoi Reconstruction (nvidia-isaac/video_to_data, 850 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deep Trajectory Analysis?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 175 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.