Agent skill

Differential Session Runner

by alchemiststudiosDOTai in alchemiststudiosDOTai/harness-engineering

Run or continue a differential debugging session between two implementations, traces, captures, or outputs.

MITAuto-check: notesDevelopment

Install Differential Session Runner

skills CLI
$ npx skills add alchemiststudiosDOTai/harness-engineering --skill differential-session-runner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alchemiststudiosDOTai/harness-engineering differential-session-runner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alchemiststudiosDOTai/harness-engineering.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/differential-session-runner .claude/skills/differential-session-runner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
differential-session-runner
GitHub stars
105
Token cost
~1.7k tokens
SKILL.md length
709 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Run or continue a differential debugging session between two implementations, traces, captures, or outputs.

  • Works in 9 steps: Load the repo's evidence convention first → Identify the artifact under investigation → Decide whether to continue or create → …
  • Tasks that involve Debugging
  • SKILL.md covers When to Use, Core Principle, Preferred Storage and Workflow, plus 5 more sections
  • Calls python and pytest

What it does

Differential Session Runner is an agent skill from alchemiststudiosDOTai/harness-engineering. Run or continue a differential debugging session between two implementations, traces, captures, or outputs. Record artifact identity, exact commands, first mismatch progression, findings, validation, and next probe in a durable session log.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Debugging. The repository describes itself as: harness-engineering discussion of shortcuts, automation, hacks and overall productivity with code agents like claude code, codex, and other harness. The licence is MIT.

When your agent uses it

  • Tasks that involve Debugging

Example prompts

  • “/differential-session-runner”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Load the repo's evidence convention first
  2. Identify the artifact under investigation
  3. Decide whether to continue or create
  4. Record baseline comparison commands
  5. Record first mismatch progression
  6. Record key findings
  7. Record landed changes separately
  8. Record validation
  9. Set the next probe

What it can do on your machine

Read from SKILL.md and the folder at commit 7a9fa15. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Differential Session Runner loads about 1.7k tokens when it runs. Until then it costs about 67 tokens; SKILL.md has 709 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from alchemiststudiosDOTai/harness-engineering at commit 7a9fa15, republished under its MIT licence (© alchemiststudiosDOTai). 709 words, ~1,676 tokens.

Download SKILL.mdSave it as .claude/skills/differential-session-runner/SKILL.md (or your agent's skills folder).
name
differential-session-runner
description
Run or continue a differential debugging session between two implementations, traces, captures, or outputs. Record artifact identity, exact commands, first mismatch progression, findings, validation, and next probe in a durable session log.
allowed-tools
Read, Write, Edit, Bash
writes-to
memory-bank/evidence/
hard-guards
Record exact commands, artifact paths, and hashes whenever available, Prefer continuing an existing session for the same artifact identity instead of creating…

Differential Session Runner

Use this skill when debugging requires a durable evidence trail rather than ad hoc notes.

This skill is for workflows where you are comparing:

  • original vs rewrite
  • implementation A vs implementation B
  • baseline trace vs candidate trace
  • before vs after replay output
  • captured artifact vs regenerated artifact

The goal is not only to investigate. The goal is to leave behind a session artifact another operator or agent can continue.

When to Use

Use this skill when the user asks to:

  • continue a differential debugging session
  • compare two implementations and record mismatches
  • create or update a debugging session log
  • track replay / trace / capture divergence over time
  • document what changed between mismatched and cleared runs

Core Principle

Every differential investigation should produce a reusable evidence packet.

A good session artifact lets another operator answer:

  • what exact artifact was investigated?
  • how was it identified?
  • what commands were run?
  • where did the first mismatch appear?
  • what was learned?
  • what changed?
  • what validation proves the current state?
  • what should happen next?

Preferred Storage

If the repo already has a native evidence location, use it.

Examples:

  • docs/.../sessions/
  • docs/chunks/
  • analysis/.../reports/
  • existing session indexes

If the repo does not already have a native convention, write to:

  • memory-bank/evidence/YYYY-MM-DD_HH-MM-SS_<topic>-session.md
  • optional index: memory-bank/evidence/index.md

Workflow

1. Load the repo's evidence convention first

Before creating anything, search for:

  • playbooks / runbooks
  • session indexes
  • prior session files
  • report directories
  • chunk/evidence templates

Read the relevant guidance and continue the repo's existing pattern.

2. Identify the artifact under investigation

Capture the strongest available identity for the artifact:

  • artifact path
  • SHA256 / content hash
  • commit SHA
  • case name / run id / replay id
  • timestamp if needed

If a content hash is possible, record it early and use it as the main session identity.

3. Decide whether to continue or create

Search existing sessions for the artifact identity.

  • If a session already exists for the same artifact/hash, append to that session.
  • If the artifact identity is new, create a new session and update the session index if the repo has one.
4. Record baseline comparison commands

Capture the exact commands used for the first comparison step.

Examples:

bash
python compare.py --baseline out/a.json --candidate out/b.json
pytest tests/test_replay.py -k case_17
mytool diff trace_a.cdt trace_b.cdt

Never summarize commands loosely. Record them exactly.

5. Record first mismatch progression

Capture the first relevant divergence and, if applicable, how it moved over time.

Examples:

  • first bad tick
  • first mismatched field
  • first unexpected output line
  • first extra/missing RNG draw
  • first snapshot diff

If later fixes move the mismatch frontier, append the new progression rather than deleting the old one.

Show full SKILL.md (300 more words)Show less
6. Record key findings

Write findings as evidence-backed observations, not guesses.

Good findings:

  • identify the mismatched subsystem or callsite
  • name the field or branch that diverged
  • note whether the issue self-heals or persists
  • connect the mismatch to a concrete code path or state transition
7. Record landed changes separately

If code changes are made during the investigation, capture them in a separate section:

  • files changed
  • short description of each change
  • tests added or updated

If no changes were made, state that explicitly.

8. Record validation

List the validation commands and their results.

Examples:

  • replay passes end-to-end
  • diff result is now clean
  • targeted tests pass
  • health check passes

Do not write "fixed" without a validation section.

9. Set the next probe

Every session should end with one of:

  • cleared / no next probe required
  • one explicit next investigation target
  • one blocked dependency or missing tool/input
markdown
---
title: "<topic> – Differential Session"
phase: Evidence
date: "YYYY-MM-DD HH:MM:SS"
owner: "<agent_or_user>"
tags: [evidence, differential, <topic>]
---

## Artifact
- Path: `<path>`
- Identity: `<sha256|commit|case-id>`

## Baseline Commands
- `<exact command 1>`
- `<exact command 2>`

## First Mismatch Progression
- baseline: `<first mismatch>`
- after fix 1: `<new frontier>`
- after fix 2: `<cleared|new mismatch>`

## Key Findings
- finding 1
- finding 2

## Landed Changes
- `path/to/file` → change summary
- `tests/...` → validation coverage added

## Validation
- `<command>` → `<result>`
- `<command>` → `<result>`

## Outcome / Next Probe
- `<cleared | next probe | blocked reason>`

Good Session Behavior

  • preserve earlier mismatch states instead of overwriting them
  • use hashes or artifact IDs to avoid ambiguous session naming
  • record exact commands so another operator can replay the same evidence path
  • separate findings from fixes
  • separate fixes from validation

Bad Session Behavior

  • "Investigated replay issue and fixed some stuff"
  • omitting the artifact identity
  • omitting commands
  • replacing prior mismatch history with the latest state only
  • claiming success without validation evidence

Output Requirements

A completed session artifact should make handoff possible with no hidden context.

It must include:

  • artifact identity
  • exact commands
  • first mismatch progression
  • key findings
  • landed changes or explicit no-change note
  • validation evidence
  • next probe or clear outcome

Handoff

After updating the session artifact:

  • if the investigation uncovered code work, hand off to plan-phase
  • if code work is already scoped in a plan, hand off to execute-phase
  • if the investigation is complete, send the user the session path and a one-line status

© alchemiststudiosDOTai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/differential-session-runner of alchemiststudiosDOTai/harness-engineering.

Open the folder on GitHubat commit 7a9fa15

Compare with similar skills

Differential Session Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Differential Session Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Differential Session Runner this skillalchemiststudiosDOTai/harness-engineering105—~1.7kAutomated safety check: NotesMIT
Trellis Session Insightmindfold-ai/Trellis15k4 repos~1.7kAutomated safety check: PassAGPL-3.0
Native Data FetchingCherryHQ/cherry-studio-app4k6 repos~2.9kAutomated safety check: NotesMIT
Debugging Executionsn8n-io/n8n207k—~2.6kAutomated safety check: PassCustom licence
Aoti Debugpytorch/pytorch104k1 repos~1.7kAutomated safety check: PassCustom licence
Herdr Throwaway Reproductionherdrdev/herdr43k—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Trellis Session Insight

    mindfold-ai/Trellis

    Reach into past AI conversation history through the trellis mem CLI.

    15k GitHub starsUsed in 4 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Native Data Fetching

    CherryHQ/cherry-studio-app

    A skill your agent uses when implementing or debugging ANY network request, API call, or data fetching.

    4k GitHub starsUsed in 6 repos~2.9k tokens
    DevelopmentAuto-check: notes
  • Official

    Debug failed or wrong-output workflow executions using executions tools.

    207k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • Runs a disposable, uniquely named Herdr session inside an existing one so runtime, pane, terminal or API bugs can be reproduced without touching the main session.

    43k GitHub stars~2.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Systematic Debugging

    ultralisp/ultralisp

    A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes

    258 GitHub starsUsed in 51 repos~2.4k tokens
    DevelopmentAuto-check passed

More from alchemiststudiosDOTai/harness-engineering

All 8 skills in this repo
  • Ast Grep Setup

    alchemiststudiosDOTai/harness-engineering

    Set up ast-grep for a codebase with common TypeScript rules for detecting anti-patterns, enforcing best practices, and preventing bugs.

    105 GitHub stars~4.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Research Phase

    alchemiststudiosDOTai/harness-engineering

    This skill should be used when mapping or researching a codebase to understand its structure, patterns, and architecture.

    105 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check: notes
  • Execute Phase

    alchemiststudiosDOTai/harness-engineering

    Execute implementation plans from .artifacts/plan/. An agent skill from alchemiststudiosDOTai/harness-engineering.

    105 GitHub stars~1.8k tokensUpdated 6 mo ago
    Auto-check passed
  • Harness Map

    alchemiststudiosDOTai/harness-engineering

    Map a repository's mechanical harness layers: canonical check command, local and CI gates, architecture boundaries, structural rules, behavioral verification, docs ratchets, evidence workflows, and…

    105 GitHub stars~1.8k tokensUpdated 6 mo ago
    Auto-check passed
  • Agents Md Mapper

    alchemiststudiosDOTai/harness-engineering

    This skill should be used when creating, refreshing, or validating a repository AGENTS.md so it stays concise, current, and grounded in repository evidence.

    105 GitHub stars~1.8k tokensUpdated 6 mo ago
    Auto-check passed
  • Plan Phase

    alchemiststudiosDOTai/harness-engineering

    Generate execution-ready implementation plans from research docs - planning ONLY, no fixing or verifying.

    105 GitHub stars~1.8k tokensUpdated 6 mo ago
    Auto-check passed

Categories

Questions about Differential Session Runner

What does Differential Session Runner do?

Run or continue a differential debugging session between two implementations, traces, captures, or outputs. Differential Session Runner is an agent skill from alchemiststudiosDOTai/harness-engineering. Run or continue a differential debugging session between two implementations, traces, captures, or outputs.

When should I use Differential Session Runner?

Differential Session Runner fits situations like: tasks that involve Debugging.

How do I install Differential Session Runner in Claude Code?

Run `npx skills add alchemiststudiosDOTai/harness-engineering --skill differential-session-runner -a claude-code`. Or copy the skill folder (skills/differential-session-runner in alchemiststudiosDOTai/harness-engineering) into .claude/skills/differential-session-runner in your project. Claude Code loads it when a task matches its description.

How do I install Differential Session Runner in Codex?

Run `npx skills add alchemiststudiosDOTai/harness-engineering --skill differential-session-runner -a codex`. Or copy the skill folder (skills/differential-session-runner in alchemiststudiosDOTai/harness-engineering) into .agents/skills/differential-session-runner in your project. Codex loads it when a task matches its description.

Can I use Differential Session Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alchemiststudiosDOTai/harness-engineering --skill differential-session-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/differential-session-runner, .gemini/skills/differential-session-runner, .github/skills/differential-session-runner and .opencode/skills/differential-session-runner in your project.

What does Differential Session Runner need to run?

Going by SKILL.md and its folder, Differential Session Runner needs the command-line tools its instructions call (python and pytest). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash.

Does Differential Session Runner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Differential Session Runner safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Differential Session Runner use?

Differential Session Runner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Differential Session Runner use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Differential Session Runner?

Skills that share tags, products or a category with Differential Session Runner: Trellis Session Insight (mindfold-ai/Trellis, 15k stars), Native Data Fetching (CherryHQ/cherry-studio-app, 4k stars), Debugging Executions (n8n-io/n8n, 207k stars) and Aoti Debug (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Differential Session Runner?

alchemiststudiosDOTai (a GitHub organization) maintains it in alchemiststudiosDOTai/harness-engineering, which has 105 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on March 17, 2026.

Source: alchemiststudiosDOTai/harness-engineering on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.