Agent skill

Dr Drive Harness

by AHepi in AHepi/DeepReason

The driving manual for DeepReason - how to run the harness properly (session preflight, the public CLI lifecycle, live-run ladders) and where to look before modifying anything or when diagnosing a…

MITAuto-check passedAgent Workflows

Install Dr Drive Harness

skills CLI
$ npx skills add AHepi/DeepReason --skill dr-drive-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AHepi/DeepReason dr-drive-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AHepi/DeepReason.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/dr-drive-harness .claude/skills/dr-drive-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dr-drive-harness
GitHub stars
141
Token cost
~3.4k tokens
SKILL.md length
1,689 words
Files
1
Skills in repo
30
Repo updated
First seen
Licence
MIT

At a glance

The driving manual for DeepReason - how to run the harness properly (session preflight, the public CLI lifecycle, live-run ladders) and where to look before modifying anything or when diagnosing a…

  • Works in 6 steps: Session preflight (before anything else) → Running it — the public lifecycle → Running it — live experiment ladders → …
  • Tasks that involve Agent instruction files
  • SKILL.md covers 1. Session preflight (before…, 2. Running it — the public…, 3. Running it — live… and 4. Where to look BEFORE…, plus 3 more sections
  • Calls python and git

What it does

Dr Drive Harness is an agent skill from AHepi/DeepReason. The driving manual for DeepReason - how to run the harness properly (session preflight, the public CLI lifecycle, live-run ladders) and where to look before modifying anything or when diagnosing a problem. An index over the owning authorities (CLAUDE.md, docs/map, the workflow skills), not a replacement for them. Load at the start of any session that will run, modify, or diagnose the harness, especially a first session in this repo.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Agent instruction files. It works with Git and Python. The licence is MIT.

When your agent uses it

  • Tasks that involve Agent instruction files

Example prompts

  • “/dr-drive-harness”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Session preflight (before anything else)
  2. Running it — the public lifecycle
  3. Running it — live experiment ladders
  4. Where to look BEFORE modifying anything
  5. Where to look WHEN something breaks
  6. Routing to the workflows

What it can do on your machine

Read from SKILL.md and the folder at commit 9607fba. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dr Drive Harness loads about 3.4k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 1,689 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AHepi/DeepReason at commit 9607fba, republished under its MIT licence (© AHepi). 1,689 words, ~3,377 tokens.

Download SKILL.mdSave it as .claude/skills/dr-drive-harness/SKILL.md (or your agent's skills folder).
name
dr-drive-harness
description
The driving manual for DeepReason - how to run the harness properly (session preflight, the public CLI lifecycle, live-run ladders) and where to look before modifying anything or when diagnosing a problem. An index over the owning authorities (CLAUDE.md, docs/map, the workflow skills), not a replacement for them. Load at the start of any session that will run, modify, or diagnose the harness, especially a first session in this repo.

Drive the harness

You are operating a Popperian reasoning harness whose entire epistemology rests on one rule: the typed record is the only admissible evidence. log.jsonl, objects/, progress.jsonl, run-status.json, REPLAY_VALIDATION.json, verify_root — those are evidence. Model prose, including yours, is not. Every section below is an index: it tells you the load-bearing command and WHERE the full authority lives, so you never operate from a half-remembered copy.

1. Session preflight (before anything else)

The cloud container rolls back silently — stale checkout, dead processes, deleted gitignored files. CLAUDE.md's "Environment" section is the authority; the sequence is:

git log --oneline -1                      # stale head? resync:
git fetch origin <branch> && git checkout -B <branch> origin/<branch>
python -c "import deepreason" || pip install -e . --break-system-packages -q
ls experiments/*/env 2>/dev/null          # gitignored credentials survive?

No live launch without a green soak on the launch config: run python -u scripts/cycle_soak.py --case <case> before any ladder launch. It drives the managed path to cycle 8 on the launch configuration's own shape against the deterministic stub, and carries a named assertion for each of the four 2026-08-22 cycle-0-to-2 operational deaths. It has REPRODUCED one of them offline (the reservation-bound seam); the other three are asserted, not demonstrated — read that tranche's RESULTS.md before treating a green soak as full coverage (experiments/2026-08-23-change-cycle-soak-instrument/).

Always python -m pytest, never bare pytest (PATH shim). Credentials are recreated from the operator's handover, never committed — env files are gitignored; check with git check-ignore <path> before writing near them. Commit and push at every phase boundary — work between pushes is work at risk. Then read, in order: CLAUDE.md, the newest experiments/*/RESULTS.md segments, docs/ERRATA.md.

Where the truth lives, in reading order: CLAUDE.md (law) → docs/map/INDEX.md (navigation) → experiments/*/RESULTS.md (what is proven) → docs/ERRATA.md (what was corrected) → each tranche's PARKED.md (what is deliberately not done).

Re-entering mid-tranche needs no conversation history: every tranche is resumable from its committed artifacts alone. Read the tranche dir's CHECKLIST.md State: line, then REQUEST.md/SPEC.md, and continue. The whole fresh-window prompt is one line — "Resume tranche <dir> from its artifacts." If a session cannot resume from the artifacts, the previous session under-committed; record that gap, reconstruct, and commit.

2. Running it — the public lifecycle

The supported product surface (authority: README.md):

deepreason setup                 # one strict provider profile
deepreason qualify --yes         # explicit; tier ladder full/shallow/unqualified
deepreason status [--json]       # readiness + the one next action
                                 # NB: provider readiness, NOT a run's
                                 # outcome — for that, `results` below
deepreason results ROOT-OR-HOME [--json] [--verify]
                                 # read a run's typed results: id, state,
                                 # stop_reason, cycles, tokens vs budget,
                                 # artifact/survivor/frontier counts,
                                 # defended-trial + judge-call counts, the
                                 # STORED verify_root verdict (--verify
                                 # re-derives), amendment epochs, and
                                 # whether the root is amend-ready.
                                 # Read-only; absent facts print as typed
                                 # absences, never omitted. This is the
                                 # ONE retrieval surface — do not go
                                 # hunting through root files for it.
deepreason reason "QUESTION" [--cycles N] [--token-budget N]
deepreason reason "Q" --attach file.pdf        # frozen evidence, dossier digest
deepreason --root ROOT amend --attach f --reshape-question "Q2"
deepreason --root ROOT continue --budget cycles=N
deepreason reason --shallow "Q"  # MiniReason reduced engine
deepreason web                   # loopback-only page over the MCP facade

Facts that bite (authority: CLAUDE.md "Live runs"): qualification caches by subject digest — same home + profile + opt-ins is a ~1s cache hit, any change reruns a ~14-minute battery; qualify opt-ins must match reason opt-ins (--attached-evidence ⇔ --attach); provider reasoning must be EXPLICITLY disabled for ollama when required (unset is not off — the refusal is typed).

3. Running it — live experiment ladders

Ladders are shell scripts (experiments/*/**_run.sh): setup → qualify → reason → audit against a DEEPREASON_HOME. The rules, each learned the expensive way (authority: CLAUDE.md "Live runs"):

  • Run identity is deterministic. Same question + config → same run id; a leftover root refuses with RUN_ALREADY_STARTED. Retire by rename (git mv run-<id> <state>-epochN-run-<id>) and COMMIT THE RENAME FIRST. Never edit a committed root — to change the question or add evidence, deepreason amend then continue.
  • Launch detached, never foreground: from the ladder's directory, setsid nohup ./<ladder>.sh & disown. Arm the snapshot loop (snapshot_loop.sh) and a monitor on the newest root's progress.jsonl plus the driver log's rc= lines — alert on failure signatures, not just success.
  • Judge only typed outcomes: run state, stop_reason, the audit JSON, verify_root, FINDINGS.md. Capability-channel use is stochastic across identical runs — one live miss is inconclusive; the offline regression is the proof.

4. Where to look BEFORE modifying anything

Never scope a change by grepping 125k lines. The map is the navigation layer, and the reading order is fixed:

  1. docs/map/INDEX.md — resolve the work to ids (DR-SUB-<pkg>, DR-CON-<concept>, DR-SEAM-<a>-x-<b>).
  2. docs/map/INV-frozen-surfaces.md — first, always: five surfaces are not yours to change (state digests, harness event application, replay-validation formats, manifest schemas + validators, qualification subjects). Readers may be fixed; formats may not.
  3. If the change spans two things, the seam document BEFORE either subsystem: the file is docs/map/SEAM-<a>-x-<b>.md, sides in alphabetical order. It names the small fraction of each side actually involved. The worked recipe for any seam change is docs/map/REC-change-a-seam.md.
  4. docs/map/SCHEMA.md before writing or editing any map document. The map moves in the SAME commit as the code, or it becomes a document that lies.
  5. Record the resolved ids in the tranche's first artifact (GOAL.md or REQUEST.md) — every later phase starts from the same map. If the map has no id for something the work touches, that is a finding, not a blocker: say so, and creating the missing document becomes part of the tranche.

Instruments that prove you broke nothing: the full gate (python -m pytest tests/ -q -n 4, 0 failed only) and the root sweep (python tools/root_sweep.py — no committed root's verdict may move). Third instrument, which NO gate runs for you: the wheel smokes (python scripts/wheel_smoke.py; python -u scripts/wheel_operational_smoke.py) — build-and-operate checks over the INSTALLED package. They pin the public surface (console entry points, MCP tool set + schema sha, wheel layout), so any change to that surface updates the pins and re-runs the smoke in the SAME commit, or the instrument rots silently and nothing else will catch it. python tools/docs_verify.py is the same gate for the map — and its --fast mode reuses cached results, so it CANNOT catch a document your src/ change just broke. Iterate with --fast; run the FULL mode at least once before any commit that touches src/.

Show full SKILL.md (676 more words)Show less

5. Where to look WHEN something breaks

Record first, code second, theory last. In order:

Look atIt tells you
deepreason stop-report <root-or-home>run this first. What actually ran per seat, what qualification already knew, provider health, the stop classified into CONFIGURATION / ENVIRONMENT / MODEL / HARNESS ranked by evidence, and whether continue would be accepted. Derives the four rows below and the qualification rows hand-reading skips. dr-diagnose gates on its section 4
<root>/run-status.jsonstate, stop_reason, message — often the whole answer
<root>/progress.jsonlwhich cycle/phase/token count it died at
<root>/REPLAY_VALIDATION.json, verify_root(<root>)typed violations: check name + detail (open the root READ-ONLY — a writable open repairs, i.e. destroys, the evidence)
the violation's blob under <root>/blobs/the verbatim error and the rejected value — read this BEFORE theorizing; both recorded cycle-0 deaths were misattributed by readers who skipped it
the covering map document's Traps sectionwhether this exact failure happened before — the cheapest diagnosis available
docs/ERRATA.mdwhether the document you are trusting was already corrected

Two instruments can disagree and both be right (verify_root vs verify_root_report vs the sweep) — always cite the instrument with the number. When the cause is located, do not fix it inline: route it.

5b. Process hygiene (each rule paid for in the record)

  • Kill by PID, never by pattern. pkill -f/pgrep -f can match your own shell's command line and kill your own session.
  • Never run the full gate concurrently with docs_verify (or any other worker-spawning instrument): both fan out processes, and the contention manufactures failures. One instrument at a time, on an otherwise idle box.
  • A surprising measurement taken under load is not a measurement. Re-run idle before recording it, and say which run you recorded.
  • Long work launches detached (setsid nohup ... & disown, §3) — a foreground process dies with the session.
  • Scratch and temp files go to the session scratchpad, never the repo.

6. Routing to the workflows

All substantive work goes through a workflow family — that is repo law (CLAUDE.md), not preference. This section is the index of all of them (CLAUDE.md's "Which workflow to use" carries the same summary).

  • Something is broken or suspicious → deepreason-orchestrator (dr-set-goal → dr-diagnose → dr-reproduce → dr-propose-fix → dr-implement-fix → dr-verify-outcome). Diagnosis from the typed record BEFORE code reading.
  • The operator suggests a change → dr-change-orchestrator (dr-capture-request → dr-spec-change → dr-plan-steps → dr-execute-step → dr-validate-change → dr-deliver-change). Authority is the operator's verbatim words, ledgered in REQUEST.md.
  • The operator's message is ambiguous or terse, a phase says "stop and ask", or evidence contradicts your expectation → dr-ask-the-right-question first: route the question to the cheapest authority (record → framework → operator) before spending operator attention.

Cross-routing is strict: a defect found mid-change is PARKED, not fixed; a change wished for mid-defect is PARKED, not implemented. One tranche, one goal.

Calibration for less capable executors. The documents this manual points at are complete by design — execute them literally rather than improvising a summary of them. Never generalize an instruction beyond its stated scope; if a spec seems silent about your case, that is a question (load dr-ask-the-right-question), not an invitation to infer — an accepted, judgment-only exception to authoring-skills' GATE-every- negation rule, docs/ERRATA.md E24. A multi-step program (a handover, a checklist, a ladder) runs one step per tranche — finishing a step early is never a reason to start the next in the same tranche. Stop conditions and DESIGN-AND-STOP gates are hard stops: the deliverable at a gate is a committed document and an ended turn, not an implementation. And every stop presented to the operator leads with the decision needed in ONE sentence, the options priced, and a recommendation with its reason — the operator should be able to answer with a word. Style, per the operator's recorded preference (CLAUDE.md, Conventions): answer their actual worry first; say what a scary finding does NOT mean before what it does; own the workflow's own contribution to any confusion; and close hard explanations with one short, accurate everyday analogy.

Exit criterion. You know you are driving properly when every claim you make about a run ends in a typed artifact, every modification you plan started from INDEX.md and INV-frozen-surfaces.md, and every problem you chase entered a workflow tranche with its evidence committed.

© AHepi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/dr-drive-harness of AHepi/DeepReason.

Open the folder on GitHubat commit 9607fba

Compare with similar skills

Dr Drive Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dr Drive Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dr Drive Harness this skillAHepi/DeepReason141—~3.4kAutomated safety check: PassMIT
Commitayutaz/piper-plus234—~839Automated safety check: NotesMIT
Sync Docsayutaz/piper-plus234—~1.4kAutomated safety check: PassMIT
Neat-Freak Knowledge CloseoutKKKKhazix/khazix-skills21k—~1.9kAutomated safety check: PassMIT
Harness Evolution Feedback Looprevfactory/harness9.1k—~855Automated safety check: PassApache-2.0
Mspm0 Ccsmc3545dada/mspm0-skill374—~4.5kAutomated safety check: PassMIT

Similar skills

  • Commit

    ayutaz/piper-plus

    piper-plus のコミットルール (CLAUDE.md 準拠) でステージ済みファイルをコミットします。--no-verify 禁止、HEREDOC、適切な prefix を強制。

    234 GitHub stars~839 tokensUpdated today
    Agent WorkflowsAuto-check: notes
  • Sync Docs

    ayutaz/piper-plus

    コミット前にエージェントチームで全ドキュメント (CLAUDE.md / README / CHANGELOG / docs/) を監査し、コード変更に応じて自動更新します。大規模変更時の documentation drift を予防。

    234 GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Neat-Freak Knowledge Closeout

    KKKKhazix/khazix-skills

    Brings project docs, agent rule files, authorized memory and leftover workspace files back in line with what the code and runtime actually do at the end of a work session.

    21k GitHub stars~1.9k tokensUpdated 9 days ago
    Agent WorkflowsAuto-check passed
  • Collects feedback on how an agent harness performed, generalizes it, and updates the harness agents, skills and orchestrator along with a change-history table.

    9.1k GitHub stars~855 tokensUpdated 12 days ago
    Agent WorkflowsAuto-check passed
  • Mspm0 Ccs

    mc3545dada/mspm0-skill

    Tool-neutral CLI agent rules for TI MSPM0 development with Code Composer Studio, Keil/uVision, CMake/GCC/OpenOCD, SysConfig, and DriverLib.

    374 GitHub stars~4.5k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Generate AI Rules

    divar-ir/ai-doc-gen

    Generate AI assistant configuration files for a repository — CLAUDE.md, AGENTS.md, and Cursor rules (.cursor/rules/.mdc) — from codebase analysis.

    767 GitHub stars~1.2k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed

More from AHepi/DeepReason

All 30 skills in this repo
  • Pinker Clarity Workflow

    AHepi/DeepReason

    Orchestrate a Steven Pinker-grounded workflow for teaching, explanatory writing, or material that must do both.

    141 GitHub stars~1.5k tokensUpdated 29 days ago
    Auto-check passed
  • Design, deliver, or audit explanations and lessons with a Pinker-informed focus on phenomena, the curse of knowledge, concrete models, active reasoning, feedback, and revision.

    141 GitHub stars~1.8k tokensUpdated 29 days ago
    Auto-check passed
  • Pinker Write For Readers

    AHepi/DeepReason

    Draft, revise, teach, or audit expository prose using Pinker's cognitive approach to style: classic presentation, reader modeling, curse-of-knowledge repair, coherent information order, deliberate…

    141 GitHub stars~2.1k tokensUpdated 29 days ago
    Auto-check passed
  • Example Battery

    AHepi/DeepReason

    Build a battery of concrete instances BEFORE writing or evaluating any definition, pin, or semantic clause (Reed step 1).

    141 GitHub stars~811 tokensUpdated 29 days ago
    Auto-check passed
  • Authoring Skills

    AHepi/DeepReason

    Rules for writing, editing, and retiring skill and workflow files for LLM agents.

    141 GitHub stars~1.7k tokensUpdated 29 days ago
    Auto-check passed
  • Deepreason Orchestrator

    AHepi/DeepReason

    Entry point for any DeepReason problem. An agent skill from AHepi/DeepReason.

    141 GitHub stars~1.1k tokensUpdated 29 days ago
    Auto-check passed

Works with

Categories

Questions about Dr Drive Harness

What does Dr Drive Harness do?

The driving manual for DeepReason - how to run the harness properly (session preflight, the public CLI lifecycle, live-run ladders) and where to look before modifying anything or when diagnosing a…. Dr Drive Harness is an agent skill from AHepi/DeepReason. The driving manual for DeepReason - how to run the harness properly (session preflight, the public CLI lifecycle, live-run ladders) and where to look before modifying anything or when diagnosing a problem.

When should I use Dr Drive Harness?

Dr Drive Harness fits situations like: tasks that involve Agent instruction files.

How do I install Dr Drive Harness in Claude Code?

Run `npx skills add AHepi/DeepReason --skill dr-drive-harness -a claude-code`. Or copy the skill folder (.claude/skills/dr-drive-harness in AHepi/DeepReason) into .claude/skills/dr-drive-harness in your project. Claude Code loads it when a task matches its description.

How do I install Dr Drive Harness in Codex?

Run `npx skills add AHepi/DeepReason --skill dr-drive-harness -a codex`. Or copy the skill folder (.claude/skills/dr-drive-harness in AHepi/DeepReason) into .agents/skills/dr-drive-harness in your project. Codex loads it when a task matches its description.

Can I use Dr Drive Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AHepi/DeepReason --skill dr-drive-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dr-drive-harness, .gemini/skills/dr-drive-harness, .github/skills/dr-drive-harness and .opencode/skills/dr-drive-harness in your project.

What does Dr Drive Harness need to run?

Going by SKILL.md and its folder, Dr Drive Harness needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Dr Drive Harness access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Dr Drive Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dr Drive Harness use?

Dr Drive Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dr Drive Harness use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dr Drive Harness?

Skills that share tags, products or a category with Dr Drive Harness: Commit (ayutaz/piper-plus, 234 stars), Sync Docs (ayutaz/piper-plus, 234 stars), Neat-Freak Knowledge Closeout (KKKKhazix/khazix-skills, 21k stars) and Harness Evolution Feedback Loop (revfactory/harness, 9.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dr Drive Harness?

AHepi (a GitHub user) maintains it in AHepi/DeepReason, which has 141 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on September 10, 2026.

Source: AHepi/DeepReason on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.