Show Me Your Work Decision Log
cursor/plugins
Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.
Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read…
$ npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install krakenfx/kraken-cli kraken-lab-experiment --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/krakenfx/kraken-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/kraken-lab-experiment .claude/skills/kraken-lab-experiment && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "kraken-lab-experiment" agent skill from https://github.com/krakenfx/kraken-cli/tree/main/skills/kraken-lab-experiment into .claude/skills/kraken-lab-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kraken-lab-experiment", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/krakenfx/kraken-cli/tree/main/skills/kraken-lab-experimentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install krakenfx/kraken-cli kraken-lab-experiment --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/krakenfx/kraken-cli.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/kraken-lab-experiment .agents/skills/kraken-lab-experiment && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "kraken-lab-experiment" agent skill from https://github.com/krakenfx/kraken-cli/tree/main/skills/kraken-lab-experiment into .agents/skills/kraken-lab-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kraken-lab-experiment", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install krakenfx/kraken-cli kraken-lab-experiment --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/krakenfx/kraken-cli.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/kraken-lab-experiment .cursor/skills/kraken-lab-experiment && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "kraken-lab-experiment" agent skill from https://github.com/krakenfx/kraken-cli/tree/main/skills/kraken-lab-experiment into .cursor/skills/kraken-lab-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kraken-lab-experiment", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/krakenfx/kraken-cli.git --path skills/kraken-lab-experiment--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install krakenfx/kraken-cli kraken-lab-experiment --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/krakenfx/kraken-cli.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/kraken-lab-experiment .gemini/skills/kraken-lab-experiment && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "kraken-lab-experiment" agent skill from https://github.com/krakenfx/kraken-cli/tree/main/skills/kraken-lab-experiment into .gemini/skills/kraken-lab-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kraken-lab-experiment", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install krakenfx/kraken-cli kraken-lab-experimentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/krakenfx/kraken-cli.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/kraken-lab-experiment .github/skills/kraken-lab-experiment && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "kraken-lab-experiment" agent skill from https://github.com/krakenfx/kraken-cli/tree/main/skills/kraken-lab-experiment into .github/skills/kraken-lab-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kraken-lab-experiment", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install krakenfx/kraken-cli kraken-lab-experiment --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/krakenfx/kraken-cli.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/kraken-lab-experiment .opencode/skills/kraken-lab-experiment && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "kraken-lab-experiment" agent skill from https://github.com/krakenfx/kraken-cli/tree/main/skills/kraken-lab-experiment into .opencode/skills/kraken-lab-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kraken-lab-experiment", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
kraken-lab-experimentRun the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read…
Kraken Lab Experiment is an agent skill from krakenfx/kraken-cli. Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read the verdict — resumable from disk at every step, designed for /loop.
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Autonomous loops. The repository describes itself as: The first AI-native CLI for trading crypto, stocks, forex, and derivatives. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit aa56e59. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Kraken Lab Experiment loads about 2.7k tokens when it runs. Until then it costs about 72 tokens; SKILL.md has 1,310 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from krakenfx/kraken-cli at commit aa56e59, republished under its MIT licence (© krakenfx). 1,310 words, ~2,717 tokens.
.claude/skills/kraken-lab-experiment/SKILL.md (or your agent's skills folder).PREREQUISITE: Load
kraken-playgroundto understand what a session is and where it lives.
Turn a trading hunch into a mechanical verdict: pre-register the hypothesis and its success criteria (hash-sealed, so the goalposts cannot move), run it as recorded paper runs, score each session against its own tape, and let the frozen criteria — not vibes — say pass or fail. Every step reads and writes plain files under the active scope, so the lifecycle survives kills, restarts, and context loss: whatever step you find half-done on disk is the step you resume.
Use this skill for:
/loop (freeze → run → score →
verdict)export KRAKEN_WORKSPACE=<name>), where the order verbs fill on the
paper account and no real funds are reachable.parse envelope.lab score --experiment
or lab compare output in this session. No estimating, no "it would
probably pass".With a sealed session plan the whole table collapses to one command — run
kraken lab next <exp> and do what it says:
kraken lab next momentum-1 -o json 2>/dev/null
# state: "start" → run the exact `command` it prints, then stop+score
# state: "running" → drive or wait, then `kraken session stop` + `lab score`
# state: "plan_complete" → `kraken lab compare <exp>`, conclude (Step 5)lab next re-derives progress from the sessions tree on every call: completed
runs consume plan entries in order, aborted sessions keep their ordinals and
count for nothing, and a replay entry whose tape no longer matches its
sealed hash is refused outright. Stop and score a finished run before
asking again — an unstopped run has no summary and reads as aborted.
For a plan-less experiment (registered before run plans existed), fall back to the manual table:
| State | Evidence | Next action |
|---|---|---|
| unregistered | kraken lab show <exp> → validation error | Step 1: freeze |
| no run yet | kraken session list → no run stamped with the experiment | Step 2: start run |
| run live | kraken session show → status recording | Step 3: drive or wait, then stop |
| run stopped, unjudged | status stopped, no verdict recorded in your notes | Step 4: score + verdict |
| judged | verdict obtained | Step 5: conclude or start run n+1 |
Pick a name (path-safe segment: letters, digits, ., -, _) and state
the falsifiable claim plus at least one mechanical criterion:
kraken lab new momentum-1 \
--hypothesis "buying 15m strength beats sitting in cash" \
--strategy recipe-playground-dca \
--min-return-pct 1 --max-drawdown-pct 5 --min-fills 2 \
--session replay:tape:jun-crash --session replay:tape:jun-chop \
--session live:24h \
-o json 2>/dev/nullSeal the session plan with the spec: repeatable --session entries, in
execution order — replay:<ref> (a finalized tape from
kraken tape list; its content hash is sealed, so a re-recorded tape
can never satisfy the plan) or live:<window>. Replay first to screen
cheaply, live last: the promotion gate requires a passing live session because
replay cannot enforce lookahead. The session count is sealed too — that is the
point: no running until a pass appears.
The output echoes the sealed spec with its frozen hash (sha256:…).
Freezing is once-only: a second lab new under the same name is a
validation error saying the spec is immutable. Verify any time with:
kraken lab show momentum-1 -o json 2>/dev/nullIf show returns a parse error mentioning "modified after freezing", the
file was tampered with. Stop the experiment and report it — do not re-freeze
over it.
With a session plan, kraken lab next <exp> prints the exact start command —
use it verbatim; adjusting --speed per the reaction-time note below is
the one permitted edit. Runs are ordinary recorded sessions stamped with
--experiment — the stamp is what discovery keys on; the --label <exp>-s<n> is the human handle:
# live entry (fill in the market to record):
kraken session start --label momentum-1-s1 --experiment momentum-1 \
--symbols BTC/USD --channels ticker,trade \
-o json 2>/dev/null &
# replay entry (the tape defines the market; 10x compresses the wait):
kraken session start --label momentum-1-s2 --experiment momentum-1 \
--from tape:jun-crash --speed 10 -o json 2>/dev/null &Replay reaction-time note: at speed N you have 1/N of live reaction time —
suits slow-cadence strategies (DCA); drop to --speed 1 for reactive ones.
(At least one --channels entry is required on live entries, and the
recorder runs in the foreground, so background it with & as the
playground skill does. The session_started stdout line carries the allocated
session id.)
Drive the strategy the hypothesis names (e.g. the recipe skill from
--strategy), passing --reason on every trade so the scorecard's
decided-fill join has data — the rationale lands in the active session's
decision log automatically.
A session must stop before it can be scored — the stop summary is the anchor every later number decomposes:
kraken session stopOne command produces the honest scorecard and the mechanical verdict:
kraken lab score --session latest --experiment momentum-1 -o json 2>/dev/nullsource: what the session traded against — {kind:"replay", dataset, speed}
or {kind:"live", symbols}, derived from the session's contract, not
asserted.anchor + components[]: the stop totals and the explain-pnl waterfall,
carried verbatim — this output can never disagree with
kraken explain pnl.metrics: return_pct, max_drawdown, max_drawdown_pct, fees,
friction, turnover_pct, hit_rate, profit_factor,
slippage_sensitivity — every null is an honest "not derivable", never
zero (a buys-only DCA run has no closed round trips, so hit_rate and
profit_factor read null by design, not by breakage).verdict.pass and verdict.checks[]: one check per frozen criterion with
threshold, observed, pass. An unknown observation fails its check.caveats[]: read them; a verdict with caveats is reported with them.Digest for the loop log:
kraken lab score --session latest --experiment momentum-1 -o json 2>/dev/null \
| jq '{session, experiment, pass: .verdict.pass,
checks: [.verdict.checks[] | {criterion, threshold, observed, pass}],
caveats}'The retry cues are in the error envelopes themselves: "still recording — stop it with 'kraken session stop' first" means go back to Step 3; "was aborted before it stopped … only a stopped session can be scored" means that session is dead evidence — start a fresh one; "experiment '…' not found; run 'kraken lab new …' first" means Step 1 never happened in this scope.
The verdict is mechanical; the conclusion is yours to report, not adjust:
threshold vs observed) and what the
waterfall says the money actually went to (fees? drawdown? never traded?).
A failed experiment is a completed experiment.Promotion: when lab next says plan_complete, assemble the promotion
evidence block from kraken lab compare and hand off to the
kraken-paper-to-live skill. Its gate is mechanical: at least 2 passing
runs with a passing live run among them — replay verdicts count only
with their lookahead caveat attached. Two replay passes never promote.
More runs: repeat Steps 2–4 (each start allocates the next ordinal), then read every session side by side:
kraken lab compare momentum-1 -o json 2>/dev/nullOne verdict column per session (sessions[].verdict), plus pass_count/total
and a named error column for any run that cannot be scored (a still-running
run says so — stop it first). Only the verdicts compare across runs: they
cover different market windows, so the metric rows are per-session context.
Never average scorecards across runs into a blended number the CLI did not
produce — compare deliberately never aggregates.
Any context — including one that never saw the sessions — reproduces a verdict from disk alone:
kraken lab show momentum-1 -o json 2>/dev/null # the sealed criteria
kraken lab score --session s1 --experiment momentum-1 -o json 2>/dev/nullSame files, same exact-Decimal math, same verdict, byte for byte. If the two
disagree with a previously reported verdict, something on disk changed —
lab show tells you whether it was the spec.
Invoke as /loop with this skill and the experiment name. Each firing:
kraken lab show <exp> — missing → freeze (Step 1, with --session
entries) and return.kraken lab next <exp> -o json and act on its state. If it refuses
with a no-run-plan validation error (an experiment frozen before run
plans existed), drive the manual state table above instead.
start → run its command; running → drive/wait, then stop + score
(Steps 3–4) and log the digest; plan_complete → kraken lab compare,
conclude (Step 5), and stop the loop with the per-session verdicts (and the
promotion evidence block, if concluding toward promotion) as the final
report.Idempotent by construction: a firing that finds nothing to advance does nothing. A killed loop resumes at the same step.
If you hit a mismatch between what you are trying to do and the CLI's interface or responses — including a mismatch between this skill and the installed CLI version's contract — feel free to submit feedback with kraken feedback.
© krakenfx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/kraken-lab-experiment of krakenfx/kraken-cli.
Open the folder on GitHubat commit aa56e59
Kraken Lab Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Kraken Lab Experiment this skillkrakenfx/kraken-cli | 751 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Show Me Your Work Decision Logcursor/plugins | 11k | 8 repos | ~1.6k | Automated safety check: Pass | None | |
| Autoresearch Iteration Loopuditgoenka/autoresearch | 6.5k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Install Loop Engineeringcobusgreyling/loop-engineering | 11k | 1 repos | ~648 | Automated safety check: Pass | MIT | |
| LoopyForward-Future/loopy | 3.2k | — | ~3.9k | Automated safety check: Pass | MIT | |
| AI Performance Improvement Plantanweai/pua | 20k | 2 repos | ~6.9k | Automated safety check: Pass | MIT |
cursor/plugins
Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.
uditgoenka/autoresearch
Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.
cobusgreyling/loop-engineering
Installs Loop Engineering into a project through the single @cobusgreyling/loop CLI, scaffolding a report-only loop and a readiness score.
Forward-Future/loopy
Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication.
tanweai/pua
Pushes an agent to exhaust every option, investigate before asking and take initiative beyond the literal request, instead of giving up or waiting passively.
loopx-project/loopx
Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.
krakenfx/kraken-cli
Price alerts, threshold monitoring, and notification triggers for agents.
krakenfx/kraken-cli
Progress from manual trading to full agent autonomy with controlled risk at each level.
krakenfx/kraken-cli
Capture the spot-futures price spread with delta-neutral basis trades.
krakenfx/kraken-cli
Dollar cost averaging with scheduled buys and performance tracking.
krakenfx/kraken-cli
Discover staking strategies, allocate funds, and track earn positions.
krakenfx/kraken-cli
Handle order failures, network errors, and duplicate submissions safely.
Categories
Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read…. Kraken Lab Experiment is an agent skill from krakenfx/kraken-cli. Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read the verdict — resumable from disk at every step, designed for /loop.
Kraken Lab Experiment fits situations like: tasks that involve Autonomous loops.
Run `npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a claude-code`. Or copy the skill folder (skills/kraken-lab-experiment in krakenfx/kraken-cli) into .claude/skills/kraken-lab-experiment in your project. Claude Code loads it when a task matches its description.
Run `npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a codex`. Or copy the skill folder (skills/kraken-lab-experiment in krakenfx/kraken-cli) into .agents/skills/kraken-lab-experiment in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kraken-lab-experiment, .gemini/skills/kraken-lab-experiment, .github/skills/kraken-lab-experiment and .opencode/skills/kraken-lab-experiment in your project.
Going by SKILL.md and its folder, Kraken Lab Experiment needs the command-line tools its instructions call (jq).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Kraken Lab Experiment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Kraken Lab Experiment: Show Me Your Work Decision Log (cursor/plugins, 11k stars), Autoresearch Iteration Loop (uditgoenka/autoresearch, 6.5k stars), Install Loop Engineering (cobusgreyling/loop-engineering, 11k stars) and Loopy (Forward-Future/loopy, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
krakenfx (a GitHub organization) maintains it in krakenfx/kraken-cli, which has 751 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on August 7, 2026.
Source: krakenfx/kraken-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.