Agent skill

Kraken Lab Experiment

by krakenfx in krakenfx/kraken-cli

Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read…

MITAuto-check passedAgent Workflows

Install Kraken Lab Experiment

skills CLI
$ npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install krakenfx/kraken-cli kraken-lab-experiment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/krakenfx/kraken-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/kraken-lab-experiment .claude/skills/kraken-lab-experiment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kraken-lab-experiment
GitHub stars
751
Token cost
~2.7k tokens
SKILL.md length
1,310 words
Files
1
Skills in repo
59
Repo updated
First seen
Licence
MIT

At a glance

Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read…

  • Works in 5 steps: Freeze the hypothesis (once) → Start a session → Stop the session → …
  • Tasks that involve Autonomous loops
  • SKILL.md covers Safety rules, The state machine, Step 1: Freeze the hypothesis… and Step 2: Start a session, plus 5 more sections
  • Calls jq

What it does

Kraken Lab Experiment is an agent skill from krakenfx/kraken-cli. Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read the verdict — resumable from disk at every step, designed for /loop.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Autonomous loops. The repository describes itself as: The first AI-native CLI for trading crypto, stocks, forex, and derivatives. The licence is MIT.

When your agent uses it

  • Tasks that involve Autonomous loops

Example prompts

  • “/kraken-lab-experiment”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Freeze the hypothesis (once)
  2. Start a session
  3. Stop the session
  4. Score and judge
  5. Conclude — or run again

What it can do on your machine

Read from SKILL.md and the folder at commit aa56e59. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kraken Lab Experiment loads about 2.7k tokens when it runs. Until then it costs about 72 tokens; SKILL.md has 1,310 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from krakenfx/kraken-cli at commit aa56e59, republished under its MIT licence (© krakenfx). 1,310 words, ~2,717 tokens.

Download SKILL.mdSave it as .claude/skills/kraken-lab-experiment/SKILL.md (or your agent's skills folder).
name
kraken-lab-experiment
description
Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read the verdict — resumable from disk at every step, designed for /loop.
version
2.0.0

Lab Experiment Lifecycle

PREREQUISITE: Load kraken-playground to understand what a session is and where it lives.

Turn a trading hunch into a mechanical verdict: pre-register the hypothesis and its success criteria (hash-sealed, so the goalposts cannot move), run it as recorded paper runs, score each session against its own tape, and let the frozen criteria — not vibes — say pass or fail. Every step reads and writes plain files under the active scope, so the lifecycle survives kills, restarts, and context loss: whatever step you find half-done on disk is the step you resume.

Use this skill for:

  • driving one experiment end-to-end under /loop (freeze → run → score → verdict)
  • resuming an experiment a previous context started
  • reproducing a verdict cold, from files alone

Safety rules

  • Paper only: run experiments inside a paper workspace (export KRAKEN_WORKSPACE=<name>), where the order verbs fill on the paper account and no real funds are reachable.
  • Never edit a frozen experiment file. The seal exists to catch exactly that; a tampered spec fails every later step with a parse envelope.
  • Never report a verdict you did not obtain from lab score --experiment or lab compare output in this session. No estimating, no "it would probably pass".

The state machine

With a sealed session plan the whole table collapses to one command — run kraken lab next <exp> and do what it says:

bash
kraken lab next momentum-1 -o json 2>/dev/null
# state: "start"         → run the exact `command` it prints, then stop+score
# state: "running"       → drive or wait, then `kraken session stop` + `lab score`
# state: "plan_complete" → `kraken lab compare <exp>`, conclude (Step 5)

lab next re-derives progress from the sessions tree on every call: completed runs consume plan entries in order, aborted sessions keep their ordinals and count for nothing, and a replay entry whose tape no longer matches its sealed hash is refused outright. Stop and score a finished run before asking again — an unstopped run has no summary and reads as aborted.

For a plan-less experiment (registered before run plans existed), fall back to the manual table:

StateEvidenceNext action
unregisteredkraken lab show <exp> → validation errorStep 1: freeze
no run yetkraken session list → no run stamped with the experimentStep 2: start run
run livekraken session show → status recordingStep 3: drive or wait, then stop
run stopped, unjudgedstatus stopped, no verdict recorded in your notesStep 4: score + verdict
judgedverdict obtainedStep 5: conclude or start run n+1

Step 1: Freeze the hypothesis (once)

Pick a name (path-safe segment: letters, digits, ., -, _) and state the falsifiable claim plus at least one mechanical criterion:

bash
kraken lab new momentum-1 \
  --hypothesis "buying 15m strength beats sitting in cash" \
  --strategy recipe-playground-dca \
  --min-return-pct 1 --max-drawdown-pct 5 --min-fills 2 \
  --session replay:tape:jun-crash --session replay:tape:jun-chop \
  --session live:24h \
  -o json 2>/dev/null

Seal the session plan with the spec: repeatable --session entries, in execution order — replay:<ref> (a finalized tape from kraken tape list; its content hash is sealed, so a re-recorded tape can never satisfy the plan) or live:<window>. Replay first to screen cheaply, live last: the promotion gate requires a passing live session because replay cannot enforce lookahead. The session count is sealed too — that is the point: no running until a pass appears.

The output echoes the sealed spec with its frozen hash (sha256:…). Freezing is once-only: a second lab new under the same name is a validation error saying the spec is immutable. Verify any time with:

bash
kraken lab show momentum-1 -o json 2>/dev/null

If show returns a parse error mentioning "modified after freezing", the file was tampered with. Stop the experiment and report it — do not re-freeze over it.

Step 2: Start a session

With a session plan, kraken lab next <exp> prints the exact start command — use it verbatim; adjusting --speed per the reaction-time note below is the one permitted edit. Runs are ordinary recorded sessions stamped with --experiment — the stamp is what discovery keys on; the --label <exp>-s<n> is the human handle:

bash
# live entry (fill in the market to record):
kraken session start --label momentum-1-s1 --experiment momentum-1 \
  --symbols BTC/USD --channels ticker,trade \
  -o json 2>/dev/null &

# replay entry (the tape defines the market; 10x compresses the wait):
kraken session start --label momentum-1-s2 --experiment momentum-1 \
  --from tape:jun-crash --speed 10 -o json 2>/dev/null &

Replay reaction-time note: at speed N you have 1/N of live reaction time — suits slow-cadence strategies (DCA); drop to --speed 1 for reactive ones.

(At least one --channels entry is required on live entries, and the recorder runs in the foreground, so background it with & as the playground skill does. The session_started stdout line carries the allocated session id.)

Drive the strategy the hypothesis names (e.g. the recipe skill from --strategy), passing --reason on every trade so the scorecard's decided-fill join has data — the rationale lands in the active session's decision log automatically.

Step 3: Stop the session

A session must stop before it can be scored — the stop summary is the anchor every later number decomposes:

bash
kraken session stop
Show full SKILL.md (612 more words)Show less

Step 4: Score and judge

One command produces the honest scorecard and the mechanical verdict:

bash
kraken lab score --session latest --experiment momentum-1 -o json 2>/dev/null
  • source: what the session traded against — {kind:"replay", dataset, speed} or {kind:"live", symbols}, derived from the session's contract, not asserted.
  • anchor + components[]: the stop totals and the explain-pnl waterfall, carried verbatim — this output can never disagree with kraken explain pnl.
  • metrics: return_pct, max_drawdown, max_drawdown_pct, fees, friction, turnover_pct, hit_rate, profit_factor, slippage_sensitivity — every null is an honest "not derivable", never zero (a buys-only DCA run has no closed round trips, so hit_rate and profit_factor read null by design, not by breakage).
  • verdict.pass and verdict.checks[]: one check per frozen criterion with threshold, observed, pass. An unknown observation fails its check.
  • caveats[]: read them; a verdict with caveats is reported with them.

Digest for the loop log:

bash
kraken lab score --session latest --experiment momentum-1 -o json 2>/dev/null \
  | jq '{session, experiment, pass: .verdict.pass,
         checks: [.verdict.checks[] | {criterion, threshold, observed, pass}],
         caveats}'

The retry cues are in the error envelopes themselves: "still recording — stop it with 'kraken session stop' first" means go back to Step 3; "was aborted before it stopped … only a stopped session can be scored" means that session is dead evidence — start a fresh one; "experiment '…' not found; run 'kraken lab new …' first" means Step 1 never happened in this scope.

Step 5: Conclude — or run again

The verdict is mechanical; the conclusion is yours to report, not adjust:

  • pass: state the hypothesis, the checks that carried it, and every caveat. One passing run is evidence, not proof — follow the plan to its end before promoting anything.
  • fail: name the failing checks (threshold vs observed) and what the waterfall says the money actually went to (fees? drawdown? never traded?). A failed experiment is a completed experiment.

Promotion: when lab next says plan_complete, assemble the promotion evidence block from kraken lab compare and hand off to the kraken-paper-to-live skill. Its gate is mechanical: at least 2 passing runs with a passing live run among them — replay verdicts count only with their lookahead caveat attached. Two replay passes never promote.

More runs: repeat Steps 2–4 (each start allocates the next ordinal), then read every session side by side:

bash
kraken lab compare momentum-1 -o json 2>/dev/null

One verdict column per session (sessions[].verdict), plus pass_count/total and a named error column for any run that cannot be scored (a still-running run says so — stop it first). Only the verdicts compare across runs: they cover different market windows, so the metric rows are per-session context. Never average scorecards across runs into a blended number the CLI did not produce — compare deliberately never aggregates.

Cold-context verdict reproduction

Any context — including one that never saw the sessions — reproduces a verdict from disk alone:

bash
kraken lab show momentum-1 -o json 2>/dev/null          # the sealed criteria
kraken lab score --session s1 --experiment momentum-1 -o json 2>/dev/null

Same files, same exact-Decimal math, same verdict, byte for byte. If the two disagree with a previously reported verdict, something on disk changed — lab show tells you whether it was the spec.

Running under /loop

Invoke as /loop with this skill and the experiment name. Each firing:

  1. kraken lab show <exp> — missing → freeze (Step 1, with --session entries) and return.
  2. kraken lab next <exp> -o json and act on its state. If it refuses with a no-run-plan validation error (an experiment frozen before run plans existed), drive the manual state table above instead. start → run its command; running → drive/wait, then stop + score (Steps 3–4) and log the digest; plan_complete → kraken lab compare, conclude (Step 5), and stop the loop with the per-session verdicts (and the promotion evidence block, if concluding toward promotion) as the final report.

Idempotent by construction: a firing that finds nothing to advance does nothing. A killed loop resumes at the same step.

If you hit a mismatch between what you are trying to do and the CLI's interface or responses — including a mismatch between this skill and the installed CLI version's contract — feel free to submit feedback with kraken feedback.

© krakenfx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/kraken-lab-experiment of krakenfx/kraken-cli.

Open the folder on GitHubat commit aa56e59

Compare with similar skills

Kraken Lab Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kraken Lab Experiment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kraken Lab Experiment this skillkrakenfx/kraken-cli751—~2.7kAutomated safety check: PassMIT
Show Me Your Work Decision Logcursor/plugins11k8 repos~1.6kAutomated safety check: PassNone
Autoresearch Iteration Loopuditgoenka/autoresearch6.5k1 repos~2kAutomated safety check: PassMIT
Install Loop Engineeringcobusgreyling/loop-engineering11k1 repos~648Automated safety check: PassMIT
LoopyForward-Future/loopy3.2k—~3.9kAutomated safety check: PassMIT
AI Performance Improvement Plantanweai/pua20k2 repos~6.9kAutomated safety check: PassMIT

Similar skills

  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    11k GitHub starsUsed in 8 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • Autoresearch Iteration Loop

    uditgoenka/autoresearch

    Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.

    6.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Install Loop Engineering

    cobusgreyling/loop-engineering

    Installs Loop Engineering into a project through the single @cobusgreyling/loop CLI, scaffolding a report-only loop and a readiness score.

    11k GitHub starsUsed in 1 repo~648 tokens
    Agent WorkflowsAuto-check passed
  • Loopy

    Forward-Future/loopy

    Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication.

    3.2k GitHub stars~3.9k tokensUpdated 29 days ago
    Agent WorkflowsAuto-check passed
  • Pushes an agent to exhaust every option, investigate before asking and take initiative beyond the literal request, instead of giving up or waiting passively.

    20k GitHub starsUsed in 2 repos~6.9k tokens
    Agent WorkflowsAuto-check passed
  • LoopX Self Repair

    loopx-project/loopx

    Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.

    6.2k GitHub stars~2.2k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from krakenfx/kraken-cli

All 59 skills in this repo
  • Kraken Alert Patterns

    krakenfx/kraken-cli

    Price alerts, threshold monitoring, and notification triggers for agents.

    751 GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Kraken Autonomy Levels

    krakenfx/kraken-cli

    Progress from manual trading to full agent autonomy with controlled risk at each level.

    751 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Kraken Basis Trading

    krakenfx/kraken-cli

    Capture the spot-futures price spread with delta-neutral basis trades.

    751 GitHub stars~802 tokensUpdated 2 mo ago
    Auto-check passed
  • Kraken Dca Strategy

    krakenfx/kraken-cli

    Dollar cost averaging with scheduled buys and performance tracking.

    751 GitHub stars~1k tokensUpdated 2 mo ago
    Auto-check passed
  • Kraken Earn Staking

    krakenfx/kraken-cli

    Discover staking strategies, allocate funds, and track earn positions.

    751 GitHub stars~1k tokensUpdated 2 mo ago
    Auto-check passed
  • Kraken Error Recovery

    krakenfx/kraken-cli

    Handle order failures, network errors, and duplicate submissions safely.

    751 GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Kraken Lab Experiment

What does Kraken Lab Experiment do?

Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read…. Kraken Lab Experiment is an agent skill from krakenfx/kraken-cli. Run the Autoresearch Lab lifecycle end-to-end: freeze a hypothesis with mechanical success criteria and a sealed session plan, drive recorded and replayed paper runs, score them honestly, and read the verdict — resumable from disk at every step, designed for /loop.

When should I use Kraken Lab Experiment?

Kraken Lab Experiment fits situations like: tasks that involve Autonomous loops.

How do I install Kraken Lab Experiment in Claude Code?

Run `npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a claude-code`. Or copy the skill folder (skills/kraken-lab-experiment in krakenfx/kraken-cli) into .claude/skills/kraken-lab-experiment in your project. Claude Code loads it when a task matches its description.

How do I install Kraken Lab Experiment in Codex?

Run `npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a codex`. Or copy the skill folder (skills/kraken-lab-experiment in krakenfx/kraken-cli) into .agents/skills/kraken-lab-experiment in your project. Codex loads it when a task matches its description.

Can I use Kraken Lab Experiment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add krakenfx/kraken-cli --skill kraken-lab-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kraken-lab-experiment, .gemini/skills/kraken-lab-experiment, .github/skills/kraken-lab-experiment and .opencode/skills/kraken-lab-experiment in your project.

What does Kraken Lab Experiment need to run?

Going by SKILL.md and its folder, Kraken Lab Experiment needs the command-line tools its instructions call (jq).

Does Kraken Lab Experiment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Kraken Lab Experiment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kraken Lab Experiment use?

Kraken Lab Experiment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kraken Lab Experiment use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kraken Lab Experiment?

Skills that share tags, products or a category with Kraken Lab Experiment: Show Me Your Work Decision Log (cursor/plugins, 11k stars), Autoresearch Iteration Loop (uditgoenka/autoresearch, 6.5k stars), Install Loop Engineering (cobusgreyling/loop-engineering, 11k stars) and Loopy (Forward-Future/loopy, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kraken Lab Experiment?

krakenfx (a GitHub organization) maintains it in krakenfx/kraken-cli, which has 751 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on August 7, 2026.

Source: krakenfx/kraken-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.