Agent skill

Orca Run Replay

by iflytek in iflytek/skillhub

Answers questions about a past agent run from its recording, using causal graphs and replay, instead of reconstructing events from memory.

Apache-2.0Auto-check passedAgent Workflows

Install Orca Run Replay

skills CLI
$ npx skills add iflytek/skillhub --skill orca-replay -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install iflytek/skillhub orca-replay --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/iflytek/skillhub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/builtin-skills/skills/orca-replay .claude/skills/orca-replay && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
orca-replay
GitHub stars
5.2k
Used in
4 other repos
Token cost
~3k tokens
SKILL.md length
1,858 words
Files
3
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

Answers questions about a past agent run from its recording, using causal graphs and replay, instead of reconstructing events from memory.

  • Works in 5 steps: Find the run → Narrow to the chain that produced the… → Report recorded and inferred differently → …
  • Asking why an earlier agent run deleted, moved or overwrote a file
  • SKILL.md covers When to Use This Skill, Workflow, If there is no recording yet and Sharing a run with someone else, plus 2 more sections
  • Calls npx, npm and git

What it does

When a question is about something that already happened, such as why a file was deleted, which step broke the build or whether yesterday's failure reproduces, the agent reads the recorded trace first. It lists runs with orca_list_runs, shows the whole timeline with orca_show_run (model turns, tool calls, shell commands with exit codes and changed files), and uses orca_graph to get the causal chain that produced one event.

Edges in the graph are labeled recorded or inferred, and answers must keep that distinction and name the rule behind any inferred edge. The agent then replays the recording with orca_replay to see what diverges or could not be served, and the skill notes that replay cannot tell whether a fresh run would fail again. Forking a run to try something different is also supported. It needs the orcareplay npm package with its MCP server registered as orca, Node 20 or newer, and a recorded run under .orca/runs.

When your agent uses it

  • Asking why an earlier agent run deleted, moved or overwrote a file
  • Finding which step of a recorded run broke the build
  • Reproducing a failure from a recorded run
  • Checking whether a different model would have handled the run correctly

Example prompts

  • “Why did the last run delete the build folder? Check the recording.”
  • “Replay yesterday's failed run and tell me where it diverges.”
  • “Show the chain of events that led to the config file change.”

Requirements

  • The orcareplay npm package with its MCP server registered as orca
  • Node 20 or newer
  • At least one recorded run in .orca/runs
  • Compatibility (from SKILL.md): Requires the `orcareplay` npm package (Node 20+) with its MCP server registered as `orca`, and at least one recorded run in the project's .orca/runs directory.

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Find the run
  2. Narrow to the chain that produced the thing being asked about
  3. Report recorded and inferred differently
  4. Reproduce it before explaining it
  5. Only then consider comparing models

What it can do on your machine

Read from SKILL.md and the folder at commit 753af72. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • npm
    • git
    • tsc

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, npm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the `orcareplay` npm package (Node 20+) with its MCP server registered as `orca`, and at least one recorded run in the project's .orca/runs directory.

    From compatibility in the SKILL.md frontmatter.

Context cost

Orca Run Replay loads about 3k tokens when it runs. Until then it costs about 51 tokens; SKILL.md has 1,858 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from iflytek/skillhub at commit 753af72, republished under its Apache-2.0 licence (© iflytek). 1,858 words, ~3,048 tokens.

Download SKILL.mdSave it as .claude/skills/orca-replay/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
orca-replay
description
Answers questions about a past agent run from its recording rather than from memory, and replays or forks that run. Use when asked why an earlier run did something, or to reproduce a failure.
compatibility
Requires the `orcareplay` npm package (Node 20+) with its MCP server registered as `orca`, and at least one recorded run in the project's .orca/runs directory.
version
1.0.0
license
Apache-2.0
metadata.author
Continuum-AI-Corp
metadata.version
0.1
metadata.homepage
https://github.com/Continuum-AI-Corp/OrcaReplay

Reading a recorded agent run

A recording is evidence. Your memory of a session is not, and neither is a transcript you were handed — both are missing the tool results, the exit codes, and the files that changed without anyone mentioning it.

The rule: when a question is about something that already happened, read the trace before you answer. Do not reconstruct it. If a recording exists, guessing is the wrong move even when the guess would have been right.

When to Use This Skill

  • "Why did you delete/overwrite/move X?"
  • "What changed this file?" / "Which step broke the build?"
  • "Can you reproduce yesterday's failure?"
  • "Does this still reproduce?" (see the limit on that in step 4 — replay cannot tell you whether a fresh run would fail again)
  • "Would a different model have got this right?"

Workflow

1. Find the run

orca_list_runs — newest first, and it names the run each fork came from. Skip this only when the user clearly means the most recent one; every other tool defaults to run: "last".

2. Narrow to the chain that produced the thing being asked about

orca_show_run gives the whole timeline: model turns with token counts and stop reasons, tool calls with arguments and results, shell commands with exit codes, and every file the run changed. Good for orientation, long for a specific question.

orca_graph is usually the better tool. It returns causal edges — which event produced which. Pass to: <event seq> to get only the chain that produced one event. That is the shape of an answer to "why did this happen", where the full timeline is the shape of an answer to "what happened".

3. Report recorded and inferred differently

Every edge from orca_graph is labelled:

  • recorded — the recorder watched it happen and wrote it into the trace.
  • inferred — derived just now from a rule the edge names. The trace does not vouch for it.

Carry that distinction into your answer. "The trace shows the rm at step 14 removed it" and "this looks like the rm at step 14, going by timing" are different claims, and flattening them into one confident sentence is the specific failure this tool exists to prevent. Name the rule when you lean on an inferred edge.

4. Reproduce it before explaining it

orca_replay re-runs the recording and reports what could not be reproduced — divergences, and requests the recording could not serve.

What "offline" covers, and what it does not. Every model response comes from the trace and the proxy forwards nothing upstream, so no provider is contacted and no tokens are spent. An unmatched request halts the replay rather than falling through to the network, unless --loose was asked for.

That covers the model traffic. It does not cover the agent's own subprocesses: unless the recording used --tls-intercept — in which case replay re-establishes interception for the hosts it recorded — a curl, npm install, git push or database call inside a recorded shell command goes straight out. Replay is not a sandbox; only a network-isolated container makes it one.

What a matching replay proves, and what it does not. It shows the recorded decisions reproduce against today's environment. It cannot show the failure is deterministic, because the model is not being asked again — the same recorded responses are served back. If the user wants to know whether a fresh run would fail the same way, say that replay cannot answer it; that needs real runs.

Replay re-executes the agent, not just its model traffic. The recorded model responses are served from the trace, but the agent process runs again for real — so every shell command it issued runs again too. worktree: true isolates repository files and nothing else. Anything the run touched outside the tree — /tmp, Docker, a local database, a package manager, another host — is mutated a second time.

Hard gate before every replay. Before calling orca_replay, the agent MUST use orca_show_run (or an equivalent trace view) to enumerate the complete shell-command list, including commands that may touch /tmp, Docker, databases, package managers, or remote hosts. It MUST show that list to the user and obtain explicit approval for the exact replay. If any command reaches outside the worktree, approval MUST name those external effects or the replay MUST run inside a genuinely isolated container. worktree: true protects repository files only; it does not authorize external side effects. Do not infer approval from silence, a previous approval, or the fact that the original run was recorded. A run that only read files and edited the repository is free and repeatable, but it still requires this preview-and-confirm gate.

Pass worktree: true. It replays into a scratch copy and leaves the working tree alone.

Without it, replay is destructive for as long as it runs: it restores the recorded filesystem over the working tree and puts the tree back when the replay ends. Uncommitted work is absent in the meantime, and stays absent if the replay is interrupted before it can restore. Run an in-place replay only when the user has been told that and has agreed to it. "They do not appear to be typing" is not consent.

A replay reporting reused=3/5 on an interactive recording is not a partial failure. Harnesses make calls for themselves — a quota probe, a session-naming request — and a replay does not repeat them.

5. Only then consider comparing models

orca_compare forks one run onto several models from the same checkpoint: same files, same conversation prefix, so the model is the only variable. Pick the fork point with orca_checkpoints and pass it as from.

Grade with verify — a shell command whose exit code is the verdict. Use something the repository already declares ("npm test", "npm run typecheck") or an explicitly local binary ("./node_modules/.bin/tsc --noEmit"), not npx <tool>: with no local install, npx runs whatever the registry has under that name, and npx tsc resolves a package deprecated in 2016 that is not TypeScript.

orca_compare uploads the recording to other people's models, and spends real money doing it. Each model named receives the same files and conversation prefix the original run had — so whatever that run touched (source, prompts, configuration, anything a credential was pasted into) is sent to every provider behind those model ids.

And each fork is a live agent, not a replay. From the fork point onward the model is really being asked, and whatever it decides to do, it does — its shell commands execute for real, and so does the verify command you pass. Each fork gets its own worktree, so repository files are isolated per model; nothing outside the tree is. A fork can also take actions the original run never took, because it is a different model making fresh decisions.

So the approval has three parts, and they are not the same question:

  1. Disclosure — what context is uploaded, and to which providers. Approving a bill is not approving a disclosure, and the two need separate answers when the recording is from a private codebase. orca scrub is for when the comparison is worth running but the trace is not safe to send as-is.
  2. Side effects — what the recorded run did outside its worktree, since each fork may repeat it and may go further. Same check as step 4, orca_show_run, and the same answer if it reached Docker, a database, a deployment or another host: get approval for that specifically, or run the comparison in an isolated environment.
  3. Cost — how many models times how many forks.

Never run it to satisfy curiosity the user did not express.

Show full SKILL.md (605 more words)Show less

If there is no recording yet

Say so plainly rather than falling back to guessing, and offer to start one.

If orca is already installed:

console
orca record claude           # or codex, opencode, openclaw, grok

If it is not installed, stop and ask the user to install it separately. Do not install packages, change global state, or use a package-manager command as part of this Skill.

orca record <agent> runs the agent unmodified behind a local proxy. Nothing about the agent changes; two environment variables get set. Recording a session now is what makes the next "why did it do that" answerable.

For a run started with a prompt in argv — orca record claude -- -p "…" — the replay is exact. A session someone typed into replays approximately, because the prompts were never on the wire and are recovered from the harness's own transcript; orca replay says which is which rather than papering over it.

Sharing a run with someone else

orca export last -o run.html writes one self-contained file. A trace holds whatever the run held, so run orca scrub before sending one anywhere.

Scrubbing is best-effort, not a guarantee. It matches known key shapes and high-entropy strings; it cannot know that a particular internal hostname, customer name, or unreleased feature is confidential to this user. So scrub, then have the user look at what is actually going out, and get their agreement — do not describe a scrubbed trace as safe on the strength of the scrubber alone.

Limitations

  • It only sees what was recorded. Runs started without orca record leave no trace, and nothing here recovers them. The answer to "why did it do that" in an unrecorded session is honestly "there is no recording", not a reconstruction.
  • A typed session replays approximately, not exactly. Prompts entered at a terminal were never on the wire; orca recovers them from the harness's own transcript. Only a run started with the prompt in argv (orca record claude -- -p "…") replays byte-for-byte.
  • Some turns are not repeated. A harness makes calls for itself — a quota probe, a session-naming request — and a replay steps over them. Tools that need a person (AskUserQuestion, plan mode) are absent when the same agent runs without one, which can make a replayed request differ from the recorded one by enough to halt.
  • inferred edges are not evidence. They are derived from a named rule at query time. Treat them as a reading of the trace, never as something the recorder witnessed.
  • Not every harness is recordable. Agents that read no base-URL variable and pin their own origin need --tls-intercept, and some cannot be reached at all. A recording that came back empty means the harness was not captured, not that nothing happened.
  • Replay is not a time machine, and not a sandbox. It reproduces the agent's side of the run against today's world. External state the run depended on — a database row, a remote branch, the clock — is whatever it is now, and the run's own shell commands reach it for real.
  • A matching replay is not a determinism result. The model is not re-asked; its recorded responses are served back. Whether a fresh run would fail the same way is a different question that replay cannot answer.

Tools

toolargumentsnotes
orca_list_runs—newest first, names the parent of each fork
orca_show_runrunthe full timeline
orca_checkpointsrunwhere a fork can start
orca_graphrun, tocausal edges; to narrows to one chain
orca_replayrun, worktreeoffline, free, repeatable
orca_comparerun, models*, from, verifyspends real tokens

run accepts a run id or "last", and defaults to "last". Replay traces are skipped when resolving "last", so it means the newest run you actually recorded.

© iflytek, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in builtin-skills/skills/orca-replay of iflytek/skillhub.

  • SKILL.md
  • LICENSE.txt
  • NOTICE.md

Open the folder on GitHubat commit 753af72

Used in 4 other repositories

We found 8 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in iflytek/skillhub, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Orca Run Replay next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Orca Run Replay compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Orca Run Replay this skilliflytek/skillhub5.2k4 repos~3kAutomated safety check: PassApache-2.0
Octocode Code Researchbgauryy/octocode949—~1.5kAutomated safety check: PassMIT
Flowstudio Power Automate Debuggithub/awesome-copilot40k2 repos~5kAutomated safety check: PassMIT
QA Find Bugs MCPbex-co/beancount-io296—~3kAutomated safety check: PassMIT
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT
Debugging and Error Recoveryaddyosmani/agent-skills103k1 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • Octocode Code Research

    bgauryy/octocode

    Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label.

    949 GitHub stars~1.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Flowstudio Power Automate Debug

    github/awesome-copilot

    Official

    Debug failing Power Automate cloud flows using the FlowStudio MCP server.

    40k GitHub starsUsed in 2 repos~5k tokens
    DevelopmentAuto-check passed
  • QA Find Bugs MCP

    bex-co/beancount-io

    Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and…

    296 GitHub stars~3k tokensUpdated today
    Backend & APIsAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    103k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed
  • Debug

    agentic-community/mcp-gateway-registry

    Debug issues in the MCP Gateway Registry using first-principles thinking.

    967 GitHub stars~1.8k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes

More from iflytek/skillhub

All 29 skills in this repo
  • Zero Slop Prose Editor

    iflytek/skillhub

    Audits and rewrites formulaic, AI-sounding prose while keeping facts, voice and format, using a local Python scorer and inspect-only, rewrite or embedded-gate modes.

    5.2k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sandbase

    iflytek/skillhub

    Access 2,000+ AI models and API tools through one MCP interface for inference, media generation, search, scraping, embeddings, social data, and structured retrieval.

    5.2k GitHub starsUsed in 2 repos~2.1k tokens
    Auto-check passed
  • LinkedIn Post Formatter

    iflytek/skillhub

    Drafts a copy-paste-ready LinkedIn post from your facts and ideas, choosing the smallest structure that fits and keeping an accessible plain-text fallback for any styled text.

    5.2k GitHub stars~901 tokensUpdated today
    Auto-check passed
  • SkillHub CLI

    iflytek/skillhub

    Connects an agent to a SkillHub registry and uses the official SkillHub CLI to search, install, list and explicitly upgrade skills from that registry.

    5.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • AI Claim Checker

    iflytek/skillhub

    Breaks AI-generated text into checkable claims, verifies them against independent sources and labels each one, with an optional exercise for learners.

    5.2k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Produces short daily standups, evening reflections and weekly retrospectives for one person or a small team, kept in the session unless you name a place to save.

    5.2k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Orca Run Replay

What does Orca Run Replay do?

Answers questions about a past agent run from its recording, using causal graphs and replay, instead of reconstructing events from memory. When a question is about something that already happened, such as why a file was deleted, which step broke the build or whether yesterday's failure reproduces, the agent reads the recorded trace first. It lists runs with orca_list_runs, shows the whole timeline with orca_show_run (model turns, tool calls, shell commands with exit codes and changed files), and uses orca_graph to get the causal chain that produced one event.

When should I use Orca Run Replay?

Orca Run Replay fits situations like: asking why an earlier agent run deleted, moved or overwrote a file; finding which step of a recorded run broke the build; reproducing a failure from a recorded run; checking whether a different model would have handled the run correctly.

How do I install Orca Run Replay in Claude Code?

Run `npx skills add iflytek/skillhub --skill orca-replay -a claude-code`. Or copy the skill folder (builtin-skills/skills/orca-replay in iflytek/skillhub) into .claude/skills/orca-replay in your project. Claude Code loads it when a task matches its description.

How do I install Orca Run Replay in Codex?

Run `npx skills add iflytek/skillhub --skill orca-replay -a codex`. Or copy the skill folder (builtin-skills/skills/orca-replay in iflytek/skillhub) into .agents/skills/orca-replay in your project. Codex loads it when a task matches its description.

Can I use Orca Run Replay in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add iflytek/skillhub --skill orca-replay -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/orca-replay, .gemini/skills/orca-replay, .github/skills/orca-replay and .opencode/skills/orca-replay in your project.

What does Orca Run Replay need to run?

Going by SKILL.md and its folder, Orca Run Replay needs the command-line tools its instructions call (npx, npm, git and tsc). Our summary lists: The orcareplay npm package with its MCP server registered as orca; Node 20 or newer; At least one recorded run in .orca/runs. Compatibility (from SKILL.md): Requires the `orcareplay` npm package (Node 20+) with its MCP server registered as `orca`, and at least one recorded run in the project's .orca/runs directory..

Does Orca Run Replay access the network?

SKILL.md contains no URLs. Its commands use npx, npm and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Orca Run Replay safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Orca Run Replay use?

Orca Run Replay is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Orca Run Replay use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Orca Run Replay?

Skills that share tags, products or a category with Orca Run Replay: Octocode Code Research (bgauryy/octocode, 949 stars), Flowstudio Power Automate Debug (github/awesome-copilot, 40k stars), QA Find Bugs MCP (bex-co/beancount-io, 296 stars) and Graph-Based Bug Tracing (tirth8205/code-review-graph, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Orca Run Replay?

iflytek (a GitHub organization) maintains it in iflytek/skillhub, which has 5,162 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 9, 2026.

Source: iflytek/skillhub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.