Agent skill

Downtime Recovery

by Towow-ai in Towow-ai/Flowness

停机后复工的标准安全流程(水位线追平/积压泄流/服务分批重启)。当系统经历过 daemon 停机、性能冲刺减负、事故停摆之后要恢复常驻服务时触发;即使 owner 只说"把服务开回来"、"复工"、"追平水位线",也应触发。核心使命:绝不让"重启"变成"积压喷发"(2026-07-04 实锤:orchestrator 停机后水位线落后 3240 条,直接重启把机器负载打到 22+,owner…

Apache-2.0Auto-check passedAgent Workflows

Install Downtime Recovery

skills CLI
$ npx skills add Towow-ai/Flowness --skill downtime-recovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Towow-ai/Flowness downtime-recovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Towow-ai/Flowness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/downtime-recovery .claude/skills/downtime-recovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
downtime-recovery
GitHub stars
107
Token cost
~1.1k tokens
SKILL.md length
319 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
Apache-2.0

At a glance

停机后复工的标准安全流程(水位线追平/积压泄流/服务分批重启)。当系统经历过 daemon 停机、性能冲刺减负、事故停摆之后要恢复常驻服务时触发;即使 owner 只说"把服务开回来"、"复工"、"追平水位线",也应触发。核心使命:绝不让"重启"变成"积压喷发"(2026-07-04 实锤:orchestrator 停机后水位线落后 3240 条,直接重启把机器负载打到 22+,owner…

  • Works in 4 steps: substrate-monitor(只读巡检)→ 起后看一轮日志有 ⚑… → feishu-gateway /… → wake-watcher + watchdog /… → …
  • Agent Workflows work in your project
  • SKILL.md covers 我是谁, 我相信什么, 怎么做(六步 SOP) and 我不做什么, plus 1 more section
  • Calls python3

What it does

Downtime Recovery is an agent skill from Towow-ai/Flowness. 停机后复工的标准安全流程(水位线追平/积压泄流/服务分批重启)。当系统经历过 daemon 停机、性能冲刺减负、事故停摆之后要恢复常驻服务时触发;即使 owner 只说"把服务开回来"、"复工"、"追平水位线",也应触发。核心使命:绝不让"重启"变成"积压喷发"(2026-07-04 实锤:orchestrator 停机后水位线落后 3240 条,直接重启把机器负载打到 22+,owner 被迫立即紧急关停)。

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows. The repository describes itself as: A work-centered runtime for agentic software engineering. Work persists; agents, context, and graphs assemble around it. The licence is Apache-2.0.

When your agent uses it

  • Agent Workflows work in your project

Example prompts

  • “把服务开回来”
  • “,也应触发。核心使命:绝不让”
  • “/downtime-recovery”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. substrate-monitor(只读巡检)→ 起后看一轮日志有 ⚑ scanned 输出。
  2. feishu-gateway / feishu-adapter(事件驱动,若停过)→ 起后 launchctl 有 PID。
  3. wake-watcher + watchdog / session-reattach(会动手但有多重刹车)→ 起后各看一轮扫描日志无异常动作。
  4. orchestrator-daemon(最后,且必须先完成第 4 步的水位线决策)→ 起后头 5 分钟盯 orchestrator status 与负载;派发数异常上冲立即 pkill -9 -f "towow.cli.main orchestrator"(连它派出的…

What it can do on your machine

Read from SKILL.md and the folder at commit c9d6abe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Downtime Recovery loads about 1.1k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 319 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Towow-ai/Flowness at commit c9d6abe, republished under its Apache-2.0 licence (© Towow-ai). 319 words, ~1,123 tokens.

Download SKILL.mdSave it as .claude/skills/downtime-recovery/SKILL.md (or your agent's skills folder).
name
downtime-recovery
description
停机后复工的标准安全流程(水位线追平/积压泄流/服务分批重启)。当系统经历过 daemon 停机、性能冲刺减负、事故停摆之后要恢复常驻服务时触发;即使 owner 只说"把服务开回来"、"复工"、"追平水位线",也应触发。核心使命:绝不让"重启"变成"积压喷发"(2026-07-04 实锤:orchestrator 停机后水位线落后 3240 条,直接重启把机器负载打到 22+,owner 被迫立即紧急关停)。

停机复工(downtime-recovery)

我是谁

我是停机与复工之间的那道闸。系统停机(daemon 停摆/冲刺减负/事故)期间,账本还在被活跃会话写入、任务还在完工、升级还在堆积——停机不是暂停,是欠债。我的职责是:复工时先把债盘清楚,再按安全顺序泄流,让每个服务"温和接管"而不是"积压喷发"。

我相信什么

重启不是恢复,是一次放大器。 停机越久,水位线落得越远;直接重启,daemon 会把整段积压一口气处理掉——并发派发、并发抢锁、并发起会话,停机多久就炸多大(实锤:落后 3240 条 → load 22+ 秒炸;现在见过落后 22 万条的)。

先测量,后动手。 复工的第一动作永远是跑盘点(只读、零副作用),不是 launchctl load。不知道欠了多少债就还债,是赌博。

一次一个,动完核实。 每启动一个服务,先确认它的行为正常(日志/内存/负载)再动下一个。资源门看真信号:memory_pressure 等级 = normal 且换页速率 ≈0 且 load < 核数×1.5 才继续;不达标就停下等,不硬上。(macOS 的"空闲内存 MB"和 swap 占用量都是误导指标,别拿它们做依据——2026-07-05 内存诊断的教训。)

积压的处理是决策,不是默认。 22 万条积压里多数派发早已过期(对应任务可能已被别人做完)。"补派积压"还是"快进放弃、只管新事件",是要看着报告做的判断——常常值得派一个排序 agent 专门研究,偶尔需要 owner 拍板(放弃积压=有任务永不被自动派发,属范围决定)。

修好的代码不会自动进正在跑的旧进程。(2026-07-04 教训)性能修复合并后,停机前启动的常驻进程内存里还是旧代码——复工清单必须包含"识别在跑旧代码的进程并重启之",否则修了白修。

怎么做(六步 SOP)

第 1 步 · 盘点(只读)
bash
python3 .claude/downtime-recovery/survey.py          # 人读(从仓库根运行)
python3 .claude/downtime-recovery/survey.py --json   # 喂排序 agent(从仓库根运行)

产出六组数字:A 各水位线落后量 / B 积压队列 / C 服务实况 vs 登记态 / D 卫生(stale 锁、死 pid 持有 commit.lock)/ E 在跑旧代码的进程 / F 当前资源水位。

(登记:survey.py 是单点脚本——只 track 在仓库根 .claude/downtime-recovery/,不随任何 skill 部署面分发;本 SKILL 三面与 RUNNING-SERVICES.md 顶部指针都指这同一个绝对路径。要挪它必须所有指针同改,owner 选向之前只登记、不挪。)

第 2 步 · 分诊排序(默认派 agent)

把 --json 报告喂给一个排序 agent(Sonnet 够用),让它交回:安全复工顺序 + 每步理由 + 每步的验证方式 + 需要 owner 拍板的点。派发信要点:报告全文 + 本 SKILL 的硬规则 + RUNNING-SERVICES.md 各服务登记(暂停原因/恢复命令/踩坑记录,尤其 §4 orchestrator 的 2026-07-04 积压喷发记录)。积压小(各水位线 behind < 几百、队列个位数)时可弃权自己排,但弃权留账:在第 6 步的复工记录里写明理由与当时的积压数字(埋掉的是一个独立的排序视角)。

第 3 步 · 卫生先行(低风险、腾地方)
  • 死 pid 文件:python 内部 unlink(先 ps 验证进程真死;绝不用 shell rm 碰 .towow 路径——guard 会拦且拦得对)。
  • stale 会话锁:用 ./tw plan/goal reap-stale-session 系列(vitality 裁决),不手删。
  • commit.lock 被死进程持有:新提交会经"诚实 holder 探测"(commit 50555a102)识破,一般无需手动;若探测未上线到在跑进程,按 RUNNING-SERVICES 处置。
  • 在跑旧代码的进程(报告 E 组):逐个按其登记的停/起方式重启(launchd 的 kickstart -k;非 launchd 的按登记命令)。
第 4 步 · 水位线决策(本 SOP 的核心判断)

⚠ 先记住一个代码实锤(orchestrator.py 主循环的 E.5 安全暂停块,搜 is_orchestrator_paused):paused 状态下 daemon 不 scan、不派发、也不推进水位线,纯 idle——不存在"暂停着慢慢追平"这条路。真实可走的是:

  • 快进放弃(fast-forward,大积压默认):daemon 停着时把 watermark.json 直接写到账本头(原子写,格式同 save_watermark_atomic)+ 留痕。已派未完的任务不会丢(dispatched/ 标记有独立于水位线的重扫通道);丢的是积压区间里"从未派出的新触发"。⚠ 范围决定,需 owner 点头。
  • 离线错过清单 + 手动补派(快进的兜底,两者配合用):派一次性 agent/脚本只读扫 [watermark, head] 区间,只提取会触发派发的事件类型,产出"错过的触发清单"给人审;重要的用 orchestrator dispatch-one --while-paused 逐个补派(这是官方支持的暂停期单发通道,orchestrator.py dispatch_one_task 的裁决注释明说"暂停期手动单发恰恰是合法核心场景")。补完再快进水位线、起 daemon——此时积压=0,温和接管。
  • 直接重启硬追(仅小积压):积压 < 几百条且机器资源门全绿时,直接起 daemon 让它自己扫完。超过千条禁用——就是 2026-07-04 的炸法。

投影水位线(graph/.cursor)例外:它每次提交都自动追平且已增量化(2026-07-04 批次合并修复),落后大时跑一次任意只读 CLI 即追平,无需决策。

第 5 步 · 服务分批重启(顺序固定,每步验证)

按"只读→事件驱动→会动手"的风险递增序:

  1. substrate-monitor(只读巡检)→ 起后看一轮日志有 ⚑ scanned 输出。
  2. feishu-gateway / feishu-adapter(事件驱动,若停过)→ 起后 launchctl 有 PID。
  3. wake-watcher + watchdog / session-reattach(会动手但有多重刹车)→ 起后各看一轮扫描日志无异常动作。
  4. orchestrator-daemon(最后,且必须先完成第 4 步的水位线决策)→ 起后头 5 分钟盯 orchestrator status 与负载;派发数异常上冲立即 pkill -9 -f "towow.cli.main orchestrator"(连它派出的 dispatch 残留一起清,见 RUNNING-SERVICES §4 关停手法)。

每步之间:重跑一次 survey.py 看 F 组(memory_pressure / 换页速率 / load),不达标就等;服务行为异常就停在这一步排查,不带病继续。

第 5.5 步 · 审计 resume 保留的延迟派发队列(2026-07-11 实锤补步)

resume 的 T-FIX-COST-DROP 守卫会保留"暂停前已 deferred 的真 spawn 决策"(nonexec backlog marker)——但保留时不检查该触发器的阶段产出是否已被别人消化。实锤:brief-512d61a9 的共识触发器 7-3 被 deferred、当天已被另一会话冻结消化,7-10 复工重放白烧一个 Opus 共识席(debt-986ba644fc44)。复工起 daemon 之前(或之后第一时间):逐个审 orchestrator/nonexec_backlog/ 标记,按 dispatch_to 查阶段产出(consensus→已冻结?planning→PlanFreezed?fix→FindingResolved?)已存在的清掉留痕,别让 daemon 重放过期触发器。

第 6 步 · 复工验证与记录
  • 跑一次 survey 对比前后(水位线应追平/快进、卫生项清零、服务实况=登记态)。
  • session-reattach 的哨兵脉搏检查下一轮应全绿。
  • 工位积压若有完工待 promote 的,按既有 promote 机制逐个处理(参照 01-reconciliation/perf-sprint/worktree-backlog-cleanup-2026-07-04.md 的死规则:有完工证据+测试绿才晋升、冲突留人裁)。
  • RUNNING-SERVICES.md 状态行更新(🔴/🟡 → 🟢,注明复工时间与本次水位线决策)。

我不做什么

  • 不在盘点前启动任何服务(先测量后动手是底线)。
  • 不并发启动多个服务(一次一个,动完核实)。
  • 不替 owner 决定"放弃大额积压"(范围决定,surface 给她选项+推荐)。
  • 不用 shell 删除形命令碰 .towow 任何路径(python unlink + 既有 reap 机制)。

退出条件

  • 系统本就在跑、无停机债 → 不需要我,别为了走流程而走流程。
  • 单个服务的日常重启(无停机期积压)→ 直接按 RUNNING-SERVICES 该服务的登记命令即可。

© Towow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/downtime-recovery of Towow-ai/Flowness.

Open the folder on GitHubat commit c9d6abe

Compare with similar skills

Downtime Recovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Downtime Recovery compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Downtime Recovery this skillTowow-ai/Flowness107—~1.1kAutomated safety check: PassApache-2.0
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k10 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k36 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers297k2 repos~5.1kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79689 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 10 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 36 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    297k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    796 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed

More from Towow-ai/Flowness

All 12 skills in this repo
  • Dependency Analyze

    Towow-ai/Flowness

    从 task 的 read/write set + concept statemachine 推导 6 种依赖类型的提案。主 planner 决定边的真实性。派它时只给 read/write set 与疑点、不给预期边集;已有预判逐条标「待复核」交它取证。

    107 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Execution

    Towow-ai/Flowness

    M-1.4 execution skill — 跑 single task 产 patch + 提交 envelope。

    107 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Execution Self Check

    Towow-ai/Flowness

    Pre-submit 自检——envelope 提交 commit gate 前必跑。独立 OPUS fork 逐项判 blocking checks(清单以 dispatch prompt 注入为准),executor 不能 self-assess(运动员不当裁判)。

    107 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Fix

    Towow-ai/Flowness

    修复者 — 把一条被发现的问题(finding)按它的闭合合约修干净,修一个不制造下一个。产 FixProposed + 临时的 FixCompleted,不自判问题关闭(那是复查的权)。当 daemon 派一条 finding 来修、或需要闭合一个已发现的问题时用,即使只说"修一下这个 finding""把这个问题闭合"也触发。调用名就是 fix(Skill 工具)或 /fix(命令),没有…

    107 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Fix Self Check

    Towow-ai/Flowness

    M-1.6 envelope self-check——独立性保证不自欺欺人 (5 blockingcheck)。由 CLI ./tw fix complete --self-check-mode fork(默认即 fork)自动派起,不经 Skill 工具调用;fix 主会话产 FixCompleted 前直读本文,是为理解双层验证关系。

    107 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Meta Review

    Towow-ai/Flowness

    F-08g 元 review——审 reviewplan 自身够不够 (dimensions 覆盖/voi 具体/historical feed 漏)。用 named error patterns + 历史比对。design-time mode 调它审 reviewplancreator 的产出, critical meta-finding → orchestrator 回头让…

    107 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Downtime Recovery

What does Downtime Recovery do?

停机后复工的标准安全流程(水位线追平/积压泄流/服务分批重启)。当系统经历过 daemon 停机、性能冲刺减负、事故停摆之后要恢复常驻服务时触发;即使 owner 只说"把服务开回来"、"复工"、"追平水位线",也应触发。核心使命:绝不让"重启"变成"积压喷发"(2026-07-04 实锤:orchestrator 停机后水位线落后 3240 条,直接重启把机器负载打到 22+,owner…. Downtime Recovery is an agent skill from Towow-ai/Flowness.

When should I use Downtime Recovery?

Downtime Recovery fits situations like: agent Workflows work in your project.

How do I install Downtime Recovery in Claude Code?

Run `npx skills add Towow-ai/Flowness --skill downtime-recovery -a claude-code`. Or copy the skill folder (.claude/skills/downtime-recovery in Towow-ai/Flowness) into .claude/skills/downtime-recovery in your project. Claude Code loads it when a task matches its description.

How do I install Downtime Recovery in Codex?

Run `npx skills add Towow-ai/Flowness --skill downtime-recovery -a codex`. Or copy the skill folder (.claude/skills/downtime-recovery in Towow-ai/Flowness) into .agents/skills/downtime-recovery in your project. Codex loads it when a task matches its description.

Can I use Downtime Recovery in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Towow-ai/Flowness --skill downtime-recovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/downtime-recovery, .gemini/skills/downtime-recovery, .github/skills/downtime-recovery and .opencode/skills/downtime-recovery in your project.

What does Downtime Recovery need to run?

Going by SKILL.md and its folder, Downtime Recovery needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Downtime Recovery access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Downtime Recovery safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Downtime Recovery use?

Downtime Recovery is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Downtime Recovery use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Downtime Recovery?

Skills that share tags, products or a category with Downtime Recovery: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Downtime Recovery?

Towow-ai (a GitHub organization) maintains it in Towow-ai/Flowness, which has 107 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on August 8, 2026.

Source: Towow-ai/Flowness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.