Agent skill

Ultrawork Mode

by code-yeongyu in code-yeongyu/lazycodex

The binding directive for ultrawork mode: evidence-first delivery, a light or heavy process tier, and constant use of memory during the work.

MITAuto-check passedAgent Workflows

Install Ultrawork Mode

skills CLI
$ npx skills add code-yeongyu/lazycodex --skill ultrawork -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install code-yeongyu/lazycodex ultrawork --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/code-yeongyu/lazycodex.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/omo/skills/ultrawork .claude/skills/ultrawork && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ultrawork
GitHub stars
3.7k
Token cost
~6.9k tokens
SKILL.md length
3,919 words
Files
2
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

The binding directive for ultrawork mode: evidence-first delivery, a light or heavy process tier, and constant use of memory during the work.

  • Works in 4 steps: Survey the skills, gather context, then… → Create the goal with binding success… → Open the durable notepad → …
  • Someone asks for ultrawork mode and the directive is not yet in the conversation
  • SKILL.md covers 0. Survey the skills, gather…, 1. Create the goal with…, 2. Open the durable notepad and 3. Register obsessive todos…
  • Calls curl, git and node

What it does

Loaded only when ultrawork mode is requested and the directive is not already in the conversation, this file is the directive itself. It opens with a fixed banner line the agent must print first and then sets an outcome-first stance: deliver exactly what was asked, working end to end, proven by captured evidence from running the changed behavior through its real surface.

A tier triage step classifies the work once as light or heavy and records the choice and its justification in a notepad. Light is the default for known patterns such as a one-spot bugfix; heavy applies to new modules, auth or security code, external integrations, schema changes, concurrency or any request that signals care, and the tier can only move upward. The directive also tells the agent to read and write memory continuously and to treat a green test suite as insufficient proof of done. A reference file, `astra-variant.md`, sits alongside it.

When your agent uses it

  • Someone asks for ultrawork mode and the directive is not yet in the conversation
  • Starting a task that must be proven with real run evidence, not only passing tests
  • Sizing the process for a change as light or heavy before planning it

Example prompts

  • “Turn on ultrawork mode and fix the failing login redirect.”
  • “Use ultrawork to add a validation rule to the checkout form and prove it works.”
  • “Run this database migration in ultrawork mode and record every regression a check caught.”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Survey the skills, gather context, then size the work
  2. Create the goal with binding success criteria
  3. Open the durable notepad
  4. Register obsessive todos via update_plan

What it can do on your machine

Read from SKILL.md and the folder at commit 36ba46a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • git
    • node
    • rg
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, git and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ultrawork Mode loads about 6.9k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 3,919 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from code-yeongyu/lazycodex at commit 36ba46a, republished under its MIT licence (© code-yeongyu). 3,919 words, ~6,936 tokens.

Download SKILL.mdSave it as .claude/skills/ultrawork/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ultrawork
description
Binding ultrawork mode directive for omo on Codex. When a prompt contains ultrawork or ulw, the omo UserPromptSubmit hook injects a short bootstrap that points at this file. Read the whole file and follow every rule in it for the rest of the task.
metadata.short-description
Binding ultrawork mode directive
<ultrawork-mode>

MANDATORY: First user-visible line this turn MUST be exactly: ULTRAWORK MODE ENABLED!

[CODE RED] Maximum precision. Outcome-first. Evidence-driven.

Role

Expert coding agent. Ship verified work; report at handoffs, not between them.

Goal

Deliver EXACTLY what the user asked, end-to-end working, proven by captured evidence: the changed behavior RUN through its real surface, sized by the tier below, with the tests the repository keeps for it still green. TESTS ALONE NEVER PROVE DONE — a green suite means the unit-level contract holds, not that the user-facing behavior works.

Tier triage (classify ONCE at bootstrap; record tier + one-line

justification in the notepad; ratchet up only) Your change set is what THIS session will itself edit or execute; work handed to another session, thread, or delegated loop is payload and sizes THAT session's process, not yours. Launching it — sync, prompt, create, verify — is control-plane work: LIGHT however large the delegated project is. Default is LIGHT. Take HEAVY only when the change set hits a fact you can point to: a new module / layer / domain model / abstraction; auth, security, session-handling code, or permissions; building or changing an external integration (API, queue, payment, webhook) — calling an existing API is not one; a DB schema or migration; concurrency, transaction boundaries, or cache invalidation; a refactor crossing domain boundaries; or the user signaled care ("carefully", "thoroughly", "design first") or demanded review of this session's work. When unsure, take HEAVY. If a HEAVY fact surfaces mid-task, upgrade immediately and redo whatever the LIGHT path skipped; never downgrade mid-task. The tier sizes process, never honesty: both tiers capture evidence, record cleanup receipts, and obey the never-suppress rules.

LIGHT — the deliverable follows a known pattern with no open design decisions (one-spot bugfix, an endpoint following an existing pattern, a validation rule, a query tweak, copy/constants, launching or steering another session): plan directly in the notepad; 1-2 success criteria (happy path + the riskiest edge); one real-surface proof of the user-visible deliverable, where auxiliary surfaces are first-class for CLI- or data-shaped work; self-review recorded in the notepad instead of the reviewer loop. HEAVY — anything a fact above names: 3+ success criteria (happy, edge, regression, adversarial risk), each with its own channel scenario and both evidence pieces; when the verification gate is triggered, run the reviewer loop until unconditional approval.

Manual-QA channels

Run real-surface proof yourself through the channel that faithfully exercises the surface; capture the artifact.

  1. HTTP call — hit the live endpoint with curl -i (or an HTTP client from js eval); capture status line + headers + body.
  2. Terminal / TUI - drive a real pty and prove it through the xterm.js web terminal (see the TUI visual QA note below). tmux send-keys is fine for a boot smoke; NEVER tmux capture-pane for color / layout / CJK evidence, which degrades truecolor.
  3. Browser use — in Codex, use browser:control-in-app-browser first when available and no authenticated/persistent user browser profile is required. Otherwise drive the page with omowright (staged in the browser skill; load it through that skill's scripts/omowright.mjs from js eval): the owned engine (connectPipe on a task-owned profile, connectCloakProfile for bot-scored targets), or the attached engine (connectBrowserSkill() in the user's own signed-in browser) when the page needs their login. Capture action log + screenshot path. Never downgrade to a non-browser surface for a browser-facing criterion, and never launch a headless browser because the attached one is missing — run the browser skill's onboarding script and relay its one human step. NEVER clear cookies, cache, or site data (Network.clearBrowserCookies, Storage.clearCookies, chrome.browsingData.remove, "clear browsing data") on the user's real/main browser profile, and never clone it — it wipes or invalidates their logged-in state.
  4. Computer use — when the surface is a desktop/GUI app rather than a page, drive it via OS-level automation (a computer-use agent, AppleScript, xdotool, etc.) against the running app; capture action log + screenshot. USE THIS for any non-browser GUI criterion; do not substitute a CLI dump for it.

For EVERY scenario name the exact tool and the exact invocation upfront: the literal command / API call / page action with its concrete inputs (URL, payload, keystrokes, selectors) and the single binary observable that decides PASS vs FAIL. "run the endpoint", "open the page", "check it works" are NOT scenarios — write the curl ..., the send-keys ..., the Browser plugin action, the page.click(...), the expected status/text.

Auxiliary surfaces (CLI stdout / DB state diff / parsed config dump) are first-class evidence for CLI- or data-shaped criteria; use a channel scenario when the behavior is user-facing. --dry-run, printing the command, "should respond", and "looks correct" never count.

For TUI visual QA, render the terminal through the real xterm.js web terminal and screenshot it - never a tmux capture-pane dump, which degrades color and wide-glyph width. In this repo: node script/qa/web-terminal-visual-qa.mjs --title "<surface>" --command "<cmd>" --input "{Enter}" --evidence-dir <dir> (live pty + xterm.js in Chrome; --from-file <capture> replays a raw stream). Outside this repo, capture equivalent browser-rendered terminal evidence: screenshot + plain transcript + cleanup receipt.

Bootstrap (DO ALL FOUR BEFORE ANY OTHER WORK — NO SKIPPING)

0. Survey the skills, gather context, then size the work

First, survey the loaded skill list and read the description of each loosely relevant skill. Decide explicitly which skills this task will use and prefer using every genuinely applicable one — name them in the notepad with a one-line reason each. Skipping a skill that fits the task is a defect. Open a skill's body only when THIS session will execute its workflow; skills a delegated session needs are named in its prompt and read there, not here. Next, fire the first discovery wave under Finding things below. Then run Tier triage (above) on the change set and record the tier — tier sizes evidence and review, never who plans. Size planning by what the wave left UNDECIDED, not by how many steps you can list: spawn the plan agent only when open design decisions remain — unclear module boundaries, several viable decompositions, or a multi-file build whose dependency order is not obvious — pass it the gathered findings (file:line facts, constraints, unknowns), and follow its wave order, parallel grouping, and verification exactly. A known procedure — however many steps — and questions about work you are delegating never justify a planner: plan directly in the notepad. Never spawn plan before the discovery wave has returned.

1. Create the goal with binding success criteria

You MUST register the goal with the create_goal tool — NOT prose, NOT the notepad, NOT the plan: the registered goal is the binding contract for the whole run, and skipping it is a defect. Call it with exactly objective; do not include status. Only when no goal tool exists on this surface, open your reply with a # Goal block treated as binding. Goals are unlimited; never invent a numeric budget or limit. Check get_goal first: continue a matching active goal instead of duplicating one; surface a conflicting one. Write the objective outcome-first: the concrete thing that will be TRUE when done (an outcome, never an activity), the named deliverable surfaces, and explicit scope bounds — a vague objective produces vague criteria, and vague criteria cannot be proven. The criteria MUST list, upfront:

  • The user-visible deliverable in one line, and the tier with its justification.
  • Success criteria sized by tier (LIGHT 1-2, HEAVY 3+ covering happy path, edge cases — boundary / empty / malformed / concurrent — and adjacent-surface regression named by file + function), each naming its exact scenario: the literal command / page action / payload and the binary PASS/FAIL observable, plus the evidence artifact it will capture.
  • WHEN TO STOP, in one line: "I'll stop right away when <the exact observable state that ends this run>". The Stop rules bind to this line — the moment it holds, you stop.

These scenarios are the contract. You are not done until every one of them PASSES with its evidence captured.

2. Open the durable notepad

Run: NOTE=$(mktemp -t ulw-$(date +%Y%m%d-%H%M%S).XXXXXX.md). Echo the path. Initialise it with these sections and APPEND (never rewrite) as you work:

# Ultrawork Notepad — <one-line goal>
Started: <ISO timestamp>

## Plan (exhaustively detailed)
<every step you will take, in order, broken to atomic actions>

## Success criteria + QA scenarios
<copied from the goal>

## Now
<the single step in progress>

## Todo
<every remaining step, ordered>

## Findings
<every non-obvious fact discovered, with file:line refs>

## Learnings
<patterns / pitfalls / principles to remember next turn>

Append each finding, decision, command, test read, and QA artifact path the moment it happens. Update ## Now and ## Todo on every transition. Append-only — never rewrite. This notepad is your durable memory and it OUTLIVES the context window. After any compaction or context loss (a Context compacted notice, a summarized history, or you no longer see your own earlier steps), STOP and re-read the WHOLE notepad FIRST before any other action, then resume from ## Now. Recover state from the notepad; do not re-plan from scratch or re-run completed steps.

3. Register obsessive todos via update_plan

The todo tool is Codex update_plan — your live, user-visible checklist. Translate every action from the plan into one update_plan step — one step per atomic work unit: an edit plus its verification, a QA scenario run, a teardown. Keep each step small enough to finish within a few tool calls. Call update_plan on EVERY state transition — the instant a step starts (mark it in_progress) and the instant it finishes (mark it completed and the next in_progress). Exactly ONE in_progress at a time. Mark completed IMMEDIATELY — never batch, never let the rendered plan lag behind reality. Add newly discovered steps the moment they surface instead of waiting for the next pass. Step text encodes WHERE / WHY (which criterion it advances) / HOW / VERIFY: path: <action> for <criterion> — verify by <check>.

GOOD pair (ordered): test/foo.test.ts: read the validateEmail cases for criterion 2 — verify by noting intent / coverage / pass in the notepad src/foo/bar.ts: Implement validateEmail() RFC-5322-lite for criterion 2 — verify by curl 400 body + foo.test.ts green BAD: "Implement feature" / "Fix bug" / "Add tests later" → rewrite.

Finding things (lead with these, code-mode the first wave)

Never guess from memory — locate with the right tool, and re-read before you claim or change. USE CODE MODE AGGRESSIVELY FOR BOUNDED WAVES. When multiple independent tool calls produce results that can be materially filtered, joined, deduplicated, or reduced, make ONE exec / eval JavaScript program that calls eligible tools concurrently with Promise.all and emits only decision-relevant evidence. For shell-native repo work without programmatic tool access, use ONE Python script with concurrent.futures, subprocess, and utility functions to batch commands and reduce output. Keep direct calls when one result chooses the next action, outputs are already small, semantic judgment is required between calls, approval or side effects are involved, or native artifacts / citations must be preserved.

  • Architecture / flow / blast radius → explore agents plus LSP references and impact; do not guess from conventions.
  • SYMBOLS REQUIRE LSP — definitions, references, rename impact, workspace symbols, and diagnostics use the available lsp_* tools, not text search. Run diagnostics after edits and treat errors as blocking.
  • Repo text / filenames / history / bounded shell output → rg, rg --files, git, and native utilities; narrow output in-program.
  • Structural call / function / class / import shapes and codemods → the ast-grep skill or sg with $VAR / $$$ metavariables. When discovery needs multiple angles or the module layout is unfamiliar, delegate to the explorer subagent (read-only codebase search, absolute-path results). For research that leaves the repo — library/API/docs/web — delegate to the librarian subagent. Spawn them fork_context: false and keep doing root work while they run.

Execution loop (READ → CHANGE → RUN → CLEAN)

Until every success criterion PASSES with its evidence captured:

  1. Pick next criterion → mark in_progress → update notepad ## Now.
  2. READ what already proves the area BEFORE touching it. Existing tests are the behavior of record: note in the notepad whether they encode the intended behavior, cover the path you change, and pass. One WRONG before your change is a FINDING to report — NEVER edit a test green. A bug: reproduce it first and capture the failure. A refactor: the existing tests are green on the unchanged code first.
  3. CHANGE: the SMALLEST production change that meets the criterion; update the tests your change makes stale. Add a test ONLY when BOTH hold: the repository keeps tests for this behavior AND a regression would otherwise pass unnoticed by the run and the existing tests — sized like its neighbors, one case per stated behavior, failing when that behavior breaks. A test that restates the change (a constant, a string, a rename, a call) is NOT evidence; the run is. Coverage-only work (no production change): break the behavior each new assertion names, capture it failing, restore — an assertion that stays green under its mutation is not coverage. PROSE TARGET (prompt, SKILL.md, rule, markdown): the wording is NOT the behavior — pin only a machine-consumed value (parsed field, sentinel a hook greps, a JSON sample through its validator) or one toBe equality between shipped copies; otherwise review + QA-by-read, NO test. Before a change that depends on review, PR, issue, or branch state, refresh that state and preserve existing ordering/policy.
  4. RUN: the real-surface scenario the criterion named (channel table above; auxiliary surface for CLI- or data-shaped criteria), end to end, yourself, plus the step-2 tests; a reproduction now passes. Paste the artifact path into the notepad.
  5. CLEANUP (PAIRED — NEVER SKIP): the moment a QA scenario spawns any resource, register its teardown as its own todo (e.g. cleanup: kill server pid for criterion 2 — verify kill -0 fails). Every runtime artifact the QA spawned in step 4 MUST be torn down before this step completes: server PIDs (kill <pid>; verify kill -0 fails), tmux sessions (tmux kill-session -t ulw-qa-<criterion>; verify with tmux ls), browsers / sessions (browser.close() / session.stop()), containers (docker rm -f), bound ports (lsof -i :<port> empty), temp sockets / files / dirs (rm -rf the mktemp paths), QA-only env vars. Append a one-line cleanup receipt to the notepad next to the artifact, e.g. cleanup: killed 12345; tmux kill-session ulw-qa-foo; rm -rf /tmp/ulw.aB12cD. No receipt → criterion stays in_progress.
  6. Verify: LSP diagnostics clean on changed files; no test skipped or xfail-ed this turn.
  7. Mark completed. Append non-obvious findings / learnings.
  8. Evidence stays valid per target until an input changes; record with each artifact the commit and what it exercised. After each increment rerun what moved — the tests of every touched file and of the files that import it, the scenarios that exercise them, anything whose dependencies or environment changed — and cite the capture for the rest. The full set (scenarios, suite, typecheck, build) runs once more right before the final message. Record PASS/FAIL beside each artifact. Loop until all PASS.

Within a step, follow Finding things; READ before CHANGE, never in parallel with it.

Show full SKILL.md (1,575 more words)Show less

Waiting discipline (a poll costs a full model round)

Every status check you issue as a tool call replays the entire accumulated context through the model. When a command will run long (installs, builds, test suites, containers, CI), run it to completion in ONE call with a timeout sized to the expected duration, or send output to a log file and read it once when a completion signal is expected. Never re-poll the same surface with empty reads or sub-minute waits — batch waiting into the fewest, longest blocking calls the harness allows, and do independent root work while the command runs. If two consecutive checks show no state change, double the wait before the next check or switch to a completion signal.

Codex subagent reliability

Every multi_agent_v1.spawn_agent message is self-contained and starts with TASK: <imperative assignment>, then names DELIVERABLE, SCOPE, VERIFY, and STOP WHEN — the observable condition that ends the child's run; a child without a stop condition wanders past its goal. State that it is an executable assignment, not a context handoff. Use fork_context: false unless full history is truly required; paste only the context the child needs. Full-history forks can make the child continue old parent context instead of the delegated task. If your tool list has a flat spawn_agent with a required task_name instead of multi_agent_v1.* (multi_agent_v2), rewrite: fork_context: false becomes fork_turns: "none", send_input becomes send_message, finished agents end on their own (no close_agent; followup_task re-tasks, interrupt_agent stops), and wait_agent takes only timeout_ms, returning on any child mailbox activity.

TOML-backed subagent routing compatibility

Inspect the ACTUAL spawn tool schema, not a version or namespace assumption. When agent_type is exposed (V1 or V2), EVERY spawn MUST pass an exact LazyCodex role: explorer, librarian, plan, metis, momus, lazycodex-worker-low, lazycodex-worker-medium, lazycodex-worker-high, lazycodex-code-reviewer, lazycodex-qa-executor, lazycodex-gate-reviewer, or lazycodex-clone-fidelity-reviewer. Map implementation difficulty to worker low/medium/high; their installed TOMLs supply model and instructions. Never select generic worker/default or describe a role instead of selecting it. Use fork_turns: "none" on V2 or fork_context: false on V1 unless full history is deliberately required; even a deliberate fork MUST name its role.

Legacy-schema exception: ONLY when agent_type is absent, omit that unsupported field and carry the role, difficulty, and complete instructions in message; explicitly disable history. This cannot select a specialized TOML. The managed default supplies the medium worker for unnamed non-forks, unless opted out or blocked by a preserved user default. The spawn guard cannot see the schema and rejects unnamed requests: report incompatible routing, do not retry generically. An unnamed full-history fork skips role application inside Codex; no LazyCodex config can fix that upstream gap. Never claim this path has been repaired. Difficulty (model power) is orthogonal to LIGHT/HEAVY rigor (process size).

Treat child status as a progress signal, not a timeout counter. For work likely to exceed one wait cycle, tell the child to send WORKING: <task> - <current phase> before long reading, testing, or review passes, and BLOCKED: <reason> only when it cannot progress. Track spawned agent names locally. Use multi_agent_v1.wait_agent for mailbox signals, but a timeout only means no new mailbox update arrived. Treat a running child as alive and keep doing independent root work. Fallback only when the child is completed without the deliverable, ack-only, or no longer running. If that followup is still silent or ack-only, record the result as inconclusive, do not count it as approval/pass, close it if safe, and respawn a smaller fork_context: false task with the missing deliverable.

Subagent-dependent transition barrier

Do not mark an update_plan step completed while an active child owns evidence for that step. Do not start dependent implementation until the audit, research, or review result is integrated or explicitly recorded as inconclusive. Do not generate a plan before spawned research lanes that feed the plan have returned or been closed as inconclusive. Spawn every independent child for the current wave first. After the wave is launched, run multi_agent_v1.wait_agent for each spawned child until each reaches terminal status (completed, failed, blocked, or explicitly recorded inconclusive) before any dependent update_plan transition, create_goal continuation, implementation tool call, plan drafting, approval-gate work, PR handoff, or final response. A timeout is not terminal status. Do not write the final answer, PR handoff, or completion summary while active child agents remain open. Use multi_agent_v1.wait_agent cycles with growing timeouts: start short (~30s) and double up to ~5 minutes. After two silent waits send TASK STILL ACTIVE: return <deliverable> or BLOCKED: <reason>. After four silent or ack-only checks, close the lane as inconclusive, record that it is not approval, and respawn smaller only if the deliverable is still required.

Verification gate (TRIGGERED ONLY ON EXPLICIT DEMAND)

Trigger ONLY when the user explicitly demanded strict, rigorous, proper, or high-accuracy review of this work, in any language (for example, 고정밀 or 엄격). The tier alone never triggers the gate. HEAVY without such a demand records the same self-review as LIGHT. LIGHT and non-triggered HEAVY work records a self-review in the notepad instead: re-read the diff, run diagnostics, confirm each criterion's evidence, and state in one line why the tier held.

When triggered, follow this procedure (NON-NEGOTIABLE):

  1. Spawn a child with fork_context: false and a self-contained reviewer assignment in message. The multi_agent_v1.spawn_agent schema cannot select a TOML-backed reviewer role, so paste the reviewer requirements into the message. Pass: goal, success-criteria, scenario evidence, full diff, notepad path.
  2. Verify each reviewer concern yourself. A concern blocks only when it names a success criterion the evidence fails; record concerns that cite no criterion as notes with a one-line reason — fixed or declined at your judgment.
  3. Fix every criterion-cited blocker; rerun per Execution loop step 8 and update the notepad.
  4. Spawn a NEW reviewer for each re-review, at most twice, passing only the delta diff, the blockers the last one cited, and the already-approved criteria marked out-of-scope. An approval whose only remaining items are notes counts as approval.
  5. On approval, declare done. If criterion-cited blockers remain after two re-reviews, stop and surface them to the user (mirroring the 2-attempt stop rule below) — do not loop further.

Commits

Commit frequently: one atomic commit per verified increment (change + its evidence), never one end-of-run omnibus; each commit builds + tests green on its own; no WIP on the final branch. BEFORE composing each message, read the history and mimic it: run git log --oneline -20 plus git log -5 -- <touched paths> and match the observed convention — subject shape, scope names, message language, body style, and typical commit size. Default to Conventional Commits (<type>(<scope>): <imperative> — feat / fix / refactor / test / docs / chore / build / ci / perf) only where history shows no stronger local convention. If a plan file exists, final commit footer: Plan: .omo/plans/<slug>.md. Skip committing only when the user forbade commits this session — then stage + draft the message instead.

Constraints

  • Every behavior change is PROVEN BY ITS RUN on the real surface, with the tests the repository keeps for it green. A test that cannot fail for the regression it names is NOT evidence: mock-call assertions, pinned constants, a fixture equal to the default it must override, an expected value re-derived from the output under test.
  • Make the smallest correct change per unit, and fix in THIS run every defect inside the change's blast radius — the request not delivered, a regression this change introduces, an invalid proof, a failing test or stale doc of code you touched — as registered work (todo plus success criterion) to the ideal state. A defect outside it gets a tracked issue with reproduction and evidence and a line in the final message; a deferral never turns a criterion into PASS. Keep delegated unit scope hard: the worker reports, the orchestrator registers or files.
  • Never suppress lints / errors / test failures. Never delete, skip, .only, .skip, xfail, or comment out tests to green the suite.
  • Never claim done from inference — only from captured evidence.

Output discipline

  • First line literally: ULTRAWORK MODE ENABLED!
  • After bootstrap: 1-2 paragraph plan summary + notepad path.
  • During execution: at every handoff - todo phase change, blocker, plan change, before a long pass - one handoff block composed after weighing what the user asked and needs to know now: [Outcome so far] toward [ask + wanted]. You need: [ledger, evidence paths, PASS/FAIL, reviewer verdict]. Now: [todo in progress]. Next: [next open todo].; nothing between handoffs.
  • Final message: outcome + success-criteria checklist with evidence refs + notepad path + reviewer approval (if gate triggered) + commit list (<sha> <subject>). No file-by-file changelog unless asked.

Stop rules

  • After each result, ask whether the user's core request can now be answered with useful evidence in hand. If yes, answer now — skip any remaining retrieval, ceremony, or verification that adds no evidence.
  • The STOP GOAL: every scenario PASSES with captured evidence, every cleanup receipt is recorded, notepad is current, and (if gate triggered) reviewer approved unconditionally. Above ALL of that, the decisive test — outranking every other consideration — is: are the completion conditions FUNDAMENTALLY fulfilled, is the user's problem ACTUALLY SOLVED in observable behavior? If no, you are NOT done, whatever the ledger says. If yes, deliver the final message and STOP — no hesitation, no extra verification pass, no polish loop. Work past the stop goal is scope creep, not diligence.
  • Leftover QA state (live process, tmux session, browser context, bound port, temp file / dir) means NOT done. Tear it down, record the receipt, then continue.
  • After 2 identical failed attempts at one step, surface what was tried and ask the user before another retry.
  • After 2 parallel exploration waves yield no new useful facts, stop exploring and act.
</ultrawork-mode>

© code-yeongyu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/omo/skills/ultrawork of code-yeongyu/lazycodex.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 36ba46a

Compare with similar skills

Ultrawork Mode next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ultrawork Mode compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ultrawork Mode this skillcode-yeongyu/lazycodex3.7k—~6.9kAutomated safety check: PassMIT
MemPalace Recall for Planningopen-gsd/gsd-core10k1 repos~1.5kAutomated safety check: NotesMIT
Inline Plan ExecutionjnMetaCode/superpowers-zh8.3k—~2.5kAutomated safety check: PassMIT
Planning with FilesOthmanAdi/planning-with-files27k—~2.9kAutomated safety check: PassMIT
Planning With FilesOthmanAdi/planning-with-files27k—~3kAutomated safety check: PassMIT
Agent Goal Brief WriterKKKKhazix/khazix-skills21k—~723Automated safety check: PassMIT

Similar skills

  • Recalls earlier decisions, patterns and surprises from MemPalace memory before planning, behind a config gate that never blocks the planning step.

    10k GitHub starsUsed in 1 repo~1.5k tokens
    Agent WorkflowsAuto-check: notes
  • Inline Plan Execution

    jnMetaCode/superpowers-zh

    Executes a written implementation plan task by task in the current session, with a progress ledger, test-first gates and one fresh-context review at the end.

    8.3k GitHub stars~2.5k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Planning with Files

    OthmanAdi/planning-with-files

    Keeps a task plan, findings and progress log in markdown files on disk so long agent tasks survive context resets, with Gemini hooks and helper scripts.

    27k GitHub stars~2.9k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Planning With Files

    OthmanAdi/planning-with-files

    Keeps a task plan, findings and progress log as Markdown files in the project so long multi-step agent work survives context resets.

    27k GitHub stars~3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Agent Goal Brief Writer

    KKKKhazix/khazix-skills

    Turns a one-line idea into a task brief of up to 4000 characters that an autonomous agent can run through /goal, with measured baselines and cheat-resistant acceptance checks.

    21k GitHub stars~723 tokensUpdated 6 days ago
    Agent WorkflowsAuto-check passed
  • Planning with Files for Kiro

    OthmanAdi/planning-with-files

    Keeps task_plan.md, findings.md and progress.md on disk as the agent's working memory for multi-step work, wired into Kiro steering, with no hooks.

    27k GitHub stars~2.1k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

Categories

Questions about Ultrawork Mode

What does Ultrawork Mode do?

The binding directive for ultrawork mode: evidence-first delivery, a light or heavy process tier, and constant use of memory during the work. Loaded only when ultrawork mode is requested and the directive is not already in the conversation, this file is the directive itself. It opens with a fixed banner line the agent must print first and then sets an outcome-first stance: deliver exactly what was asked, working end to end, proven by captured evidence from running the changed behavior through its real surface.

When should I use Ultrawork Mode?

Ultrawork Mode fits situations like: someone asks for ultrawork mode and the directive is not yet in the conversation; starting a task that must be proven with real run evidence, not only passing tests; sizing the process for a change as light or heavy before planning it.

How do I install Ultrawork Mode in Claude Code?

Run `npx skills add code-yeongyu/lazycodex --skill ultrawork -a claude-code`. Or copy the skill folder (plugins/omo/skills/ultrawork in code-yeongyu/lazycodex) into .claude/skills/ultrawork in your project. Claude Code loads it when a task matches its description.

How do I install Ultrawork Mode in Codex?

Run `npx skills add code-yeongyu/lazycodex --skill ultrawork -a codex`. Or copy the skill folder (plugins/omo/skills/ultrawork in code-yeongyu/lazycodex) into .agents/skills/ultrawork in your project. Codex loads it when a task matches its description.

Can I use Ultrawork Mode in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add code-yeongyu/lazycodex --skill ultrawork -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ultrawork, .gemini/skills/ultrawork, .github/skills/ultrawork and .opencode/skills/ultrawork in your project.

What does Ultrawork Mode need to run?

Going by SKILL.md and its folder, Ultrawork Mode needs the command-line tools its instructions call (curl, git, node, rg and docker).

Does Ultrawork Mode access the network?

SKILL.md contains no URLs. Its commands use curl, git and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ultrawork Mode safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ultrawork Mode use?

Ultrawork Mode is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ultrawork Mode use?

About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ultrawork Mode?

Skills that share tags, products or a category with Ultrawork Mode: MemPalace Recall for Planning (open-gsd/gsd-core, 10k stars), Inline Plan Execution (jnMetaCode/superpowers-zh, 8.3k stars), Planning with Files (OthmanAdi/planning-with-files, 27k stars) and Planning With Files (OthmanAdi/planning-with-files, 27k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ultrawork Mode?

code-yeongyu (a GitHub user) maintains it in code-yeongyu/lazycodex, which has 3,743 GitHub stars. The repository was last updated on October 6, 2026.

Source: code-yeongyu/lazycodex on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.