QA Find Bugs MCP
bex-co/beancount-io
Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and…
A skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…
$ npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Intelligent-Internet/zenith engineering-mission-playbook --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook .claude/skills/engineering-mission-playbook && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "engineering-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook into .claude/skills/engineering-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engineering-mission-playbook", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbookType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Intelligent-Internet/zenith engineering-mission-playbook --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .agents/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook .agents/skills/engineering-mission-playbook && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "engineering-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook into .agents/skills/engineering-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engineering-mission-playbook", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Intelligent-Internet/zenith engineering-mission-playbook --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook .cursor/skills/engineering-mission-playbook && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "engineering-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook into .cursor/skills/engineering-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engineering-mission-playbook", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Intelligent-Internet/zenith.git --path zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Intelligent-Internet/zenith engineering-mission-playbook --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook .gemini/skills/engineering-mission-playbook && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "engineering-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook into .gemini/skills/engineering-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engineering-mission-playbook", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Intelligent-Internet/zenith engineering-mission-playbookInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .github/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook .github/skills/engineering-mission-playbook && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "engineering-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook into .github/skills/engineering-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engineering-mission-playbook", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Intelligent-Internet/zenith engineering-mission-playbook --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook .opencode/skills/engineering-mission-playbook && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "engineering-mission-playbook" agent skill from https://github.com/Intelligent-Internet/zenith/tree/main/zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook into .opencode/skills/engineering-mission-playbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engineering-mission-playbook", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
engineering-mission-playbookA skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…
Engineering Mission Playbook is an agent skill from Intelligent-Internet/zenith. Use when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs, data/migrations, libraries, or operator workflows. Defines investigation, scope inventory, coherent VAL- contracts, evidence floors, multi-target task topology, engineering validation, root-cause patching, and durable guidance. For pure metric search, use optimization-mission-playbook instead.
Its SKILL.md is about 8.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Backend & APIs, covering Background jobs and Root cause analysis. It works with Model Context Protocol. The repository describes itself as: Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a8d9b57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Engineering Mission Playbook loads about 8.2k tokens when it runs. Until then it costs about 125 tokens; SKILL.md has 4,072 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Intelligent-Internet/zenith at commit a8d9b57, republished under its Apache-2.0 licence (© Intelligent-Internet). 4,072 words, ~8,154 tokens.
.claude/skills/engineering-mission-playbook/SKILL.md (or your agent's skills folder).Engineering success means the accepted user, caller, operator, or consumer-visible behavior is implemented, preserved, and independently proven on the real surface. Code motion, file diffs, worker claims, green commands, or task completion are not success by themselves.
Success requires:
mission.md or the current scope charter, the scope/capability inventory, live VAL-* assertions, or explicit accepted non-goal/risk decisions;work task that completes it, not merely contributes to it;Needs and Evidence fields are binding: unmet prerequisites, missing evidence, skipped assigned targets, wrong validation surface, or incomplete source-suite parity means blocked or failed, not passed;Do not accept a mission as successful when only happy paths pass, only worker-authored tests pass, only source inspection looks correct, or only a small manually chosen subset was exercised while requested real surfaces, rare cases, compatibility behavior, or source-suite parity remain unproven.
Honest failure is allowed. Do not convert missing setup, weak oracle, broad contract, bad task topology, unavailable evidence, incomplete test coverage, repeated worker miss, or changed scope into a speculative pass, weaker validator method, vague assertion, or late scope shrink.
Use this playbook in order. Do not write contract files, task lists, project skills, or durable guidance until investigation has produced an evidence-backed engineering mission model.
On a planning wake, follow this order:
VAL-* contract assertions: convert the inventory into compact, falsifiable validation targets with Surface, Needs, Behavior, and Evidence.contract-review before task planning. Fix missing coverage, broad buckets, unverifiable evidence, shortcut paths, and inventory-to-contract mapping gaps before continuing.work, validate, and gate tasks from the reviewed contract. Preserve many atomic assertions mapped into fewer coherent work tasks when implementation boundaries are shared.submit_plan, then drive work with advance_project. Do not implement product behavior in the orchestrator.mission.md, inventory, contract files, task list, project skills, AGENTS.md, MEMORY.md, and decisions when their old content would mislead later agents.On an attention or replan wake, do not resume from the current symptom. Trace the issue back to the earliest invalid step in this order. A failed validator may mean bad implementation, but it may also mean missing inventory coverage, weak contract, missing setup, wrong oracle, wrong task topology, stale skill, weak validation method, or changed scope.
Investigation, decomposition, contract authoring, validation-readiness checks, and task design are planning responsibilities. Do not hide them as runtime work tasks merely to postpone thinking.
Investigation is planning work. Do not create runtime work tasks for generic repository analysis, architecture mapping, test discovery, verifier discovery, source-suite discovery, feature enumeration, or task decomposition. Resolve those questions before submit_plan.
Before contract authoring, build an evidence-backed engineering mission model from primary sources:
Use bounded investigator lanes for non-trivial missions. A mission is non-trivial when it spans multiple surfaces, changes public behavior, preserves compatibility, ports behavior, depends on an external oracle, needs validation setup, has unclear scope, or likely needs more than one coherent worker session. Give each lane one focused question, why it matters for contract quality, exact paths or surfaces, non-goals, evidence to collect, and a stop condition.
For broad, parity, migration, porting, drop-in replacement, product-surface, multi-flow, or ambiguous missions, create a scope/capability inventory before writing VAL-* files. The inventory is not the contract and not the task list. It is the coverage map that prevents silent scope shrinkage.
The inventory must enumerate:
For porting, parity, compatibility, or drop-in replacement missions, the inventory must name the accepted source version or commit, source-to-target surface map, original/source test suite, golden corpus, differential checks, compatibility examples, accepted divergences, and environment required to run both source and target. Missing access to the source suite, golden corpus, or differential oracle is a planning blocker unless the user explicitly accepts the validation risk.
When the workspace includes an external verifier, reward script, scoring test suite, benchmark harness, oracle config, or hidden/public evaluator wrapper, treat it as primary evidence for success. Read the allowed verifier/reward surface before freezing scope. Record scored capabilities, commands, APIs, files, weights, anti-cheat rules, required setup, and known unscored areas. Do not mark scored behavior as out of scope just because it is broad or hard.
After investigation, synthesize the inventory into a planning diagnosis:
VAL-* assertions;Do not proceed to contract authoring while required surfaces have only generic conclusions. Missing evidence, unknown oracle behavior, unavailable source suites, unclear setup, or ambiguous scope must be recorded as explicit blockers or decisions, not hidden inside vague contract wording.
Engineering contracts use VAL-* assertions. Each assertion is a compact validation target stored as contract/<ID>.md under the current runtime mission contract directory. The contract is the engineering definition of done; it is not the task list, not the implementation plan, and not a summary of work packages.
Author contracts from the accepted scope charter, scope/capability inventory, investigation evidence, and actor interaction inventory. Enumerate what the user, caller, operator, or consumer can do, see, click, type, call, run, trigger, retry, cancel, edit, delete, import, export, migrate, recover, or observe. Organize assertions by feature area, surface area, workflow, and cross-area flow when useful.
Do not let task count determine assertion count. A broad mission should normally have many more VAL-* assertions than work tasks. One work task may later own many related assertions when one coherent implementation boundary completes them. Contract boundaries are chosen by validation coherence, not task topology.
Each assertion should be compact and falsifiable:
# VAL-AREA-001: Short user/caller/operator-facing title
Surface: browser | api | cli | tui | job | artifact | data | migration | library | parity | other.
Needs: none | prerequisite assertion/setup/oracle/source baseline/fixture/service.
Behavior: one coherent actor/action/outcome promise or related scenario set.
Evidence: exact evidence a validator must collect.
Fail: optional failure or blocked condition.
Oracle: optional source baseline, golden corpus, compatibility suite, rubric, or authoritative reference.
Scope: optional non-goals, accepted divergences, or adjacent boundaries.Surface, Needs, Behavior, and Evidence are required. Use none for Needs only when no prerequisite is known. Add Fail, Oracle, or Scope when needed to prevent ambiguity, shortcut passing, or silent scope shrink.
Split assertions when grouped behavior would hide unrelated proof, owner, setup, oracle, validator method, milestone boundary, or failure diagnosis. Keep related scenarios together when they share actor, surface, setup, oracle, work owner, milestone boundary, and validators can still give clear evidence and verdicts.
Reject broad bucket assertions such as "refs work", "status works", "API works", "object database works", "migration works", or "UI works" unless the assertion defines the bounded scenario set, oracle, per-scenario evidence path, and failure conditions tightly enough for independent validation.
For cross-area behavior, write explicit VAL-CROSS-* assertions. Include first-use flows, actual reachability/navigation, persistence across surfaces, import/edit/export flows, auth/session state, generated artifact consumption, migration-to-runtime behavior, background side effects, recovery, rollback, and compatibility handoffs when relevant.
For porting, parity, compatibility, or drop-in replacement missions, assertions must preserve source behavior through named source version, source test suite, golden corpus, differential command/API/library examples, and accepted divergences. Do not replace source-suite parity with a few handpicked examples unless the user accepts that validation risk.
For each assertion, the Evidence field must say what fresh validator evidence proves it: screenshots, console/network observations, request/response traces, stdout/stderr/exit codes, generated files, DB/state checks, logs, golden diffs, source-suite output, compatibility suite output, or differential source/target results. Missing required evidence means failed or blocked, not passed.
Contract review is mandatory before task planning. Use the contract-review subagent to review the full contract set against the user request, accepted scope, scope/capability inventory, actor interaction inventory, loaded playbook, evidence floor, shortcut risk, and future task-mapping shape.
For non-trivial engineering missions, run at least two sequential contract-review passes. After each pass, synthesize findings, revise contract files, then review the revised contract again. Do not author or submit the task list while review returns missing coverage, broad buckets, unverifiable assertions, weak evidence floors, shortcut risk, or unresolved inventory-to-contract gaps.
Contract review must check:
VAL-* assertions, an accepted non-goal, or an explicit deferred-scope decision;Needs and Evidence;The review output is a contract set ready for task planning: coverage gaps resolved, broad assertions split or clarified, stale assertions removed or superseded, validation blockers named, accepted risks recorded, and every live assertion ready to map to one owning work task plus independent validation.
Validation proves VAL-* assertions through independent evidence. It does not confirm worker prose, task completion, or optimistic summaries.
Choose validation from the assertion surface and risk:
scrutiny-validator: implementation, diff, tests, fixtures, generated outputs, hard-gate commands, evidence integrity, fake-pass risk, maintainability, and responsibility drift.user-testing-validator: real user, caller, operator, or consumer surfaces: browser, API, CLI/TUI, background job, generated artifact, migration/data state, public library call, operator workflow, or parity flow.For user-visible, caller-visible, operator-visible, generated-artifact, migration/data, public API, CLI, or parity behavior, fresh real-surface evidence is the verdict source. Automated tests, worker-authored E2E specs, source inspection, and green commands are supporting evidence unless the assertion explicitly names them as the oracle.
Validators must parse and honor every assigned assertion field:
Surface: choose the real validation surface named by the contract.Needs: verify prerequisite assertions, setup, fixtures, services, credentials, source baselines, accepted decisions, and oracles before exercising behavior.Behavior: exercise the actor/action/outcome promise and relevant scenarios.Evidence: collect the exact artifacts required by the contract.Fail, Oracle, and Scope: apply them when present; do not pass by ignoring boundaries or accepted divergences.Missing required evidence, skipped assigned targets, unmet Needs, unavailable setup, missing source baseline, wrong surface, unverifiable assertion, or evidence artifacts that cannot be attributed to the target mean passed=false or blocked attention, not pass.
Use a two-lane validation model for externally observable engineering behavior:
The real-surface lane carries the verdict for real-surface behavior. The scrutiny lane may fail an assertion on integrity or implementation grounds, but it cannot pass a real-surface assertion by itself unless the contract names scrutiny as the oracle.
For porting, parity, compatibility, or drop-in replacement missions, validation must include the relevant source/original suite, golden corpus, compatibility examples, or differential source-target checks for the accepted source version. Target-only tests and handpicked examples are supporting evidence, not full parity proof, unless the user accepted that limitation.
For UI/browser assertions, evidence should include real navigation or interaction steps, screenshots, console error review, and relevant network observations. For API assertions, include real request/response traces, auth context, status/body/schema, and persistence side effects when relevant. For CLI/TUI assertions, include exact commands or interactions, stdout/stderr, exit code, cwd/env assumptions, and terminal behavior when relevant.
For generated artifacts, migrations, jobs, and data behavior, evidence must prove the real producer and consumer path where applicable: generation command, artifact path, schema/golden/checksum diff, before/after data state, idempotency, rollback/retry behavior, logs/events, and downstream consumption.
When validators use flow-validator lanes, each lane must receive exact target ids, contract bodies or paths, isolation resources, evidence directory or artifact prefix, setup/recovery limits, non-goals, and output schema. Parallel lanes must not share mutable state unless isolation is explicit. Parent validators must reject or fail lane results whose evidence cannot be attributed to exact assertions.
Validation cost is scheduling pressure, not permission to weaken the contract. Reduce cost by batching validators that share setup, surface, oracle, isolation, or evidence artifacts, but keep per-assertion verdicts and evidence. Do not broaden assertions or drop rare cases to make validation cheaper.
If required evidence cannot be collected, patch setup, oracle, fixtures, contract clarity, task topology, or validator skill. Do not downgrade the validation method to fit convenient evidence.
Build the SWE task list after the mission scope, target set, and validation floor are already clear. This section is only about arranging runtime tasks: implementation ownership, parallel execution, tester lanes, ordering, and gates. Do not use task topology to redefine scope.
Use task types deliberately:
work: a coherent implementation or setup unit.validate: an independent tester lane for a milestone.gate: a checkpoint after required validation lanes complete.Create work tasks by implementation boundary. Group work when it belongs to the same subsystem, files, setup, behavior, or handoff. Split work when chunks can be owned independently, run safely in parallel, reduce risk, or make debugging and review clearer.
Tasks that do not depend on each other may run in parallel. Use depends_on only for real ordering: setup before dependent work, one work task before another when it needs the earlier result, work before milestone validation, validation before gate, and promotion/integration before validation when a candidate was not auto-merged.
Leave work tasks independent only when they can start from the same project state and neither task needs the other's output. Add dependencies when tasks share mutable files, services, fixtures, data, generated artifacts, or behavior that must be integrated before the next task can proceed.
Create validate tasks as independent tester lanes for a milestone, not as mirrors of worker tasks. A validator receives the milestone target set and performs comprehensive adversarial testing for its assigned lane. It should act like a real tester: inspect implementation when useful, run real product flows, exercise edge cases, check regressions, verify setup, and collect evidence across the milestone scope.
At the end of a milestone, plan the tester lanes needed to prove the milestone as a whole. Split validators only when the milestone needs distinct tester perspectives, such as implementation scrutiny, real user-surface testing, benchmark/performance validation, security review, migration/data validation, or source/parity validation. Do not create one validator per worker unless the worker boundary is also genuinely the validation boundary.
Create gate tasks by milestone. A gate depends on the relevant validators and seals only the milestone targets covered by those validators.
Work task bodies must be comprehensive and detailed enough for a worker to execute without guessing. Include assigned target ids, milestone or slice name, required setup, relevant surfaces/files, behavior to change or preserve, constraints, non-goals, expected output, self-checks, and handoff requirements. Add implementation guidance when it reduces ambiguity.
Validate task bodies must be comprehensive and detailed enough for a tester to verify the milestone independently. Include assigned target ids, tester lane purpose, setup, surfaces or flows to exercise, edge cases and regressions to probe, evidence artifacts to collect, blocked/fail policy, and required per-target verdicts. Make the tester's job clear enough that they can attack the milestone without reconstructing the plan from prior context.
Targetless setup work is allowed only when it enables later implementation work and cannot honestly be folded into a concrete work task. Keep it few, explicit, and dependency-linked.
Avoid one-task-per-target planning unless the implementation boundary is truly that small. Avoid one-validator-per-worker planning unless the validation boundary truly matches the worker boundary. Avoid catch-all work, validator, or gate tasks whose scope is too broad for clear ownership, evidence, or failure diagnosis.
Before submit_plan, check:
work task has a coherent implementation boundary;depends_on;Use durable memory to preserve operational facts that later workers and validators need without depending on session memory. Engineering missions often discover setup, services, commands, source baselines, oracles, and validation procedures during investigation or execution. Promote the durable parts into the right carrier before more work depends on them.
When reproducing or migrating behavior from an older harness, inspect old operational artifacts such as init.sh, services.yaml, service manifests, validation scripts, source-suite commands, browser setup notes, fixture setup, and library notes. Do not copy them blindly. Distill the reusable facts and preserve pointers to the original artifacts when useful.
Use carriers by purpose:
MEMORY.md: reusable engineering facts such as package manager, setup sequence, service map, ports, commands, healthchecks, source-suite command, golden corpus path, fixture setup, oracle locations, baseline refs, known flakes, accepted divergences, and validation limitations.AGENTS.md: normative guidance that all workers and validators must obey, such as port boundaries, off-limits resources, cleanup rules, required command discipline, source-suite requirements, and validation constraints.skills/: reusable worker or validator procedures when setup, oracle use, source-suite execution, browser validation, migration checks, or artifact generation would otherwise be repeated in task bodies.contract/<VAL-ID>.md: mission-specific done criteria and evidence requirements.decisions/: rationale for accepted scope cuts, accepted validation risk, changed setup, deferred source-suite coverage, retry/patch choices, or aborts.attempts/, regressions/, and evidence/: raw command output, logs, screenshots, traces, source-suite output, failed reproductions, and forensic records.Do not paste full scripts, raw logs, or long one-off attempt summaries into MEMORY.md. Promote only curated facts, procedures, constraints, and failure lessons that future agents need to act correctly.
If a discovered fact changes how future work should be planned, implemented, or validated, update memory before dispatching more work. If old memory conflicts with current setup, update or supersede it instead of leaving two truths alive.
Avoid these failure modes:
work tasks.mission.md or the current scope charter and scope/capability inventory preserve accepted scope.VAL-* assertion count.exactly one active owning work task per assertion as one assertion per work task.Needs, missing Evidence, skipped assigned targets, wrong surface, unverifiable setup, unattributed artifacts, or missing per-assertion verdicts.optimization-mission-playbook.© Intelligent-Internet, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook of Intelligent-Internet/zenith.
Open the folder on GitHubat commit a8d9b57
Engineering Mission Playbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Engineering Mission Playbook this skillIntelligent-Internet/zenith | 338 | — | ~8.2k | Automated safety check: Pass | Apache-2.0 | |
| QA Find Bugs MCPbex-co/beancount-io | 297 | — | ~3.1k | Automated safety check: Pass | MIT | |
| Edt MCP Project YaxunitDitriXNew/EDT-MCP | 296 | — | ~1.1k | Automated safety check: Pass | AGPL-3.0 | |
| FoundatioFoundatioFx/Foundatio | 2.1k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| NubaseOtterMind/Nubase | 622 | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | |
| Dv Adminmicrosoft/Dataverse-skills | 243 | — | ~5.1k | Automated safety check: Pass | MIT |
bex-co/beancount-io
Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and…
DitriXNew/EDT-MCP
Discover, run, debug, poll, and cancel YAXUnit tests through current EDT-MCP launch and background-job contracts.
FoundatioFx/Foundatio
A skill your agent uses when working with Foundatio infrastructure abstractions for .NET -- caching, queuing, messaging, file storage, distributed locking, or background jobs.
OtterMind/Nubase
A skill your agent uses when the user mentions Nubase broadly, wants a backend for an AI-generated app, or needs to deploy/publish generated code online — across Database, Auth, Storage, Assets…
microsoft/Dataverse-skills
Environment-level Dataverse administration — bulk delete, retention/archival, organization settings, OrgDB settings, recycle bin, audit, and the 37 allowlisted PPAC toggles.
modiqo/skillspec
A skill your agent uses when the task needs to run a local command and remember the result, inspect CLI output with provenance, follow a log or process stream, start or observe a background job…
Intelligent-Internet/zenith
Benchmark validation procedure for one assigned benchmark-related target.
Intelligent-Internet/zenith
Adversarial scrutiny procedure for engineering validation assignments.
Intelligent-Internet/zenith
Real-surface validation coordinator for engineering validation assignments.
Intelligent-Internet/zenith
Automates browser and Electron app interactions for user-flow validation.
Intelligent-Internet/zenith
Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…
Works with
Categories
A skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…. Engineering Mission Playbook is an agent skill from Intelligent-Internet/zenith. Use when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs, data/migrations, libraries, or operator workflows.
Engineering Mission Playbook fits situations like: replanning engineering missions that create; preserve durable codebase behavior across UI; background jobs; data/migrations.
Run `npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a claude-code`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook in Intelligent-Internet/zenith) into .claude/skills/engineering-mission-playbook in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a codex`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/engineering-mission-playbook in Intelligent-Internet/zenith) into .agents/skills/engineering-mission-playbook in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Intelligent-Internet/zenith --skill engineering-mission-playbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/engineering-mission-playbook, .gemini/skills/engineering-mission-playbook, .github/skills/engineering-mission-playbook and .opencode/skills/engineering-mission-playbook in your project.
SKILL.md names no scripts, command-line tools or credentials: Engineering Mission Playbook is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Engineering Mission Playbook is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.2k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Engineering Mission Playbook: QA Find Bugs MCP (bex-co/beancount-io, 297 stars), Edt MCP Project Yaxunit (DitriXNew/EDT-MCP, 296 stars), Foundatio (FoundatioFx/Foundatio, 2.1k stars) and Nubase (OtterMind/Nubase, 622 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Intelligent-Internet (a GitHub organization) maintains it in Intelligent-Internet/zenith, which has 338 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 6, 2026.
Source: Intelligent-Internet/zenith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.