Agent skill

Goal Lock

by AlexZio00 in AlexZio00/sovereign-skills

Agent Discipline Engine — lock the goal, run PLAN→DO→VERIFY→FINALIZE→OUTPUT loop, detect success masquerading.

MITAuto-check passed

Install Goal Lock

skills CLI
$ npx skills add AlexZio00/sovereign-skills --skill goal-lock -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlexZio00/sovereign-skills goal-lock --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlexZio00/sovereign-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/goal-lock .claude/skills/goal-lock && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
goal-lock
GitHub stars
140
Token cost
~9.4k tokens
SKILL.md length
4,420 words
Files
3
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Agent Discipline Engine — lock the goal, run PLAN→DO→VERIFY→FINALIZE→OUTPUT loop, detect success masquerading.

  • Works in 4 steps: CRITIQUE — re-read the artifact and… → REWRITE — rewrite only the identified… → DELTA CHECK — compare original vs rewrite → …
  • SKILL.md covers Dominant Variable, Trigger, Discard If and Architecture: 2 Layers, plus 8 more sections
  • Calls npm, npx and go

What it does

Goal Lock is an agent skill from AlexZio00/sovereign-skills. Agent Discipline Engine — lock the goal, run PLAN→DO→VERIFY→FINALIZE→OUTPUT loop, detect success masquerading. Triggers: '/goal-lock', '/goal-lock quick', 'goal lock', 'task harness'.

Its SKILL.md is about 9.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `.claude-plugin/plugin.json` and `agents/openai.yaml`).

The repository describes itself as: 20 production-grade skills for AI coding agents — setup, scope, discipline, code review, security, session management, governance, ops, and quality audits (eval-leakage… The licence is MIT.

Example prompts

  • “/goal-lock”
  • “/goal-lock quick”
  • “goal lock”
  • “/goal-lock”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. CRITIQUE — re-read the artifact and identify the "3 weakest points."
  2. REWRITE — rewrite only the identified weaknesses, once. Leave strong
  3. DELTA CHECK — compare original vs rewrite
  4. 1-round limit — REFINE runs at most once. A second round has

What it can do on your machine

Read from SKILL.md and the folder at commit c062683. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • npx
    • go
    • cargo
    • git
    • pytest
    • ruff

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, npx and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Goal Lock loads about 9.4k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 4,420 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~9.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AlexZio00/sovereign-skills at commit c062683, republished under its MIT licence (© AlexZio00). 4,420 words, ~9,413 tokens.

Download SKILL.mdSave it as .claude/skills/goal-lock/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
goal-lock
description
Agent Discipline Engine — lock the goal, run PLAN→DO→VERIFY→FINALIZE→OUTPUT loop, detect success masquerading. Triggers: '/goal-lock', '/goal-lock quick', 'goal lock', 'task harness'.
user-invocable
true
not_for
Simple questions/conversation (no code changes), Single file 1-line fix

/goal-lock — Agent Discipline Engine v1.0

Lock the goal. Run the loop. Ship clean.

Prevents agents from drifting off target, masquerading success, or creeping scope. Quality through enforced loops, not prompt obedience.

Dominant Variable

Is DONE EVIDENCE verified by actual execution? — What the agent says is done vs what is actually done. Closing this gap to zero is the purpose of this skill.

Trigger

  • /goal-lock
  • /goal-lock quick (Quick mode)
  • "goal lock"
  • "task harness"

Discard If

  • Simple question/conversation (no code changes)
  • goal-lock already active in this session
  • Single-file 1-line fix — goal-lock overhead > the work itself

Architecture: 2 Layers

[A] GOAL Input Sheet — fill per task (goal definition)
[B] Fixed Loop — same for every task (execution discipline)

Missing/contradictory input → STOP. Conflicts → PRIORITY. STOP RULES → halt.


Mode Selection

ModeConditionInput SheetLoop
Quick1 file, clear change, ≤10 lines3 fields (GOAL/DONE/SCOPE)DO→VERIFY only
FullEverything elseAll 7 fieldsB1~B5 full

User specifies /goal-lock quick, or change fits Quick criteria. When unsure, use Full.


[A] GOAL Input Sheet

Full Mode (7 fields)
markdown
## GOAL Input Sheet

### 1. GOAL
[Single measurable goal. No expansion.]

### 2. DONE EVIDENCE
[Completion proof. The evidence contract branches by artifact type —
don't force one shape onto both:]
- **Code artifact** → command to run + expected result. No subjective
  criteria.
  e.g.: `pytest tests/test_X.py -v` → 5 passed
  e.g.: `curl localhost:3000/api/health` → 200 OK
- **Non-code artifact** (writing, analysis, reports, designs, prompts, spec
  docs) → no exit code exists to demand. State the review contract instead:
  what a reviewer checks off, or what a named approver signs off on (e.g.
  "reviewer confirms the 3 required sections are present and each claim
  cites a source" or "user approves the draft"). This feeds directly into
  the REFINE loop below (VERIFY/REFINE split) rather than VERIFY's execution
  path.

**Adversarial criteria design**: when setting DONE EVIDENCE, ask first "how
could an agent game this criterion." An unblocked loophole tends to get
found eventually — threshold relaxation, mock wrapping, hardcoding all
exploit a DONE EVIDENCE that was underspecified to begin with. Check for
loopholes at design time, especially on long-running or repeated tasks.

**Evidence-Rigor Pre-spec** [borrowed from ultraprompt]: when DONE EVIDENCE
includes concurrency, benchmark, p99-style statistics, or long-running-process
claims, pre-check the verification agent's evidence-rigor rules (N≥5 repeats,
before/after symmetry, evidence-scope matching, flaky-means-new-bug,
positive-signal-required) and write DONE EVIDENCE to already satisfy them —
this prevents a later insufficient-evidence rejection at the verify step by
fixing the design at spec time instead.

### 3. CONTEXT
[Current state · existing structure · prior decisions · dependencies · known constraints]

### 4. STARTING POINT
[Files/logs/tests to look at first. Start here, no broad exploration.]

### 5. SCOPE
- **Include**: [Editable area + required work]
- **Exclude**: [Out of bounds · unrelated refactors · new features · production behavior changes]
  - **Capability-spillover (flag, don't fix)**: other bugs, design/structural
    improvement ideas, or similar edge cases noticed mid-task all stay in
    Exclude. Report them separately (one inline line, or a follow-up task)
    and return to the current GOAL. Stronger models trend toward "fixing it
    all while I'm in here" — scope is a lock, not a ceiling.

### 6. CONSTRAINTS
- New dependencies: allow/forbid
- Network/API calls: allow/forbid
- Commit/PR/push: allow/forbid
- Migration/DB changes: allow/forbid
- Destructive actions: allow/forbid

### 7. BUDGET
[Time/token/call/cost limits. Follow if given, don't invent if not.]

### 8. EVAL TYPE (optional — only for tasks measuring a skill/hook/gate's own reliability)
[yes — this GOAL measures whether the verification logic itself actually works]
[no or omit — regular implementation. Normal DO→VERIFY iteration is allowed]
Quick Mode (3 fields)
markdown
## GOAL (Quick)

### 1. GOAL
[One-line goal]

### 2. DONE EVIDENCE
[One verification command]

### 3. SCOPE
- **Include**: [Files to modify]
- **Exclude**: [Don't touch]
Scope Check Surface

SCOPE Include naming a file is not blanket permission for everything inside it. The scope check surface is file changes + interface/functionality surface — an unrequested CLI flag, a new public API parameter, or a test scenario broader than what was asked for is scope creep even when the file it lives in sits squarely inside SCOPE Include. Applies to both Quick and Full mode: "the file is in scope" answers a different question than "was this specific change asked for."

Auto-fill Rules

Fields extractable from conversation context are auto-filled and shown for user confirmation:

  • GOAL: from user's last request
  • CONTEXT: current git status + recent commits + related files
  • STARTING POINT: files mentioned in GOAL, or git diff files

After auto-fill: "Input sheet filled. Proceed if correct, or tell me what to change." — never proceed without confirmation.

Input Validation

Any of 7 fields (Quick: 3) missing or contradictory → don't guess, STOP.

  • GOAL contains 2+ independent goals → STOP ("Please narrow to one goal")
  • DONE EVIDENCE is subjective ("if it works well") → STOP ("Replace with executable verification command")
  • SCOPE Include and Exclude overlap → STOP ("X appears in both Include and Exclude")
  • BUDGET missing → OK (just don't invent one)

[B] Fixed Loop

B1. Always-ON Guards (never violate)

Honesty (highest priority):

  • Never hide failure, uncertainty, or unverified results
  • Mark uncertain judgments as uncertain

Success Masquerading Blocklist:

PatternDescription
Test deletion/disableDelete failing tests or neutralize with skip/xfail
Mock wrappingAssert only mock return values and report "passed"
Threshold relaxationLower the bar to pass
Benchmark condition changeChange measurement conditions to improve numbers
Failure log suppressionSend error output to /dev/null or catch-all
HardcodingHardcoded results matching test inputs
Requirement reinterpretation"Actually this was supposed to work like..."
Acceptance criteria weakeningSubtly lower DONE EVIDENCE standards
Production behavior changeModify production logic in a way that contradicts the actual requirement/spec to make a test pass — normal RED→GREEN (writing the minimal production code a correct failing test demands) is not this pattern; the violation is the direction of the change, not the fact that production code changed after a test
Goal-lock declaration ignoredDeclare "proceeding with goal-lock" then skip the input sheet
Structural fix reported as upgradeReport boilerplate additions as "substantive improvements"
Ralph Wiggum (early completion)Skip VERIFY or run it partially, then jump to OUTPUT. Emit completion signal from an incomplete state
CEF Thanatosis (external failure fabrication)Evading constraints via unverified failure claims like "API error"/"file not found"/"permission denied". Failure reports must be accompanied by actual Bash/Read execution results
Post-hoc abstentionExecute an irreversible action first, then declare "failed"/"on hold" after the fact. Abstention judgment is only valid before the commit-point gate — declaring it after the action has already landed is still success masquerading
Layer launderingNarrating a unit-test pass as if it proves the user-facing feature actually works — laundering one evidence layer as a higher one
Silent self-correctionOn an EVAL TYPE=yes task, quietly re-running DONE EVIDENCE multiple times off the record to hide failures, then reporting only the last (passing) run

Language-specific patterns:

  • Python: @pytest.mark.skip, @pytest.mark.xfail, mock.return_value abuse
  • JavaScript: test.skip, .only left in, jest.fn() chains bypassing real logic
  • Go: t.Skip(), //go:build ignore
  • Rust: #[ignore], #[should_panic] misuse

Judgment-reversal discipline (arXiv 2608.11624, 2608.21377): when user pushback lands mid-loop, sort it into one of three buckets before reacting — (a) a new fact (a new log, a new requirement, a new constraint) → incorporate it, (b) a pointed-out reasoning error (they name which premise or step is wrong) → re-examine that specific point, (c) pressure or preference with no new information ("are you sure?", "look again", "that doesn't seem right") → restate your original reasoning once, and if nothing they said actually invalidates a specific premise, keep the original judgment. When a judgment does change, log it in OUTPUT as prior conclusion → new conclusion (reason: ...). Reversing a conclusion under pressure alone, with no new information, is success masquerading of the same weight as the patterns above.

B1.1 Evidence-Rigor Ladder + Reporting Order [borrowed from ultraprompt]

5-tier evidence ladder — no claim can outrank this ladder: executed (actually observed running) > integration-tested > unit-tested > typed (type-checked only) > reasoned (reasoning only)

Every claim must state its tier: verified: {concrete evidence} or unverified: {reason}. An unlabeled "it's done"-type claim is not allowed.

Failure-first reporting order: describing successes first and only mentioning failures afterward is itself an anti-pattern ("burying the failure") — report failed/unverified items first, successes after.

Banned hedge phrases: "should work", "probably fine", "this looks right" and similar are explicitly banned. If unverified, write unverified: {reason} instead.

B2. PRIORITY (on conflict)
0 Honesty → 1 Stability → 2 Preserve existing behavior → 3 Verifiability → 4 Performance → 5 Code cleanup
B3. Execution Loop
PLAN → DO → VERIFY(code) → FINALIZE → OUTPUT
              ↘ REFINE(non-code) ↗
PLAN GATE
  • No immediate fixes. Identify root cause + short plan first.
  • Big changes, schema changes, dependency additions, production behavior changes → stop and get approval.
  • Plan is 3 lines max. Steps, not documents.
  • If several root-cause candidates exist and there isn't enough information to rank them: don't just list the candidates and move on — first run the single cheapest check that discriminates between them, within this skill's own tools (Read, Bash, etc.). If you are still uncertain afterwards, STOP at S5.
DO
  • Minimum change to achieve GOAL. Don't touch SCOPE Exclude.
  • Before starting, check RISKS (applicable items only):
RiskCheck
Breaking changeWill existing callers break?
Race conditionConcurrent access to shared resource?
Stale stateCache/state might not update?
Data lossIrreversible deletion/overwrite?
SecurityInput validation, permissions, secret exposure?
Perf regressionO(n²) introduction?
Backward compatExisting API/interface changing?
Chain lengthCan the total number of steps/tool calls be cut before optimizing any single step?

Chain length as a dominant variable: reliability tracks step count, not just each step's individual correctness. Benchmarks show tool-chain accuracy falling from roughly 39% to 13% as chains lengthen, and sequential-turn-depth scores dropping from 82.3 to 51.2 over comparable depth increases — longer chains fail more often even when every individual step looks reasonable. Before tuning how a step is done, ask whether it needs to exist at all; fewer, more consequential steps beat more, smaller ones.

Risk detected → return to PLAN with avoidance strategy.

ATTACK (mandatory sub-step in Full mode): if any RISKS item above is checked "yes," attack the plan yourself before executing it — enumerating risks and actually trying to break the plan are not the same exercise.

Self-assessed depth: informally gauge your own reasoning capability tier and scale ATTACK's depth to it — a lighter-capability tier warrants the full RISKS list plus multiple lenses below, a stronger tier can compress this to a one-line self-check. When unsure which tier applies, default to the more thorough end rather than assuming a strong tier.

Tier-0 (always — reuses only the lens concept from an independent adversarial-review skill, not its full machinery): apply as many of these lenses as the self-assessed depth calls for — decompose / invert / draw an analogy / push to the extreme / follow the incentive / check for grandfathered assumptions — to attack your own plan. Don't import a full independent reviewer's claim/evidence separation or ground-truth testing here — doing so just turns this into a shrunken copy of that skill for no net benefit. Sort findings into four buckets — actionable / tradeoff / contract-misread / noise — and tag them [self-attack, non-independent]: this is goal-lock interrogating itself, not an independent check, and the label says so. An actionable finding sends you back to PLAN for an avoidance strategy. Tradeoff/contract-misread findings get logged in CONTEXT only. Discard noise.

Tier-1 (conditional — defer to doubt-reviewer, or an equivalent adversarial pre-implementation review skill, instead of judging further yourself): if GOAL/SCOPE/CONTEXT match any of the following, stop self-judging and halt with S8: a change to branching or module boundaries · a property the type system can't verify · an irreversible blast radius · a change to a core parameter · a change to data-collection logic · confidence that outruns the certainty actually behind it. (These six are a subset of a broader trigger set such a review skill might use on its own — two related conditions are deliberately handled elsewhere instead: the same approach failing repeatedly is S6's job, and "right before declaring a large task complete" is a known gap goal-lock doesn't cover by itself — the calling session should judge whether that review is warranted at FINALIZE time for large tasks.) goal-lock is a single-process skill with no sub-agent dispatch of its own (see B5) — it cannot invoke doubt-reviewer directly. The calling session dispatches it and resumes goal-lock with the resulting verdict.

First-Attempt Ledger: before making any changes, run the DONE EVIDENCE command once and record the raw result under a ## First run (raw) field in .goal-lock-progress.md. Root-causing, fixing, and re-running proceed as normal after that. OUTPUT reports the first-run result side by side with the final result — hiding the first failure and reporting only the final pass is success masquerading.

Pre-DO failure baseline: at that same first run, record the items that already fail (test names, lint rules, error lines) as the baseline. VERIFY then judges only failures that are not in the baseline as regressions of this change; failures that exist only in the baseline are reported in OUTPUT as "pre-existing, out of scope" (never silently ignored). If the baseline failure is the GOAL itself (a bug you were asked to fix), it is the goal, not a baseline. Without a baseline, pre-existing failures burn retries or get misreported as your own regressions.

VERIFY (Code) / REFINE (Non-code)

Branches by artifact type:

Code artifacts → VERIFY: actually execute the verification specified in DONE EVIDENCE.

DONE EVIDENCE ↔ GOAL alignment check (Building to the Test): Before running verification, confirm: does DONE EVIDENCE actually prove GOAL? Agents tend to "build what gets checked, not what was asked." Passing DONE EVIDENCE while GOAL remains unmet is still a FAIL.

  • GOAL: "add search feature" / DONE EVIDENCE: "pytest passes" → confirm the test actually validates search functionality
  • If DONE EVIDENCE measures something unrelated to GOAL → STOP + rewrite DONE EVIDENCE

Verification recipes (auto-detect stack):

StackCommand
Python (pytest)pytest -q + ruff check (if available)
JavaScript (jest)npm test + npx eslint . (if available)
TypeScriptnpx tsc --noEmit + npm test
Gogo test ./... + go vet ./...
Rustcargo test + cargo clippy
Generalgit diff --stat (verify change scope)

Items not verified: NOT RUN: {label} using the failure-label enum below, plus a one-line detail. Never "it should be fine."

GroundEval: verify that verification tool calls actually executed. If the OUTPUT claims "tests passed" but no pytest/npm test Bash call exists in the tool history, the claim is ungrounded. Every verification claim in OUTPUT must trace back to an actual tool invocation.

Failure-label enum: when a verification claim can't be grounded, classify it with one of these fixed labels instead of a free-text reason — labels are greppable and comparable across runs, prose isn't:

  • no_fetch — no verification tool call exists in the tool history at all
  • irrelevant_fetch — a tool call happened, but it checked something other than DONE EVIDENCE
  • checker_overfit — the check ran, passed, and traces back to a real tool call, but is narrow enough that a broken implementation would pass it too

Evidence channel branching: not every DONE EVIDENCE produces an exit code. If DONE EVIDENCE is a visual artifact, confirm via rendered output (screenshot / extracted page text); if it's a read-only analysis, confirm via artifact comparison. Absence of an exit code is not an automatic FAIL — but regardless of channel, "no evidence produced" is still NOT RUN.

Non-code artifacts (writing, analysis, reports, designs, prompts, spec docs) → REFINE: artifacts without an executable verification command are validated through a self-review loop.

  1. CRITIQUE — re-read the artifact and identify the "3 weakest points." Be specific about where and why each is weak. Weakness types: insufficient evidence / logical leap / vague wording / missing perspective / structural imbalance / potential for reader misunderstanding
  2. REWRITE — rewrite only the identified weaknesses, once. Leave strong parts untouched.
  3. DELTA CHECK — compare original vs rewrite:
    • improved → adopt rewrite → FINALIZE
    • negligible difference or worse → keep original → FINALIZE
    • can't tell → note "REFINE performed, improvement uncertain" → FINALIZE
    • Self-judgment caveat: CRITIQUE → REWRITE → DELTA CHECK is the same agent grading its own work — a self-report, not an independent check. Treat "improved" as a working judgment, not proof. For high-stakes artifacts (specs, published content, anything a real decision gets made from), route the result through a separate reviewer pass after FINALIZE instead of trusting DELTA CHECK alone.
  4. 1-round limit — REFINE runs at most once. A second round has diminishing returns. No infinite self-correction loops.

REFINE eligibility criteria:

  • DONE EVIDENCE has an executable command → VERIFY
  • DONE EVIDENCE is a content criterion ("includes X", "explains Y", "analyzes Z") → REFINE
  • Both present → VERIFY first, then REFINE after passing (code + docs delivered together)
FINALIZE
  • After goal achieved, no additional refactoring.
  • Clean up: temp code, debug prints, failed experiments, junk files.
  • Before reporting: re-check scope + verification — didn't touch Exclude, met DONE EVIDENCE.
  • Comprehension check: can the key change be explained in one sentence? If not, flag ⚠️ High complexity — review recommended in OUTPUT. Code you can't explain is debt.
OUTPUT
markdown
## Result

**Changed files**: [list]
**Key changes**: [what and why]
**Completion evidence**: [commands run + results] or [REFINE: original→rewrite DELTA summary]
**Verification tool calls**: [actual tools invoked during VERIFY — GroundEval principle] or [REFINE: 3 CRITIQUE weaknesses + REWRITE fixes]
**Verification**: [passed/failed/not run — each with reason] or [REFINE: adopted/kept original/uncertain]
**Risks/trade-offs**: [if any]
**Remaining known issues**: [if any]
**Follow-up work**: [if any]

**Final status**: WORKING / PARTIAL / BROKEN / BLOCKED
  • PARTIAL: partially working, list specific defects
  • BROKEN: core functionality not working, state cause
  • BLOCKED: cannot proceed due to an external unresolved dependency or pending user approval — not a code defect. List exactly what's being waited on (which dependency, which decision, from whom)
  • Claiming "done" while PARTIAL/BROKEN/BLOCKED = success masquerading (B1 violation)
B4. STOP RULES (halt and ask — no progress until answered)
#ConditionAction
S1Goal splits into 2+ independent goals"Goal is branching. Which one first?"
S2Input missing/contradictorySpecify exactly what's ambiguous
S3Need to change SCOPE Exclude area"Need to modify X but it's Excluded. Allow?"
S4Destructive / external side effect needed"DB deletion/API call/push needed. Proceed?"
S5Insufficient confidence in root cause after the PLAN GATE discriminating check"Not sure if cause is A or B"
S6Same blocker repeated (2+ times) — stagnation circuit breakerAsk one question before escalating: "Is this repetition a problem with how DO is attempting it, or was the GOAL input sheet (GOAL / DONE EVIDENCE / SCOPE) set up wrongly from the start?" If it looks like an execution error (approach problem), report "Same problem repeating. Need to change the DO approach" and escalate to a human. If it looks like a design error (the input sheet itself), report "This repetition looks like a problem with the GOAL input sheet design, not the execution approach — the input sheet needs rewriting" and propose rewriting the input sheet to a human instead of retrying DO. In both cases there is no auto-retry. What counts as a repeat: count a repeat only when the command, input, working tree and environment are unchanged and no new information arrived. If the cause is outside the code (missing dependency or command, permissions, service down, port in use, auth, disk), report it separately as an environment error — distinct from execution and design errors — and give the human one command to run instead of retrying source edits. A timeout or forced kill is not automatically classified as an environment error
S7Already aware that execution evidence (a deterministic oracle — a failing test, a broken existing contract) contradicts an explicit user instruction — an awareness-is-not-resistance response [borrowed from Blind Obedience 07385]STOP before forcing the implementation through: "The instruction contradicts execution evidence: [evidence]. Proceed anyway?" Even after approval, do not paper over it with a later self-directed autonomous fix (a Ghost Error cannot be recovered by iterative post-hoc correction) — report the outcome exactly as it is
S8ATTACK Tier-1 matches one of its six escalation conditions"This change is a candidate for adversarial pre-implementation review: [matching condition]. Dispatch doubt-reviewer (or equivalent), then resume with its verdict (proceed / revise first / escalate)." If the user explicitly says "just proceed," continue on the Tier-0 result alone — that's an intentional override, not a bypass

S7 scope: "execution evidence" applies only to a code context where a deterministic oracle exists (tests, type checker, an existing API/contract). The REFINE path (non-code deliverables — prose, analysis, design docs) is out of scope, since there is no objective answer to contradict. S7 is the epistemic opposite of S5 (uncertain root cause) — S7 blocks a model that is confident yet wrong from complying anyway.

Show full SKILL.md (1,632 more words)Show less
B5. Long-running Tasks
  • Short status reports at major milestones — separate completed vs incomplete.
  • Save current state in .goal-lock-progress.md (session crash protection).
  • Resume from last verification point. Never restart from scratch.
  • At BUDGET 80% or extended stall → report status, ask whether to continue.
  • Early self-doubt boundary: long-reasoning models have been shown to misjudge remaining budget by up to 24%, triggering premature self-doubt ("I'm probably about to run out") that causes early abandonment or inefficient hedging — measured at -15pp accuracy and +54% token use on eventual success. When a hard counter like budget.remaining() is available, trust that measurement over your own felt pressure — don't shrink and quit early when the actual budget is still fine.
  • Constraint re-echo: at each BUDGET-80% checkpoint or progress-resume point, echo the GOAL input sheet's CONSTRAINTS/SCOPE-Exclude verbatim, separately from the status report. This is a static check against constraints quietly falling out of view during long tasks as attention shifts to raw progress — a full separate memory agent or a learned injection-timing policy would be overkill for this scale of harness.
B5.1 Physical Completion Gate (Stop Hook)

Separately from self-reported progress via .goal-lock-progress.md, a Stop hook can be registered to intervene physically at session-end time — checking whether a progress file's "current step" is still non-empty when the agent attempts to end its response, and blocking once per session if so (cap=1, to avoid infinite re-blocking).

  • Trigger: session Stop event (when the agent attempts to end its response)
  • cap=1: gate intervention limited to once per session — prevents infinite re-blocking
  • What it checks: whether a .goal-lock-progress.md-style progress file exists with a non-empty "current step," meaning DONE EVIDENCE verification may not have run before the agent tried to end the session
  • L1 (prompt) vs L2 (physical): B1/B3 VERIFY is an L1 discipline (the model is expected to follow it). A Stop hook is L2 (tool/hook-level) enforcement — it blocks session termination itself even if the model forgets the discipline.
  • On failure: fail-open — if the progress file can't be read or the block-count file can't be written, let termination proceed without blocking (avoid false blocks).

Order gate (loophole closure): even when the check above passes (progress file's current step is empty), a second pass can reconstruct the deterministic order of tool calls from the session transcript — if no verification-class command (test runner / linter / typechecker / diff) appears after the last file-modifying edit, block once as UNVERIFIED-CHANGE. This closes the loophole where verification passes, the agent makes one more edit, and then declares completion without re-verifying.

Document-edit exception: the gate's own progress-tracking file self-updating its "current step" doesn't count as a real change requiring re-verification — otherwise the gate would perpetually re-trigger on its own bookkeeping writes. Other document edits (a report, a spec, a design doc) do still count as a change, but once that document has been re-read afterward (mirroring the REFINE loop's CRITIQUE re-read), the gate can pass without requiring a verification-class Bash command — code files get no such exception and still need one. Recognize wrapped verification commands: a verification-class command run through a package-manager wrapper (poetry run pytest, uv run pytest, pipenv run pytest, npx jest) or a named script segment (npm run test:unit) still counts as verification — don't require the bare binary name.

Status: reference implementation not shipped. This section specifies the intended behavior a Stop hook of this kind should have — this repo does not ship the hook script, the settings.json wiring that registers it, or a test file for it. A team adopting B5.1 has to write and test that hook itself; until then, treat every claim below as a design spec, not evidence that the gate is running. The hook should block termination only when all four of the following hold (AND, not OR):

  1. A Stop event has actually fired for this session.
  2. The re-entrancy flag (e.g. stop_hook_active) is not already true — infinite-loop guard: without this, the hook's own block can trigger another Stop event and re-block itself forever.
  3. The per-session block cap (cap=1) has not already been spent.
  4. Transcript reconstruction shows no verification-class command after the last file-modifying edit.

If any one of the four is false, termination proceeds unblocked. Known limitation (document it, don't silently claim full coverage): file changes made through shell text-editing commands (e.g. sed -i) or custom verification scripts outside the standard test/lint/typecheck vocabulary aren't detected. A gate like this must carry a small self-test suite (synthetic transcripts exercising each of the 4 conditions, pass and fail cases both) that is re-run whenever the gate's logic is modified, to confirm the change didn't silently disable detection.

This pattern implements the L2 (no tool provided / physical block) layer of a 4-level safety framework: prompt rules alone (L1) can be forgotten by the model; a hook enforced at the tool/session layer (L2) cannot.

B5.2 Ralph Mode — Context-Isolation Alternative

The default B5 loop assumes context continuity — the same session carries forward, still referencing the accumulated conversation. In an unattended, long-running autonomous loop (an overnight unmanned batch, a pipeline that auto-retries N times with no approval gate between rounds), that continuity itself becomes the risk — a wrong assumption, an accumulated rationalization, or drift from one round carries straight into the next with nothing there to interrupt it. Ralph Mode [the "run fresh agent instances in a loop" pattern popularized as the Ralph Wiggum technique; not to be confused with the B1 "Ralph Wiggum" masquerading pattern above, which is about premature completion signaling, not context isolation] is a structural alternative for exactly that situation.

  • Each round inherits nothing from prior rounds — no carried-over parent/child session context. Every round starts in a genuinely fresh context.
  • State crosses rounds through exactly two channels: (1) the shared workspace itself (the actual filesystem artifacts — a code/doc diff is its own evidence), and (2) a single bounded, structured handoff — not a free-text summary but fixed fields: status (active|blocked|complete) / summary (1-3 sentences) / evidence (commands run + results) / next_steps (what the next round should do) / blocker (if any). The handoff supplements the workspace, it doesn't replace it — it's an instruction about what's left, not a narrative of what happened.
  • When to use this instead of default B5: only for unattended execution stretches where no human is present between rounds to correct direction (an autonomous overnight batch, an auto-retry pipeline running N times without approval). In an ordinary session where a human is watching every turn, default B5 stays more efficient (no context-rebuild cost) — don't switch just because Ralph Mode is available.
  • Relationship to B4 S6 (same blocker repeated 2+ times → escalate): S6 still applies under Ralph Mode. If the handoff's blocker field stays the same across rounds — as long as a short history of recent handoffs is kept — the repetition is detectable even with no memory carried forward, and S6 still routes to human escalation instead of letting memoryless restarts continue indefinitely. The same repeat definition applies: a round only counts as a repeat if nothing (input, tree, environment, information) changed between rounds.

Safety Layers

goal-lock's own actions differ in how reversible they are. Match the defense to the reversibility tier, and never let a single layer carry anything above "trivially reversible" alone:

ReversibilityExample goal-lock actionRequired defense
Easy — local, no external effectLocal file edit, .goal-lock-progress.md checkpoint writeLoop discipline (B1–B3) is sufficient
Costly — local but expensive to redoLocal deletion, force-terminating a stalled sub-stepLoop discipline + explicit stop-and-ask (S3/S4)
Hard — touches external/shared historygit push, remote branch changes, migrationsLoop discipline + STOP RULE + explicit user confirmation before executing
Unrecoverable — external side effect, no undoDB DROP/TRUNCATE, a live API call with real-world effect, secret exposureLoop discipline + STOP RULE + user confirmation + an independent post-hoc review pass

Layers stack upward only — moving to a higher tier adds a layer, it never drops a lower one. An action dressed up as "just this once, low risk" is still evaluated at its actual reversibility tier, not the tier that's convenient for the agent in the moment.


Task Templates (optional — quick start)

bug-fix
GOAL: Fix [symptom]
DONE EVIDENCE: Reproduction test → PASS + all existing tests PASS
SCOPE Exclude: No API signature changes, no new features
feature
GOAL: Implement [feature]
DONE EVIDENCE: N new tests PASS + all existing tests PASS + render/behavior confirmed
SCOPE Exclude: No changes to existing feature behavior
refactor
GOAL: Refactor [target] — no behavior change
DONE EVIDENCE: All existing tests PASS (same test count) + before/after diff scope confirmed
SCOPE Exclude: No new features, no API signature changes
migration
GOAL: Migrate [target]
DONE EVIDENCE: Migration succeeds both up and down + all existing tests PASS
CONSTRAINTS: Existing data must be preserved, destructive changes require approval

Scope Boundary

DoesDoes NOT
Auto-fill input sheet + get user confirmationStart work without user confirmation
Enforce PLAN→DO→VERIFY→FINALIZE→OUTPUT loopSkip loop steps
Halt immediately on STOP RULESContinue with "probably fine"
Detect and block success masquerading patternsDesign code logic (that's the developer/agent's role)
Save .goal-lock-progress.md checkpointsManage memory/handoff systems

Invariants (never violate)

  1. DONE EVIDENCE must be actually executed: Run all items before OUTPUT. "It should pass" is not verification. Violation → unverified code reported as "done."

  2. SCOPE Exclude is absolute: Need to touch Exclude → S3 STOP. Only proceed after user approval. Violation → unintended changes reach production.

  3. Success masquerading detected → OUTPUT PARTIAL/BROKEN: If B1 pattern found in code, mark that verification as FAIL and downgrade OUTPUT to PARTIAL/BROKEN. Violation → false success report.

  4. Incomplete input sheet → no work: Any of 7 (Quick: 3) fields missing/contradictory → STOP. Don't guess. Violation → unclear goal → rework.

Error Recovery

Failure TypeRecovery
VERIFY failureReturn to PLAN for root cause analysis → re-DO. Same approach fails twice → S6 STOP
Tool failure (Bash/Edit)1 retry → report "tool failure" + suggest alternative
BUDGET exceededReport status + clearly separate done/not-done → user decides

Rationalization Table

RationalizationCounter
"Simple change, don't need the input sheet"Quick mode exists. Can't fill 3 fields → goal is unclear
"Most of VERIFY passed so it's WORKING"One FAIL = PARTIAL. "Most" ≠ "all"
"Test was too strict so I skipped it"Success masquerading B1 violation. If test is strict, fix the code
"Doing the refactor together is more efficient"SCOPE Exclude violation. Achieve goal first, then separate goal-lock for refactor
"This should be fine"DONE EVIDENCE not executed = Invariant 1 violation
"I'll add tests later"If DONE EVIDENCE includes tests, now. If not, they were never needed
"Goal-lock format is overhead, just look at the result"Format IS discipline. Without the input sheet, scope drift and masquerading detection opportunities vanish. Quick mode is 10 seconds

© AlexZio00, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in goal-lock of AlexZio00/sovereign-skills.

  • SKILL.md
  • .claude-plugin/plugin.json
  • agents/openai.yaml

Open the folder on GitHubat commit c062683

Compare with similar skills

Goal Lock next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Goal Lock compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Goal Lock this skillAlexZio00/sovereign-skills140—~9.4kAutomated safety check: PassMIT
Goalscodewhale-hq/Codewhale41k—~273Automated safety check: PassMIT
Agent Goal Plannerruvnet/ruflo74k2 repos~842Automated safety check: PassMIT
Goal Planruvnet/ruflo74k—~807Automated safety check: NotesMIT
Goal Loopsickn33/agentic-awesome-skills47k1 repos~2.6kAutomated safety check: PassMIT
Agent Code Goal Plannerruvnet/ruflo74k2 repos~3.6kAutomated safety check: PassMIT

Similar skills

  • Goals

    codewhale-hq/Codewhale

    Set, review, and update the user's goals. An agent skill from codewhale-hq/Codewhale.

    41k GitHub stars~273 tokensUpdated today
    Auto-check passed
  • Agent Goal Planner

    ruvnet/ruflo

    Agent skill for goal-planner - invoke with $agent-goal-planner

    74k GitHub starsUsed in 2 repos~842 tokens
    Auto-check passed
  • Goal Plan

    ruvnet/ruflo

    Create and execute Goal-Oriented Action Plans (GOAP) with precondition analysis, cost optimization, and adaptive replanning

    74k GitHub stars~807 tokensUpdated yesterday
    Auto-check: notes
  • Goal Loop

    sickn33/agentic-awesome-skills

    Draft and explain persistent goal-loop prompts for long-running agent work with clear stop conditions.

    47k GitHub starsUsed in 1 repo~2.6k tokens
    Agent WorkflowsAuto-check passed
  • Agent skill for code-goal-planner - invoke with $agent-code-goal-planner

    74k GitHub starsUsed in 2 repos~3.6k tokens
    Agent WorkflowsAuto-check passed
  • Fable Goal

    alirezarezvani/claude-skills

    Convert a rambling description of a desired outcome into one polished, autonomous /goal prompt ready to paste into a fresh session.

    28k GitHub stars~2.4k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed

More from AlexZio00/sovereign-skills

All 17 skills in this repo
  • Project Overview

    AlexZio00/sovereign-skills

    A skill your agent uses when the user wants a deterministic cross-project status map generated from registered projects' session handoffs.

    140 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Scope

    AlexZio00/sovereign-skills

    Scope definition before implementation — two modes. An agent skill from AlexZio00/sovereign-skills.

    140 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Project Init

    AlexZio00/sovereign-skills

    Interview-based project setup — generates CLAUDE.md, ROADMAP, .gitignore, .env.example from scratch.

    140 GitHub stars~4.1k tokensUpdated 2 days ago
    Auto-check: notes
  • Collab Audit

    AlexZio00/sovereign-skills

    This skill should be used when the user types /collab-audit or requests AI collaboration diagnosis.

    140 GitHub stars~8k tokensUpdated 2 days ago
    Auto-check passed
  • Doc Drift

    AlexZio00/sovereign-skills

    A skill your agent uses when the user wants to audit the memory and documents Claude Code loads into context — CLAUDE.md (user global + project + nested), MEMORY.md, @imports, .claude/skills…

    140 GitHub stars~6.2k tokensUpdated 2 days ago
    Auto-check passed
  • Session Checkpoint

    AlexZio00/sovereign-skills

    A skill your agent uses when saving session state before context compaction, switching tasks, or ending a session.

    140 GitHub stars~14k tokensUpdated 2 days ago
    Auto-check passed

Questions about Goal Lock

What does Goal Lock do?

Agent Discipline Engine — lock the goal, run PLAN→DO→VERIFY→FINALIZE→OUTPUT loop, detect success masquerading. Goal Lock is an agent skill from AlexZio00/sovereign-skills. Agent Discipline Engine — lock the goal, run PLAN→DO→VERIFY→FINALIZE→OUTPUT loop, detect success masquerading.

How do I install Goal Lock in Claude Code?

Run `npx skills add AlexZio00/sovereign-skills --skill goal-lock -a claude-code`. Or copy the skill folder (goal-lock in AlexZio00/sovereign-skills) into .claude/skills/goal-lock in your project. Claude Code loads it when a task matches its description.

How do I install Goal Lock in Codex?

Run `npx skills add AlexZio00/sovereign-skills --skill goal-lock -a codex`. Or copy the skill folder (goal-lock in AlexZio00/sovereign-skills) into .agents/skills/goal-lock in your project. Codex loads it when a task matches its description.

Can I use Goal Lock in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexZio00/sovereign-skills --skill goal-lock -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/goal-lock, .gemini/skills/goal-lock, .github/skills/goal-lock and .opencode/skills/goal-lock in your project.

What does Goal Lock need to run?

Going by SKILL.md and its folder, Goal Lock needs the command-line tools its instructions call (npm, npx, go, cargo, git and pytest).

Does Goal Lock access the network?

SKILL.md contains no URLs. Its commands use npm, npx and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Goal Lock safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Goal Lock use?

Goal Lock is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Goal Lock use?

About 9.4k tokens (SKILL.md is roughly 38k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Goal Lock?

Skills that share tags, products or a category with Goal Lock: Goals (codewhale-hq/Codewhale, 41k stars), Agent Goal Planner (ruvnet/ruflo, 74k stars), Goal Plan (ruvnet/ruflo, 74k stars) and Goal Loop (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Goal Lock?

AlexZio00 (a GitHub user) maintains it in AlexZio00/sovereign-skills, which has 140 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: AlexZio00/sovereign-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.