Agent skill

Agent Prompt Patterns

by LeoYeAI in LeoYeAI/openclaw-master-skills

Battle-tested prompt patterns for production AI agents. An agent skill from LeoYeAI/openclaw-master-skills.

MITAuto-check passedAI & LLM Engineering

Install Agent Prompt Patterns

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill agent-prompt-patterns -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills agent-prompt-patterns --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-prompt-patterns .claude/skills/agent-prompt-patterns && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-prompt-patterns
GitHub stars
2.2k
Token cost
~6.8k tokens
SKILL.md length
1,479 words
Files
2
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

Battle-tested prompt patterns for production AI agents. An agent skill from LeoYeAI/openclaw-master-skills.

  • Works in 10 steps: Consumer-First Design → Proof-of-Work Enforcement → Cascading Validation → …
  • Designing agent behavior
  • SKILL.md covers When to Use, When NOT to Use, 1. Consumer-First Design and 2. Proof-of-Work Enforcement, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Prompt Patterns is an agent skill from LeoYeAI/openclaw-master-skills. Battle-tested prompt patterns for production AI agents. Covers consumer-first design, deletion test, cascading validation, advisory mode tiers, proof-of-work enforcement, heartbeat protocol, contradiction detection, WAL protocol, rule escalation ladder, and cross-validation patterns. Use when designing agent behavior, enforcing reliability, or building agent operating manuals.

Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `_meta.json`).

It sits in AI & LLM Engineering, covering Prompt engineering, Verification before completion and Machine learning. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • Designing agent behavior
  • Enforcing reliability
  • Building agent operating manuals

Example prompts

  • “/agent-prompt-patterns”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Consumer-First Design
  2. Proof-of-Work Enforcement
  3. Cascading Validation
  4. Advisory Mode Tiers
  5. Completion Contracts
  6. Cross-Validation
  7. Rule Escalation Ladder
  8. Heartbeat Protocol
  9. Contradiction Detection
  10. Tight Harness Principle

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown, bash and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Prompt Patterns loads about 6.8k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 1,479 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 1,479 words, ~6,779 tokens.

Download SKILL.mdSave it as .claude/skills/agent-prompt-patterns/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
agent-prompt-patterns
description
Battle-tested prompt patterns for production AI agents. Covers consumer-first design, deletion test, cascading validation, advisory mode tiers, proof-of-work enforcement, heartbeat protocol, contradiction detection, WAL protocol, rule escalation ladder, and cross-validation patterns. Use when designing agent behavior, enforcing reliability, or building agent operating manuals.
license
MIT

Agent Prompt Patterns

Battle-tested patterns for agents that ship, not agents that demo. If your agent works in a live-fire notebook but breaks in production, you have a demo, not an agent.


When to Use

  • Designing a new agent's behavioral rules and operating manual
  • An agent is hallucinating completions, skipping steps, or claiming work it didn't do
  • Building multi-agent pipelines where output quality compounds (or collapses)
  • Setting up human-in-the-loop approval tiers for different risk levels
  • Enforcing reliability in automated workflows (cron jobs, scheduled tasks, pipelines)
  • Writing AGENTS.md or operating manuals for production agent workspaces
  • Debugging why an agent keeps violating rules you've already stated
  • Evaluating whether an agent should exist at all (deletion test)
  • Building harnesses that make autonomy safe and useful

When NOT to Use

  • One-shot prompts with no agent persistence — these patterns assume continuity
  • Pure chatbot / conversational UX with no action-taking capability
  • Academic prompt engineering research — these are production patterns, not benchmarks
  • Agents with no filesystem, no tool access, and no side effects — nothing to harness
  • You're still in the "make it work at all" phase — get basic functionality first, then harden

1. Consumer-First Design

Principle: Every agent output must have a named consumer. If nobody uses the output, the agent shouldn't exist.

This is the most important pattern because it kills bloat before it starts. Agents proliferate. Each one feels useful when you build it. Six months later you have 14 agents and can't remember what half of them do.

The Deletion Test

Ask: If I delete this agent, which other agent's work breaks?

If the answer is "nothing" or "I'm not sure," the agent is a vanity project.

markdown
# Agent Registry (in AGENTS.md)

## daily-digest
- **Consumers:** Sam (morning briefing), weekly-report agent (aggregation)
- **Deletion impact:** Sam loses morning summary, weekly-report loses daily inputs
- **Verdict:** KEEP

## inbox-sorter
- **Consumers:** None identified
- **Deletion impact:** Unknown
- **Verdict:** CANDIDATE FOR REMOVAL — validate or kill within 7 days
How to Apply

Every agent entry in your operating manual should answer:

  1. Who consumes this output? (name the human or agent)
  2. What format do they need? (not what's convenient to produce)
  3. What breaks if this stops? (the deletion test)
  4. What's the feedback loop? (how does the consumer signal quality issues?)

If an agent produces beautiful summaries that nobody reads, it's burning tokens for nothing.

Anti-Pattern: The "Nice to Have" Agent
markdown
# BAD: No consumer, no deletion impact
## sentiment-tracker
Monitors social media sentiment about our brand.
Runs daily. Outputs to sentiment-log.md.

# GOOD: Named consumer, clear dependency
## sentiment-tracker
Monitors social media sentiment for weekly-report.
Consumer: weekly-report agent (pulls sentiment delta for executive summary)
Deletion impact: weekly-report loses sentiment section; Sam must manually check socials
Format: JSON with {platform, score_delta, top_mentions[3]}

2. Proof-of-Work Enforcement

Principle: Never claim done unless the action actually started. Every status update needs proof — PID, file path, URL, command output. No proof = didn't happen. Write first, speak second.

This pattern exists because LLMs are pathological completers. They want to say "Done!" because that's the satisfying end of a sequence. The problem is they'll say "Done!" before doing anything, or after attempting something that silently failed.

The Rule
STATUS UPDATE FORMAT:
- "Started X" → must include: PID, command, or file path
- "Completed X" → must include: output snippet, file path, or URL
- "Failed X" → must include: error message, what was tried
- "Skipped X" → must include: reason with evidence
Examples
markdown
# BAD: No proof
✅ Backed up database
✅ Sent daily digest email
✅ Rotated API keys

# GOOD: Every claim has evidence
✅ Backed up database → /backups/2026-03-15-db.sql.gz (43MB, sha256: a1b2c3...)
✅ Sent daily digest → Message-ID: <abc123@mail.example.com>, 3 recipients
✅ Rotated API keys → new key fingerprint: sk-...x4f2, old key revoked at 14:32 UTC
Implementation Pattern
bash
# In a script gate or agent wrapper:
run_with_proof() {
  local task="$1"
  shift
  local output
  output=$("$@" 2>&1)
  local exit_code=$?

  if [ $exit_code -eq 0 ]; then
    echo "DONE: $task | proof: $(echo "$output" | tail -3)"
  else
    echo "FAIL: $task | exit=$exit_code | error: $(echo "$output" | tail -5)"
  fi
  return $exit_code
}

# Usage:
run_with_proof "database backup" pg_dump -Fc mydb -f /backups/latest.dump
Agent Operating Manual Rule
markdown
## Proof-of-Work (AGENTS.md entry)

NEVER say "done" without evidence. For every completed action, include at least one of:
- File path of output produced
- PID of process started
- URL of resource created/modified
- Command output (truncated to last 5 lines)
- Screenshot or hash of artifact

If you cannot produce proof, say "ATTEMPTED but cannot verify" and explain why.

3. Cascading Validation

Principle: Dependent sequential steps — each task validates the previous output before starting its own work. Failures loop back with fix instructions, not silent continuations.

Cascading validation prevents the "garbage in, garbage out" problem in multi-step pipelines. Without it, step 3 happily processes the corrupt output of step 2, and you don't discover the problem until step 7.

The Pattern
Step 1: Produce output A
Step 2: Validate A meets spec → if invalid, return to Step 1 with fix instructions
Step 3: Use validated A to produce B
Step 4: Validate B meets spec → if invalid, return to Step 3 with fix instructions
...
Example: Content Pipeline
markdown
## Newsletter Pipeline (cascading validation)

### Step 1: Research
- Output: research-notes.md
- Validation: must contain ≥ 3 sources, each with URL and date
- Failure: "Research incomplete — need 3+ sourced items. Currently have {n}. Add more."

### Step 2: Draft
- Input: validated research-notes.md
- Pre-check: verify research-notes.md passes Step 1 validation (don't trust upstream)
- Output: draft.md
- Validation: 400-800 words, includes all research items, no placeholder text
- Failure: "Draft {issue}. Fix and resubmit. Do not proceed to editing."

### Step 3: Edit
- Input: validated draft.md
- Pre-check: verify draft.md passes Step 2 validation
- Output: final.md
- Validation: grammar check passes, links resolve, formatting correct
- Failure: "Edit issues found: {list}. Return to editing. Do not publish."

### Step 4: Publish
- Input: validated final.md
- Pre-check: verify final.md passes Step 3 validation
- Gate: HUMAN APPROVAL REQUIRED before publish
Key Rule: Never Trust Upstream

Even if Step 1 "passed," Step 2 should re-validate Step 1's output before proceeding. This catches:

  • Race conditions (output modified between steps)
  • Silent corruption (file written but content wrong)
  • Upstream validation bugs (Step 1's validator had a gap)
Implementation
python
def cascading_step(input_path, input_validator, processor, output_validator, max_retries=3):
    """Each step validates its input AND its output."""
    # Validate input (don't trust upstream)
    input_valid, input_errors = input_validator(input_path)
    if not input_valid:
        return {"status": "BLOCKED", "reason": f"Input validation failed: {input_errors}"}

    for attempt in range(max_retries):
        output = processor(input_path)
        output_valid, output_errors = output_validator(output)
        if output_valid:
            return {"status": "DONE", "output": output, "attempts": attempt + 1}
        # Loop back with fix instructions
        processor = make_fix_processor(processor, output_errors)

    return {"status": "FAILED", "reason": f"Failed after {max_retries} attempts", "last_errors": output_errors}

4. Advisory Mode Tiers

Principle: Not all actions carry the same risk. Categorize agent capabilities into tiers with different autonomy levels and approval requirements.

The mistake people make is binary: either the agent can do everything, or it can do nothing. Tiers let you give autonomy where it's safe and require approval where it's not.

The Four Tiers
TierRiskProbationGraduationExample
LowReversible, internal only3 daysSelf-promote after clean streakRead files, search, summarize
MediumVisible to user, recoverable2 weeksHuman approves promotionCreate files, edit code, run tests
HighVisible to others, hard to reverse2 weeks minimumNever fully unsupervisedGit push, create PRs, post to Slack
RestrictedIrreversible or impersonation riskPermanentAlways draft-onlySend email from user's account, delete data, financial transactions
Critical Rule: Email = Restricted

Sending email from a user's account is always Restricted tier. No exceptions. No graduation. Always draft-only with human send.

Why: Email is identity. An AI sending email "as you" creates legal, professional, and trust risks that no amount of testing eliminates.

markdown
## Advisory Mode Configuration (AGENTS.md)

### Tier: Low (auto-approve after 3-day probation)
- Read any file in workspace
- Search codebase
- Generate summaries to memory files
- Run read-only API calls

### Tier: Medium (human approves after 2-week probation)
- Create/edit files in workspace
- Run test suites
- Generate reports
- Schedule cron jobs (read-only actions only)

### Tier: High (2-week probation, never fully autonomous)
- Git commit and push
- Create pull requests
- Post to Slack channels
- Modify cron jobs

### Tier: Restricted (always draft-only, human executes)
- Send email from user's account
- Delete files/data outside workspace
- Financial transactions (invoice, payment)
- Modify access controls or permissions
- Post to social media as user
Probation Protocol
markdown
## Probation Rules

1. New capability starts at its tier's probation period
2. During probation: agent proposes action, human approves/denies
3. Clean streak = no denials or corrections for full probation period
4. After clean streak:
   - Low: auto-promotes, agent logs the promotion
   - Medium: agent requests promotion, human approves
   - High: agent requests promotion, human approves, but spot-checks continue
   - Restricted: never promotes — always draft-only
5. Any denial during probation resets the probation clock
6. Graduated capability can be demoted if quality degrades

5. Completion Contracts

Principle: Every automated workflow needs binary done-criteria, observable evidence, staged approval, and timeout bounds. No "probably done" — it's done or it's not.

The Contract Template
markdown
## Completion Contract: {workflow name}

### Done Criteria (all must be true)
- [ ] {criterion 1} — verified by: {method}
- [ ] {criterion 2} — verified by: {method}
- [ ] {criterion 3} — verified by: {method}

### Evidence Required
- {artifact 1}: {location/format}
- {artifact 2}: {location/format}

### Approval Stages
1. Automated validation passes (criteria above)
2. Agent self-review (checklist)
3. Human approval (if tier requires it)

### Timeout
- Maximum duration: {time}
- On timeout: {action — alert human, retry once, abort}
- Escalation: {who gets notified}

### Rollback
- Revert procedure: {steps}
- Rollback trigger: {conditions}
Example: Deployment Contract
markdown
## Completion Contract: Production Deploy

### Done Criteria
- [ ] All tests pass on deploy branch — verified by: CI green check
- [ ] Docker image builds successfully — verified by: image SHA in registry
- [ ] Health check returns 200 — verified by: curl to /health within 60s
- [ ] No error spike in first 5 minutes — verified by: error rate < 0.1%

### Evidence Required
- CI run URL with green status
- Docker image SHA256
- Health check response (timestamp + status code)
- Error rate dashboard screenshot at T+5min

### Approval Stages
1. CI passes automatically
2. Agent verifies health check and error rate
3. Human confirms "deploy complete" in Slack

### Timeout
- Maximum duration: 15 minutes from deploy start
- On timeout: auto-rollback to previous version, alert #ops channel
- Escalation: page on-call engineer if rollback also fails

### Rollback
- Revert: deploy previous Docker image SHA
- Trigger: health check fails OR error rate > 0.5% OR human says "rollback"

6. Cross-Validation

Principle: Generate with one model, review with another. Different architectures catch different blind spots.

Single-model pipelines have correlated failure modes. If Claude hallucinates a fact, Claude reviewing its own work will often confirm the hallucination. A different model (or a human) brings uncorrelated errors.

The Sub-Agent QC Workflow
Produce (Sonnet) → Review (Sam) → Cross-check (GPT) → Incorporate → Deliver

This isn't about which model is "better." It's about error decorrelation. Each reviewer catches things the others miss.

Implementation
markdown
## Cross-Validation Protocol

### Step 1: Produce (Primary Model)
- Model: Claude Sonnet (fast, cost-effective for drafts)
- Output: first draft with citations

### Step 2: Review (Human)
- Reviewer: Sam
- Focus: factual accuracy, tone, strategic alignment
- Output: annotated draft with corrections

### Step 3: Cross-Check (Secondary Model)
- Model: GPT-4 or different Claude variant
- Prompt: "Review this document for factual errors, logical inconsistencies,
  and unsupported claims. Do not rewrite — only flag issues with explanations."
- Focus: catch blind spots the primary model and human missed
- Output: issue list with severity ratings

### Step 4: Incorporate
- Primary model incorporates human + cross-check feedback
- Changes tracked and justified

### Step 5: Deliver
- Final version with revision history
- Confidence rating based on number of issues found and fixed
When Cross-Validation Matters Most
  • Legal or compliance content — different models interpret regulations differently
  • Financial calculations — arithmetic errors are model-specific
  • Factual claims — hallucination patterns differ across architectures
  • Security reviews — different models catch different vulnerability classes
When It's Overkill
  • Internal notes nobody else will read
  • Ephemeral content (daily logs, scratch work)
  • Tasks where speed matters more than correctness
  • Outputs with automated validation (tests, linters) that catch errors mechanically

Show full SKILL.md (604 more words)Show less

7. Rule Escalation Ladder

Principle: Rules start as prose. If violated, they escalate to loaded rules. If violated again, they become script gates. Critical rules skip the ladder entirely.

The problem with prose rules is enforcement. An agent "knows" the rule but still violates it under pressure (long context, competing instructions, ambiguous situations). The escalation ladder adds mechanical enforcement for rules that matter.

The Three Levels
Level 1: Prose Rule (in AGENTS.md)
  "Don't send emails without approval"
  → Relies on agent reading and following the rule
  → Appropriate for: new rules, low-risk guidelines

Level 2: Loaded Rule (in decisions.md, checked at session start)
  "EMAIL_SENDING: RESTRICTED — always draft-only, never auto-send"
  → Agent must load and acknowledge before acting
  → Appropriate for: rules violated once, medium-risk operations

Level 3: Script Gate (mechanical enforcement)
  Pre-send hook checks for human approval token
  → Agent literally cannot bypass the rule
  → Appropriate for: rules violated twice, high-risk operations, critical rules
Escalation Protocol
markdown
## Rule Escalation (AGENTS.md)

### Escalation Triggers
- First violation of a prose rule → add to decisions.md as loaded rule
- Second violation (now a loaded rule) → implement as script gate
- Any violation of a critical rule → skip to script gate immediately

### Critical Rules (always script-gated)
- Sending email from user's account
- Deleting files outside workspace
- Financial transactions
- Modifying access controls
- Publishing to external platforms

### Currently Loaded Rules (decisions.md)
- See decisions.md for the current set — these are checked every session start
Script Gate Example
bash
#!/bin/bash
# scripts/gate-email-send.sh — mechanical enforcement of email restriction

APPROVAL_TOKEN_FILE="/tmp/.email-approval-$(date +%Y%m%d)"

if [ ! -f "$APPROVAL_TOKEN_FILE" ]; then
  echo "BLOCKED: Email sending requires human approval."
  echo "Human: run 'echo APPROVED > $APPROVAL_TOKEN_FILE' to authorize."
  exit 1
fi

APPROVAL=$(cat "$APPROVAL_TOKEN_FILE")
if [ "$APPROVAL" != "APPROVED" ]; then
  echo "BLOCKED: Approval token invalid."
  exit 1
fi

echo "GATE PASSED: Email send authorized for today."
# Proceed with email send
"$@"

# Consume the token (one-time use)
rm "$APPROVAL_TOKEN_FILE"

8. Heartbeat Protocol

Principle: Periodic health checks batched together. Context monitor, system health, memory maintenance — all in one scheduled pulse, not scattered across individual crons.

Heartbeat vs. Cron Decision
Use Heartbeat WhenUse Individual Cron When
Check is lightweight (< 30 seconds)Task is heavyweight (minutes)
Multiple checks share contextTask is completely independent
Failure in one check should inform othersTask has its own retry/error handling
You want a single "system status" viewTask needs its own schedule (not aligned)
Heartbeat Structure
markdown
## Heartbeat Protocol (runs every 4 hours)

### Phase 1: Context Monitor (5 seconds)
- Check MEMORY.md size (warn if > 200 lines)
- Check daily note exists for today
- Verify SOUL.md and AGENTS.md haven't been modified unexpectedly

### Phase 2: System Health (10 seconds)
- Disk space check (warn if < 10% free)
- Check if critical services are running (by PID file)
- Verify cron jobs are registered and last-ran within expected windows

### Phase 3: Memory Maintenance (15 seconds)
- Scan for contradictions (see Pattern 9)
- Archive daily notes older than 7 days
- Update system-health.json with current status

### Output
- Write to HEARTBEAT.md: timestamp, all-clear or issues found
- If issues found: list them with severity and suggested fix
- If critical issue: alert human immediately (don't wait for next heartbeat)
Implementation
bash
#!/bin/bash
# scripts/heartbeat.sh

HEARTBEAT_FILE="HEARTBEAT.md"
TIMESTAMP=$(date -u +"%Y-%m-%dT%H:%M:%SZ")
STATUS="ALL CLEAR"
ISSUES=""

# Phase 1: Context Monitor
MEMORY_LINES=$(wc -l < MEMORY.md 2>/dev/null || echo "0")
if [ "$MEMORY_LINES" -gt 200 ]; then
  ISSUES="$ISSUES\n- WARN: MEMORY.md is $MEMORY_LINES lines (limit: 200)"
  STATUS="ISSUES FOUND"
fi

if [ ! -f "memory/$(date +%Y-%m-%d).md" ]; then
  ISSUES="$ISSUES\n- INFO: No daily note for today"
fi

# Phase 2: System Health
DISK_FREE=$(df -h . | tail -1 | awk '{print $5}' | tr -d '%')
if [ "$DISK_FREE" -gt 90 ]; then
  ISSUES="$ISSUES\n- CRITICAL: Disk usage at ${DISK_FREE}%"
  STATUS="CRITICAL"
fi

# Phase 3: Memory Maintenance
# (Contradiction detection delegated to agent — see Pattern 9)

# Write heartbeat
cat > "$HEARTBEAT_FILE" << EOF
# Heartbeat
Last check: $TIMESTAMP
Status: $STATUS
$(if [ -n "$ISSUES" ]; then echo -e "\n## Issues\n$ISSUES"; fi)
EOF

echo "Heartbeat complete: $STATUS"

9. Contradiction Detection

Principle: Actively scan for conflicts between memory entries, between memory and SOUL, stale facts, and decision reversals. Don't wait for contradictions to cause errors — find them during maintenance.

Contradiction Types
TypeDescriptionExample
Memory-MemoryTwo memory entries say opposite things"Client prefers email" vs "Client prefers Slack"
Memory-SOULMemory contradicts core identity/rulesSOUL says "never auto-send email" but memory says "auto-send enabled for digest"
Stale FactsMemory entry is outdated"API endpoint: api.v1.example.com" when v1 was deprecated
Decision Reversaldecisions.md contradicts earlier decision without noting the change"Use PostgreSQL" then later "Use SQLite" with no migration note
Scan Protocol
markdown
## Contradiction Detection (run during heartbeat Phase 3)

### Scan Checklist
1. Load all memory entries with type=project and type=reference
2. For each pair, check for semantic conflicts:
   - Same topic, different conclusions
   - Same entity, different attributes
   - Same process, different steps
3. Load SOUL.md rules and check memory entries against each rule
4. Check decisions.md for entries that reverse previous decisions without rationale
5. Flag entries older than 30 days for staleness review

### Output Format
- If contradictions found: write to memory/contradictions-{date}.md
- Each entry: the two conflicting sources, the conflict, suggested resolution
- Critical contradictions (SOUL violations): alert immediately

### Resolution
- Human reviews contradictions list
- For each: keep A, keep B, merge, or delete both
- Update affected files
- Log resolution in decisions.md
Example Contradiction Report
markdown
# Contradictions Found — 2026-03-15

## CRITICAL: Memory-SOUL Conflict
- **SOUL.md line 23:** "Never auto-send email from user's account"
- **memory/gmail-daily-summary.md:** "Auto-send daily digest at 7am"
- **Resolution needed:** Either update SOUL or disable auto-send
- **Severity:** CRITICAL — active violation of core rule

## WARN: Memory-Memory Conflict
- **memory/wallets-onchain-identity.md:** "Primary wallet: 0xABC..."
- **memory/2026-03-12.md:** "Migrated primary wallet to 0xDEF..."
- **Resolution needed:** Update wallets file with new primary address
- **Severity:** MEDIUM — stale reference may cause wrong wallet usage

## INFO: Stale Entry
- **memory/starred-repos.md:** Last updated 45 days ago
- **Suggestion:** Review and refresh or archive
- **Severity:** LOW

10. Tight Harness Principle

Principle: Autonomy gets useful when the harness is tight. Don't sell agents — sell harnesses. An agent without a harness is a liability. A harness without an agent is just a script.

The Five Harness Components

Every autonomous agent operation needs all five:

ComponentQuestionExample
Objective MetricHow do we know it worked?"Test suite passes" not "code looks good"
Bounded ScopeWhat can it touch?"Only files in /src/api/" not "any file"
Time BudgetWhen does it stop?"15 minutes max" not "when it's done"
ReversibilityCan we undo it?"Git branch, not direct commit to main"
ObservabilityCan we see what it did?"Full command log" not "trust me"
The Key Insight

Most agent failures aren't capability failures — they're harness failures. The agent could do the task, but:

  • Nobody defined "done" objectively (no metric)
  • It modified files it shouldn't have (no scope bound)
  • It ran for 3 hours burning tokens (no time budget)
  • It pushed directly to main (no reversibility)
  • Nobody could tell what it did (no observability)
Harness Configuration Example
markdown
## Harness: Automated PR Review Agent

### Objective Metric
- All review comments reference specific code lines
- No false positive rate > 10% (tracked over 2-week window)
- Review completed within 5 minutes of PR open

### Bounded Scope
- READ: any file in the repository
- WRITE: only PR comments via GitHub API
- CANNOT: approve PRs, merge PRs, modify code, close PRs

### Time Budget
- Maximum 5 minutes per PR
- Maximum 20 PRs per day
- On budget exceeded: skip PR, log reason, alert human

### Reversibility
- All comments can be deleted
- No permanent actions taken
- Human can dismiss any comment

### Observability
- Every review logged to reviews/{date}-{pr-number}.md
- Includes: files reviewed, issues found, comments posted, time taken
- Weekly accuracy report generated automatically
Selling Harnesses, Not Agents

When someone asks "can your agent do X?" — the right answer is "here's the harness that makes X safe":

BAD:  "Yes, our agent can deploy to production!"
GOOD: "Yes, with this harness: deploys only to staging first, requires health
       check pass, auto-rollback on error spike, human approval for prod
       promotion, full audit log, 15-minute timeout."

The harness IS the product. The agent is just the engine inside it.


Quick Reference: Pattern Selection Guide

SituationPattern
"Should this agent exist?"Consumer-First Design (#1)
"Agent says it's done but I don't believe it"Proof-of-Work (#2)
"Multi-step pipeline keeps producing garbage"Cascading Validation (#3)
"How much autonomy should this agent have?"Advisory Mode Tiers (#4)
"When is this workflow actually done?"Completion Contracts (#5)
"Agent keeps making the same kind of error"Cross-Validation (#6)
"Agent keeps violating a rule"Rule Escalation Ladder (#7)
"How do I monitor agent health?"Heartbeat Protocol (#8)
"Agent's memory is inconsistent"Contradiction Detection (#9)
"How do I make autonomy safe?"Tight Harness Principle (#10)

Combining Patterns

These patterns are composable. A production agent typically uses several together:

Consumer-First Design    → Does this agent need to exist?
  ↓ yes
Advisory Mode Tiers      → What can it do autonomously?
  ↓ configured
Completion Contracts     → How do we know each task is done?
  ↓ defined
Cascading Validation     → How do multi-step tasks flow?
  ↓ piped
Proof-of-Work            → How do we verify claims?
  ↓ enforced
Cross-Validation         → How do we catch blind spots?
  ↓ reviewed
Rule Escalation Ladder   → How do we handle violations?
  ↓ gated
Heartbeat Protocol       → How do we monitor ongoing health?
  ↓ pulsing
Contradiction Detection  → How do we keep memory consistent?
  ↓ clean
Tight Harness            → How do we keep all of this safe?

Start with #1 (does this agent need to exist?) and #10 (is the harness tight?). Add the others as complexity demands.

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/agent-prompt-patterns of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • _meta.json

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Agent Prompt Patterns next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Prompt Patterns compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Prompt Patterns this skillLeoYeAI/openclaw-master-skills2.2k—~6.8kAutomated safety check: PassMIT
Senior Prompt Engineermaslennikov-ig/claude-code-orchestrator-kit2603 repos~1.4kAutomated safety check: PassCustom licence
Create System Promptpnp/copilot-prompts892—~3.1kAutomated safety check: PassMIT
DSPy Language Model ProgrammingOrchestra-Research/AI-Research-SKILLs13k9 repos~3.8kAutomated safety check: PassMIT
Building Agent Systemstelagod/code-abyss243—~691Automated safety check: PassMIT
Agentsop Dspyagentsope/SkillAlchemy466—~7kAutomated safety check: PassMIT

Similar skills

  • Senior Prompt Engineer

    maslennikov-ig/claude-code-orchestrator-kit

    Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.

    260 GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Create System Prompt

    pnp/copilot-prompts

    This skill should be used when the user asks to "create an agent instruction", "add agent instructions", "scaffold an agent sample", "create a system prompt sample", "add a system prompt", "create a…

    892 GitHub stars~3.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • DSPy Language Model Programming

    Orchestra-Research/AI-Research-SKILLs

    Teaches an agent to build LM pipelines, RAG systems and agents in DSPy using signatures, modules and optimizers instead of hand-tuned prompts.

    13k GitHub starsUsed in 9 repos~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Building Agent Systems

    telagod/code-abyss

    AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt…

    243 GitHub stars~691 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agentsop Dspy

    agentsope/SkillAlchemy

    Operating SOP for DSPy (Stanford NLP) — the declarative framework for "programming, not prompting" language models.

    466 GitHub stars~7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • AI ML

    aiskillstore/marketplace

    AI and machine learning workflow covering LLM application development, RAG implementation, agent architecture, ML pipelines, and AI-powered features.

    430 GitHub starsUsed in 3 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed

More from LeoYeAI/openclaw-master-skills

All 1,235 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Agent Prompt Patterns

What does Agent Prompt Patterns do?

Battle-tested prompt patterns for production AI agents. An agent skill from LeoYeAI/openclaw-master-skills. Agent Prompt Patterns is an agent skill from LeoYeAI/openclaw-master-skills. Battle-tested prompt patterns for production AI agents.

When should I use Agent Prompt Patterns?

Agent Prompt Patterns fits situations like: designing agent behavior; enforcing reliability; building agent operating manuals.

How do I install Agent Prompt Patterns in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill agent-prompt-patterns -a claude-code`. Or copy the skill folder (skills/agent-prompt-patterns in LeoYeAI/openclaw-master-skills) into .claude/skills/agent-prompt-patterns in your project. Claude Code loads it when a task matches its description.

How do I install Agent Prompt Patterns in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill agent-prompt-patterns -a codex`. Or copy the skill folder (skills/agent-prompt-patterns in LeoYeAI/openclaw-master-skills) into .agents/skills/agent-prompt-patterns in your project. Codex loads it when a task matches its description.

Can I use Agent Prompt Patterns in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill agent-prompt-patterns -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-prompt-patterns, .gemini/skills/agent-prompt-patterns, .github/skills/agent-prompt-patterns and .opencode/skills/agent-prompt-patterns in your project.

What does Agent Prompt Patterns need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Prompt Patterns is instructions for the agent only. Our summary lists: Python 3.

Does Agent Prompt Patterns access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Prompt Patterns safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Prompt Patterns use?

Agent Prompt Patterns is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Prompt Patterns use?

About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Prompt Patterns?

Skills that share tags, products or a category with Agent Prompt Patterns: Senior Prompt Engineer (maslennikov-ig/claude-code-orchestrator-kit, 260 stars), Create System Prompt (pnp/copilot-prompts, 892 stars), DSPy Language Model Programming (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Building Agent Systems (telagod/code-abyss, 243 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Prompt Patterns?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,160 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.