Agent skill

Realize

by jongwony in jongwony/epistemic-protocols

This skill should be used when the user asks to "run the eval", "test whether the protocol actually works at runtime", "check type realization", "measure protocol fulfillment", "run the…

MITAuto-check: notes

Install Realize

skills CLI
$ npx skills add jongwony/epistemic-protocols --skill realize -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jongwony/epistemic-protocols realize --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jongwony/epistemic-protocols.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/realize .claude/skills/realize && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
realize
GitHub stars
173
Token cost
~3.3k tokens
SKILL.md length
1,746 words
Files
106 (incl. scripts, references)
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when the user asks to "run the eval", "test whether the protocol actually works at runtime", "check type realization", "measure protocol fulfillment", "run the…

  • Asks to run the eval
  • SKILL.md covers Purpose, Judgment boundary, When it applies and Workflow, plus 6 more sections
  • Runs Shell scripts from its folder; calls codex; needs CODEX_API_KEY and CLAUDE_CODE_OAUTH_TOKEN
  • Test whether the protocol actually works at runtime

What it does

Realize is an agent skill from jongwony/epistemic-protocols. This skill should be used when the user asks to "run the eval", "test whether the protocol actually works at runtime", "check type realization", "measure protocol fulfillment", "run the type-realization suite", "does the gate actually stop", or wants runtime evidence that a protocol's formal transition contract is realized rather than an assessment of downstream artifact quality. Invoke explicitly with /realize and a skill argument.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 112 other files, including scripts and reference files (for example `evals/conduct-map-gate/case.yaml`, `evals/conduct-map-gate/graders/contrary-grounds-shown.md` and `evals/conduct-map-gate/graders/method-written-out.md`).

The repository describes itself as: Epistemic protocols for Claude Code — structure human-AI interaction quality at every decision point - https://epistemic-protocols.com. The licence is MIT.

When your agent uses it

  • Asks to run the eval
  • Test whether the protocol actually works at runtime
  • Check type realization
  • Measure protocol fulfillment

Example prompts

  • “run the eval”
  • “test whether the protocol actually works at runtime”
  • “check type realization”
  • “/realize”

Requirements

  • A Bash shell
  • A credential in CLAUDE_CODE_OAUTH_TOKEN
  • A credential in CODEX_API_KEY
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash

What it can do on your machine

Read from SKILL.md and the folder at commit 69bb95b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • codex

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CODEX_API_KEY
    • CLAUDE_CODE_OAUTH_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Realize loads about 3.3k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 1,746 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jongwony/epistemic-protocols at commit 69bb95b, republished under its MIT licence (© jongwony). 1,746 words, ~3,270 tokens.

Download SKILL.mdSave it as .claude/skills/realize/SKILL.md (or your agent's skills folder). This skill also uses 105 other files; get the full folder from GitHub.
name
realize
description
This skill should be used when the user asks to "run the eval", "test whether the protocol actually works at runtime", "check type realization", "measure protocol fulfillment", "run the type-realization suite", "does the gate actually stop", or wants runtime evidence that a protocol's formal transition contract is realized rather than an assessment of downstream artifact quality. Invoke explicitly with /realize and a skill argument.
allowed-tools
Read, Grep, Glob, Bash

Type Realization

Measure whether one named protocol skill's declared contract is realized during an actual run. The first argument is required; /realize inquire selects only the registered inquire cases and graders. An unknown or omitted target fails before setup.

Purpose

A protocol's formal blocks are runtime-normative: TYPES, PHASE TRANSITIONS, TOOL GROUNDING and the Rules type the prose and carry the operational contract. Static checks establish that a SKILL.md contains those blocks and enforce a bounded set of structural invariants. Nothing establishes that a run follows their transitions.

This skill supplies evidence for that gap. It drives Claude Code or Codex across a matrix of models and treatments, captures the JSONL trace, grades the deterministically observable subset, and names the transcript judgments still owed. A complete automatic cell can establish that collection occurred, the declared branch reached Stop, or the zero-uncertainty branch reached Proceed; semantic ordering and rendered type coverage remain manual until their specified judges are built.

Judgment boundary

The object of judgment is the formal transition trace:

text
X → context sufficiency → evidence collection as needed → residual deficit
  → formal branch → Stop | Proceed

Every score attaches to a node or edge in that trace. Once the selected branch is observable, judgment ends. A plan, implementation, analysis, or other downstream artifact is evidence that Proceed occurred; its quality, correctness, completeness, and usefulness are outside /realize. Likewise, an unchanged substrate is evidence of Stop only where the case granted the capability to proceed and the formal branch required stopping. references/grader-design.md carries the witness design and the separate oracle that artifact-quality evaluation would require.

When it applies

Reach for this after changing a protocol's phase structure, terminal conditions, gate definition, or answer types — the places where a contract can become unrealizable without any static check noticing. Reach for it also when adopting a new model, where the question is whether the contract still holds on it.

Do not reach for it to check that a file is well-formed. That is /verify, it is free, and it runs on every commit.

Workflow

bash
cd .claude/skills/realize/scripts

./setup.sh inquire                           # default Claude runner
export CLAUDE_CODE_OAUTH_TOKEN="$(...)" && ./run.sh inquire

REALIZE_RUNNER=codex ./setup.sh inquire      # no credential consumed or stored
CODEX_API_KEY="$(...)" REALIZE_RUNNER=codex ./run.sh inquire
REALIZE_CODEX_AUTH=login REALIZE_RUNNER=codex ./run.sh inquire   # or: this machine's codex login
REALIZE_RUNNER=codex ./teardown.sh inquire

./teardown.sh inquire                        # Claude volatile state

For Claude, setup.sh prints the one interactive step — obtaining a token against the target-specific isolated config directory. For Codex, setup creates bare/protocol homes and installs the local protocol plugin only into the protocol home without reading or writing a credential. run.sh authenticates each codex exec child one of two ways, and codex login is never called by either:

  • API key (default): CODEX_API_KEY from the run process, forwarded only to each codex exec child.
  • Login (REALIZE_CODEX_AUTH=login): the codex login already on this machine, symlinked into the arm's disposable home only for the span of each codex exec — never copied, because a ChatGPT login rotates its refresh token and a copy would strand the real one. references/runbook.md carries why that keeps the homes isolated.

Read references/runbook.md before the first run. It records where runner isolation lives, where the budget floor sits, and what each column of the report asserts.

From CI

.github/workflows/type-realization.yml runs the Claude path on manual dispatch and comments the report on the PR. Codex remains local-only until the matrix can run behind the official action's credential proxy; repository-controlled harness code does not receive an OpenAI key in Actions.

Configuration

harness.config.json carries shared runner settings and a target registry. Each target owns its plugin directory, skill id, invocation, and case set. REALIZE_RUNNER=codex selects the committed Luna xhigh profile. Results are keyed by target, runner, and a hash of the actual protocol/style treatment, so one skill or ablation cannot reuse another's cache.

Prefer a capable model for the primary measurement. The weakest available one exercises the safeguards but not the protocol, so a failure there cannot separate a defect in the contract from a limit of the model.

Arms

Claude has four arms crossing the protocol against the Epistemic Ink output style, and one that removes the protocol's formal blocks. The style ships in the cc-plugin marketplace's ink-figure plugin; styleSource in harness.config.json names the file the style arms read, so those arms need that marketplace installed:

armprotocolstyleanswers
bare——baseline
style—✓sham — is the structure coming from form alone?
protocol✓—is the SKILL.md self-contained, as required?
protocol+style✓✓the deployed configuration
protocol-prose✓, every ```lean block removed—what the formal blocks add over the same SKILL.md prose

protocol-prose runs only when named in REALIZE_ARMS. It loads a copy of the plugin rebuilt at each run with the Lean blocks removed from its SKILL.md files, and refuses to run when the copy lost none. Read against protocol: a transition realized in both arms is realized without the formal blocks, and one realized only under protocol points to what they add.

The sham arm is not a construction. Published work on rule files for coding agents found random rules helping as much as curated ones, which makes "a long structured instruction is present" a live alternative explanation for any positive result. The output style is that alternative made available as a control: it fixes gate shape and observer markers while fixing none of a protocol's own obligations.

Codex has bare and protocol arms. Codex has no equivalent of Claude's shipped output-style treatment, so requesting style or protocol+style fails rather than simulating a different deployment condition; protocol-prose is likewise Claude-only. The Codex protocol treatment is present only when codex plugin list reports that plugin installed and enabled in the isolated protocol home while the bare home reports it absent. Codex JSONL currently exposes no separate skill-invocation event, so its skill report cell states that limitation rather than inferring invocation from the model's prose.

Show full SKILL.md (844 more words)Show less

Cases

Every case runs under explicit invocation. Each registered target needs at least a case where the protocol's obligations must be realized, and one that holds the counterpart the protocol must not fabricate — for /inquire, a collection that leaves nothing open and relays; for /grasp, an answer with nothing to check it against. Whether a protocol is selected at all, or stays silent, is measured by route's own eval, not here. The registered targets are the keys of targets in harness.config.json; no result is implied for a protocol absent there.

A target's paired cases mount the same scaffold, deliberately: for /inquire, one observes whether a file-discoverable fact was asked, the other whether a supplied parameter was re-asked, and differing substrates would let a run pass one by luck. A case names its scaffold in case.yaml (scaffold_script); one that names none mounts evals/scaffold.sh.

Write the prompt as the task alone. The line that invokes the protocol lives in harness.config.json and reaches only the arms that have it — a prompt naming the command makes an arm without the plugin gate on the missing tool instead of on the task.

Read every clause of a negative case as an adversary would. A specification that looks exhaustive can still leave a term underdetermined, and a correct protocol will open a gate on it — recording a failure that belongs to the case author.

evals/inquire-underspecified/ and evals/inquire-fully-specified/ are the worked pair for /inquire. Follow their shape when adding a protocol: a prompt.md carrying frontmatter and the user's words, and one grader per obligation under graders/. A grader checks that the fields a judgment produces exist and are faithful to their sources; it does not grade the judgment itself, such as whether a ground is sufficient.

A protocol whose obligations fire only after the user answers needs turns past the first Stop. A case reaches them by declaring multi_turn in case.yaml. Where its user side can be written as a fixed script — driver: harness, the replies in reply-1.md, reply-2.md, … — the harness sends them itself, on either runner, resuming one session; evals/grasp-adjudicable/ and evals/grasp-unattachable/ are that pair for /grasp, and their oracle.md says why every reply is written to stand at whichever gate it lands on. An oracle that must read the subject's turn to compose a reply, as /elicit's does, is walked by hand (references/runbook.md) and refused if registered.

Reading results

Read integrity first. It reports whether each arm's treatment actually applied — both dimensions of it, the plugin and the output style, since either can fail silently and leave two arms running the same treatment. A row whose integrity falls short of its run count is not evidence about the protocol, and the report reprints those rows separately so they are not mistaken for findings.

pass_k is one only when every repetition passed its deterministic transition predicates. The manual column counts scenario-specific transcript judgments excluded from that composite; the report names them. For inquire, collection-before-surfacing order, unasked cheap evidence, faithful basis, kept ownership, stated answer openings, the relay of a collection that left nothing open, and the absence of a design gate remain manual observations grounded by the grader files. For grasp, the automatic set is what both cases share — the target read in the first turn, the tree unchanged after every turn, every turn reported — and the quoted correction, the withheld verdict with its named need, the stop at each gate, and closure on the user's word are manual. For conduct, the only automatic predicate is every turn reported: whether the method's work started — the map's stop, the taking's proceed, the relay's proceed — is read from the transcript, so each transition is manual along with what its turn presents. Whether the tree differed from the scaffold after each turn is recorded in the cell's sidecar as an observation those graders may read, not as a verdict. The predicates column breaks pass_k down by predicate; turns shows how many scripted turns a multi-turn cell reached.

On Claude, skill says whether the protocol fired where it was available, and n/a where there was no plugin to fire. Codex reports trace-unavailable for that column and keeps plugin installation integrity separate from behavioral fulfillment.

Absolute transition-fulfilment rates are the reportable automatic quantity here, not deltas. A baseline arm has no gate to stop at, so most predicates have no counterpart there to subtract.

For Codex, the report records token use rather than inventing a dollar amount the CLI did not emit. Every requested cell must produce a complete thread.started → turn.completed pair. A launch failure, missing cell, unreadable predicate, or treatment-integrity failure makes run or report exit non-zero while preserving partial transcripts for inspection.

Additional Resources

  • references/runbook.md — the workflow, where isolation lives, how to widen the matrix, how to read the report.
  • references/grader-design.md — why the deterministic axis is behaviour rather than wording, the obligation-to-predicate mapping, what the arms answer, and the structured-extraction and metamorphic-validation passes that are specified but not yet built.
  • scripts/harness.mjs — the runner and the deterministic graders. Node standard library only.
  • evals/*.sh — the fixtures; each case selects its own through scaffold_script in case.yaml.

© jongwony, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 105 other files (scripts, references) in .claude/skills/realize of jongwony/epistemic-protocols.

  • SKILL.md
  • evals/conduct-map-gate/case.yaml
  • evals/conduct-map-gate/graders/contrary-grounds-shown.md
  • evals/conduct-map-gate/graders/method-written-out.md
  • evals/conduct-map-gate/graders/turn-ends-at-gate.md
  • evals/conduct-map-gate/prompt.md
  • evals/conduct-relay/case.yaml
  • evals/conduct-relay/graders/map-relayed-before-dispatch.md
  • evals/conduct-relay/graders/proceed-observed.md
  • evals/conduct-relay/prompt.md
  • evals/conduct-scaffold.sh
  • evals/conduct-taking-with-change/case.yaml
  • evals/conduct-taking-with-change/graders/map-relayed-before-dispatch.md
  • evals/conduct-taking-with-change/graders/relayed-not-gated.md
  • … and 92 more

Open the folder on GitHubat commit 69bb95b

Compare with similar skills

Realize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Realize compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Realize this skilljongwony/epistemic-protocols173—~3.3kAutomated safety check: NotesMIT
Benefits Realizationcbrock84/headcount2k—~861Automated safety check: PassMIT
Value RealizationDone-0/value-realization531—~11kAutomated safety check: PassMIT
Modeling Competition Computation StageWuXinbo-bo/Math-model-skills111—~4.5kAutomated safety check: PassMIT
HTML Ppt Zhangzara Cobalt Gridnexu-io/open-design100k—~1.2kAutomated safety check: PassMIT
Gtm Enterprise Onboardinggithub/awesome-copilot40k1 repos~3.6kAutomated safety check: PassMIT

Similar skills

  • Benefits Realization

    cbrock84/headcount

    Ensures projects deliver the value they were approved on — defining measurable benefits, baselining, tracking after delivery, and honest post-implementation review.

    2k GitHub stars~861 tokensUpdated 20 days ago
    Auto-check passed
  • Value Realization

    Done-0/value-realization

    Judge whether something actually produces value in a concrete scenario.

    531 GitHub stars~11k tokensUpdated 2 mo ago
    Marketing & SEOAuto-check passed
  • Modeling Competition Computation Stage

    WuXinbo-bo/Math-model-skills

    Pipeline stage that turns a mathematical modeling report into runnable programs per sub-question, frozen numerical results and reviewable evidence files.

    111 GitHub stars~4.5k tokensUpdated 16 days ago
    Research & ScienceAuto-check passed
  • OpenDesign renewal + seat-expansion business case for a growing customer: realized value, usage proof, and the expansion ROI.

    100k GitHub stars~1.2k tokensUpdated today
    Sales & SupportAuto-check passed
  • Gtm Enterprise Onboarding

    github/awesome-copilot

    Official

    Four-phase framework for onboarding enterprise customers from contract to value realization.

    40k GitHub starsUsed in 1 repo~3.6k tokens
    Marketing & SEOAuto-check passed
  • Okx Dex Market

    LeoYeAI/openclaw-master-skills

    A skill your agent uses for on-chain market data: token prices/价格, K-line/OHLC charts, index prices, and wallet PnL/盈亏分析 (win rate, my DEX trade history, realized/unrealized PnL per token).

    2.2k GitHub stars~5.9k tokensUpdated 2 mo ago
    Backend & APIsAuto-check: notes

More from jongwony/epistemic-protocols

All 29 skills in this repo
  • Outcome

    jongwony/epistemic-protocols

    This skill should be used when the user asks to "run the outcome eval", "paired bare vs protocol", "which decisions did the protocol surface", "count what the AI asked or presented", "does /inquire…

    173 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check: notes
  • Verify

    jongwony/epistemic-protocols

    This skill should be used when the user asks to "verify protocols", "check consistency before commit", "validate definitions", "run pre-commit checks", "verify soundness", or wants to ensure…

    173 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Encapsulation

    jongwony/epistemic-protocols

    This skill should be used when the user asks to "audit plugin encapsulation", "check self-containment semantics", "find contributor-knowledge assumptions", or invokes /encapsulation.

    173 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Formal Review

    jongwony/epistemic-protocols

    This skill should be used when the user asks to "formal review", "formal lens review", or invokes /formal-review.

    173 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check: notes
  • Recollect

    jongwony/epistemic-protocols

    The user vaguely recalls something discussed before but cannot name it — one session, or a line of work, topic, or settled concept across several: find it in past records to recognize.

    173 GitHub stars~8.7k tokensUpdated yesterday
    Auto-check passed
  • White Bear

    jongwony/epistemic-protocols

    A skill your agent uses when the user asks to "check white bear", "audit prohibitions", "find negative framing", or invokes /white-bear.

    173 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed

Questions about Realize

What does Realize do?

This skill should be used when the user asks to "run the eval", "test whether the protocol actually works at runtime", "check type realization", "measure protocol fulfillment", "run the…. Realize is an agent skill from jongwony/epistemic-protocols. This skill should be used when the user asks to "run the eval", "test whether the protocol actually works at runtime", "check type realization", "measure protocol fulfillment", "run the type-realization suite", "does the gate actually stop", or wants runtime evidence that a protocol's formal transition contract is realized rather than an assessment of downstream artifact quality.

When should I use Realize?

Realize fits situations like: asks to run the eval; test whether the protocol actually works at runtime; check type realization; measure protocol fulfillment.

How do I install Realize in Claude Code?

Run `npx skills add jongwony/epistemic-protocols --skill realize -a claude-code`. Or copy the skill folder (.claude/skills/realize in jongwony/epistemic-protocols) into .claude/skills/realize in your project. Claude Code loads it when a task matches its description.

How do I install Realize in Codex?

Run `npx skills add jongwony/epistemic-protocols --skill realize -a codex`. Or copy the skill folder (.claude/skills/realize in jongwony/epistemic-protocols) into .agents/skills/realize in your project. Codex loads it when a task matches its description.

Can I use Realize in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jongwony/epistemic-protocols --skill realize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/realize, .gemini/skills/realize, .github/skills/realize and .opencode/skills/realize in your project.

What does Realize need to run?

Going by SKILL.md and its folder, Realize needs a shell for the scripts in its folder, the command-line tools its instructions call (codex) and credentials named CODEX_API_KEY and CLAUDE_CODE_OAUTH_TOKEN. Our summary lists: A Bash shell; A credential in CLAUDE_CODE_OAUTH_TOKEN; A credential in CODEX_API_KEY. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash.

Does Realize access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Realize safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Realize use?

Realize is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Realize use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.3k tokens, read only when the agent opens those files.

What are the alternatives to Realize?

Skills that share tags, products or a category with Realize: Benefits Realization (cbrock84/headcount, 2k stars), Value Realization (Done-0/value-realization, 531 stars), Modeling Competition Computation Stage (WuXinbo-bo/Math-model-skills, 111 stars) and HTML Ppt Zhangzara Cobalt Grid (nexu-io/open-design, 100k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Realize?

jongwony (a GitHub user) maintains it in jongwony/epistemic-protocols, which has 173 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 6, 2026.

Source: jongwony/epistemic-protocols on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.