Caveman Gateway Setup
JuliusBrussee/caveman
Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.
Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Eigenwise/eigenwise-toolshed model-gateway --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Eigenwise/eigenwise-toolshed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/model-gateway/skills/model-gateway .claude/skills/model-gateway && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "model-gateway" agent skill from https://github.com/Eigenwise/eigenwise-toolshed/tree/main/plugins/model-gateway/skills/model-gateway into .claude/skills/model-gateway/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-gateway", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Eigenwise/eigenwise-toolshed/tree/main/plugins/model-gateway/skills/model-gatewayType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Eigenwise/eigenwise-toolshed model-gateway --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Eigenwise/eigenwise-toolshed.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/model-gateway/skills/model-gateway .agents/skills/model-gateway && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "model-gateway" agent skill from https://github.com/Eigenwise/eigenwise-toolshed/tree/main/plugins/model-gateway/skills/model-gateway into .agents/skills/model-gateway/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-gateway", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Eigenwise/eigenwise-toolshed model-gateway --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Eigenwise/eigenwise-toolshed.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/model-gateway/skills/model-gateway .cursor/skills/model-gateway && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "model-gateway" agent skill from https://github.com/Eigenwise/eigenwise-toolshed/tree/main/plugins/model-gateway/skills/model-gateway into .cursor/skills/model-gateway/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-gateway", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Eigenwise/eigenwise-toolshed.git --path plugins/model-gateway/skills/model-gateway--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Eigenwise/eigenwise-toolshed model-gateway --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Eigenwise/eigenwise-toolshed.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/model-gateway/skills/model-gateway .gemini/skills/model-gateway && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "model-gateway" agent skill from https://github.com/Eigenwise/eigenwise-toolshed/tree/main/plugins/model-gateway/skills/model-gateway into .gemini/skills/model-gateway/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-gateway", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Eigenwise/eigenwise-toolshed model-gatewayInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Eigenwise/eigenwise-toolshed.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/model-gateway/skills/model-gateway .github/skills/model-gateway && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "model-gateway" agent skill from https://github.com/Eigenwise/eigenwise-toolshed/tree/main/plugins/model-gateway/skills/model-gateway into .github/skills/model-gateway/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-gateway", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Eigenwise/eigenwise-toolshed model-gateway --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Eigenwise/eigenwise-toolshed.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/model-gateway/skills/model-gateway .opencode/skills/model-gateway && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "model-gateway" agent skill from https://github.com/Eigenwise/eigenwise-toolshed/tree/main/plugins/model-gateway/skills/model-gateway into .opencode/skills/model-gateway/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-gateway", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
model-gatewaySet up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker.
Model Gateway is an agent skill from Eigenwise/eigenwise-toolshed. Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker. Use for gateway setup, login, model visibility, routing, or failures.
Its SKILL.md is about 8.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud. It works with OpenAI. The repository describes itself as: Six Claude Code plugins for the work that keeps coming back: repo maps, conditional rules, ticketed parallel work, extra subscription models, local usage metrics, and guided setup. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0140ec1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.anthropic.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Model Gateway loads about 8.7k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 4,800 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
ot reload settings or the picker cache. Do not tell the user to select a new row until that full restart is complete.g Codex subagents through the gateway): do not tell the user to restart Claude Code just toAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Eigenwise/eigenwise-toolshed at commit 0140ec1, republished under its MIT licence (© Eigenwise). 4,800 words, ~8,651 tokens.
.claude/skills/model-gateway/SKILL.md (or your agent's skills folder).Two local processes give Claude Code native access to the user's ChatGPT subscription models:
claude-code-proxy (does OpenAI OAuth and translates Anthropic Messages API to the Codex
backend) and a shim router this plugin owns. ANTHROPIC_BASE_URL points at the shim: requests
for claude-gpt-* models are un-prefixed and go to the proxy, everything else passes through
to api.anthropic.com with the user's normal claude.ai login. The shim's /v1/models advertises
Codex models with a claude- prefix because Claude Code's model discovery drops ids that don't
start with claude/anthropic. That prefix is shared with the real Anthropic ids, so the route
is decided by the backend family segment (claude-gpt-*, claude-grok-*), never by the prefix.
All commands: node ~/.claude/model-gateway/model-gateway.js <command>. SessionStart recreates this stable launcher from Claude Code's installed-plugin registry, so it follows upgrades and falls back to an installed downgrade; it reports a missing install directly.
Project-local wiring is the standard setup. env --write-project writes the current project's .claude/settings.local.json, so the gateway stays configured for this project's sessions and executor worktrees without putting a machine-local endpoint in a committed file. setup uses the same project-local target by default.
env --write-user remains an opt-in shared fallback for people who deliberately want one gateway URL in ~/.claude/settings.json across every project. Claude Code gives a current project's settings.local.json higher precedence, so doctor marks the winner [effective], names both files, and says that project-local wiring wins when their gateway modes disagree.
The first-run order matters:
setup.login, then run setup again. The second setup finishes the download, wiring, and project confirmation.A plugin reload alone does not reload settings or the picker cache. Do not tell the user to select a new row until that full restart is complete.
The SessionStart hook injects a one-line nudge while the gateway is in any half-configured
state; act on it. The user sees that same line in the transcript, because a state only they can fix used to
reach the model alone. Anything routine stays out of it, and the hook always exits 0 so the line survives:
run ensure yourself when you need an exit code. SessionStart waits at most 12 seconds for a newly started
gateway, then leaves its supervisor to finish in the background so it stays inside Claude Code's hook budget. setup is one-shot and idempotent: it downloads the claude-code-proxy binary
(sha256-verified) and starts everything. Re-running it later is also the upgrade path. Worker hot-replacement accepts the current CLI or a non-older version-shaped sibling under the same lexical and resolved parent directory. This checks filesystem layout, not marketplace registration; a development layout with the same structure can also satisfy it.
Setup fails before extraction when the required checksum asset is absent, its download fails, its contents are invalid or name a different archive, or the archive digest does not match. Report the verification failure; never skip the checksum check or run the unverified download.
Use the bundled lifecycle commands for restarts and drains. They authenticate with the per-user ~/.claude/model-gateway/control-token file. Never print or copy this secret into a conversation, request body, or settings file. Browser-origin lifecycle requests are refused, and a replacement worker must be this CLI or a non-older sibling within the same resolved plugin cache. Development checkouts cannot hot-switch to unrelated installations. After updating, restart the calling Claude Code process to load the current control protocol. These controls do not authenticate inference requests to the separate proxy or protect against a process with the same OS-user access.
node ~/.claude/model-gateway/model-gateway.js setup
# only if setup says sign-in is needed:
node ~/.claude/model-gateway/model-gateway.js login # browser OAuth; --device for headless
node ~/.claude/model-gateway/model-gateway.js setup # finishes the wiringlogin opens the user's browser; they complete it themselves (suggest ! node ... login if it
needs a real TTY). env --write-project writes this project's .claude/settings.local.json.
Use env --write-user only when the user wants a shared fallback across projects. If both are
present, project-local wiring wins: doctor marks it [effective], names the shadowed user file,
and fails when their gateway modes differ. All wiring changes apply to new Claude Code sessions,
so restart after the write. The Codex rows appear in /model labeled "From gateway".
The winning ANTHROPIC_BASE_URL source follows Claude Code's precedence: process environment,
current-project .claude/settings.local.json, project .claude/settings.json, then user
~/.claude/settings.json. A process export always wins, so settings writes cannot replace it.
When doctor or SessionStart says that process env shadows a wired settings file, Model Gateway is
bypassed. If the user controls the Claude Code CLI launch, they can correct or unset
ANTHROPIC_BASE_URL, then restart. If the host replaces that value, use the supported Claude Code CLI
on the wired project instead. Model Gateway does not support Desktop routing under forced overrides on
Windows or macOS, and settings, parent, or User-scope edits cannot be promised to win.
env --write-user --reconcile is confirmation-gated. Plain env --write-user writes the shared
user fallback, then lists recorded projects whose local URL differs without changing their files.
The confirmed command removes only Model Gateway-owned keys from those other projects'
.claude/settings.local.json files: its base URL, the three Claude alias pins, and static gateway
flags whose values equal plugin defaults. It leaves unrelated settings alone, skips projects already
agreeing, cannot change process.env, and needs a restart to affect a new session.
Discovery needs Claude Code v2.1.129+; models shows exactly what the shim advertises. Claude Code
only refetches gateway discovery when it has an API-key credential. OAuth subscriptions do not give it
one, so Model Gateway writes Claude Code's discovery cache whenever its advertised list changes, but
only from a list the proxy answered. While the proxy is unreachable the shim serves models.json or its
built-in list, keeps the previous cache, retries the proxy on its next refresh tick, and status says
fallback catalog (proxy unreachable). setup --preserve-wiring (what the Toolshed updater runs) never
wires the directory it runs from.
Restart remains necessary to surface new rows in /model: Claude Code reads the picker cache once at
session start. /reload-plugins does not reload it. Restoring or refreshing auth on an already-wired
install needs no restart of the current Claude Code process: the proxy is a separate process, so once login + setup re-authenticate it,
the next request routes through cleanly. Settings, discovery-cache, plugin, or model-row changes do need a full restart of the affected project process. Keep these two recovery paths separate. The shim supervisor also probes the proxy's /v1/models endpoint
while it runs, confirming a failed probe through a fresh connection before restarting an unavailable proxy with single-flight bounded backoff. It leaves a healthy proxy
alone. Recovery output remains in ~/.claude/model-gateway/logs/guardian.log; bounded lifecycle records in
~/.claude/model-gateway/logs/lifecycle.jsonl identify supervisor, worker, and proxy PIDs, orderly
stop/restart requests, observed exits, and recovery outcomes. Use doctor to print the evidence path and
the last observed exit. An OS termination or force-killed supervisor may leave no final record, so treat an
absent exit record as absence of evidence, not a clean shutdown. Cleanup accepts a PID record whose start time matches even when the live command line is unavailable, and deletes a record only when a readable command line contradicts it; it never stops a reused PID. On Windows, an unavailable command line can mean the process is elevated, so doctor names the condition and stop or setup must run from a session with the same privileges. Proxy recovery stops a listener using the shared proxy binary only
when the live process tree proves it descends from the recovering supervisor. A matching shared binary alone
never proves ownership. A failed /v1/models check gets one fresh-connection confirmation before recovery can stop an owned listener; a healthy confirmation resets recovery without stopping or starting the proxy. This matters when an agent is mid-orchestration
(e.g. dispatching Codex subagents through the gateway): do not tell the user to restart Claude Code just to
bring auth back, or you kill the session that was about to use it.
/model picker: rows like "GPT-6.1-sol (Codex)" and "Grok 4.5".
Typed: /model claude-gpt-6.1-sol[1m] or /model claude-grok-4.5[1m]. The picker and Sidequest catalog emit those exact ids. The suffix is stripped before routing to Codex or Grok.
lib/runtime.js's exported MODEL_WINDOW_POLICY is the sole authority for gateway backend windows, picker aliases, advertised windows, and sentry mode. GPT-6.1 Sol, GPT-5.6 Sol, Terra, Luna, and GPT-6 Astra are measured rows. GPT-6.1 Sol is the default Codex model; GPT-6 Astra is reserved for frontier or high-stakes tickets. GPT ids absent from the table are deliberately advertised through its explicitly unmeasured 920k default, rather than silently inheriting a window. Grok 4.5 is a measured 500k row with a [1m] picker alias and the shared synthetic-413 sentry.
Codex GPT-5.6 through the ChatGPT Codex product (the subscription login this gateway routes to, not the pay-per-token API) accepted 920,012 input tokens and refused 935,012 on 2026-09-05 through claude-code-proxy 0.1.35 (upstream 55bf0b58). Every synthetic-413 policy row keeps its backend-window-minus-40000 ceiling. With no saved compactAt, the lowest legacy environment trigger or cap-minus-85000 wins. With compactAt.codex or compactAt.grok saved, the explicit maximum replaces those legacy candidates, is clamped to the cap, and ignores CODEX_GATEWAY_COMPACT_TRIGGER without mutating it. Claude Code 2.1.261 ignores a settings-file CLAUDE_CODE_MAX_CONTEXT_TOKENS value for its own unrecognized-model resolver, so rows above 200k use their policy's recognized [1m] alias. That alias gives Claude Code a 1M client window, the closest available setting to the verified 920k backend window; it does not promise a 1M backend input limit. A lower explicit autoCompactWindow still wins. Use /context to inspect the selected model and effective cap.
Context window and cost: context-window saves backend windows in ~/.claude/model-gateway/context-window.json (--claude <tokens|full> --codex <tokens|full> --grok <tokens|full>). Defaults stay Claude full, Codex 272000, Grok full; existing installs are never auto-migrated. /v1/models advertises the cap and keeps [1m] ids. With no explicit maximum, the gateway's legacy cap policy subtracts 85000, so Codex still compacts past 187000.
Direct gateway compaction maximum: context-window --codex-compact-at <tokens|cap> and --grok-compact-at <tokens|cap> save positive whole counts under compactAt; cap removes the explicit field and restores the legacy policy. These flags have no 100000 floor. Window-only updates preserve saved maxima. An explicit maximum has no hidden 85000 subtraction; the cap and actual backend window minus 40000 remain authoritative. status, doctor, and catalog notes report the requested and effective trigger, actual backend window, limiting source, and any ignored CODEX_GATEWAY_COMPACT_TRIGGER unchanged. --codex-compact-at 242000 is an opt-in example that leaves the 272000 cap unchanged. OpenAI bills input above 272k at 2x; the crossing turn and compaction request can still exceed 272k and pay double. Never promise this example keeps every request below the billing boundary.
Native Claude window: preserve --claude <tokens|full> and existing gateway-owned autoCompactWindow synchronization to project-wired .claude/settings.local.json; numeric windows require project-scoped wiring. Native engine headroom and exact trigger are unverified. --claude-compact-at and compactAt.claude are unsupported and refuse without writing. Same-model main/subagent thresholds are unverified; use one backend policy with no role inference. A lower native session window can bound gateway models too. After deliberately adopting a change, restart the gateway (stop, then ensure) and restart open Claude Code sessions for native-window or picker changes. No automatic 242000 adoption, machine environment mutation, or worker restart during implementation. CODEX_GATEWAY_CONTEXT_WINDOW still applies only when no Codex window is saved.
Claude models (opus/sonnet/fable, with or without [1m]) keep their OWN separate native windows
and compaction limits: the shim forwards their requests byte-identically to Anthropic and never
applies Codex window advertisement or error rewriting to them. The env block pins the current
real 1M aliases (Opus, Sonnet, Fable) to [1m] ids so a gateway session on one gets its full 1M
window instead of the 200k gateway default; Haiku stays unpinned (it's 200k). An env --write-*
command resolves those aliases through the installed Claude CLI's credential-free headless probe;
SessionStart refreshes its cache after the CLI changes or the cache ages out, and rewrites stale pins
it wrote, including in a project where the gateway was turned off for Remote Control. A failed probe keeps
the last good pin, then a shipped safe default. A detected alias older than the shipped default, or
one the plugin retired, loses to the shipped default. Set a persistent per-alias override with
pin --opus claude-opus-5-5[1m] (same for --sonnet and --fable), or use pin --opus default
to return to auto-detection. Overrides always win. pin with no arguments and doctor show each
effective pin, whether it is overridden, and which lagging CLI alias a shipped pin replaced. Overrides live in
~/.claude/model-gateway/pins.json, outside
the plugin cache. A pin change, and every SessionStart pin refresh, updates every registered wired
project's gateway-owned pins (so a new shipped default reaches them too) and skips any project with a
user-owned pin value. Restart every open Claude Code session in an
affected project; changing a saved value cannot alter an open session.
Do NOT set a
global CLAUDE_CODE_AUTO_COMPACT_WINDOW: it applies to both providers and can make Codex
/compact fail after history already exceeds the Codex limit.
Caution: loading a huge reference skill (e.g. claude-api, ~800k chars) in a single turn can
spike Codex context past the point proactive compaction can recover from. Prefer pulling large
references incrementally on Codex models.
The advertised catalog comes from the proxy's /v1/models; models.json and the built-in list only stand in while the proxy is unreachable. A models.json file cannot add a backend that the claude-code-proxy allowlist does not support; update the proxy through setup instead.
Claude Desktop cannot use Codex/Grok models in this version: Desktop has its own native Gateway
configuration, separate from Claude Code CLI settings, and can point at this shim's endpoint. But
installed Desktop 1.49585.0 validates every Gateway model ID client-side and rejects any
gpt/codex/non-Anthropic family marker before it reaches the picker or a session, whether the ID
came from an explicit config entry or from discovery; a terminal [1m] is stripped first and does
not change the outcome. Only a genuinely Anthropic-backed route (a real Claude alias) is usable
there. This is a Desktop-side restriction, not a Model Gateway bug, and there is no supported
workaround: do not suggest an Anthropic-named alias to disguise a Codex/Grok route, patch the
Desktop binary, substitute credentials or auth, intercept TLS, or use a global env or hosts trick.
Tell the user Desktop end-to-end Codex/Grok support is not available on this version; the Claude
Code CLI is the verified path. VS Code success has been reported by users but is not independently
verified here. This is version-specific and can be revisited if a future Desktop release removes
the model-family filter.
RC-compat and missing Codex rows: On inspected Claude Code 2.1.267, RC-compatibility points ANTHROPIC_BASE_URL at api.anthropic.com, which disables gateway model discovery. The rows disappear and cache refresh cannot restore them. Claude Code can accept and persist /model claude-gpt-5.6-terra[1m], but that client-side action does not prove a later request reaches the gateway. The current cleaner preserves canonical [1m] ids; do not present it as a fix for a reported request error. RC-compatibility makes Claude Code treat the gateway as first-party and can enable experimental message threading. The shim refuses Codex/Grok thread continuations locally with HTTP 400 because those backends have no conversation state. Nothing is forwarded or rerouted. A recognizing client drops the threading beta and resends the full message history. This is version- and experiment-limited, so do not promise one refusal per session. Normal gateway mode is the verified inference path. The reported 2.1.259 client and inspected 2.1.267 client have no verified end-to-end RC result: inspected session creation uses HTTPS while compatibility transport is HTTP. A detected hosts entry or bound listener proves local HTTP transport only. Sidequest dispatch is unaffected because it resolves its explicit route marker and never uses picker discovery.
Claude models keep working normally at the same time (passthrough path). That does not mean a
subagent invocation can mix in a gateway id freely: the Agent tool's own model parameter is a
fixed host enum (sonnet/opus/haiku/fable), independent of gateway or plugin state, and a
gateway id passed there is refused. To route one subagent to a specific gateway model, give it a
definition file instead, with a full gateway id in model: frontmatter, and omit model from the
invocation so that frontmatter applies:
---
name: luna-reviewer
description: Reviews code changes using the GPT-5.6 Luna gateway model.
model: claude-gpt-5.6-luna[1m]
tools: Read, Grep, Glob
---
Review the diff for correctness and report findings.Save that as .claude/agents/luna-reviewer.md and start a new session before invoking it —
definitions are read at session start, so a running session won't see one just added. Invoke with
subagent_type: "luna-reviewer" and no model argument. This needs only Model Gateway, not
Sidequest: a concrete gateway id needs no route marker, only Sidequest's own claude-codex-auto id
does. Never edit Claude Code's built-in model aliases, the Agent tool schema, or the host binary —
none of that is supported or necessary. This frontmatter contract and Model Gateway's routing for a
concrete id are confirmed; an end-to-end custom-agent spawn through this path has not been verified
here.
Codex schema compatibility: Codex rejects some Unicode property escapes such as \p{Cc} and \P{Cf}. After deferred hydration, the shim changes only a Codex-bound provider hint, never Claude Code's host schema or an Anthropic request. It admits a missing dialect or Draft 2020-12 and only pattern instances reached through properties, compatible patternProperties values, additionalProperties, items, prefixItems, allOf, anyOf, dependentSchemas, propertyNames, or unevaluatedProperties/unevaluatedItems. Each real property atom must stand alone in a negated character class without ranges, set syntax, captures, backreferences, or a negative regex context. The shim removes only that atom and preserves every other regex byte. not, conditionals, oneOf, contains, references, definitions, content schemas, affected pattern-property keys, unknown containers, and unsafe regexes refuse locally with HTTP 400 naming the tool, JSON Pointer, and reason code. Tell the user it was not forwarded or rerouted. Do not claim arbitrary schemas are preserved or try to bypass that diagnostic.
Request-route logging is enabled by default. It writes metadata-only JSONL records to
~/.claude/model-gateway/logs/request-routes.jsonl: timestamp, backend, model, request path, route and
effort when present, safe session and agent correlation ids, and dispatch-marker length when present. It
never writes request bodies, prompts, messages, tools, authentication, or arbitrary headers. Honor a user
request to disable it by setting CODEX_GATEWAY_REQUEST_LOG=0 before the shim starts, then restart the shim
through setup or ensure. The value is read when the shim process starts, so /reload-plugins does not
change an already-running shim. CODEX_GATEWAY_REQUEST_LOG_PATH changes the file location.
Usage observability also writes one high-water JSON file per valid session under
~/.claude/model-gateway/request-body/. The filename is derived from the session id. Its contents are the
largest forwarded request-body byte count observed for that session and an observation timestamp. It does
not contain the request body. No retention period is promised for either local record.
When the user asks about a failed Codex compaction and consent-gated route telemetry is available, inspect fixed compaction trace metadata only: selected/effective model, backend, upstream status, elapsed time, outcome, terminal/error code, and observed usage counters. completed and empty_summary both mean the generation finished: either a message_stop frame arrived, or a message_delta declared an end_turn/stop_sequence stop reason with no content after it and no content block left open. empty_summary is that finish with no visible non-whitespace text. incomplete means the HTTP-200 SSE stream never showed a finished generation: no stop reason, a stop reason that truncates (max_tokens, refusal, pause_turn, tool_use, anything unrecognized), content still arriving after the stop, an open block, or an upstream error frame. compaction_terminal_code is message_stop only when that frame really arrived, so a completed outcome with no terminal code means the frame was lost and the shim closed the turn itself instead of re-running the compaction. Unknown provider errors appear only as unknown_error. Never infer summary content or repeat prompts, response text, raw errors, headers, credentials, or environment values from diagnostics.
Claude Code's /remote-control only lights up when ANTHROPIC_BASE_URL is exactly the real
Anthropic host, which conflicts with gateway model discovery. RC-compatibility is a reversible,
opt-in local HTTP transport configuration, not a verified end-to-end Remote Control solution. At
inspected Claude Code 2.1.267 its session creation uses HTTPS while the compatibility listener is
HTTP; the reported 2.1.259 client and inspected client have no verified end-to-end RC result. A
bound port or detected hosts entry proves only that transport state. Before enabling, tell the user
that the picker rows disappear and an explicit [1m] id may be accepted client-side without proving
its later inference request reaches the gateway. Keep normal gateway mode as the verified inference
path. Do not suggest a cache refresh as a fix for a reported request error; the current cleaner
preserves canonical [1m] ids.
For the confirmation-gated procedure, use the remote-control-compatibility skill. It manages the
plugin-marked hosts block, creates a backup before an elevated write, reconciles gateway mode, and
checks the final transport state. Do not edit the hosts file outside that procedure. If effective
process env ANTHROPIC_BASE_URL is HTTPS api.anthropic.com (including port 443), enabling is
refused before any backup, hosts write, startup, or reconciliation because the loopback mapping
cannot serve TLS. A user-controlled Claude Code CLI launch can correct or unset that value, then
restart. If a host replaces it, use the supported Claude Code CLI on the wired project instead.
Desktop routing is unsupported under forced overrides on Windows and macOS, and settings, parent,
or User-scope edits cannot be promised to win. Disabling stays available.
remote-control enable --confirm, the plugin creates a backup and writes its marked hosts block mapping api.anthropic.com to loopback: 127.0.0.1 api.anthropic.com on Windows (C:\Windows\System32\drivers\etc\hosts, needs Administrator), macOS, and Linux (/etc/hosts, needs sudo). Do not edit the hosts file outside that procedure.ensure/setup/doctor detect the entry (read-only) and, only after confirming the shim can
actually bind loopback port 80, switch ANTHROPIC_BASE_URL to http://api.anthropic.com and
start a second listener on port 80 next to the usual 127.0.0.1:18764. Exactly one line tells
the user to restart Claude Code when the mode changes either direction.doctor reports the hosts entry (if any), whether port 80 actually bound (and why not if it
didn't), and which mode each settings scope (user/project) is wired to.CODEX_GATEWAY_HOSTS_FILE (custom hosts path), CODEX_GATEWAY_COMPAT_PORT
(port other than 80). Neither is needed for normal use.... status # what's running
... doctor # binary, auth, ports, model count, settings wiring
... ensure # start whatever is down (SessionStart hook runs this with --quiet)
... stop
... env --remove # unwire Claude Code (do this BEFORE uninstalling the plugin)status, doctor and ensure read the shim through one shared probe and print the same line,
shim (model router) on :<port>: <state>, where the state is running-ours (serving <version>),
running-foreign (PID <pid>, <install root or owner unidentified>), starting (PID <pid> since <time>) or
stopped. ensure leaves running-ours at the installed version alone and succeeds, waits up to its startup
window for starting before replacing anything, and refuses running-foreign.
doctor prints the full model-window table: backend and picker ids, backend and advertised windows,
Claude Code's resolved client window and compaction point, sentry mode and trigger, and the measurement
date. It includes Codex, Grok, and native Claude pin rows. Its model-id check is useful for stale shim ids,
but a PASS does not prove every supported model is present. A proxy from 0.1.14 through 0.1.35 can
omit GPT-6 Astra while this check passes. If Astra is missing, check the installed and serving proxy
version, rerun setup to fetch the latest release, and fully restart Claude Code. Astra requires
claude-code-proxy 0.1.36 or newer. A models.json edit cannot add a backend that the proxy allowlist
does not support. A FAIL naming missing and extra ids means the shim is stale even when its version
matches: restart it through the normal ensure or setup path, then restart Claude Code sessions so the
picker re-discovers the rows.
Logs live in ~/.claude/model-gateway/logs/. guardian.log has recovery output; lifecycle.jsonl
has bounded process evidence that doctor summarizes. Ports: shim 18764, proxy 18765 (override with
CODEX_GATEWAY_PORT / CODEX_GATEWAY_PROXY_PORT, but the env block and running processes must
agree).
doctor, check logs. Worst case env --remove restores stock behavior instantly.doctor and include its supervisor conflict line and lifecycle evidence.login state
(doctor shows auth), then proxy log. OpenAI gates non-Codex clients by request fingerprint;
when they tighten it, requests die mid-stream until claude-code-proxy ships a fix, so
suggest re-running setup (it fetches the latest release).doctor says upstream-unavailable: a final Codex inference failed in the last 30 seconds; the message names its status, time, and hold end.
It records completed request outcomes, not /v1/models or a health check, and clears only after
a completed successful Codex response. The 30-second expiry means there is no recent failure
evidence, not that Codex is live. An attributed OpenAI 401, 403, or 429 rejection enters
upstream-blocked. A 401 or 403 stays until setup or a completed successful Codex response
clears it. A 429 block expires: upstreamBlocked.expiresAt comes from the 429's Retry-After,
else claude-code-proxy's usage-limit reset header, else 60 seconds, and the doctor message
names that time. It lifts by itself then, or sooner on a completed successful Codex response,
and a later rejected request can latch it again. Sidequest consumes a cached catalog and can lag
this state by up to five minutes.end_turn instead, so the session ends the turn rather than retrying
into the same empty answer; the shim log records each one.catalog --refresh --json.
Run that command by hand and read stderr plus the exit code. It exits non-zero and names the reason
when it declines to write (shim not answering /healthz, /v1/models erroring, or a model list
with no gateway ids in it), leaving the stored catalog and its timestamp untouched. Exit 0 with no
diagnostic means it did write, so compare the printed updatedAt with the stored file.CODEX_GATEWAY_PROBE_TIMEOUT_MS budget (2 seconds by default, 8 seconds on Windows, where the Win32_Process
lookup itself typically takes 1.8-2.4 seconds). When that budget expires, the refusal says so, names the elapsed
budget, and points to CODEX_GATEWAY_PROBE_TIMEOUT_MS as the override. A malformed process result or
unrecognized command remains an ownership-unknown refusal without the timeout guidance. Startup records
owner-unknown and leaves that listener untouched; recovery leaves it for the next tick. A confirmed foreign
owner is also left untouched. Probe children are stopped with the supervisor, so they cannot keep a test fixture home open.doctor shows Not authenticated right after an upgrade: bumping the proxy binary (e.g.
0.1.10 → 0.1.17 via setup) can invalidate the credential the old version accepted — the new
binary reads it as not authenticated and setup stops before wiring. Fix: re-run login, then
setup again to finish. Until then every Codex model is down, so any run that routes to
Codex (a whole sidequest board of Codex-tier tickets, for one) stalls entirely./model: do not diagnose account access first. Check the installed and serving claude-code-proxy version. Astra requires 0.1.36 or newer; 0.1.35 does not include its backend allowlist, while the current doctor floor can still pass. Re-run setup to fetch the latest GitHub release, then fully restart Claude Code. Do not propose models.json: it cannot add a backend the proxy does not allow.CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY
missing), Claude Code < v2.1.129, CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC is set (it
disables discovery), or RC-compatibility is active. Claude Code only refetches discovery with an
API-key credential, so OAuth users rely on Model Gateway's cache write. Run doctor to check the
cache, then restart Claude Code after it updates; /reload-plugins does not reload picker rows.stop + start). Shift+Tab restores the mode in an affected session. Escape
hatch to re-enable plan tools: CODEX_GATEWAY_KEEP_PLAN_TOOLS=1.© Eigenwise, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/model-gateway/skills/model-gateway of Eigenwise/eigenwise-toolshed.
Open the folder on GitHubat commit 0140ec1
Model Gateway next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Model Gateway this skillEigenwise/eigenwise-toolshed | 277 | — | ~8.7k | Automated safety check: Warn | MIT | |
| Caveman Gateway SetupJuliusBrussee/caveman | 110k | 1 repos | ~2.6k | Automated safety check: Warn | Apache-2.0 | |
| Azure Architecture Autopilotgithub/awesome-copilot | 40k | 1 repos | ~1.9k | Automated safety check: Pass | MIT | |
| Capacitymicrosoft/GitHub-Copilot-for-Azure | 255 | 2 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Youtubeeat-pray-ai/yutu | 696 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Local Stack RuntimeOpenHands/OpenHands | 90k | — | ~375 | Automated safety check: Pass | MIT |
JuliusBrussee/caveman
Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.
github/awesome-copilot
Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.
microsoft/GitHub-Copilot-for-Azure
Discovers available Azure OpenAI model capacity across regions and projects.
eat-pray-ai/yutu
A skill your agent uses whenever the user mentions YouTube, video uploads, channel management, playlists, video SEO, or any YouTube Data API operation.
OpenHands/OpenHands
This skill should be used when the user asks to "change the dev stack", "add a runtime service", "change the launcher", "update Docker", "bump Agent Server", "change ingress routing", or changes…
aafqaq/codex-lb-enhanced
A skill your agent uses when the user asks how to build with OpenAI products or APIs and needs up-to-date official documentation with citations (for example: Codex, Responses API, Chat Completions…
Eigenwise/eigenwise-toolshed
Create or edit a live-rules instruction in the project's atomic rule set.
Eigenwise/eigenwise-toolshed
Create a self-maintaining codebase map in .claude/.codebase-info/.
Eigenwise/eigenwise-toolshed
Set up a Claude Code workspace for a new or existing project, informed by hindsight from the user's whole session history.
Eigenwise/eigenwise-toolshed
Audit a Sidequest board for completed, stale, duplicate, or superseded tickets, then safely close clear cases.
Eigenwise/eigenwise-toolshed
Inspect, audit, enable, or disable project live-rules. An agent skill from Eigenwise/eigenwise-toolshed.
Eigenwise/eigenwise-toolshed
Run a read-only health check for Quartermaster and installed Toolshed plugins.
Works with
Categories
Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker. Model Gateway is an agent skill from Eigenwise/eigenwise-toolshed. Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker.
Model Gateway fits situations like: model visibility.
Run `npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a claude-code`. Or copy the skill folder (plugins/model-gateway/skills/model-gateway in Eigenwise/eigenwise-toolshed) into .claude/skills/model-gateway in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a codex`. Or copy the skill folder (plugins/model-gateway/skills/model-gateway in Eigenwise/eigenwise-toolshed) into .agents/skills/model-gateway in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-gateway, .gemini/skills/model-gateway, .github/skills/model-gateway and .opencode/skills/model-gateway in your project.
Going by SKILL.md and its folder, Model Gateway needs the command-line tools its instructions call (node).
SKILL.md names 1 domain. In commands or code: api.anthropic.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 2 warning(s): contains instruction-override wording (e.g. “without asking the user”). Read the flagged lines before installing; the check is not a guarantee either way.
Model Gateway is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.7k tokens (SKILL.md is roughly 35k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Model Gateway: Caveman Gateway Setup (JuliusBrussee/caveman, 110k stars), Azure Architecture Autopilot (github/awesome-copilot, 40k stars), Capacity (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Youtube (eat-pray-ai/yutu, 696 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Eigenwise (a GitHub user) maintains it in Eigenwise/eigenwise-toolshed, which has 277 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.
Source: Eigenwise/eigenwise-toolshed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.