Agent skill

Model Gateway

by Eigenwise in Eigenwise/eigenwise-toolshed

Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker.

MITAuto-check: warningsDevOps & Cloud

Install Model Gateway

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Eigenwise/eigenwise-toolshed model-gateway --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Eigenwise/eigenwise-toolshed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/model-gateway/skills/model-gateway .claude/skills/model-gateway && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-gateway
GitHub stars
277
Token cost
~8.7k tokens
SKILL.md length
4,800 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker.

  • Works in 5 steps: Install Model Gateway at the recommended… → Reload plugins so the new skill is… → Invoke this skill and run setup. → …
  • Model visibility
  • SKILL.md covers First-time setup, Selecting models, Local gateway records and Day-2 operations, plus 1 more section
  • Calls node; reaches api.anthropic.com

What it does

Model Gateway is an agent skill from Eigenwise/eigenwise-toolshed. Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker. Use for gateway setup, login, model visibility, routing, or failures.

Its SKILL.md is about 8.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud. It works with OpenAI. The repository describes itself as: Six Claude Code plugins for the work that keeps coming back: repo maps, conditional rules, ticketed parallel work, extra subscription models, local usage metrics, and guided setup. The licence is MIT.

When your agent uses it

  • Model visibility

Example prompts

  • “/model-gateway”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Install Model Gateway at the recommended project scope.
  2. Reload plugins so the new skill is available.
  3. Invoke this skill and run setup.
  4. If setup says sign-in is needed, have the user complete login, then run setup again. The second setup finishes the download, wiring, and…
  5. After the project wiring is confirmed, tell the user to fully restart the Claude Code process for this same project before selecting a…

What it can do on your machine

Read from SKILL.md and the folder at commit 0140ec1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.anthropic.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Gateway loads about 8.7k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 4,800 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~8.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningContains instruction-override wording (e.g. “without asking the user”)SKILL.md:35
    ot reload settings or the picker cache. Do not tell the user to select a new row until that full restart is complete.
  • WarningContains instruction-override wording (e.g. “without asking the user”)SKILL.md:98
    g Codex subagents through the gateway): do not tell the user to restart Claude Code just to

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Eigenwise/eigenwise-toolshed at commit 0140ec1, republished under its MIT licence (© Eigenwise). 4,800 words, ~8,651 tokens.

Download SKILL.mdSave it as .claude/skills/model-gateway/SKILL.md (or your agent's skills folder).
name
model-gateway
description
Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker. Use for gateway setup, login, model visibility, routing, or failures.

model-gateway

Two local processes give Claude Code native access to the user's ChatGPT subscription models: claude-code-proxy (does OpenAI OAuth and translates Anthropic Messages API to the Codex backend) and a shim router this plugin owns. ANTHROPIC_BASE_URL points at the shim: requests for claude-gpt-* models are un-prefixed and go to the proxy, everything else passes through to api.anthropic.com with the user's normal claude.ai login. The shim's /v1/models advertises Codex models with a claude- prefix because Claude Code's model discovery drops ids that don't start with claude/anthropic. That prefix is shared with the real Anthropic ids, so the route is decided by the backend family segment (claude-gpt-*, claude-grok-*), never by the prefix.

All commands: node ~/.claude/model-gateway/model-gateway.js <command>. SessionStart recreates this stable launcher from Claude Code's installed-plugin registry, so it follows upgrades and falls back to an installed downgrade; it reports a missing install directly.

First-time setup

Project-local wiring is the standard setup. env --write-project writes the current project's .claude/settings.local.json, so the gateway stays configured for this project's sessions and executor worktrees without putting a machine-local endpoint in a committed file. setup uses the same project-local target by default.

env --write-user remains an opt-in shared fallback for people who deliberately want one gateway URL in ~/.claude/settings.json across every project. Claude Code gives a current project's settings.local.json higher precedence, so doctor marks the winner [effective], names both files, and says that project-local wiring wins when their gateway modes disagree.

The first-run order matters:

  1. Install Model Gateway at the recommended project scope.
  2. Reload plugins so the new skill is available.
  3. Invoke this skill and run setup.
  4. If setup says sign-in is needed, have the user complete login, then run setup again. The second setup finishes the download, wiring, and project confirmation.
  5. After the project wiring is confirmed, tell the user to fully restart the Claude Code process for this same project before selecting a model.

A plugin reload alone does not reload settings or the picker cache. Do not tell the user to select a new row until that full restart is complete.

The SessionStart hook injects a one-line nudge while the gateway is in any half-configured state; act on it. The user sees that same line in the transcript, because a state only they can fix used to reach the model alone. Anything routine stays out of it, and the hook always exits 0 so the line survives: run ensure yourself when you need an exit code. SessionStart waits at most 12 seconds for a newly started gateway, then leaves its supervisor to finish in the background so it stays inside Claude Code's hook budget. setup is one-shot and idempotent: it downloads the claude-code-proxy binary (sha256-verified) and starts everything. Re-running it later is also the upgrade path. Worker hot-replacement accepts the current CLI or a non-older version-shaped sibling under the same lexical and resolved parent directory. This checks filesystem layout, not marketplace registration; a development layout with the same structure can also satisfy it.

Setup fails before extraction when the required checksum asset is absent, its download fails, its contents are invalid or name a different archive, or the archive digest does not match. Report the verification failure; never skip the checksum check or run the unverified download.

Use the bundled lifecycle commands for restarts and drains. They authenticate with the per-user ~/.claude/model-gateway/control-token file. Never print or copy this secret into a conversation, request body, or settings file. Browser-origin lifecycle requests are refused, and a replacement worker must be this CLI or a non-older sibling within the same resolved plugin cache. Development checkouts cannot hot-switch to unrelated installations. After updating, restart the calling Claude Code process to load the current control protocol. These controls do not authenticate inference requests to the separate proxy or protect against a process with the same OS-user access.

bash
node ~/.claude/model-gateway/model-gateway.js setup
# only if setup says sign-in is needed:
node ~/.claude/model-gateway/model-gateway.js login    # browser OAuth; --device for headless
node ~/.claude/model-gateway/model-gateway.js setup    # finishes the wiring

login opens the user's browser; they complete it themselves (suggest ! node ... login if it needs a real TTY). env --write-project writes this project's .claude/settings.local.json. Use env --write-user only when the user wants a shared fallback across projects. If both are present, project-local wiring wins: doctor marks it [effective], names the shadowed user file, and fails when their gateway modes differ. All wiring changes apply to new Claude Code sessions, so restart after the write. The Codex rows appear in /model labeled "From gateway".

The winning ANTHROPIC_BASE_URL source follows Claude Code's precedence: process environment, current-project .claude/settings.local.json, project .claude/settings.json, then user ~/.claude/settings.json. A process export always wins, so settings writes cannot replace it.

When doctor or SessionStart says that process env shadows a wired settings file, Model Gateway is bypassed. If the user controls the Claude Code CLI launch, they can correct or unset ANTHROPIC_BASE_URL, then restart. If the host replaces that value, use the supported Claude Code CLI on the wired project instead. Model Gateway does not support Desktop routing under forced overrides on Windows or macOS, and settings, parent, or User-scope edits cannot be promised to win.

env --write-user --reconcile is confirmation-gated. Plain env --write-user writes the shared user fallback, then lists recorded projects whose local URL differs without changing their files. The confirmed command removes only Model Gateway-owned keys from those other projects' .claude/settings.local.json files: its base URL, the three Claude alias pins, and static gateway flags whose values equal plugin defaults. It leaves unrelated settings alone, skips projects already agreeing, cannot change process.env, and needs a restart to affect a new session. Discovery needs Claude Code v2.1.129+; models shows exactly what the shim advertises. Claude Code only refetches gateway discovery when it has an API-key credential. OAuth subscriptions do not give it one, so Model Gateway writes Claude Code's discovery cache whenever its advertised list changes, but only from a list the proxy answered. While the proxy is unreachable the shim serves models.json or its built-in list, keeps the previous cache, retries the proxy on its next refresh tick, and status says fallback catalog (proxy unreachable). setup --preserve-wiring (what the Toolshed updater runs) never wires the directory it runs from.

Restart remains necessary to surface new rows in /model: Claude Code reads the picker cache once at session start. /reload-plugins does not reload it. Restoring or refreshing auth on an already-wired install needs no restart of the current Claude Code process: the proxy is a separate process, so once login + setup re-authenticate it, the next request routes through cleanly. Settings, discovery-cache, plugin, or model-row changes do need a full restart of the affected project process. Keep these two recovery paths separate. The shim supervisor also probes the proxy's /v1/models endpoint while it runs, confirming a failed probe through a fresh connection before restarting an unavailable proxy with single-flight bounded backoff. It leaves a healthy proxy alone. Recovery output remains in ~/.claude/model-gateway/logs/guardian.log; bounded lifecycle records in ~/.claude/model-gateway/logs/lifecycle.jsonl identify supervisor, worker, and proxy PIDs, orderly stop/restart requests, observed exits, and recovery outcomes. Use doctor to print the evidence path and the last observed exit. An OS termination or force-killed supervisor may leave no final record, so treat an absent exit record as absence of evidence, not a clean shutdown. Cleanup accepts a PID record whose start time matches even when the live command line is unavailable, and deletes a record only when a readable command line contradicts it; it never stops a reused PID. On Windows, an unavailable command line can mean the process is elevated, so doctor names the condition and stop or setup must run from a session with the same privileges. Proxy recovery stops a listener using the shared proxy binary only when the live process tree proves it descends from the recovering supervisor. A matching shared binary alone never proves ownership. A failed /v1/models check gets one fresh-connection confirmation before recovery can stop an owned listener; a healthy confirmation resets recovery without stopping or starting the proxy. This matters when an agent is mid-orchestration (e.g. dispatching Codex subagents through the gateway): do not tell the user to restart Claude Code just to bring auth back, or you kill the session that was about to use it.

Selecting models

  • /model picker: rows like "GPT-6.1-sol (Codex)" and "Grok 4.5".

  • Typed: /model claude-gpt-6.1-sol[1m] or /model claude-grok-4.5[1m]. The picker and Sidequest catalog emit those exact ids. The suffix is stripped before routing to Codex or Grok.

  • lib/runtime.js's exported MODEL_WINDOW_POLICY is the sole authority for gateway backend windows, picker aliases, advertised windows, and sentry mode. GPT-6.1 Sol, GPT-5.6 Sol, Terra, Luna, and GPT-6 Astra are measured rows. GPT-6.1 Sol is the default Codex model; GPT-6 Astra is reserved for frontier or high-stakes tickets. GPT ids absent from the table are deliberately advertised through its explicitly unmeasured 920k default, rather than silently inheriting a window. Grok 4.5 is a measured 500k row with a [1m] picker alias and the shared synthetic-413 sentry.

  • Codex GPT-5.6 through the ChatGPT Codex product (the subscription login this gateway routes to, not the pay-per-token API) accepted 920,012 input tokens and refused 935,012 on 2026-09-05 through claude-code-proxy 0.1.35 (upstream 55bf0b58). Every synthetic-413 policy row keeps its backend-window-minus-40000 ceiling. With no saved compactAt, the lowest legacy environment trigger or cap-minus-85000 wins. With compactAt.codex or compactAt.grok saved, the explicit maximum replaces those legacy candidates, is clamped to the cap, and ignores CODEX_GATEWAY_COMPACT_TRIGGER without mutating it. Claude Code 2.1.261 ignores a settings-file CLAUDE_CODE_MAX_CONTEXT_TOKENS value for its own unrecognized-model resolver, so rows above 200k use their policy's recognized [1m] alias. That alias gives Claude Code a 1M client window, the closest available setting to the verified 920k backend window; it does not promise a 1M backend input limit. A lower explicit autoCompactWindow still wins. Use /context to inspect the selected model and effective cap.

  • Context window and cost: context-window saves backend windows in ~/.claude/model-gateway/context-window.json (--claude <tokens|full> --codex <tokens|full> --grok <tokens|full>). Defaults stay Claude full, Codex 272000, Grok full; existing installs are never auto-migrated. /v1/models advertises the cap and keeps [1m] ids. With no explicit maximum, the gateway's legacy cap policy subtracts 85000, so Codex still compacts past 187000.

  • Direct gateway compaction maximum: context-window --codex-compact-at <tokens|cap> and --grok-compact-at <tokens|cap> save positive whole counts under compactAt; cap removes the explicit field and restores the legacy policy. These flags have no 100000 floor. Window-only updates preserve saved maxima. An explicit maximum has no hidden 85000 subtraction; the cap and actual backend window minus 40000 remain authoritative. status, doctor, and catalog notes report the requested and effective trigger, actual backend window, limiting source, and any ignored CODEX_GATEWAY_COMPACT_TRIGGER unchanged. --codex-compact-at 242000 is an opt-in example that leaves the 272000 cap unchanged. OpenAI bills input above 272k at 2x; the crossing turn and compaction request can still exceed 272k and pay double. Never promise this example keeps every request below the billing boundary.

  • Native Claude window: preserve --claude <tokens|full> and existing gateway-owned autoCompactWindow synchronization to project-wired .claude/settings.local.json; numeric windows require project-scoped wiring. Native engine headroom and exact trigger are unverified. --claude-compact-at and compactAt.claude are unsupported and refuse without writing. Same-model main/subagent thresholds are unverified; use one backend policy with no role inference. A lower native session window can bound gateway models too. After deliberately adopting a change, restart the gateway (stop, then ensure) and restart open Claude Code sessions for native-window or picker changes. No automatic 242000 adoption, machine environment mutation, or worker restart during implementation. CODEX_GATEWAY_CONTEXT_WINDOW still applies only when no Codex window is saved.

  • Claude models (opus/sonnet/fable, with or without [1m]) keep their OWN separate native windows and compaction limits: the shim forwards their requests byte-identically to Anthropic and never applies Codex window advertisement or error rewriting to them. The env block pins the current real 1M aliases (Opus, Sonnet, Fable) to [1m] ids so a gateway session on one gets its full 1M window instead of the 200k gateway default; Haiku stays unpinned (it's 200k). An env --write-* command resolves those aliases through the installed Claude CLI's credential-free headless probe; SessionStart refreshes its cache after the CLI changes or the cache ages out, and rewrites stale pins it wrote, including in a project where the gateway was turned off for Remote Control. A failed probe keeps the last good pin, then a shipped safe default. A detected alias older than the shipped default, or one the plugin retired, loses to the shipped default. Set a persistent per-alias override with pin --opus claude-opus-5-5[1m] (same for --sonnet and --fable), or use pin --opus default to return to auto-detection. Overrides always win. pin with no arguments and doctor show each effective pin, whether it is overridden, and which lagging CLI alias a shipped pin replaced. Overrides live in ~/.claude/model-gateway/pins.json, outside the plugin cache. A pin change, and every SessionStart pin refresh, updates every registered wired project's gateway-owned pins (so a new shipped default reaches them too) and skips any project with a user-owned pin value. Restart every open Claude Code session in an affected project; changing a saved value cannot alter an open session.

  • Do NOT set a global CLAUDE_CODE_AUTO_COMPACT_WINDOW: it applies to both providers and can make Codex /compact fail after history already exceeds the Codex limit.

  • Caution: loading a huge reference skill (e.g. claude-api, ~800k chars) in a single turn can spike Codex context past the point proactive compaction can recover from. Prefer pulling large references incrementally on Codex models.

  • The advertised catalog comes from the proxy's /v1/models; models.json and the built-in list only stand in while the proxy is unreachable. A models.json file cannot add a backend that the claude-code-proxy allowlist does not support; update the proxy through setup instead.

  • Claude Desktop cannot use Codex/Grok models in this version: Desktop has its own native Gateway configuration, separate from Claude Code CLI settings, and can point at this shim's endpoint. But installed Desktop 1.49585.0 validates every Gateway model ID client-side and rejects any gpt/codex/non-Anthropic family marker before it reaches the picker or a session, whether the ID came from an explicit config entry or from discovery; a terminal [1m] is stripped first and does not change the outcome. Only a genuinely Anthropic-backed route (a real Claude alias) is usable there. This is a Desktop-side restriction, not a Model Gateway bug, and there is no supported workaround: do not suggest an Anthropic-named alias to disguise a Codex/Grok route, patch the Desktop binary, substitute credentials or auth, intercept TLS, or use a global env or hosts trick. Tell the user Desktop end-to-end Codex/Grok support is not available on this version; the Claude Code CLI is the verified path. VS Code success has been reported by users but is not independently verified here. This is version-specific and can be revisited if a future Desktop release removes the model-family filter.

  • RC-compat and missing Codex rows: On inspected Claude Code 2.1.267, RC-compatibility points ANTHROPIC_BASE_URL at api.anthropic.com, which disables gateway model discovery. The rows disappear and cache refresh cannot restore them. Claude Code can accept and persist /model claude-gpt-5.6-terra[1m], but that client-side action does not prove a later request reaches the gateway. The current cleaner preserves canonical [1m] ids; do not present it as a fix for a reported request error. RC-compatibility makes Claude Code treat the gateway as first-party and can enable experimental message threading. The shim refuses Codex/Grok thread continuations locally with HTTP 400 because those backends have no conversation state. Nothing is forwarded or rerouted. A recognizing client drops the threading beta and resends the full message history. This is version- and experiment-limited, so do not promise one refusal per session. Normal gateway mode is the verified inference path. The reported 2.1.259 client and inspected 2.1.267 client have no verified end-to-end RC result: inspected session creation uses HTTPS while compatibility transport is HTTP. A detected hosts entry or bound listener proves local HTTP transport only. Sidequest dispatch is unaffected because it resolves its explicit route marker and never uses picker discovery.

  • Claude models keep working normally at the same time (passthrough path). That does not mean a subagent invocation can mix in a gateway id freely: the Agent tool's own model parameter is a fixed host enum (sonnet/opus/haiku/fable), independent of gateway or plugin state, and a gateway id passed there is refused. To route one subagent to a specific gateway model, give it a definition file instead, with a full gateway id in model: frontmatter, and omit model from the invocation so that frontmatter applies:

    markdown
    ---
    name: luna-reviewer
    description: Reviews code changes using the GPT-5.6 Luna gateway model.
    model: claude-gpt-5.6-luna[1m]
    tools: Read, Grep, Glob
    ---
    
    Review the diff for correctness and report findings.

    Save that as .claude/agents/luna-reviewer.md and start a new session before invoking it — definitions are read at session start, so a running session won't see one just added. Invoke with subagent_type: "luna-reviewer" and no model argument. This needs only Model Gateway, not Sidequest: a concrete gateway id needs no route marker, only Sidequest's own claude-codex-auto id does. Never edit Claude Code's built-in model aliases, the Agent tool schema, or the host binary — none of that is supported or necessary. This frontmatter contract and Model Gateway's routing for a concrete id are confirmed; an end-to-end custom-agent spawn through this path has not been verified here.

  • Codex schema compatibility: Codex rejects some Unicode property escapes such as \p{Cc} and \P{Cf}. After deferred hydration, the shim changes only a Codex-bound provider hint, never Claude Code's host schema or an Anthropic request. It admits a missing dialect or Draft 2020-12 and only pattern instances reached through properties, compatible patternProperties values, additionalProperties, items, prefixItems, allOf, anyOf, dependentSchemas, propertyNames, or unevaluatedProperties/unevaluatedItems. Each real property atom must stand alone in a negated character class without ranges, set syntax, captures, backreferences, or a negative regex context. The shim removes only that atom and preserves every other regex byte. not, conditionals, oneOf, contains, references, definitions, content schemas, affected pattern-property keys, unknown containers, and unsafe regexes refuse locally with HTTP 400 naming the tool, JSON Pointer, and reason code. Tell the user it was not forwarded or rerouted. Do not claim arbitrary schemas are preserved or try to bypass that diagnostic.

Show full SKILL.md (1,860 more words)Show less

Local gateway records

Request-route logging is enabled by default. It writes metadata-only JSONL records to ~/.claude/model-gateway/logs/request-routes.jsonl: timestamp, backend, model, request path, route and effort when present, safe session and agent correlation ids, and dispatch-marker length when present. It never writes request bodies, prompts, messages, tools, authentication, or arbitrary headers. Honor a user request to disable it by setting CODEX_GATEWAY_REQUEST_LOG=0 before the shim starts, then restart the shim through setup or ensure. The value is read when the shim process starts, so /reload-plugins does not change an already-running shim. CODEX_GATEWAY_REQUEST_LOG_PATH changes the file location.

Usage observability also writes one high-water JSON file per valid session under ~/.claude/model-gateway/request-body/. The filename is derived from the session id. Its contents are the largest forwarded request-body byte count observed for that session and an observation timestamp. It does not contain the request body. No retention period is promised for either local record.

When the user asks about a failed Codex compaction and consent-gated route telemetry is available, inspect fixed compaction trace metadata only: selected/effective model, backend, upstream status, elapsed time, outcome, terminal/error code, and observed usage counters. completed and empty_summary both mean the generation finished: either a message_stop frame arrived, or a message_delta declared an end_turn/stop_sequence stop reason with no content after it and no content block left open. empty_summary is that finish with no visible non-whitespace text. incomplete means the HTTP-200 SSE stream never showed a finished generation: no stop reason, a stop reason that truncates (max_tokens, refusal, pause_turn, tool_use, anything unrecognized), content still arriving after the stop, an open block, or an upstream error frame. compaction_terminal_code is message_stop only when that frame really arrived, so a completed outcome with no terminal code means the frame was lost and the shim closed the turn itself instead of re-running the compaction. Unknown provider errors appear only as unknown_error. Never infer summary content or repeat prompts, response text, raw errors, headers, credentials, or environment values from diagnostics.

Claude Code's /remote-control only lights up when ANTHROPIC_BASE_URL is exactly the real Anthropic host, which conflicts with gateway model discovery. RC-compatibility is a reversible, opt-in local HTTP transport configuration, not a verified end-to-end Remote Control solution. At inspected Claude Code 2.1.267 its session creation uses HTTPS while the compatibility listener is HTTP; the reported 2.1.259 client and inspected client have no verified end-to-end RC result. A bound port or detected hosts entry proves only that transport state. Before enabling, tell the user that the picker rows disappear and an explicit [1m] id may be accepted client-side without proving its later inference request reaches the gateway. Keep normal gateway mode as the verified inference path. Do not suggest a cache refresh as a fix for a reported request error; the current cleaner preserves canonical [1m] ids.

For the confirmation-gated procedure, use the remote-control-compatibility skill. It manages the plugin-marked hosts block, creates a backup before an elevated write, reconciles gateway mode, and checks the final transport state. Do not edit the hosts file outside that procedure. If effective process env ANTHROPIC_BASE_URL is HTTPS api.anthropic.com (including port 443), enabling is refused before any backup, hosts write, startup, or reconciliation because the loopback mapping cannot serve TLS. A user-controlled Claude Code CLI launch can correct or unset that value, then restart. If a host replaces it, use the supported Claude Code CLI on the wired project instead. Desktop routing is unsupported under forced overrides on Windows and macOS, and settings, parent, or User-scope edits cannot be promised to win. Disabling stays available.

  • After the user directly confirms remote-control enable --confirm, the plugin creates a backup and writes its marked hosts block mapping api.anthropic.com to loopback: 127.0.0.1 api.anthropic.com on Windows (C:\Windows\System32\drivers\etc\hosts, needs Administrator), macOS, and Linux (/etc/hosts, needs sudo). Do not edit the hosts file outside that procedure.
  • ensure/setup/doctor detect the entry (read-only) and, only after confirming the shim can actually bind loopback port 80, switch ANTHROPIC_BASE_URL to http://api.anthropic.com and start a second listener on port 80 next to the usual 127.0.0.1:18764. Exactly one line tells the user to restart Claude Code when the mode changes either direction.
  • Removing the hosts entry, or port 80 becoming unavailable (no permission, or something else is using it), reverts to default mode automatically, again with one restart line.
  • doctor reports the hosts entry (if any), whether port 80 actually bound (and why not if it didn't), and which mode each settings scope (user/project) is wired to.
  • Test/advanced overrides: CODEX_GATEWAY_HOSTS_FILE (custom hosts path), CODEX_GATEWAY_COMPAT_PORT (port other than 80). Neither is needed for normal use.

Day-2 operations

bash
... status      # what's running
... doctor      # binary, auth, ports, model count, settings wiring
... ensure      # start whatever is down (SessionStart hook runs this with --quiet)
... stop
... env --remove   # unwire Claude Code (do this BEFORE uninstalling the plugin)

status, doctor and ensure read the shim through one shared probe and print the same line, shim (model router) on :<port>: <state>, where the state is running-ours (serving <version>), running-foreign (PID <pid>, <install root or owner unidentified>), starting (PID <pid> since <time>) or stopped. ensure leaves running-ours at the installed version alone and succeeds, waits up to its startup window for starting before replacing anything, and refuses running-foreign.

doctor prints the full model-window table: backend and picker ids, backend and advertised windows, Claude Code's resolved client window and compaction point, sentry mode and trigger, and the measurement date. It includes Codex, Grok, and native Claude pin rows. Its model-id check is useful for stale shim ids, but a PASS does not prove every supported model is present. A proxy from 0.1.14 through 0.1.35 can omit GPT-6 Astra while this check passes. If Astra is missing, check the installed and serving proxy version, rerun setup to fetch the latest release, and fully restart Claude Code. Astra requires claude-code-proxy 0.1.36 or newer. A models.json edit cannot add a backend that the proxy allowlist does not support. A FAIL naming missing and extra ids means the shim is stale even when its version matches: restart it through the normal ensure or setup path, then restart Claude Code sessions so the picker re-discovers the rows.

Logs live in ~/.claude/model-gateway/logs/. guardian.log has recovery output; lifecycle.jsonl has bounded process evidence that doctor summarizes. Ports: shim 18764, proxy 18765 (override with CODEX_GATEWAY_PORT / CODEX_GATEWAY_PROXY_PORT, but the env block and running processes must agree).

Failure modes worth knowing

  • Every request fails after wiring: a SessionStart hook can time out while it starts the shim. Claude Code cancels that hook, and a directly spawned supervisor can die with its process tree before it records an exit. Model Gateway launches the supervisor outside that tree and stops waiting before the hook budget, but if this is an older install or it still repeats, run doctor, check logs. Worst case env --remove restores stock behavior instantly.
  • Codex sessions drop while this plugin's suite runs: this version scopes fixture cleanup to the test gateway home, so the suite never touches the installed gateway. If it happens after updating, run doctor and include its supervisor conflict line and lifecycle evidence.
  • Codex models error, Claude models fine: proxy or OpenAI side. Check login state (doctor shows auth), then proxy log. OpenAI gates non-Codex clients by request fingerprint; when they tighten it, requests die mid-stream until claude-code-proxy ships a fix, so suggest re-running setup (it fetches the latest release).
  • doctor says upstream-unavailable: a final Codex inference failed in the last 30 seconds; the message names its status, time, and hold end. It records completed request outcomes, not /v1/models or a health check, and clears only after a completed successful Codex response. The 30-second expiry means there is no recent failure evidence, not that Codex is live. An attributed OpenAI 401, 403, or 429 rejection enters upstream-blocked. A 401 or 403 stays until setup or a completed successful Codex response clears it. A 429 block expires: upstreamBlocked.expiresAt comes from the 429's Retry-After, else claude-code-proxy's usage-limit reset header, else 60 seconds, and the doctor message names that time. It lifts by itself then, or sooner on a completed successful Codex response, and a later rejected request can latch it again. Sidequest consumes a cached catalog and can lag this state by up to five minutes.
  • Codex turn with no output: claude-code-proxy answers a Codex turn that completed with no text, tool call, or thinking as a 503 "Codex completed without producing output". The shim answers it as an empty end_turn instead, so the session ends the turn rather than retrying into the same empty answer; the shim log records each one.
  • Gateway models vanish from a Sidequest board a few minutes after the shim starts: Sidequest discards a catalog older than five minutes and refreshes it by running catalog --refresh --json. Run that command by hand and read stderr plus the exit code. It exits non-zero and names the reason when it declines to write (shim not answering /healthz, /v1/models erroring, or a model list with no gateway ids in it), leaving the stored catalog and its timestamp untouched. Exit 0 with no diagnostic means it did write, so compare the printed updatedAt with the stored file.
  • Startup, recovery, restart, or drain refuses to touch a listener: each ownership probe shares one CODEX_GATEWAY_PROBE_TIMEOUT_MS budget (2 seconds by default, 8 seconds on Windows, where the Win32_Process lookup itself typically takes 1.8-2.4 seconds). When that budget expires, the refusal says so, names the elapsed budget, and points to CODEX_GATEWAY_PROBE_TIMEOUT_MS as the override. A malformed process result or unrecognized command remains an ownership-unknown refusal without the timeout guidance. Startup records owner-unknown and leaves that listener untouched; recovery leaves it for the next tick. A confirmed foreign owner is also left untouched. Probe children are stopped with the supervisor, so they cannot keep a test fixture home open.
  • doctor shows Not authenticated right after an upgrade: bumping the proxy binary (e.g. 0.1.10 → 0.1.17 via setup) can invalidate the credential the old version accepted — the new binary reads it as not authenticated and setup stops before wiring. Fix: re-run login, then setup again to finish. Until then every Codex model is down, so any run that routes to Codex (a whole sidequest board of Codex-tier tickets, for one) stalls entirely.
  • GPT-6 Astra is missing from /model: do not diagnose account access first. Check the installed and serving claude-code-proxy version. Astra requires 0.1.36 or newer; 0.1.35 does not include its backend allowlist, while the current doctor floor can still pass. Re-run setup to fetch the latest GitHub release, then fully restart Claude Code. Do not propose models.json: it cannot add a backend the proxy does not allow.
  • No "From gateway" rows in /model: discovery is off (CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY missing), Claude Code < v2.1.129, CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC is set (it disables discovery), or RC-compatibility is active. Claude Code only refetches discovery with an API-key credential, so OAuth users rely on Model Gateway's cache write. Run doctor to check the cache, then restart Claude Code after it updates; /reload-plugins does not reload picker rows.
  • Thinking/reasoning: the Codex backend doesn't return thinking blocks into Claude Code's UI; that's an upstream limitation, not a bug here.
  • Permission mode flips to "accept edits on" during Codex sessions: caused by GPT models calling the plan-mode tools; an approved ExitPlanMode downgrades the mode instead of restoring it (anthropics/claude-code#39973). The shim strips EnterPlanMode/ExitPlanMode from Codex-bound requests since 0.2.1, so this shouldn't recur; if it does, make sure the shim was restarted (stop + start). Shift+Tab restores the mode in an affected session. Escape hatch to re-enable plan tools: CODEX_GATEWAY_KEEP_PLAN_TOOLS=1.

© Eigenwise, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/model-gateway/skills/model-gateway of Eigenwise/eigenwise-toolshed.

Open the folder on GitHubat commit 0140ec1

Compare with similar skills

Model Gateway next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Gateway compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Gateway this skillEigenwise/eigenwise-toolshed277—~8.7kAutomated safety check: WarnMIT
Caveman Gateway SetupJuliusBrussee/caveman110k1 repos~2.6kAutomated safety check: WarnApache-2.0
Azure Architecture Autopilotgithub/awesome-copilot40k1 repos~1.9kAutomated safety check: PassMIT
Capacitymicrosoft/GitHub-Copilot-for-Azure2552 repos~1.7kAutomated safety check: PassMIT
Youtubeeat-pray-ai/yutu696—~1.1kAutomated safety check: PassMIT
Local Stack RuntimeOpenHands/OpenHands90k—~375Automated safety check: PassMIT

Similar skills

  • Caveman Gateway Setup

    JuliusBrussee/caveman

    Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

    110k GitHub starsUsed in 1 repo~2.6k tokens
    DevOps & CloudAuto-check: warnings
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Capacity

    microsoft/GitHub-Copilot-for-Azure

    Official

    Discovers available Azure OpenAI model capacity across regions and projects.

    255 GitHub starsUsed in 2 repos~1.7k tokens
    DevOps & CloudAuto-check passed
  • Youtube

    eat-pray-ai/yutu

    A skill your agent uses whenever the user mentions YouTube, video uploads, channel management, playlists, video SEO, or any YouTube Data API operation.

    696 GitHub stars~1.1k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Local Stack Runtime

    OpenHands/OpenHands

    This skill should be used when the user asks to "change the dev stack", "add a runtime service", "change the launcher", "update Docker", "bump Agent Server", "change ingress routing", or changes…

    90k GitHub stars~375 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Openai Docs

    aafqaq/codex-lb-enhanced

    A skill your agent uses when the user asks how to build with OpenAI products or APIs and needs up-to-date official documentation with citations (for example: Codex, Responses API, Chat Completions…

    102 GitHub starsUsed in 3 repos~861 tokens
    DevOps & CloudAuto-check passed

More from Eigenwise/eigenwise-toolshed

All 14 skills in this repo
  • Add Rule

    Eigenwise/eigenwise-toolshed

    Create or edit a live-rules instruction in the project's atomic rule set.

    277 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Map Codebase

    Eigenwise/eigenwise-toolshed

    Create a self-maintaining codebase map in .claude/.codebase-info/.

    277 GitHub stars~4.2k tokensUpdated 2 days ago
    Auto-check passed
  • Setup

    Eigenwise/eigenwise-toolshed

    Set up a Claude Code workspace for a new or existing project, informed by hindsight from the user's whole session history.

    277 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Groom

    Eigenwise/eigenwise-toolshed

    Audit a Sidequest board for completed, stale, duplicate, or superseded tickets, then safely close clear cases.

    277 GitHub stars~3.4k tokensUpdated 2 days ago
    Auto-check passed
  • Manage Rules

    Eigenwise/eigenwise-toolshed

    Inspect, audit, enable, or disable project live-rules. An agent skill from Eigenwise/eigenwise-toolshed.

    277 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Toolshed Doctor

    Eigenwise/eigenwise-toolshed

    Run a read-only health check for Quartermaster and installed Toolshed plugins.

    277 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed

Works with

Categories

Questions about Model Gateway

What does Model Gateway do?

Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker. Model Gateway is an agent skill from Eigenwise/eigenwise-toolshed. Set up, update, or diagnose the local ChatGPT/Codex gateway for Claude Code's /model picker.

When should I use Model Gateway?

Model Gateway fits situations like: model visibility.

How do I install Model Gateway in Claude Code?

Run `npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a claude-code`. Or copy the skill folder (plugins/model-gateway/skills/model-gateway in Eigenwise/eigenwise-toolshed) into .claude/skills/model-gateway in your project. Claude Code loads it when a task matches its description.

How do I install Model Gateway in Codex?

Run `npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a codex`. Or copy the skill folder (plugins/model-gateway/skills/model-gateway in Eigenwise/eigenwise-toolshed) into .agents/skills/model-gateway in your project. Codex loads it when a task matches its description.

Can I use Model Gateway in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Eigenwise/eigenwise-toolshed --skill model-gateway -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-gateway, .gemini/skills/model-gateway, .github/skills/model-gateway and .opencode/skills/model-gateway in your project.

What does Model Gateway need to run?

Going by SKILL.md and its folder, Model Gateway needs the command-line tools its instructions call (node).

Does Model Gateway access the network?

SKILL.md names 1 domain. In commands or code: api.anthropic.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Model Gateway safe to install?

Our automated static check of SKILL.md flagged 2 warning(s): contains instruction-override wording (e.g. “without asking the user”). Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Model Gateway use?

Model Gateway is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Gateway use?

About 8.7k tokens (SKILL.md is roughly 35k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Gateway?

Skills that share tags, products or a category with Model Gateway: Caveman Gateway Setup (JuliusBrussee/caveman, 110k stars), Azure Architecture Autopilot (github/awesome-copilot, 40k stars), Capacity (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Youtube (eat-pray-ai/yutu, 696 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Gateway?

Eigenwise (a GitHub user) maintains it in Eigenwise/eigenwise-toolshed, which has 277 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.

Source: Eigenwise/eigenwise-toolshed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.