Agent skill

Model Routing Patterns

by softspark in softspark/ai-toolkit

Capability-aware model routing for Codex, Copilot, Claude and provider APIs.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Model Routing Patterns

skills CLI
$ npx skills add softspark/ai-toolkit --skill model-routing-patterns -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install softspark/ai-toolkit model-routing-patterns --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/app/skills/model-routing-patterns .claude/skills/model-routing-patterns && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-routing-patterns
GitHub stars
179
Token cost
~2.4k tokens
SKILL.md length
1,190 words
Files
1
Skills in repo
112
Repo updated
First seen
Licence
Apache-2.0

At a glance

Capability-aware model routing for Codex, Copilot, Claude and provider APIs.

  • Tasks that involve Model routing and gateways
  • SKILL.md covers Claude Code: available…, Reviewed model reference…, Effort and caching and Pattern 1: explicit task routing, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Model Routing Patterns is an agent skill from softspark/ai-toolkit. Capability-aware model routing for Codex, Copilot, Claude and provider APIs. Triggers: model routing, model selection, reasoning effort, approved fallback, cost, escalation.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model routing and gateways. The repository describes itself as: Professional-grade AI coding toolkit: 94 skills, 44 agents, multi-platform (Claude, Cursor, Windsurf, Copilot, Gemini, Cline, Roo Code, Aider, Augment, Antigravity, Codex CLI… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Model routing and gateways

Example prompts

  • “/model-routing-patterns”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read

What it can do on your machine

Read from SKILL.md and the folder at commit d64db2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com
    • learn.chatgpt.com
    • developers.openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Routing Patterns loads about 2.4k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 1,190 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from softspark/ai-toolkit at commit d64db2b, republished under its Apache-2.0 licence (© softspark). 1,190 words, ~2,439 tokens.

Download SKILL.mdSave it as .claude/skills/model-routing-patterns/SKILL.md (or your agent's skills folder).
name
model-routing-patterns
description
Capability-aware model routing for Codex, Copilot, Claude and provider APIs. Triggers: model routing, model selection, reasoning effort, approved fallback, cost, escalation.
allowed-tools
Read
effort
medium
user-invocable
false

Model Routing Patterns

Choose routes from measured quality, latency and cost under the user's approved model and spending policy. An explicit model choice overrides automatic routing. This skill never authorizes a model-tier, permission or budget change.

Apply the same policy across Codex, Copilot and Claude: retain the model selected by the native client. Provider API IDs, editor picker labels and agent aliases are separate namespaces; do not translate between them by resemblance. For a cross-provider application route, validate capabilities and availability on each provider, and obtain authorization before transferring data or changing spend. The project inventory is kb/reference/model-compatibility.md.

<!-- CLAUDE_CODE_ONLY_START -->

Claude Code: available executors and supervised routing

Apply this section only in Claude Code. Other clients use their native agents and keep their configured models and reasoning effort. Never call the Claude Codex plugin from Codex, Antigravity or another client.

Check availability before choosing an executor. The Codex route requires all three: the codex@openai-codex plugin is installed, effectively enabled for the current project, and codex:codex-rescue is callable through the current Agent tool. The toolkit SessionStart hint checks local metadata only; it does not prove runtime availability, authentication, model access or security access. An existing cache directory alone is insufficient. If the hint is absent, verify those conditions from the client's plugin status and agent catalog.

When unavailable or disabled, use installed native agents with their configured models and effort. Do not install or enable a plugin, force a Codex command, or change the current session's model to satisfy this policy. If delegation itself is unavailable, do the authorized work in the current session and report that limitation. Explicit user model choices always take precedence.

The toolkit's Claude defaults use these roles when available:

WorkExecutor
Planning, coordination and acceptanceNative orchestrator, Opus high
Bounded implementation, tests and routine fixesNative implementation agent, Sonnet high, or the available Codex plugin
Security workPrefer Codex gpt-6-astra, xhigh; otherwise the installed native security agent and its configured model
Hard debuggingAvailable Codex at xhigh, or the native debugger, Opus xhigh

Use Codex proactively for substantial independent implementation or diagnosis; do not require a failure first, and keep trivial edits with the current executor. Retain the native domain role, file ownership and review criteria when selecting the Codex executor. If Astra is unavailable, report that and use the configured native security agent. Do not silently select a different Codex model.

Dispatch with Agent(subagent_type="codex:codex-rescue", prompt="--fresh ..."). Include the task, working directory, owned paths, relevant evidence, acceptance criteria and verification commands. Pass --model gpt-6-astra --effort xhigh for the security route, and --effort xhigh for hard debugging. Ordinary coding keeps the configured Codex model and effort unless the user chose otherwise. These flags control the Codex worker, not the Sonnet forwarding wrapper. Do not use the retired Spark alias or substitute a plugin command for the agent.

Request read-only work explicitly for audits or diagnosis; write access is appropriate only for an authorized implementation task. Codex's sandbox and approval restrictions still apply. Permission or authentication failures return to the supervisor; they do not authorize bypasses or a new account/provider. Selecting Astra does not load a security package or confer Trusted Access. Use an existing security package only when it is available in the delegated runtime and within the user's task; report which workflow actually ran.

The supervisor owns lifecycle and verification. A background job identifier is not completion: collect the final result using the plugin's documented status and result controls, keyed to that job. Resume only the identified prior task; use a fresh task for independent work. Empty output, failed startup and partial results are failures to resolve, not successful delegation. Do not redispatch failed writes until their partial changes have been inspected. Verify the diff, tests and required review before acceptance, then finish or cancel owned work. The forwarding wrapper must only forward; do not ask it to inspect or monitor.

When creating agents, keep native Claude model fields valid. Codex is a separate executor, not a Claude model: value or an automatic Claude Team member. Reuse the installed plugin instead of generating another wrapper or editing its cache.

<!-- CLAUDE_CODE_ONLY_END -->
Show full SKILL.md (510 more words)Show less

Reviewed model reference (2026-10-01)

ModelClaude API IDAPI effort default
Claude Opus 5.5claude-opus-5-5medium
Claude Fable 5.1claude-fable-5-1high
Claude Sonnet 5.5claude-sonnet-5-5high
Claude Haiku 4.5claude-haiku-4-5-20251001Effort unsupported

These are dated identifiers, not a runtime upgrade policy. Check the provider's model availability and current pricing before estimating costs. Do not encode universal cost ratios or declare a model best for every workload. Preserve a user-specified older model while supported; surface retirement or availability problems explicitly.

Effort and caching

Opus 5.5, Fable 5.1 and Sonnet 5.5 support low, medium, high, xhigh, and max. Effort is a behavior control, not a hard spending cap. Opus 5.5 and Fable 5.1 use always-on adaptive thinking; a small output limit can truncate the answer.

Changing top-level output_config.effort invalidates message cache blocks, with model-dependent effects on earlier caches. Supported per-message effort changes can preserve the prefix. Do not assume effort tuning is cache-neutral.

Keep the configured agent/skill effort. An approved application experiment may compare effort settings, recording total thinking/output usage and completion quality at the same task budget.

Pattern 1: explicit task routing

Use application configuration reviewed for the workload. Labels such as "classification" or "architecture" are evaluation slices, not proof that one family is sufficient or necessary.

python
def choose_model(task, routes, allowed_models, explicit_model=None):
    candidate = explicit_model if explicit_model is not None else routes.get(task)
    if candidate is None or candidate not in allowed_models:
        raise ValueError("No approved model for this request")
    return candidate

Start with the selected model. Add a separate classification call only when measured routing savings exceed its latency and token cost.

Pattern 2: validation-based escalation

Evaluate a result with task-specific checks: schema validation, failing tests, retrieval evidence, or human labels. A model's self-reported confidence is not a calibrated probability. Do not pass hidden reasoning between models; pass the problem, relevant evidence, and a short failure summary.

Escalate only along an approved route with a bounded attempt count. If no approved route remains, report failure or send the item for human review.

Pattern 3: delegation within configured roles

A planner can split independent tasks between workers when the task and client permit it. Use each agent's configured model and tools. Do not rewrite frontmatter or force a cheaper worker because a generic diagram suggests it.

Compare end-to-end quality and cost, including planning, handoffs and synthesis. More agents do not inherently save tokens.

Pattern 4: resilience fallback

Retry transient failures within the existing retry policy before considering a different model. The official SDK may already retry requests; avoid multiplying its retries with another unbounded loop.

A fallback must preserve the user's model requirement, context limits, structured output support and tool permissions. If changing models is not authorized, stop with the original model's error. Record every actual fallback and its reason.

Measuring

Track model ID, effort, policy version, attempts, latency, cache reads/writes and total billable tokens. Evaluate quality per task type and language using held-out examples. Set acceptance criteria before changing the route; do not use fixed confidence thresholds, traffic percentages or cost multipliers as universal rules.

Reviewed 2026-10-01:

Use prompt-caching-patterns for cache design and json-mode-patterns for structured results. Use the llm-ops-engineer agent for application routing.

© softspark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in app/skills/model-routing-patterns of softspark/ai-toolkit.

Open the folder on GitHubat commit d64db2b

Compare with similar skills

Model Routing Patterns next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Routing Patterns compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Routing Patterns this skillsoftspark/ai-toolkit179—~2.4kAutomated safety check: PassApache-2.0
Shogun Bloom Configyohey-w/multi-agent-shogun1.4k—~3.1kAutomated safety check: PassMIT
Codemie Analyticscodemie-ai/codemie-code294—~7.5kAutomated safety check: PassApache-2.0
Model Routernidhi-singh02/agent-router110—~1.2kAutomated safety check: PassMIT
Codex Model Routing Teamzjp1997720/codex-model-routing-team158—~736Automated safety check: PassMIT
Add Modelget-convex/convex-evals129—~1.5kAutomated safety check: NotesApache-2.0

Similar skills

  • Shogun Bloom Config

    yohey-w/multi-agent-shogun

    Interactive wizard: guided questions with multiple-choice options about subscriptions, then outputs a ready-to-paste capabilitytiers YAML + fixed agent model assignments.

    1.4k GitHub stars~3.1k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Codemie Analytics

    codemie-ai/codemie-code

    CodeMie Analytics expert — use this skill whenever the user asks about CodeMie usage data, AI adoption metrics, user leaderboards, CLI insights, spending, LiteLLM costs, token usage, or wants to…

    294 GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Model Router

    nidhi-singh02/agent-router

    A skill your agent uses when the user asks to pick a model, subscription, or reasoning effort, or to run router status, usage refresh, or resume a router session.

    110 GitHub stars~1.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Codex Model Routing Team

    zjp1997720/codex-model-routing-team

    在 Codex App 中为复杂、可并行的知识工作或编程任务自动创建多个可指定模型与推理强度的后台任务,由主 Agent 负责规划、分工、集成和验收。用于多来源调研、多章节内容、复杂 Skill/PPT、跨模块开发、独立验证或 2 个以上互不依赖工作流;也用于用户明确要求模型路由、后台 Worker、Agents Team…

    158 GitHub stars~736 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    get-convex/convex-evals

    Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

    129 GitHub stars~1.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Jev Route

    cobusgreyling/Jev

    Route with TypeSafe Jev — map Choice plus confidence to act/confirm/human, or pick a coding-agent model tier.

    134 GitHub stars~481 tokensUpdated 17 days ago
    AI & LLM EngineeringAuto-check passed

More from softspark/ai-toolkit

All 112 skills in this repo
  • Prepare Test Env

    softspark/ai-toolkit

    Prepare or verify a project QA environment with source identity, readiness, browser access, evidence paths and owned cleanup.

    179 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check: notes
  • A11y Validate

    softspark/ai-toolkit

    Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check: notes
  • Analyze

    softspark/ai-toolkit

    Analyzes code quality, complexity, patterns across codebase.

    179 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Autonomous Dev

    softspark/ai-toolkit

    Drives a brief, specification, issue or existing PR through implementation, review, tests and QA to a ready PR.

    179 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Brand Voice

    softspark/ai-toolkit

    Direct technical voice for docs, README, user-facing text. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • CI

    softspark/ai-toolkit

    Detect/generate/debug CI pipeline config (GitHub Actions, GitLab CI).

    179 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes

Questions about Model Routing Patterns

What does Model Routing Patterns do?

Capability-aware model routing for Codex, Copilot, Claude and provider APIs. Model Routing Patterns is an agent skill from softspark/ai-toolkit. Capability-aware model routing for Codex, Copilot, Claude and provider APIs.

When should I use Model Routing Patterns?

Model Routing Patterns fits situations like: tasks that involve Model routing and gateways.

How do I install Model Routing Patterns in Claude Code?

Run `npx skills add softspark/ai-toolkit --skill model-routing-patterns -a claude-code`. Or copy the skill folder (app/skills/model-routing-patterns in softspark/ai-toolkit) into .claude/skills/model-routing-patterns in your project. Claude Code loads it when a task matches its description.

How do I install Model Routing Patterns in Codex?

Run `npx skills add softspark/ai-toolkit --skill model-routing-patterns -a codex`. Or copy the skill folder (app/skills/model-routing-patterns in softspark/ai-toolkit) into .agents/skills/model-routing-patterns in your project. Codex loads it when a task matches its description.

Can I use Model Routing Patterns in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softspark/ai-toolkit --skill model-routing-patterns -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-routing-patterns, .gemini/skills/model-routing-patterns, .github/skills/model-routing-patterns and .opencode/skills/model-routing-patterns in your project.

What does Model Routing Patterns need to run?

SKILL.md names no scripts, command-line tools or credentials: Model Routing Patterns is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read.

Does Model Routing Patterns access the network?

SKILL.md names 3 domains. As links in the text: platform.claude.com, learn.chatgpt.com and developers.openai.com. This is read from the text; nothing was executed.

Is Model Routing Patterns safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Routing Patterns use?

Model Routing Patterns is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Routing Patterns use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Routing Patterns?

Skills that share tags, products or a category with Model Routing Patterns: Shogun Bloom Config (yohey-w/multi-agent-shogun, 1.4k stars), Codemie Analytics (codemie-ai/codemie-code, 294 stars), Model Router (nidhi-singh02/agent-router, 110 stars) and Codex Model Routing Team (zjp1997720/codex-model-routing-team, 158 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Routing Patterns?

softspark (a GitHub user) maintains it in softspark/ai-toolkit, which has 179 GitHub stars. The repository holds 112 skills in this directory. The repository was last updated on October 7, 2026.

Source: softspark/ai-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.