Agent skill

Production Model Router

by LeoYeAI in LeoYeAI/openclaw-master-skills

Route each user request to the most cost-effective model or multi-model workflow based on task type, complexity, risk, latency, budget, tool needs, and verification requirements.

MITAuto-check passedAI & LLM Engineering

Install Production Model Router

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill production-model-router -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills production-model-router --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-routing-orchestrator .claude/skills/production-model-router && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
production-model-router
GitHub stars
2.2k
Token cost
~4.9k tokens
SKILL.md length
2,504 words
Files
2
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

Route each user request to the most cost-effective model or multi-model workflow based on task type, complexity, risk, latency, budget, tool needs, and verification requirements.

  • Works in 12 steps: Complexity → Exactness requirement → Risk level → …
  • Tasks that involve Model routing and gateways
  • SKILL.md covers Overview, Primary objective, When to use and Do not use, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Production Model Router is an agent skill from LeoYeAI/openclaw-master-skills. Route each user request to the most cost-effective model or multi-model workflow based on task type, complexity, risk, latency, budget, tool needs, and verification requirements.

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `_meta.json`).

It sits in AI & LLM Engineering, covering Model routing and gateways. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • Tasks that involve Model routing and gateways

Example prompts

  • “/production-model-router”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Complexity
  2. Exactness requirement
  3. Risk level
  4. Latency priority
  5. Budget strategy
  6. Context burden
  7. Modality burden
  8. Single economy
  9. Single balanced
  10. Single premium
  11. Tool-assisted reasoning
  12. Staged pipeline

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Production Model Router loads about 4.9k tokens when it runs. Until then it costs about 51 tokens; SKILL.md has 2,504 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 2,504 words, ~4,874 tokens.

Download SKILL.mdSave it as .claude/skills/production-model-router/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
production-model-router
description
Route each user request to the most cost-effective model or multi-model workflow based on task type, complexity, risk, latency, budget, tool needs, and verification requirements.
version
2026.3.27

Production Model Router

Overview

Use this skill to decide which model tier, workflow shape, and verification strategy should handle a user's request.

The goal is to maximize cost-effectiveness without sacrificing task fit, correctness, or operational reliability.

This skill does not blindly choose the strongest model. It chooses the cheapest safe path that still meets the quality bar for the task.

It may recommend:

  • a single low-cost model
  • a single balanced model
  • a single premium model
  • a tool-assisted model workflow
  • a staged multi-model pipeline
  • a parallel comparison workflow
  • a draft-and-review workflow
  • a consensus or verifier workflow

Primary objective

For every request, choose the minimum-cost execution path that can still satisfy:

  • task quality
  • correctness requirements
  • latency expectations
  • safety or risk constraints
  • output format needs
  • tool and modality requirements

When to use

Use this skill when you need to decide:

  • which model should answer a given user request
  • whether a cheap model is enough
  • when to escalate to a stronger reasoning model
  • when to use one model versus multiple models
  • when to use tools instead of relying on pure model reasoning
  • how to handle complex calculations, code, multimodal input, long context, or high-risk tasks
  • how to balance cost, speed, and answer quality in production

Do not use

Do not use this skill to:

  • answer the original business question directly
  • fabricate model capabilities without evidence from the environment or configuration
  • assume the most expensive model is always the best choice
  • route high-risk exact tasks to a cheap model without verification
  • rely on pure language generation for exact arithmetic when tools are available

Inputs to collect

Collect or infer the following from the request and system context:

Request characteristics
  • task type
  • domain
  • expected output type
  • presence of images, files, tables, code, or long documents
  • need for exactness versus approximate usefulness
  • whether the request is open-ended or precision-critical
Execution constraints
  • budget sensitivity
  • latency sensitivity
  • quality expectation
  • token or context size pressure
  • tool availability
  • need for citations or traceability
  • need for reproducibility
Risk profile
  • low-risk
  • medium-risk
  • high-risk
Failure tolerance
  • whether a rough answer is acceptable
  • whether the answer must be verified
  • whether disagreement between models would be valuable

Task taxonomy

Classify the request into one or more of these categories:

  1. Simple generation

    • rewrite
    • summarization
    • formatting
    • light translation
    • basic brainstorming
  2. General reasoning

    • explanation
    • comparison
    • concept mapping
    • normal business analysis
  3. Deep reasoning

    • multi-step planning
    • tradeoff analysis
    • architecture design
    • ambiguous decision support
    • chain-dependent reasoning
  4. Exact calculation or formal logic

    • arithmetic
    • financial calculations
    • unit conversion
    • spreadsheet-like reasoning
    • symbolic or step-sensitive math
    • combinatorics or logic puzzles where exactness matters
  5. Coding and technical execution

    • code generation
    • debugging
    • refactoring
    • test generation
    • query writing
    • infrastructure or API design
  6. Long-context synthesis

    • large documents
    • multiple files
    • multi-source comparison
    • transcript or contract review
  7. Multi-modal tasks

    • image understanding
    • diagram interpretation
    • PDF with layout-heavy content
    • video or audio related tasks if supported
  8. High-risk tasks

    • medical
    • legal
    • financial decisions
    • compliance
    • security-sensitive operations
    • anything where incorrect advice has material consequences

Core routing principle

Always prefer the cheapest path that can safely succeed.

Apply this order of preference:

  1. Cheap single-model path
  2. Balanced single-model path
  3. Premium single-model path
  4. Tool-assisted path
  5. Staged multi-model path
  6. Parallel multi-model comparison
  7. Premium plus verifier or consensus workflow

Do not escalate unless the task characteristics justify it.

Model tiers

Use abstract capability tiers unless the deployment specifies exact providers.

Economy tier

Use for:

  • simple rewriting
  • formatting
  • low-risk classification
  • short summaries
  • lightweight extraction
  • first-pass triage

Strengths:

  • lowest cost
  • fast response
  • good for straightforward tasks

Weaknesses:

  • weaker deep reasoning
  • more brittle on ambiguity
  • worse on exactness-critical tasks
Balanced tier

Use for:

  • everyday product and engineering work
  • standard reasoning
  • moderate code tasks
  • moderate document analysis
  • most business and writing tasks

Strengths:

  • solid quality-cost tradeoff
  • handles most normal production traffic
  • reasonable speed and robustness

Weaknesses:

  • may still fail on highly ambiguous or exacting tasks
  • not always enough for hard reasoning or high-risk requests
Premium tier

Use for:

  • deep reasoning
  • difficult code and architecture problems
  • long-context synthesis with subtle dependencies
  • high-value outputs
  • high-risk tasks requiring stronger judgment

Strengths:

  • strongest reasoning
  • better ambiguity handling
  • better synthesis quality

Weaknesses:

  • highest cost
  • often slower
  • overkill for simple tasks
Tool-assisted tier

Use when exactness matters more than fluent wording.

Use this path for:

  • arithmetic
  • deterministic calculations
  • spreadsheet operations
  • formula application
  • structured data transformation
  • exact code execution or testing if available
  • retrieval-backed factual tasks

Rule: When a task requires exact numeric correctness, prefer tools plus model orchestration over pure model reasoning.

Decision dimensions

Score the request across these dimensions:

1. Complexity
  • low
  • medium
  • high
  • very high
2. Exactness requirement
  • low: approximate answer is acceptable
  • medium: mostly correct is acceptable
  • high: exact result expected
  • critical: exact result plus verification required
3. Risk level
  • low
  • medium
  • high
4. Latency priority
  • urgent
  • normal
  • relaxed
5. Budget strategy
  • minimize cost
  • balanced
  • quality-first
6. Context burden
  • short
  • moderate
  • long
  • extreme
7. Modality burden
  • text only
  • image or PDF
  • mixed inputs

Hard routing rules

Apply these rules before any soft optimization.

Exact calculation rule

If the task involves exact arithmetic, formulas, tables, accounting-like operations, unit-sensitive conversions, or step-sensitive logic:

  • do not rely on a pure language-only route when tools are available
  • prefer tool-assisted execution
  • use a balanced or premium model only to interpret the task and explain results
  • add a verification step for high-impact numeric outputs
High-risk rule

If the task is high-risk:

  • do not use economy-only routing as the final path
  • require either premium single-model reasoning with grounding or a model plus verifier workflow
  • add citations, checks, or a review pass when possible
Ambiguity rule

If the task is materially ambiguous and the answer quality depends on interpretation:

  • use a stronger reasoning tier or a two-stage workflow
  • do not finalize on a cheap first-pass answer without clarification or review
Long-context rule

If the input is large or multi-document:

  • prefer staged processing
  • use extraction or chunk summarization first
  • then use a stronger model for synthesis if needed
  • avoid sending everything to the strongest model by default if staged reduction is cheaper and safe
Multimodal rule

If the task includes images, diagrams, PDFs with layout dependence, or visual interpretation:

  • use a model path that actually supports the required modality
  • do not route to a text-only path
Coding rule

For code tasks:

  • simple boilerplate or syntax transforms may use balanced or economy tiers
  • debugging, architecture, concurrency, performance, or tricky refactors should escalate to balanced or premium tiers
  • if execution, linting, tests, or static analysis tools are available, prefer tool-assisted validation

Choose one of these workflow shapes.

1. Single economy

Use when:

  • low complexity
  • low risk
  • low exactness requirement
  • low business impact
  • latency and cost matter more than polish

Examples:

  • rewrite text
  • generate short summaries
  • classify intent
  • format content
2. Single balanced

Use when:

  • the task is typical production traffic
  • moderate reasoning is needed
  • quality matters but premium is not justified

Examples:

  • standard technical Q&A
  • ordinary product copy
  • moderate coding tasks
  • document understanding with limited ambiguity
3. Single premium

Use when:

  • the task needs strong reasoning
  • the output is strategically important
  • ambiguity is high
  • long dependency chains matter

Examples:

  • system design
  • complex debugging
  • nuanced tradeoff analysis
  • sensitive writing requiring higher judgment
4. Tool-assisted reasoning

Use when:

  • exactness matters
  • calculations are required
  • data must be transformed reliably
  • code can be executed or checked
  • retrieval is needed for factual grounding

Pattern:

  • model interprets the request
  • tools compute, retrieve, or validate
  • model explains and formats the result
5. Staged pipeline

Use when:

  • the request is large, expensive, or decomposable
  • cheap preprocessing can reduce downstream cost

Pattern:

  1. economy or balanced model for triage or extraction
  2. balanced or premium model for synthesis
  3. optional verifier pass

Examples:

  • long-document analysis
  • large support threads
  • multi-file engineering review
6. Draft and review

Use when:

  • low-cost drafting is possible but final quality matters

Pattern:

  1. cheaper model drafts
  2. stronger model critiques, corrects, or upgrades

Best for:

  • writing
  • technical explanations
  • proposal drafting
  • code review style tasks
7. Parallel comparison

Use when:

  • model disagreement is informative
  • solution diversity is valuable
  • the task is comparative or open-ended

Pattern:

  1. two models produce independent answers
  2. a stronger model or rule layer compares and merges

Best for:

  • architecture options
  • planning alternatives
  • ambiguous recommendations
8. Consensus or verifier workflow

Use when:

  • correctness matters enough to justify extra cost
  • false confidence is dangerous

Pattern:

  1. primary model produces answer
  2. verifier model checks logic, calculations, or policy fit
  3. disagreements trigger escalation or explicit uncertainty

Best for:

  • high-risk reasoning
  • important financial outputs
  • compliance-sensitive content
  • high-value technical decisions

Cost-control strategy

Use these strategies to keep cost high-value.

Default strategy
  • start cheap when safe
  • escalate only on signals of failure risk
  • avoid premium for routine tasks
  • reuse extracted structure instead of repeating full-context calls
Escalation triggers

Escalate to a stronger model or multi-step workflow when any of these appear:

  • multiple dependent reasoning steps
  • ambiguous user intent with multiple plausible interpretations
  • repeated self-contradiction in draft output
  • failure to follow structure or constraints
  • long context with subtle dependencies
  • code correctness matters beyond surface syntax
  • exactness-critical math or finance output
  • high-risk domain or high business impact
Show full SKILL.md (1,007 more words)Show less
De-escalation triggers

Use a cheaper path when:

  • the task is mostly formatting or rewriting
  • the answer can be approximate
  • the task is repetitive and pattern-based
  • first-pass triage is enough
  • premium capabilities would not materially improve the outcome

Complex calculation policy

When the request includes complex calculations or formal reasoning:

  1. Separate interpretation from computation.
  2. Use the model to parse the problem and define the method.
  3. Use a deterministic tool or calculational path when available.
  4. Ask a verifier layer to check assumptions, formulas, units, and edge cases for high-impact outputs.
  5. Present the final answer with explicit assumptions and, when relevant, step order.

Never use a fluent but non-verified freeform model answer as the final authority for exact numeric work when a deterministic path exists.

Long-context policy

When the request includes large context:

  • first extract relevant segments, summaries, or structured facts
  • reduce duplication
  • preserve citations or pointers when possible
  • synthesize only after reduction
  • use premium synthesis only if the reduced problem still demands it

Output format

Return exactly this structure:

Routing Decision: <single-economy | single-balanced | single-premium | tool-assisted | staged-pipeline | draft-and-review | parallel-comparison | consensus-check>

Primary Reason: <one-sentence explanation of why this route is the best cost-effective choice>

Task Profile:

  • taskType: <task category>
  • complexity: <low | medium | high | very high>
  • exactness: <low | medium | high | critical>
  • risk: <low | medium | high>
  • latency: <urgent | normal | relaxed>
  • budget: <minimize-cost | balanced | quality-first>
  • contextLoad: <short | moderate | long | extreme>
  • modality: <text | multimodal>

Recommended Execution Plan:

  1. <step 1>
  2. <step 2>
  3. <step 3>
  4. <step 4 if needed>

Model Role Assignment:

  • planner: <economy | balanced | premium | not-needed>
  • generator: <economy | balanced | premium | tool-assisted>
  • verifier: <none | balanced | premium | tool>

Why This Is Cost-Effective:

  • <reason 1>
  • <reason 2>
  • <reason 3>

Escalation Rule: <when to move to a stronger model or more expensive workflow>

Fallback Rule: <what to do if the chosen path fails or produces uncertainty>

Examples

Example 1: Simple rewrite

Request: "Rewrite this email to sound more professional."

Output: Routing Decision: single-economy

Primary Reason: This is a low-risk rewriting task where a low-cost model should meet the quality bar.

Task Profile:

  • taskType: simple generation
  • complexity: low
  • exactness: low
  • risk: low
  • latency: urgent
  • budget: minimize-cost
  • contextLoad: short
  • modality: text

Recommended Execution Plan:

  1. Use an economy model to rewrite the email.
  2. Preserve intent and tone constraints.
  3. Return one polished version and one shorter variant.

Model Role Assignment:

  • planner: not-needed
  • generator: economy
  • verifier: none

Why This Is Cost-Effective:

  • The task is structurally simple.
  • Premium reasoning is unnecessary.
  • Fast low-cost generation is sufficient.

Escalation Rule: Escalate only if the user requests highly nuanced persuasion, legal sensitivity, or multiple tone variants.

Fallback Rule: If the first rewrite is too generic, retry once with stricter tone instructions using a balanced model.

Example 2: Complex system design

Request: "Design an AI agent architecture for long-form writing with memory, retrieval, conflict resolution, and chapter planning."

Output: Routing Decision: single-premium

Primary Reason: This task requires deep reasoning, abstraction, and multi-part architecture tradeoff analysis.

Task Profile:

  • taskType: deep reasoning
  • complexity: high
  • exactness: medium
  • risk: medium
  • latency: normal
  • budget: balanced
  • contextLoad: moderate
  • modality: text

Recommended Execution Plan:

  1. Use a premium model to identify major architecture components and constraints.
  2. Structure the answer into memory layers, retrieval flow, consistency controls, and orchestration.
  3. Include tradeoffs and failure modes.
  4. Add implementation recommendations only after the design is coherent.

Model Role Assignment:

  • planner: premium
  • generator: premium
  • verifier: none

Why This Is Cost-Effective:

  • A weaker model may produce shallow or inconsistent architecture.
  • One premium pass is cheaper than multiple failed lower-tier retries.
  • The task value justifies stronger reasoning.

Escalation Rule: Escalate to parallel-comparison only if multiple competing architecture strategies are explicitly required.

Fallback Rule: If the first answer is too generic, rerun with a staged outline-first then full-design workflow.

Example 3: Financial calculation with policy impact

Request: "Calculate the projected monthly burn and runway under three hiring scenarios and explain the assumptions."

Output: Routing Decision: tool-assisted

Primary Reason: This request depends on exact calculations, so deterministic computation is safer and more cost-effective than pure model reasoning.

Task Profile:

  • taskType: exact calculation or formal logic
  • complexity: high
  • exactness: critical
  • risk: high
  • latency: normal
  • budget: balanced
  • contextLoad: moderate
  • modality: text

Recommended Execution Plan:

  1. Use a balanced model to extract the variables, hiring scenarios, and required formulas.
  2. Use a deterministic calculator or spreadsheet-capable path to compute burn and runway.
  3. Run a verification check on formulas, assumptions, and units.
  4. Use the model to present the results clearly with assumptions and scenario breakdowns.

Model Role Assignment:

  • planner: balanced
  • generator: tool-assisted
  • verifier: tool

Why This Is Cost-Effective:

  • Exact arithmetic should not depend on freeform model reasoning.
  • Deterministic calculation reduces error risk.
  • A smaller model can still explain the results after computation.

Escalation Rule: Escalate to premium review only if the scenario assumptions are ambiguous or strategically sensitive.

Fallback Rule: If tool-based computation is unavailable, return assumptions explicitly and mark the result as unverified instead of presenting certainty.

Example 4: Multi-document review

Request: "Compare these five candidate proposals and recommend the best one based on cost, timeline, technical fit, and risk."

Output: Routing Decision: staged-pipeline

Primary Reason: The request is large and decomposable, so staged extraction and synthesis is more cost-effective than sending everything directly to a premium model.

Task Profile:

  • taskType: long-context synthesis
  • complexity: high
  • exactness: medium
  • risk: medium
  • latency: relaxed
  • budget: balanced
  • contextLoad: long
  • modality: text

Recommended Execution Plan:

  1. Use an economy or balanced model to extract structured facts from each proposal.
  2. Normalize the proposals into a common comparison table.
  3. Use a stronger model to synthesize tradeoffs and recommend the best option.
  4. Add a brief verifier pass if the recommendation is high stakes.

Model Role Assignment:

  • planner: balanced
  • generator: staged-pipeline
  • verifier: balanced

Why This Is Cost-Effective:

  • Cheap extraction lowers total token cost.
  • Structured normalization improves synthesis quality.
  • Premium reasoning is reserved for the part that truly needs it.

Escalation Rule: Escalate to consensus-check if the recommendation will drive a major decision or if proposal differences are subtle.

Fallback Rule: If extraction quality is poor, rerun the extraction stage with a stronger model before recomputing the final recommendation.

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/model-routing-orchestrator of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • _meta.json

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Production Model Router next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Production Model Router compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Production Model Router this skillLeoYeAI/openclaw-master-skills2.2k—~4.9kAutomated safety check: PassMIT
Shogun Bloom Configyohey-w/multi-agent-shogun1.4k—~3.1kAutomated safety check: PassMIT
Codemie Analyticscodemie-ai/codemie-code294—~7.5kAutomated safety check: PassApache-2.0
Model Routernidhi-singh02/agent-router112—~1.2kAutomated safety check: PassMIT
Codex Model Routing Teamzjp1997720/codex-model-routing-team158—~736Automated safety check: PassMIT
Add Modelget-convex/convex-evals130—~1.5kAutomated safety check: NotesApache-2.0

Similar skills

  • Shogun Bloom Config

    yohey-w/multi-agent-shogun

    Interactive wizard: guided questions with multiple-choice options about subscriptions, then outputs a ready-to-paste capabilitytiers YAML + fixed agent model assignments.

    1.4k GitHub stars~3.1k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Codemie Analytics

    codemie-ai/codemie-code

    CodeMie Analytics expert — use this skill whenever the user asks about CodeMie usage data, AI adoption metrics, user leaderboards, CLI insights, spending, LiteLLM costs, token usage, or wants to…

    294 GitHub stars~7.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Model Router

    nidhi-singh02/agent-router

    A skill your agent uses when the user asks to pick a model, subscription, or reasoning effort, or to run router status, usage refresh, or resume a router session.

    112 GitHub stars~1.2k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Codex Model Routing Team

    zjp1997720/codex-model-routing-team

    在 Codex App 中为复杂、可并行的知识工作或编程任务自动创建多个可指定模型与推理强度的后台任务,由主 Agent 负责规划、分工、集成和验收。用于多来源调研、多章节内容、复杂 Skill/PPT、跨模块开发、独立验证或 2 个以上互不依赖工作流;也用于用户明确要求模型路由、后台 Worker、Agents Team…

    158 GitHub stars~736 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    get-convex/convex-evals

    Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

    130 GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • OmniRoute CLI Evals

    diegosouzapw/OmniRoute

    Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

    75k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from LeoYeAI/openclaw-master-skills

All 1,235 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Production Model Router

What does Production Model Router do?

Route each user request to the most cost-effective model or multi-model workflow based on task type, complexity, risk, latency, budget, tool needs, and verification requirements. Production Model Router is an agent skill from LeoYeAI/openclaw-master-skills. Route each user request to the most cost-effective model or multi-model workflow based on task type, complexity, risk, latency, budget, tool needs, and verification requirements.

When should I use Production Model Router?

Production Model Router fits situations like: tasks that involve Model routing and gateways.

How do I install Production Model Router in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill production-model-router -a claude-code`. Or copy the skill folder (skills/model-routing-orchestrator in LeoYeAI/openclaw-master-skills) into .claude/skills/production-model-router in your project. Claude Code loads it when a task matches its description.

How do I install Production Model Router in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill production-model-router -a codex`. Or copy the skill folder (skills/model-routing-orchestrator in LeoYeAI/openclaw-master-skills) into .agents/skills/production-model-router in your project. Codex loads it when a task matches its description.

Can I use Production Model Router in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill production-model-router -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/production-model-router, .gemini/skills/production-model-router, .github/skills/production-model-router and .opencode/skills/production-model-router in your project.

What does Production Model Router need to run?

SKILL.md names no scripts, command-line tools or credentials: Production Model Router is instructions for the agent only.

Does Production Model Router access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Production Model Router safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Production Model Router use?

Production Model Router is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Production Model Router use?

About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Production Model Router?

Skills that share tags, products or a category with Production Model Router: Shogun Bloom Config (yohey-w/multi-agent-shogun, 1.4k stars), Codemie Analytics (codemie-ai/codemie-code, 294 stars), Model Router (nidhi-singh02/agent-router, 112 stars) and Codex Model Routing Team (zjp1997720/codex-model-routing-team, 158 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Production Model Router?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.