Agent skill

Model Hierarchy

by zscole in zscole/model-hierarchy-skill

Cost-optimize AI agent operations by routing tasks to appropriate models based on complexity.

MITAuto-check passedAI & LLM Engineering

Install Model Hierarchy

skills CLI
$ npx skills add zscole/model-hierarchy-skill --skill model-hierarchy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zscole/model-hierarchy-skill model-hierarchy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-hierarchy
GitHub stars
346
Token cost
~2.3k tokens
SKILL.md length
838 words
Files
10
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Cost-optimize AI agent operations by routing tasks to appropriate models based on complexity.

  • Works in 3 steps: Default to Tier 2 for interactive work → Suggest downgrade when doing routine… → Request upgrade when stuck: "This needs…
  • Deciding which model to use for a task
  • SKILL.md covers Core Principle, Model Tiers, Task Classification and Decision Algorithm, plus 6 more sections
  • Runs Python scripts from its folder

What it does

Model Hierarchy is an agent skill from zscole/model-hierarchy-skill. Cost-optimize AI agent operations by routing tasks to appropriate models based on complexity. Use this skill when: (1) deciding which model to use for a task, (2) spawning sub-agents, (3) considering cost efficiency, (4) the current model feels like overkill for the task. Triggers: "model routing", "cost optimization", "which model", "too expensive", "spawn agent".

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files (for example `README.md`, `examples/claude-code.md` and `examples/openclaw.md`).

It sits in AI & LLM Engineering, covering Model routing and gateways and Subagents. It works with Zhipu GLM, Kimi and OpenAI. The repository describes itself as: OpenClaw skill for cost-optimized model routing based on task complexity. The licence is MIT.

When your agent uses it

  • Deciding which model to use for a task
  • Spawning sub-agents
  • Considering cost efficiency
  • The current model feels like overkill for the task

Example prompts

  • “model routing”
  • “cost optimization”
  • “which model”
  • “/model-hierarchy”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Default to Tier 2 for interactive work
  2. Suggest downgrade when doing routine work: "This is routine - I can handle this on a cheaper model or spawn a sub-agent."
  3. Request upgrade when stuck: "This needs more reasoning power. Switching to [premium model]."

What it can do on your machine

Read from SKILL.md and the folder at commit 9095f83. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Hierarchy loads about 2.3k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 838 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from zscole/model-hierarchy-skill at commit 9095f83, republished under its MIT licence (© zscole). 838 words, ~2,281 tokens.

Download SKILL.mdSave it as .claude/skills/model-hierarchy/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
model-hierarchy
description
Cost-optimize AI agent operations by routing tasks to appropriate models based on complexity. Use this skill when: (1) deciding which model to use for a task, (2) spawning sub-agents, (3) considering cost efficiency, (4) the current model feels like overkill for the task. Triggers: "model routing", "cost optimization", "which model", "too expensive", "spawn agent".

Model Hierarchy

Route tasks to the cheapest model that can handle them. Most agent work is routine.

Core Principle

80% of agent tasks are janitorial. File reads, status checks, formatting, simple Q&A. These don't need expensive models. Reserve premium models for problems that actually require deep reasoning.

Model Tiers

Tier 1: Cheap ($0.10-0.50/M tokens)
ModelInputOutputBest For
DeepSeek V3$0.14$0.28General routine work
GPT-4o-mini$0.15$0.60Quick responses
Claude Haiku$0.25$1.25Fast tool use
Gemini Flash$0.075$0.30High volume
GLM 5 (Zhipu)(OpenRouter Z.AI)(OpenRouter Z.AI)Routine + moderate text; 200K context; text-only — do not use for image/vision
Kimi K2.5 (Moonshot)$0.45$2.25Routine + moderate; 262K context; multimodal (text + image + video)

Text-only models (e.g. GLM 5): Do not use for any task that requires image input or vision — no photo analysis, screenshots, image-generation tools, or document/chart vision. Route to a vision-capable model (e.g. Kimi K2.5, GPT-4o, Gemini, Claude with vision, GLM-4.5V/4.6V).

Vision-capable Tier 1/2 (e.g. Kimi K2.5): Use for routine or moderate tasks that may involve images — screenshots, photo analysis, docs, image-generation orchestration — without moving to premium vision models.

Tier 2: Mid ($1-5/M tokens)
ModelInputOutputBest For
Claude Sonnet$3.00$15.00Balanced performance
GPT-4o$2.50$10.00Multimodal tasks
Gemini Pro$1.25$5.00Long context
Tier 3: Premium ($10-75/M tokens)
ModelInputOutputBest For
Claude Opus$15.00$75.00Complex reasoning
GPT-4.5$75.00$150.00Frontier tasks
o1$15.00$60.00Multi-step reasoning
o3-mini$1.10$4.40Reasoning on budget

Prices as of Feb 2026. Check provider docs for current rates.

Task Classification

Before executing any task, classify it:

ROUTINE → Use Tier 1

Requires image/vision → Do not assign to text-only models (GLM 5, etc.). Use a vision-capable model from Tier 1/2 or 3 (e.g. Kimi K2.5, GPT-4o, Gemini, Claude, GLM-4.5V).

Characteristics:

  • Single-step operations
  • Clear, unambiguous instructions
  • No judgment required
  • Deterministic output expected

Examples:

  • File read/write operations
  • Status checks and health monitoring
  • Simple lookups (time, weather, definitions)
  • Formatting and restructuring text
  • List operations (filter, sort, transform)
  • API calls with known parameters
  • Heartbeat and cron tasks
  • URL fetching and basic parsing
MODERATE → Use Tier 2

Characteristics:

  • Multi-step but well-defined
  • Some synthesis required
  • Standard patterns apply
  • Quality matters but isn't critical

Examples:

  • Code generation (standard patterns)
  • Summarization and synthesis
  • Draft writing (emails, docs, messages)
  • Data analysis and transformation
  • Multi-file operations
  • Tool orchestration
  • Code review (non-security)
  • Search and research tasks
COMPLEX → Use Tier 3

Characteristics:

  • Novel problem solving required
  • Multiple valid approaches
  • Nuanced judgment calls
  • High stakes or irreversible
  • Previous attempts failed

Examples:

  • Multi-step debugging
  • Architecture and design decisions
  • Security-sensitive code review
  • Tasks where cheaper model already failed
  • Ambiguous requirements needing interpretation
  • Long-context reasoning (>50K tokens)
  • Creative work requiring originality
  • Adversarial or edge-case handling

Decision Algorithm

function selectModel(task):
    # Rule 1: Vision override (Tier 1/2 includes text-only models)
    if task.requiresImageInput or task.requiresVision:
        return VISION_CAPABLE_MODEL  # e.g. Kimi K2.5, GPT-4o, Gemini, Claude; do not use GLM 5 or other text-only
    
    # Rule 2: Escalation override
    if task.previousAttemptFailed:
        return nextTierUp(task.previousModel)
    
    # Rule 3: Explicit complexity signals
    if task.hasSignal("debug", "architect", "design", "security"):
        return TIER_3
    
    if task.hasSignal("write", "code", "summarize", "analyze"):
        return TIER_2
    
    # Rule 4: Default classification
    complexity = classifyTask(task)
    
    if complexity == ROUTINE:
        return TIER_1
    elif complexity == MODERATE:
        return TIER_2
    else:
        return TIER_3

Behavioral Rules

For Main Session
  1. Default to Tier 2 for interactive work
  2. Suggest downgrade when doing routine work: "This is routine - I can handle this on a cheaper model or spawn a sub-agent."
  3. Request upgrade when stuck: "This needs more reasoning power. Switching to [premium model]."
Show full SKILL.md (345 more words)Show less
For Sub-Agents
  1. Default to Tier 1 unless task is clearly moderate+
  2. Batch similar tasks to amortize overhead
  3. Report failures back to parent for escalation
For Automated Tasks
  1. Heartbeats/monitoring → Always Tier 1
  2. Scheduled reports → Tier 1 or 2 based on complexity
  3. Alert responses → Start Tier 2, escalate if needed

Communication Patterns

When suggesting model changes, use clear language:

Downgrade suggestion:

"This looks like routine file work. Want me to spawn a sub-agent on DeepSeek for this? Same result, fraction of the cost."

Upgrade request:

"I'm hitting the limits of what I can figure out here. This needs Opus-level reasoning. Switching up."

Explaining hierarchy:

"I'm running the heavy analysis on Sonnet while sub-agents fetch the data on DeepSeek. Keeps costs down without sacrificing quality where it matters."

Cost Impact

Assuming 100K tokens/day average usage:

StrategyMonthly CostNotes
Pure Opus~$225Maximum capability, maximum spend
Pure Sonnet~$45Good default for most work
Pure DeepSeek~$8Cheap but limited on hard problems
Hierarchy (80/15/5)~$19Best of all worlds

The 80/15/5 split:

  • 80% routine tasks on Tier 1 (~$6)
  • 15% moderate tasks on Tier 2 (~$7)
  • 5% complex tasks on Tier 3 (~$6)

Result: 10x cost reduction vs pure premium, with equivalent quality on complex tasks.

Integration Examples

OpenClaw
yaml
# config.yml - set default model
model: anthropic/claude-sonnet-4

# In session, switch models
/model opus  # upgrade for complex task
/model deepseek  # downgrade for routine

# Spawn sub-agent on cheap model
sessions_spawn:
  task: "Fetch and parse these 50 URLs"
  model: deepseek

OpenRouter (Tier 1 with vision or text-only):

yaml
# Tier 1 with vision — Kimi K2.5 (multimodal)
model: openrouter/moonshotai/kimi-k2.5
# Heartbeats, cron, image-involving tasks: K2.5 handles text and vision.

# Tier 1 text-only — GLM 5 (no vision)
# model: openrouter/z-ai/glm-5  # exact ID TBD on OpenRouter Z.AI
# Routine text-only only; for image tasks use Kimi K2.5 or another vision-capable model.
Claude Code
# In CLAUDE.md or project instructions
When spawning background agents, use claude-3-haiku for:
- File operations
- Simple searches  
- Status checks

Reserve claude-sonnet-4 for:
- Code generation
- Analysis tasks
General Agent Systems
python
def get_model_for_task(task_description: str) -> str:
    routine_signals = ['read', 'fetch', 'check', 'list', 'format', 'status']
    complex_signals = ['debug', 'architect', 'design', 'security', 'why']
    
    desc_lower = task_description.lower()
    
    if any(signal in desc_lower for signal in complex_signals):
        return "claude-opus-4"
    elif any(signal in desc_lower for signal in routine_signals):
        return "deepseek-v3"
    else:
        return "claude-sonnet-4"

Anti-Patterns

DON'T:

  • Run heartbeats on Opus
  • Use premium models for file I/O
  • Keep expensive model when task is clearly routine
  • Spawn sub-agents on premium models by default
  • Use GLM 5 (or any text-only Tier 1/2 model) for image/vision tasks — e.g. photo analysis, screenshot understanding, image-generation skills, or any tool that takes image input

DO:

  • Start mid-tier, adjust based on task
  • Spawn helpers on cheapest viable model
  • Escalate explicitly when stuck
  • Track cost per task type to optimize further

Extending This Skill

To customize for your use case:

  1. Adjust tier definitions based on your provider/budget
  2. Add domain-specific signals to classification rules
  3. Track actual complexity vs predicted to improve heuristics
  4. Set budget alerts to catch runaway premium usage

© zscole, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files in the repository root of zscole/model-hierarchy-skill.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • examples/claude-code.md
  • examples/openclaw.md
  • pyproject.toml
  • tests/conftest.py
  • tests/scenarios.json
  • tests/test_classification.py

Open the folder on GitHubat commit 9095f83

Compare with similar skills

Model Hierarchy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Hierarchy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Hierarchy this skillzscole/model-hierarchy-skill346—~2.3kAutomated safety check: PassMIT
Provider Integrationhex/claude-council857—~635Automated safety check: PassMIT
LLM Council on Fireworks AIdair-ai/dair-academy-plugins614—~5kAutomated safety check: NotesMIT
Model Routersundial-org/awesome-openclaw-skills663—~2.2kAutomated safety check: PassNone
Proxy Mode ReferenceMadAppGang/claude-code285—~1.3kAutomated safety check: PassMIT
Claudish UsageMadAppGang/claudish1k—~9kAutomated safety check: PassNone

Similar skills

  • Provider Integration

    hex/claude-council

    Adds new AI providers to claude-council, configures provider API settings, troubleshoots provider connections, and documents the provider script interface.

    857 GitHub stars~635 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Council on Fireworks AI

    dair-ai/dair-academy-plugins

    Has several open-weight models answer a question, rank each other's anonymized answers, then lets a chairman model write the final response through Fireworks AI.

    614 GitHub stars~5k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Model Router

    sundial-org/awesome-openclaw-skills

    A comprehensive AI model routing system that automatically selects the optimal model for any task.

    663 GitHub stars~2.2k tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Proxy Mode Reference

    MadAppGang/claude-code

    Reference guide for using external AI models via claudish CLI.

    285 GitHub stars~1.3k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Claudish Usage

    MadAppGang/claudish

    CRITICAL - Guide for using Claudish CLI ONLY through sub-agents to run Claude Code with any AI model (OpenRouter, Gemini, OpenAI, local models).

    1k GitHub stars~9k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Model-routing orchestration for any expensive frontier model (Fable, Opus, GPT-5.x) - the main model keeps judgment and Q&A review, mechanical subagent dispatches carry an explicit cheaper model…

    115 GitHub stars~1.9k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed

Questions about Model Hierarchy

What does Model Hierarchy do?

Cost-optimize AI agent operations by routing tasks to appropriate models based on complexity. Model Hierarchy is an agent skill from zscole/model-hierarchy-skill. Cost-optimize AI agent operations by routing tasks to appropriate models based on complexity.

When should I use Model Hierarchy?

Model Hierarchy fits situations like: deciding which model to use for a task; spawning sub-agents; considering cost efficiency; the current model feels like overkill for the task.

How do I install Model Hierarchy in Claude Code?

Run `npx skills add zscole/model-hierarchy-skill --skill model-hierarchy -a claude-code`. Or copy the skill folder (the zscole/model-hierarchy-skill repository) into .claude/skills/model-hierarchy in your project. Claude Code loads it when a task matches its description.

How do I install Model Hierarchy in Codex?

Run `npx skills add zscole/model-hierarchy-skill --skill model-hierarchy -a codex`. Or copy the skill folder (the zscole/model-hierarchy-skill repository) into .agents/skills/model-hierarchy in your project. Codex loads it when a task matches its description.

Can I use Model Hierarchy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zscole/model-hierarchy-skill --skill model-hierarchy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-hierarchy, .gemini/skills/model-hierarchy, .github/skills/model-hierarchy and .opencode/skills/model-hierarchy in your project.

What does Model Hierarchy need to run?

Going by SKILL.md and its folder, Model Hierarchy needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Model Hierarchy access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model Hierarchy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Hierarchy use?

Model Hierarchy is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Hierarchy use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Hierarchy?

Skills that share tags, products or a category with Model Hierarchy: Provider Integration (hex/claude-council, 857 stars), LLM Council on Fireworks AI (dair-ai/dair-academy-plugins, 614 stars), Model Router (sundial-org/awesome-openclaw-skills, 663 stars) and Proxy Mode Reference (MadAppGang/claude-code, 285 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Hierarchy?

zscole (a GitHub user) maintains it in zscole/model-hierarchy-skill, which has 346 GitHub stars. The repository was last updated on February 16, 2026.

Source: zscole/model-hierarchy-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.