Agent skill

Model Routing

by alinaqi in alinaqi/maggy

9-tier model routing system with cascading classifier fallback and result auto-evaluation

MITAuto-check passedAI & LLM Engineering

Install Model Routing

skills CLI
$ npx skills add alinaqi/maggy --skill model-routing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alinaqi/maggy model-routing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alinaqi/maggy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-routing .claude/skills/model-routing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-routing
GitHub stars
707
Token cost
~1.5k tokens
SKILL.md length
472 words
Files
1
Skills in repo
71
Repo updated
First seen
Licence
MIT

At a glance

9-tier model routing system with cascading classifier fallback and result auto-evaluation

  • Works in 3 steps: Which model should handle this? — 9-tier… → Is the classifier itself working? —… → Can we verify the result? — Tool-level…
  • Tasks that involve Model routing and gateways
  • SKILL.md covers How Routing Decisions Are Made, 9-Tier Routing Table, Delegation Commands and Delegation Script Contract, plus 5 more sections
  • Calls gemini, claude and codex; needs DEEPSEEK_API_KEY and GEMINI_API_KEY

What it does

Model Routing is an agent skill from alinaqi/maggy. 9-tier model routing system with cascading classifier fallback and result auto-evaluation

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model routing and gateways. It works with Kimi, DeepSeek and Qwen. The repository describes itself as: What started as an opinionated Claude Code setup kit is now an autonomous AI engineering command center. The licence is MIT.

When your agent uses it

  • Tasks that involve Model routing and gateways

Example prompts

  • “/model-routing”

Requirements

  • A credential in DEEPSEEK_API_KEY
  • A credential in GEMINI_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Which model should handle this? — 9-tier cost/complexity classification
  2. Is the classifier itself working? — Cascading fallback (qwen3 → kimi → deepseek → cache)
  3. Can we verify the result? — Tool-level fallback + auto-evaluation

What it can do on your machine

Read from SKILL.md and the folder at commit 72a456e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gemini
    • claude
    • codex
    • ollama

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DEEPSEEK_API_KEY
    • GEMINI_API_KEY
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Routing loads about 1.5k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 472 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from alinaqi/maggy at commit 72a456e, republished under its MIT licence (© alinaqi). 472 words, ~1,508 tokens.

Download SKILL.mdSave it as .claude/skills/model-routing/SKILL.md (or your agent's skills folder).
name
model-routing
description
9-tier model routing system with cascading classifier fallback and result auto-evaluation
when-to-use
When configuring or debugging how prompts are classified and routed to the cheapest capable model
user-invocable
false
effort
medium

Model Routing System

How Routing Decisions Are Made

Every user prompt goes through a 9-tier classification pipeline before any AI model processes it. The system answers three questions:

  1. Which model should handle this? — 9-tier cost/complexity classification
  2. Is the classifier itself working? — Cascading fallback (qwen3 → kimi → deepseek → cache)
  3. Can we verify the result? — Tool-level fallback + auto-evaluation
The Pipeline
User types prompt
    ↓
UserPromptSubmit hook fires (~/.claude/hooks/route-task-hook)
    ↓
Classifier: qwen3 (local, free) classifies into tier
    ↓  (fails?)
Classifier: kimi (local, free) retries
    ↓  (fails?)
Classifier: deepseek-flash (~$0.0001) retries
    ↓  (fails?)
Classifier: cached tier from last success
    ↓
Hook injects routing decision into Claude's context
    ↓
Claude delegates to the right model or handles directly

9-Tier Routing Table

TierModelInput (per M)Output (per M)Handles
0Qwen3 (local)$0$0grep, find, shell, syntax, log reading
1Gemini 2.5 Flash-Lite$0.10$0.40Bulk extraction, classification, CIG pipelines
2DeepSeek V4 Flash$0.14$0.28Simple code, CRUD, test writing, small fixes
3DeepSeek V4 Pro$0.44$0.87Multi-file features, refactors, debugging (~80% of work)
4Gemini 2.5 Flash$0.15$0.60Multimodal (images, video, audio), brand analysis
5Kimi K2.6$0.60$2.50Code review, commit messages, diff summaries
6Gemini 3.1 Pro + Search$1.25$10.00Deep research, Google grounding, 2M context
7CodexvariesvariesBulk generation, code review
8Claude Sonnet/Opus$3-5$15-25Architecture, security, quality-critical

Delegation Commands

When the hook says "delegate to X", run the matching command and return its output:

bash
# Tier 0 — Qwen3
~/bin/qwen3 "prompt"

# Tier 1 — Gemini Flash-Lite
~/bin/gemini --flash-lite "prompt"

# Tier 2 — DeepSeek Flash
~/bin/deepseek --flash "prompt"

# Tier 3 — DeepSeek Pro
~/bin/deepseek --pro "prompt"

# Tier 4 — Gemini Flash
~/bin/gemini --flash "prompt"

# Tier 5 — Kimi
~/bin/kimi --quiet -p "prompt"

# Tier 6 — Gemini Pro Search
~/bin/gemini --pro-search "prompt"

# Tier 7 — Codex
codex exec "prompt"

# Tier 8 — Claude
# Handle directly (no delegation)

Delegation Script Contract

Every ~/bin/ script follows the same pattern:

  1. Accepts prompt as argument: script "what is 2+2"
  2. Model flags: --flash, --pro, --flash-lite, --pro-search
  3. Quiet mode: --quiet (where applicable)
  4. Output: writes response to stdout, errors to stderr
  5. Exit codes: 0 on success, non-zero on failure
Available Scripts
~/bin/
├── qwen3       # Shell: curl to local Ollama API
├── kimi        # Shell: execs Kimi CLI binary
├── deepseek    # Python: httpx to DeepSeek Anthropic-compat API
├── gemini      # Python: httpx to Gemini OpenAI-compat API
├── research    # Python: multi-backend research with auto-evaluation
└── route-task  # Shell: qwen3-powered task classification

Classifier Fallback Chain

The classifier itself can fail. When it does, cascading fallback kicks in:

LevelClassifierCostThreshold
1qwen3 (Ollama)$02s connect, 8s classify
2kimi CLI$0Local process
3deepseek-flash~$0.0001API call
4Cached tier$0From ~/.claude/routing-cache.json

The cache (~/.claude/routing-cache.json) saves the last successful tier and timestamp. After compaction, when Ollama may be briefly unreachable, the cache ensures routing continues without dropping to CLAUDE by default.

Show full SKILL.md (158 more words)Show less

Tool Fallback Protocol

When Claude's built-in tools fail, external backends take over:

Failed ToolFallback 1Fallback 2
WebSearch / WebFetch~/bin/research "query"~/bin/deepseek --pro "query"
Read / file accesscat via Bash—
Grepgrep -r via Bash—
Research Tool (~/bin/research)

Multi-backend research with auto-evaluation:

  • Tries deepseek-flash → deepseek-pro in sequence
  • Scores results 0-10 on content quality, structure, length
  • Auto-adjusts preferred backend based on evaluation scores
  • View stats: ~/bin/research --eval
  • Score log: ~/.claude/research-eval.jsonl

Maggy Integration

Maggy's model_router.py mirrors the same 9-tier structure in DEFAULT_TIERS. The PiAdapter uses the same delegation scripts for execution. Task type overrides in routing_rules_defaults.py ensure:

  • research, competitor → Gemini Pro Search (Google grounding)
  • bulk → Gemini Flash-Lite (cheapest)
  • security, architecture, planning → Claude (quality-critical)
  • docs, tests → DeepSeek Pro (cost-efficient)
  • review → Claude (security + architecture depth)

Environment

bash
# Required for delegation scripts (in ~/.zshrc)
export DEEPSEEK_API_KEY="sk-..."
export GEMINI_API_KEY="..."       # For gemini delegator
export OPENAI_API_KEY="sk-..."    # For codex CLI

# Ollama must be running locally for qwen3
ollama serve  # or launch at startup

Observability

  • Routing log: ~/.claude/routing-log.jsonl — every classification with tier, classifier used, tokens saved
  • Routing cache: ~/.claude/routing-cache.json — last tier for post-compact recovery
  • Research eval: ~/.claude/research-eval.jsonl — per-query backend scoring
  • Maggy routing heatmap: Dashboard → Models tab → per-model reward scores

© alinaqi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/model-routing of alinaqi/maggy.

Open the folder on GitHubat commit 72a456e

Compare with similar skills

Model Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Routing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Routing this skillalinaqi/maggy707—~1.5kAutomated safety check: PassMIT
LLM Council on Fireworks AIdair-ai/dair-academy-plugins614—~5kAutomated safety check: NotesMIT
Claude Maintain ModelsKiln-AI/Kiln5.2k—~15kAutomated safety check: NotesCustom licence
LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~3.9kAutomated safety check: PassNone
Update Ollama Cloud Modelsheypinchy/pinchy182—~3.9kAutomated safety check: NotesAGPL-3.0
OmniRoute Chat CLIdiegosouzapw/OmniRoute75k—~345Automated safety check: PassMIT

Similar skills

  • LLM Council on Fireworks AI

    dair-ai/dair-academy-plugins

    Has several open-weight models answer a question, rank each other's anonymized answers, then lets a chairman model write the final response through Fireworks AI.

    614 GitHub stars~5k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Add new AI models to Kiln's mlmodellist.py and produce a Discord announcement.

    5.2k GitHub stars~15k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    938 GitHub stars~3.9k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.

    182 GitHub stars~3.9k tokensUpdated 19 days ago
    AI & LLM EngineeringAuto-check: notes
  • OmniRoute Chat CLI

    diegosouzapw/OmniRoute

    Sends chat completions, streams responses, and opens an interactive REPL against any OmniRoute-routed model provider.

    75k GitHub stars~345 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Nowait Reasoning Optimizer

    davila7/claude-code-templates

    Implements the NOWAIT technique for efficient reasoning in R1-style LLMs.

    33k GitHub starsUsed in 2 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed

More from alinaqi/maggy

All 71 skills in this repo
  • Aeo Optimization

    alinaqi/maggy

    AI Engine Optimization - semantic triples, page templates, content clusters for AI citations

    707 GitHub stars~3.7k tokensUpdated 16 days ago
    Auto-check passed
  • Agent Teams

    alinaqi/maggy

    Claude Code Agent Teams - default team-based development with strict TDD pipeline enforcement

    707 GitHub stars~5k tokensUpdated 16 days ago
    Auto-check: notes
  • AI Models

    alinaqi/maggy

    Latest AI models reference - Claude, OpenAI, Gemini, Eleven Labs, Replicate

    707 GitHub stars~4.1k tokensUpdated 16 days ago
    Auto-check passed
  • Android Java

    alinaqi/maggy

    Android Java development with MVVM, ViewBinding, and Espresso testing

    707 GitHub stars~3.9k tokensUpdated 16 days ago
    Auto-check: notes
  • Android Kotlin

    alinaqi/maggy

    Android Kotlin development with Coroutines, Jetpack Compose, Hilt, and MockK testing

    707 GitHub stars~3k tokensUpdated 16 days ago
    Auto-check passed
  • Autonomous Testing

    alinaqi/maggy

    AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type

    707 GitHub stars~1.1k tokensUpdated 16 days ago
    Auto-check passed

Questions about Model Routing

What does Model Routing do?

9-tier model routing system with cascading classifier fallback and result auto-evaluation. Model Routing is an agent skill from alinaqi/maggy.

When should I use Model Routing?

Model Routing fits situations like: tasks that involve Model routing and gateways.

How do I install Model Routing in Claude Code?

Run `npx skills add alinaqi/maggy --skill model-routing -a claude-code`. Or copy the skill folder (skills/model-routing in alinaqi/maggy) into .claude/skills/model-routing in your project. Claude Code loads it when a task matches its description.

How do I install Model Routing in Codex?

Run `npx skills add alinaqi/maggy --skill model-routing -a codex`. Or copy the skill folder (skills/model-routing in alinaqi/maggy) into .agents/skills/model-routing in your project. Codex loads it when a task matches its description.

Can I use Model Routing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alinaqi/maggy --skill model-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-routing, .gemini/skills/model-routing, .github/skills/model-routing and .opencode/skills/model-routing in your project.

What does Model Routing need to run?

Going by SKILL.md and its folder, Model Routing needs the command-line tools its instructions call (gemini, claude, codex and ollama) and credentials named DEEPSEEK_API_KEY, GEMINI_API_KEY and OPENAI_API_KEY. Our summary lists: A credential in DEEPSEEK_API_KEY; A credential in GEMINI_API_KEY.

Does Model Routing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model Routing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Routing use?

Model Routing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Routing use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Routing?

Skills that share tags, products or a category with Model Routing: LLM Council on Fireworks AI (dair-ai/dair-academy-plugins, 614 stars), Claude Maintain Models (Kiln-AI/Kiln, 5.2k stars), LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars) and Update Ollama Cloud Models (heypinchy/pinchy, 182 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Routing?

alinaqi (a GitHub user) maintains it in alinaqi/maggy, which has 707 GitHub stars. The repository holds 71 skills in this directory. The repository was last updated on September 24, 2026.

Source: alinaqi/maggy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.