Agent skill

Jev Model Routing

by kerpopule in kerpopule/hermes-jev-skills

Routes a turn or delegated task to the cheapest model and effort lane that will still do it right, using the Jev decision model to classify difficulty and escalate only when needed.

MITAuto-check passedAI & LLM Engineering

Install Jev Model Routing

skills CLI
$ npx skills add kerpopule/hermes-jev-skills --skill jev-model-routing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kerpopule/hermes-jev-skills jev-model-routing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kerpopule/hermes-jev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/jev-model-routing .claude/skills/jev-model-routing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
jev-model-routing
GitHub stars
1.1k
Token cost
~2.7k tokens
SKILL.md length
1,468 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Routes a turn or delegated task to the cheapest model and effort lane that will still do it right, using the Jev decision model to classify difficulty and escalate only when needed.

  • Picking the cheapest model that can still handle a given turn or task
  • SKILL.md covers On Hermes it is automatic, Asking directly (any agent), Lanes: delegating a task, and… and The pools, plus 2 more sections
  • Calls git
  • Deciding whether a sub-agent's work should continue, retry or escalate after a cycle

What it does

Jev answers three questions about a turn in about 0.4 seconds: how hard it is, what kind of work it is, and how costly a mistake would be. Code then walks the configured model pool for that tier and specialty and picks the first model that fits constraints like image support or context size, so model choice is asked for rather than guessed. On Hermes, the hermes-jev plugin routes each fresh user turn automatically before the first model call, with /jev commands to check status or switch between shadow mode, which logs decisions without switching, and routing on, which actually switches models; running /model yourself overrides Jev.

For a Hermes custom provider the plugin cannot infer the backing models.dev provider automatically, so it keeps the current model and logs a message asking for an explicit provider_aliases.custom entry in routing.json rather than guessing, and that alias must be checked against the real endpoint and every pool model before enabling routing.

Any agent can ask directly with jev route, passing the task in the user's own words and the agent's current provider and model, and using the returned model_id, or staying put when routed is false. For delegated work, jev lane classify picks a starting lane and model or effort level, and jev lane step re-evaluates after each cycle using test and lint results and a scope glob, deciding whether to continue, retry, verify, escalate or complete; Jev only decides, it never writes code, patches or designs itself.

When your agent uses it

  • Picking the cheapest model that can still handle a given turn or task
  • Deciding whether a sub-agent's work should continue, retry or escalate after a cycle
  • Setting up provider aliases for a custom model pool in routing.json

Example prompts

  • “Route this task to the cheapest model that can handle it: summarize these 50 support tickets.”
  • “Classify the lane for this delegated refactor task before I hand it to a sub-agent.”
  • “After this test run, decide whether the sub-agent should retry or escalate.”

Requirements

  • The Hermes agent runtime with the hermes-jev plugin, or the jev CLI directly

What it can do on your machine

Read from SKILL.md and the folder at commit a26dad0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Jev Model Routing loads about 2.7k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 1,468 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kerpopule/hermes-jev-skills at commit a26dad0, republished under its MIT licence (© kerpopule). 1,468 words, ~2,722 tokens.

Download SKILL.mdSave it as .claude/skills/jev-model-routing/SKILL.md (or your agent's skills folder).
name
jev-model-routing
description
Use to pick the cheapest good-enough model or effort for a turn or a delegated task (lanes small to escalate), to decide continue/retry/verify/escalate/complete after each cycle, or to tune routing.
version
0.2.0
license
MIT

Model routing with Jev

Jev reads a turn and answers three questions in one ~0.4 s request: how hard is it, what kind of work is it, and would a mistake be costly. Code then walks your pool for that tier and specialty and takes the first model that fits (images, context size). You do not pick models by feel; you ask.

On Hermes it is automatic

With the hermes-jev plugin enabled, each fresh user turn is routed once, before the first model call. Tool-loop follow-ups reuse that decision. Switches, per profile:

/jev                    status
/jev routing shadow     decide and log, but do not switch (start here)
/jev routing on         switch models
/jev routing off
/jev notice on          show "[Jev] medium · coding → kimi-k2.7-code · confidence 0.97" on routed replies

A plugin can swap the model, not the provider connection. On OpenRouter that still means every vendor (DeepSeek, GLM, Kimi, MiniMax, Grok, Qwen, Gemini, GPT). If you run /model yourself, your choice wins and Jev stays out of the way.

For a Hermes custom provider, the plugin cannot infer the backing models.dev provider. It now keeps the current model and logs custom provider needs an explicit provider_aliases.custom instead of blaming an unrelated pool. If and only if that endpoint actually serves the pool's models, set "provider_aliases": {"custom": "venice"} (replace venice with the real pool prefix) in routing.json. Check the endpoint and every pool model before enabling routing; an alias is an operator assertion, not cross-provider discovery. This does not edit any live routing mode.

Asking directly (any agent)

Before delegating a task or spawning a sub-agent, ask which model should get it:

bash
jev route --prompt "<the task, in the person's words>" --current "<provider:model you are on>"

Use model_id from the reply. routed: false means stay where you are; reason says why. Relay notice if the person likes to see routing.

Lanes: delegating a task, and every step after it

For work you hand to a sub-agent or worker, finish with the smallest model and lowest effort that still gets it right. Jev decides; it never writes code, patches or designs.

bash
jev lane classify --task "<the work, in the person's words>"        # first lane + model/effort
jev lane step --task "..." --lane <lane> --attempt <n> \
    --run "<test cmd>" --run "<lint/typecheck cmd>" --scope "<path glob>"   # after each cycle
LaneClaude Code (subagent)Hermes Kanban card (default map)
smalljev-lane-small: Haiku, lowgpt-5.6-luna, medium
mediumjev-lane-medium: Sonnet, mediumgpt-6-sol, medium (today's default)
highjev-lane-high: Opus, mediumgpt-6-sol, medium
escalatejev-lane-escalate: Opus, highgpt-6-astra, high

jev lane targets --host hermes shows the map in force; <hermes root>/jev/lanes.json (or ~/.config/jev/lanes.json) overrides any field. The Hermes map was calibrated on one fleet's own history (see docs/lanes.md); re-measure yours with jev lane replay-build / replay-report.

  • One request, all questions. classify asks the lane (with an other escape: work a person should see first), security sensitivity and underspecification together. Code applies the thresholds: a small pick needs 0.7 confidence; a medium pick below 0.5 goes to high; security ≥ 0.7 is at least high. keep_current means keep the model you had (do it yourself, or ask).
  • Code first. A model the person named, two failed attempts, or your own security-path check decide without asking Jev.
  • Deterministic checks first. step runs the tests, compiler, type checker and linter you name and reads git diff. A failing check is retry (and escalate once the same lane failed twice); files outside --scope are retry; unrun checks are verify; security files changed on small/medium are escalate. Jev is asked only what is left: is it implemented, is it in scope, what next.
  • Escalate one lane at a time, on evidence only. escalate from the top lane returns person.
  • Complete is earned. complete is refused (becomes verify, with complete_refused) unless the checks ran and passed and the diff stayed in scope. Say when a check failed; never hide it.
  • Only the tail of each long check output goes to Jev. Never compact or filter the agent's own reasoning.
  • Jev down: classify keeps the current model, step says verify.
  • Make --run a script, not a one-liner. It runs under a shell, so pipes and && work, but the evidence prints the command cut short and a reader cannot tell what passed. A small script that checks the exact commit and each exit code, and prints ok:/FAIL: per gate, keeps the evidence legible. Make sure it reads only this cycle's results: an earlier failed attempt left in the same log will fail (or pass) the wrong run.
  • Read-only work: --no-changes-expected. For a review or report, say so; otherwise an empty diff reads as nothing done.
  • escalate with a decision still open means person. After a review whose facts are verified but which leaves the owner a choice (a risk to accept, an approach to pick), a stronger model cannot settle it. Put the decision to the person rather than re-running the work a lane up.

On Hermes, jev lane shadow (from cron) classifies new Kanban cards and logs what it would choose; /jev lanes shadow|on|off is the switch and <hermes root>/jev/LANES_OFF wins. on sets the card's model and effort before dispatch; turn it on only after jev lane shadow-report shows fewer tokens at the same first-try success, and with the owner's yes.

The pools

jev models list shows every model this machine can call (the models.dev catalog, filtered to providers you hold a key or login for) with price, context and abilities. Pools live in ~/.hermes/jev/routing.json (or ~/.config/jev/routing.json):

json
{"tiers": {"simple": {"general": ["openrouter:deepseek/deepseek-v4.1-flash"], "coding": ["..."]},
           "medium": {"general": ["..."], "coding": ["..."], "research": ["..."], "writing": ["..."], "vision": ["..."]},
           "hard":   {"general": ["..."], "coding": ["..."]}},
 "exclude": ["*:free"], "private_profiles": ["billing"], "mode": "redacted-text"}
  • jev models suggest --write creates a first draft from price bands. Then edit: order matters, first fit wins.
  • Specialties are general, coding, writing, research, vision. A missing specialty falls back to general. A pool never falls down a tier, only up.
  • When the person names a model they like for something, put it first in that pool. Do not invent model ids: copy them from jev models list --search <name>.
Show full SKILL.md (576 more words)Show less

Guarantees you can rely on

  • Hard is earned: it needs real probability mass on "substantial" or "expert" (0.6 by default), read from the per-level spread Jev returns, never from an averaged score.
  • Unsure is not hard. An unsure answer about a harmless turn keeps the current model; about a risky turn it picks medium.
  • Risk words (production, delete, migration, security, payment, legal…) set a floor of medium, however short the prompt. They do not buy the hard tier on their own.
  • Jev judges the ask: a long turn is read as its opening plus, mostly, its end (ask_chars). Boilerplate in the middle is not what gets scored.
  • Reasoning effort is off by default and only writes the standard reasoning_effort field when routing is on. Opt in with an exact provider:model capability map, for example "effort": {"enabled": true, "levels": ["low", "medium", "high", "high"], "models": {"openrouter:your-verified-model-id": ["low", "medium", "high"]}}. Replace the example ID with a model actually verified to accept those levels; xhigh is not presumed supported. The requested level must appear in the exact model's allowed list, and an existing reasoning_effort or extra_body.reasoning always wins. Shadow/off never mutate requests. The pick reuses the routing difficulty answer without another Jev call; unsupported or malformed configuration fails open. No fleet effort setting or live routing is activated by installation.
  • Template turns are not routed: anything starting with a skip_prefixes entry ([kanban], [SESSION HANDOFF…) or from a skip_session_prefixes session (cron) keeps the model its profile or job was configured with.
  • Large context (strictly above sticky_context_tokens, ~32k by default): never switches to a cheaper model, because rebuilding the prompt cache costs more than it saves. If either model has no catalog price, it conservatively keeps the current model too. Cache keys include the exact side of this guard, not only a coarse context bucket.
  • Opt-in effort routing cannot undercut the resolved routing tier: medium floors the difficulty bucket at routine, hard at substantial. An unsure kept decision also gets at least routine. A large-context keep preserves the resolved floor through effort_tier without triggering a model switch or escalation. Explicit caller effort and exact-model capability checks still win.
  • Turns that look like they contain secrets, and any profile listed in private_profiles, send Jev only coarse features (length, code present, risk words), never text. Those turns, and a profile with mode: features, also opt out of the merged request below.
  • Its three questions normally travel in the same request as skill selection's stage 1 (jevkit/turn.py), because Jev charges per request and not per question, and the connection underneath is pooled (a fresh TLS session per call used to be ~275 ms of the ~520 ms a decision cost). Measured live 2026-09-21/22: one question ~180-250 ms warm, and 1784 ms → 672 ms per turn that needs both, 3 requests → 2. Each feature still reads its own answers through its own thresholds. /jev merge_requests off separates them again.
  • Jev down, slow (2.5 s budget) or malformed: current model, no delay beyond the budget. An answer that contradicts itself — a spread that does not cover the options, mass that does not sum to one, a chosen option that is not the maximum, a score that disagrees with its own distribution — is refused as invalid_response and lands here too.

Tuning

Decisions are logged without prompt text to <hermes home>/logs/jev-decisions.jsonl. Run in shadow for a day, read which tier real turns land in, then move models between pools. Change thresholds from your own traces, never from a hunch.

© kerpopule, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/jev-model-routing of kerpopule/hermes-jev-skills.

Open the folder on GitHubat commit a26dad0

Compare with similar skills

Jev Model Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Jev Model Routing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Jev Model Routing this skillkerpopule/hermes-jev-skills1.1k—~2.7kAutomated safety check: PassMIT
FreeRide Free Model ManagerShaivpidadi/FreeRide2372 repos~1.1kAutomated safety check: PassNone
Hyper Jevdisler/ten-levels-of-jev213—~1.7kAutomated safety check: PassMIT
Openrouter Context Optimizationjeremylongshore/tons-of-skills-marketplace2.8k—~2.4kAutomated safety check: PassMIT
Embeddings via 9Routerdecolua/9router30k—~604Automated safety check: PassMIT
Using Ccproxy Inspectorstarbaser/ccproxy350—~2.7kAutomated safety check: PassCustom licence

Similar skills

  • FreeRide Free Model Manager

    Shaivpidadi/FreeRide

    Configures OpenClaw to use free OpenRouter models, setting the best one as primary and adding ranked fallbacks so rate limits do not interrupt work.

    237 GitHub starsUsed in 2 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hyper Jev

    disler/ten-levels-of-jev

    Integrate and use Jev, TypeSafe AI's System One decision model, in production codebases.

    213 GitHub stars~1.7k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Openrouter Context Optimization

    jeremylongshore/tons-of-skills-marketplace

    Optimize context window usage for OpenRouter models to reduce cost and improve quality.

    2.8k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy Inspector

    starbaser/ccproxy

    Operates the ccproxy inspector MITM system for intercepting, inspecting, and transforming LLM API traffic.

    350 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Outsourcerer

    alexgreensh/outsourcerer

    Cross-harness orchestrator for AI coding work. An agent skill from alexgreensh/outsourcerer.

    170 GitHub stars~5.5k tokensUpdated 20 days ago
    AI & LLM EngineeringAuto-check: notes

More from kerpopule/hermes-jev-skills

All 10 skills in this repo
  • Jev Browser Use

    kerpopule/hermes-jev-skills

    Drives web pages that need interaction, letting Jev choose one action at a time from observed page elements under a host allowlist and step budget.

    1.1k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Jev Desktop Computer Use

    kerpopule/hermes-jev-skills

    Drives desktop GUI apps and OS dialogs by letting Jev pick the next action from a menu of safe actions the agent built, with a Mac Co-Agent shortcut.

    1.1k GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Jev Transcript Compaction

    kerpopule/hermes-jev-skills

    Uses Jev to mark each transcript turn keep, summarize or drop when cutting a conversation to a fixed size, with measured results on handoff quality.

    1.1k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Jev Key Setup

    kerpopule/hermes-jev-skills

    Connects the Jev decision model by storing a TypeSafe, OpenRouter, Venice or OpenCode Zen key with jev setup-key, so the key never passes through the agent.

    1.1k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check: notes
  • Jev Skill Selector

    kerpopule/hermes-jev-skills

    Ranks a large catalog of installed skills against the current request through the Jev service, and can conclude that no skill applies.

    1.1k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Frontier Model Handoff

    kerpopule/hermes-jev-skills

    Chooses which paid frontier model seat should take a task already judged hard, hands it off with proper context, and keeps a watch on the delegated run.

    1.1k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check: warnings

Works with

Questions about Jev Model Routing

What does Jev Model Routing do?

Routes a turn or delegated task to the cheapest model and effort lane that will still do it right, using the Jev decision model to classify difficulty and escalate only when needed. 4 seconds: how hard it is, what kind of work it is, and how costly a mistake would be. Code then walks the configured model pool for that tier and specialty and picks the first model that fits constraints like image support or context size, so model choice is asked for rather than guessed.

When should I use Jev Model Routing?

Jev Model Routing fits situations like: picking the cheapest model that can still handle a given turn or task; deciding whether a sub-agent's work should continue, retry or escalate after a cycle; setting up provider aliases for a custom model pool in routing.json.

How do I install Jev Model Routing in Claude Code?

Run `npx skills add kerpopule/hermes-jev-skills --skill jev-model-routing -a claude-code`. Or copy the skill folder (skills/jev-model-routing in kerpopule/hermes-jev-skills) into .claude/skills/jev-model-routing in your project. Claude Code loads it when a task matches its description.

How do I install Jev Model Routing in Codex?

Run `npx skills add kerpopule/hermes-jev-skills --skill jev-model-routing -a codex`. Or copy the skill folder (skills/jev-model-routing in kerpopule/hermes-jev-skills) into .agents/skills/jev-model-routing in your project. Codex loads it when a task matches its description.

Can I use Jev Model Routing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kerpopule/hermes-jev-skills --skill jev-model-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/jev-model-routing, .gemini/skills/jev-model-routing, .github/skills/jev-model-routing and .opencode/skills/jev-model-routing in your project.

What does Jev Model Routing need to run?

Going by SKILL.md and its folder, Jev Model Routing needs the command-line tools its instructions call (git). Our summary lists: The Hermes agent runtime with the hermes-jev plugin, or the jev CLI directly.

Does Jev Model Routing access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Jev Model Routing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Jev Model Routing use?

Jev Model Routing is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Jev Model Routing use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Jev Model Routing?

Skills that share tags, products or a category with Jev Model Routing: FreeRide Free Model Manager (Shaivpidadi/FreeRide, 237 stars), Hyper Jev (disler/ten-levels-of-jev, 213 stars), Openrouter Context Optimization (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Embeddings via 9Router (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Jev Model Routing?

kerpopule (a GitHub user) maintains it in kerpopule/hermes-jev-skills, which has 1,056 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 8, 2026.

Source: kerpopule/hermes-jev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.