Agent skill

OmniRoute CLI Evals

by diegosouzapw in diegosouzapw/OmniRoute

Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

MITAuto-check passedAI & LLM Engineering

Install OmniRoute CLI Evals

skills CLI
$ npx skills add diegosouzapw/OmniRoute --skill cli-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install diegosouzapw/OmniRoute cli-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/diegosouzapw/OmniRoute.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cli-eval .claude/skills/cli-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cli-eval
GitHub stars
74k
Token cost
~1.3k tokens
SKILL.md length
267 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

  • Building an eval suite with exact-match or LLM-judge rubrics
  • SKILL.md covers Overview, Quick install, Subcommands and What are evals?, plus 7 more sections
  • Calls jq, npm and python3
  • Running an eval against a model and watching progress live

What it does

Evals here are automated suites that score model outputs against expected answers or rubrics, stored with their run results in OmniRoute's local database. The command list covers eval suites list and create, eval run with model, combo, concurrency, tag and watch options, eval list with suite, status, since and limit filters, eval get, eval results with a failed-only flag, eval cancel, eval scorecard and a simulate command for a single prompt.

Suites are created with a name, a rubric of exact-match, contains, llm-judge or regex, and a samples file in JSON Lines with an input and expected output per row. Install with npm install -g omniroute or run it through npx, and the CLI entry-point skill covers global flags. The excerpt is truncated before the CI integration details, so those are not described here.

When your agent uses it

  • Building an eval suite with exact-match or LLM-judge rubrics
  • Running an eval against a model and watching progress live
  • Comparing scorecards across models
  • Reviewing only the failed results of an eval run

Example prompts

  • “Create an eval suite named code-quality with the exact-match rubric from samples.jsonl.”
  • “Run the suite against the default combo with watch mode on.”
  • “Show the scorecard for the last run and list only the failed results.”
  • “Simulate this prompt and explain how it would be routed.”

Requirements

  • The `omniroute` CLI, installed with `npm install -g omniroute` or run through npx
  • Node.js and npm

What it can do on your machine

Read from SKILL.md and the folder at commit 8ad6b1c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq
    • npm
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OmniRoute CLI Evals loads about 1.3k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 267 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from diegosouzapw/OmniRoute at commit 8ad6b1c, republished under its MIT licence (© diegosouzapw). 267 words, ~1,322 tokens.

Download SKILL.mdSave it as .claude/skills/cli-eval/SKILL.md (or your agent's skills folder).
name
cli-eval
description
Create and run evaluation suites, watch live benchmark progress, view scorecards, compare model performance, and integrate eval runs with CI workflows from the CLI.
<!-- generated by src/lib/agentSkills/generator.ts; manual edits will be overwritten -->

Overview

Create and run evaluation suites, watch live benchmark progress, view scorecards, compare model performance, and integrate eval runs with CI workflows from the CLI.

Quick install

bash
npm install -g omniroute   # or: npx omniroute
omniroute --version

Subcommands

eval

Example:

bash
omniroute eval
eval suites

Example:

bash
omniroute eval suites
eval list

Example:

bash
omniroute eval list
eval get <suiteId>

Example:

bash
omniroute eval get <suiteId>
eval create

Flags:

  • --file <path>

Example:

bash
omniroute eval create
eval run <suiteId>

Flags:

  • -m, --model <id>
  • --combo <name>
  • --concurrency <n>
  • --tag <tag>
  • --watch

Example:

bash
omniroute eval run <suiteId>
eval list

Flags:

  • --suite <id>
  • --status <s>
  • --since <ts>
  • --limit <n>

Example:

bash
omniroute eval list
eval get <runId>

Example:

bash
omniroute eval get <runId>
eval results <runId>

Flags:

  • --failed

Example:

bash
omniroute eval results <runId>
eval cancel <runId>

Flags:

  • --yes

Example:

bash
omniroute eval cancel <runId>
eval scorecard <runId>

Example:

bash
omniroute eval scorecard <runId>
simulate [prompt]

Flags:

  • --file <path>
  • -m, --model <id>
  • --combo <name>
  • --reasoning-effort <level>
  • --thinking-budget <n>
  • --explain

Example:

bash
omniroute simulate [prompt]
<!-- skill:custom-start -->
<!-- Migrated from skills/omniroute-cli-eval/SKILL.md (preserved curated content) -->

OmniRoute — CLI Evals

Requires the omniroute CLI. See CLI entry-point skill for install + global flags.

What are evals?

Evals are automated test suites that score LLM outputs against expected answers or rubrics. OmniRoute stores suites and run results in its local database.

Eval suites

bash
omniroute eval suites list                       # List all eval suites
omniroute eval suites list --json                # JSON output

omniroute eval suites get <suiteId>              # Full suite definition
Create a suite
bash
omniroute eval suites create \
  --name "code-quality" \
  --rubric "exact-match" \
  --samples-file ./samples.jsonl                 # JSONL: {input, expected_output}

Rubric options: exact-match, contains, llm-judge, regex.

--samples-file format (one JSON object per line):

jsonl
{"input": "What is 2+2?", "expected_output": "4"}
{"input": "Translate 'hello' to Spanish", "expected_output": "hola"}

Run an eval

bash
omniroute eval suites run <suiteId> \
  --model claude-sonnet-4-6                      # Run suite against a specific model

omniroute eval suites run <suiteId> \
  --model gpt-4o \
  --watch                                        # Live TUI progress (EvalWatch)

The run is asynchronous. Use --watch for a live terminal dashboard or poll manually:

bash
RUN_ID=$(omniroute eval suites run <suiteId> --model claude-sonnet-4-6 --output json | jq -r '.id')
omniroute eval get $RUN_ID

Manage runs

bash
omniroute eval list                              # List all eval runs
omniroute eval list --json

omniroute eval get <runId>                       # Run details (status, model, score)
omniroute eval results <runId>                   # Per-sample results
omniroute eval scorecard <runId>                 # Full scorecard with pass/fail per sample
omniroute eval cancel <runId>                    # Cancel a running eval

Scorecard output

bash
omniroute eval scorecard <runId> --output json

Response fields per sample:

json
{
  "id": "sample-1",
  "score": 0.95,
  "passed": true,
  "input": "What is 2+2?",
  "output": "4",
  "expected": "4"
}

Comparing models

Run the same suite against multiple models and compare:

bash
for MODEL in claude-sonnet-4-6 gpt-4o gemini-2.0-flash; do
  omniroute eval suites run $SUITE_ID --model $MODEL --output json | jq '{model: .model, score: .score}'
done

CI integration

bash
# Run and fail CI if score drops below threshold
SCORE=$(omniroute eval suites run $SUITE_ID --model claude-sonnet-4-6 --output json | jq -r '.score')
python3 -c "import sys; score=float('$SCORE'); sys.exit(0 if score >= 0.90 else 1)"

Errors

  • suites create fails with invalid rubric → use one of: exact-match, contains, llm-judge, regex
  • suites run returns model not found → verify model ID with omniroute models --search <name>
  • eval get shows status: failed → check omniroute logs --search eval for error details
  • scorecard returns empty results → the run may still be running; poll omniroute eval get <runId> until status is completed
<!-- skill:custom-end -->

© diegosouzapw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cli-eval of diegosouzapw/OmniRoute.

Open the folder on GitHubat commit 8ad6b1c

Compare with similar skills

OmniRoute CLI Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OmniRoute CLI Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OmniRoute CLI Evals this skilldiegosouzapw/OmniRoute74k—~1.3kAutomated safety check: PassMIT
Duel Agents2aronS/Duel-Agents746—~228Automated safety check: PassMIT
Duel Agents2aronS/Duel-Agents746—~298Automated safety check: NotesMIT
Add Modelget-convex/convex-evals129—~1.5kAutomated safety check: NotesApache-2.0
Phoenix Typescript Package DocsArize-ai/phoenix12k—~2.2kAutomated safety check: PassApache-2.0
Agent Harness DesignAnastasiyaW/codex-claude-code-config154—~764Automated safety check: PassMIT

Similar skills

  • Duel Agents

    2aronS/Duel-Agents

    Use Duel Agents model routing in Cursor. An agent skill from 2aronS/Duel-Agents.

    746 GitHub stars~228 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Duel Agents

    2aronS/Duel-Agents

    Configure and use Duel Agents model routing in Claude Code. An agent skill from 2aronS/Duel-Agents.

    746 GitHub stars~298 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Add Model

    get-convex/convex-evals

    Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

    129 GitHub stars~1.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Maintain the bundled TypeScript package docs that ship inside Phoenix npm packages.

    12k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Agent Harness Design

    AnastasiyaW/codex-claude-code-config

    Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…

    154 GitHub stars~764 tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from diegosouzapw/OmniRoute

All 50 skills in this repo
  • OmniRoute Backup and Sync CLI

    diegosouzapw/OmniRoute

    Backup and restore OmniRoute data from the CLI. Trigger incremental snapshots, sync to cloud storage, manage backup schedules, and restore from archive files.

    74k GitHub stars~948 tokensUpdated today
    Auto-check passed
  • OmniRoute Settings API

    diegosouzapw/OmniRoute

    Read and update global application settings: system prompts, thinking budget, IP filters, payload rules, combo defaults, and require-login configuration.

    74k GitHub starsUsed in 1 repo~3.6k tokens
    Auto-check passed
  • Quality Scan

    diegosouzapw/OmniRoute

    Runs a scoped, read-only quality scan on a repository candidate and reports exact evidence, failures and frozen debt, without treating a static scan as release acceptance.

    74k GitHub stars~748 tokensUpdated today
    Auto-check passed
  • OmniRoute Database Backups

    diegosouzapw/OmniRoute

    Trigger system backups, restore from backup files, and manage the SQLite database lifecycle. Supports export, import, and incremental snapshot strategies.

    74k GitHub starsUsed in 1 repo~395 tokens
    Auto-check passed
  • OmniRoute Provider Management

    diegosouzapw/OmniRoute

    Manages AI provider connections, API keys, OAuth flows and connection tests through OmniRoute's REST API across its 327-provider catalog.

    74k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • OmniRoute LLM Cache

    diegosouzapw/OmniRoute

    Documents OmniRoute's cache endpoints for reading cache statistics and clearing entries, statistics or the reasoning cache, with notes on TTL and similarity settings.

    74k GitHub starsUsed in 1 repo~529 tokens
    Auto-check passed

Works with

Questions about OmniRoute CLI Evals

What does OmniRoute CLI Evals do?

Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI. Evals here are automated suites that score model outputs against expected answers or rubrics, stored with their run results in OmniRoute's local database. The command list covers eval suites list and create, eval run with model, combo, concurrency, tag and watch options, eval list with suite, status, since and limit filters, eval get, eval results with a failed-only flag, eval cancel, eval scorecard and a simulate command for a single prompt.

When should I use OmniRoute CLI Evals?

OmniRoute CLI Evals fits situations like: building an eval suite with exact-match or LLM-judge rubrics; running an eval against a model and watching progress live; comparing scorecards across models; reviewing only the failed results of an eval run.

How do I install OmniRoute CLI Evals in Claude Code?

Run `npx skills add diegosouzapw/OmniRoute --skill cli-eval -a claude-code`. Or copy the skill folder (skills/cli-eval in diegosouzapw/OmniRoute) into .claude/skills/cli-eval in your project. Claude Code loads it when a task matches its description.

How do I install OmniRoute CLI Evals in Codex?

Run `npx skills add diegosouzapw/OmniRoute --skill cli-eval -a codex`. Or copy the skill folder (skills/cli-eval in diegosouzapw/OmniRoute) into .agents/skills/cli-eval in your project. Codex loads it when a task matches its description.

Can I use OmniRoute CLI Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add diegosouzapw/OmniRoute --skill cli-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cli-eval, .gemini/skills/cli-eval, .github/skills/cli-eval and .opencode/skills/cli-eval in your project.

What does OmniRoute CLI Evals need to run?

Going by SKILL.md and its folder, OmniRoute CLI Evals needs the command-line tools its instructions call (jq, npm and python3). Our summary lists: The `omniroute` CLI, installed with `npm install -g omniroute` or run through npx; Node.js and npm.

Does OmniRoute CLI Evals access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is OmniRoute CLI Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does OmniRoute CLI Evals use?

OmniRoute CLI Evals is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does OmniRoute CLI Evals use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to OmniRoute CLI Evals?

Skills that share tags, products or a category with OmniRoute CLI Evals: Duel Agents (2aronS/Duel-Agents, 746 stars), Duel Agents (2aronS/Duel-Agents, 746 stars), Add Model (get-convex/convex-evals, 129 stars) and Phoenix Typescript Package Docs (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OmniRoute CLI Evals?

diegosouzapw (a GitHub user) maintains it in diegosouzapw/OmniRoute, which has 73,701 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 6, 2026.

Source: diegosouzapw/OmniRoute on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.