Duel Agents
2aronS/Duel-Agents
Use Duel Agents model routing in Cursor. An agent skill from 2aronS/Duel-Agents.
Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.
$ npx skills add diegosouzapw/OmniRoute --skill cli-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install diegosouzapw/OmniRoute cli-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/diegosouzapw/OmniRoute.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cli-eval .claude/skills/cli-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cli-eval" agent skill from https://github.com/diegosouzapw/OmniRoute/tree/release%2Fv3.8.52/skills/cli-eval into .claude/skills/cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cli-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/diegosouzapw/OmniRoute/tree/release%2Fv3.8.52/skills/cli-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add diegosouzapw/OmniRoute --skill cli-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install diegosouzapw/OmniRoute cli-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/diegosouzapw/OmniRoute.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cli-eval .agents/skills/cli-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cli-eval" agent skill from https://github.com/diegosouzapw/OmniRoute/tree/release%2Fv3.8.52/skills/cli-eval into .agents/skills/cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cli-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add diegosouzapw/OmniRoute --skill cli-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install diegosouzapw/OmniRoute cli-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/diegosouzapw/OmniRoute.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cli-eval .cursor/skills/cli-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cli-eval" agent skill from https://github.com/diegosouzapw/OmniRoute/tree/release%2Fv3.8.52/skills/cli-eval into .cursor/skills/cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cli-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/diegosouzapw/OmniRoute.git --path skills/cli-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add diegosouzapw/OmniRoute --skill cli-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install diegosouzapw/OmniRoute cli-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/diegosouzapw/OmniRoute.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cli-eval .gemini/skills/cli-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cli-eval" agent skill from https://github.com/diegosouzapw/OmniRoute/tree/release%2Fv3.8.52/skills/cli-eval into .gemini/skills/cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cli-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install diegosouzapw/OmniRoute cli-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add diegosouzapw/OmniRoute --skill cli-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/diegosouzapw/OmniRoute.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cli-eval .github/skills/cli-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cli-eval" agent skill from https://github.com/diegosouzapw/OmniRoute/tree/release%2Fv3.8.52/skills/cli-eval into .github/skills/cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cli-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add diegosouzapw/OmniRoute --skill cli-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install diegosouzapw/OmniRoute cli-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/diegosouzapw/OmniRoute.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cli-eval .opencode/skills/cli-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cli-eval" agent skill from https://github.com/diegosouzapw/OmniRoute/tree/release%2Fv3.8.52/skills/cli-eval into .opencode/skills/cli-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cli-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cli-evalCreates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.
Evals here are automated suites that score model outputs against expected answers or rubrics, stored with their run results in OmniRoute's local database. The command list covers eval suites list and create, eval run with model, combo, concurrency, tag and watch options, eval list with suite, status, since and limit filters, eval get, eval results with a failed-only flag, eval cancel, eval scorecard and a simulate command for a single prompt.
Suites are created with a name, a rubric of exact-match, contains, llm-judge or regex, and a samples file in JSON Lines with an input and expected output per row. Install with npm install -g omniroute or run it through npx, and the CLI entry-point skill covers global flags. The excerpt is truncated before the CI integration details, so those are not described here.
Read from SKILL.md and the folder at commit 8ad6b1c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqnpmpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
OmniRoute CLI Evals loads about 1.3k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 267 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from diegosouzapw/OmniRoute at commit 8ad6b1c, republished under its MIT licence (© diegosouzapw). 267 words, ~1,322 tokens.
.claude/skills/cli-eval/SKILL.md (or your agent's skills folder).<!-- generated by src/lib/agentSkills/generator.ts; manual edits will be overwritten -->
Create and run evaluation suites, watch live benchmark progress, view scorecards, compare model performance, and integrate eval runs with CI workflows from the CLI.
npm install -g omniroute # or: npx omniroute
omniroute --versionevalExample:
omniroute evaleval suitesExample:
omniroute eval suiteseval listExample:
omniroute eval listeval get <suiteId>Example:
omniroute eval get <suiteId>eval createFlags:
--file <path>Example:
omniroute eval createeval run <suiteId>Flags:
-m, --model <id>--combo <name>--concurrency <n>--tag <tag>--watchExample:
omniroute eval run <suiteId>eval listFlags:
--suite <id>--status <s>--since <ts>--limit <n>Example:
omniroute eval listeval get <runId>Example:
omniroute eval get <runId>eval results <runId>Flags:
--failedExample:
omniroute eval results <runId>eval cancel <runId>Flags:
--yesExample:
omniroute eval cancel <runId>eval scorecard <runId>Example:
omniroute eval scorecard <runId>simulate [prompt]Flags:
--file <path>-m, --model <id>--combo <name>--reasoning-effort <level>--thinking-budget <n>--explainExample:
omniroute simulate [prompt]<!-- skill:custom-start -->
<!-- Migrated from skills/omniroute-cli-eval/SKILL.md (preserved curated content) -->
Requires the omniroute CLI. See CLI entry-point skill for install + global flags.
Evals are automated test suites that score LLM outputs against expected answers or rubrics. OmniRoute stores suites and run results in its local database.
omniroute eval suites list # List all eval suites
omniroute eval suites list --json # JSON output
omniroute eval suites get <suiteId> # Full suite definitionomniroute eval suites create \
--name "code-quality" \
--rubric "exact-match" \
--samples-file ./samples.jsonl # JSONL: {input, expected_output}Rubric options: exact-match, contains, llm-judge, regex.
--samples-file format (one JSON object per line):
{"input": "What is 2+2?", "expected_output": "4"}
{"input": "Translate 'hello' to Spanish", "expected_output": "hola"}omniroute eval suites run <suiteId> \
--model claude-sonnet-4-6 # Run suite against a specific model
omniroute eval suites run <suiteId> \
--model gpt-4o \
--watch # Live TUI progress (EvalWatch)The run is asynchronous. Use --watch for a live terminal dashboard or poll manually:
RUN_ID=$(omniroute eval suites run <suiteId> --model claude-sonnet-4-6 --output json | jq -r '.id')
omniroute eval get $RUN_IDomniroute eval list # List all eval runs
omniroute eval list --json
omniroute eval get <runId> # Run details (status, model, score)
omniroute eval results <runId> # Per-sample results
omniroute eval scorecard <runId> # Full scorecard with pass/fail per sample
omniroute eval cancel <runId> # Cancel a running evalomniroute eval scorecard <runId> --output jsonResponse fields per sample:
{
"id": "sample-1",
"score": 0.95,
"passed": true,
"input": "What is 2+2?",
"output": "4",
"expected": "4"
}Run the same suite against multiple models and compare:
for MODEL in claude-sonnet-4-6 gpt-4o gemini-2.0-flash; do
omniroute eval suites run $SUITE_ID --model $MODEL --output json | jq '{model: .model, score: .score}'
done# Run and fail CI if score drops below threshold
SCORE=$(omniroute eval suites run $SUITE_ID --model claude-sonnet-4-6 --output json | jq -r '.score')
python3 -c "import sys; score=float('$SCORE'); sys.exit(0 if score >= 0.90 else 1)"suites create fails with invalid rubric → use one of: exact-match, contains, llm-judge, regexsuites run returns model not found → verify model ID with omniroute models --search <name>eval get shows status: failed → check omniroute logs --search eval for error detailsscorecard returns empty results → the run may still be running; poll omniroute eval get <runId> until status is completed<!-- skill:custom-end -->
© diegosouzapw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/cli-eval of diegosouzapw/OmniRoute.
Open the folder on GitHubat commit 8ad6b1c
OmniRoute CLI Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| OmniRoute CLI Evals this skilldiegosouzapw/OmniRoute | 74k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Duel Agents2aronS/Duel-Agents | 746 | — | ~228 | Automated safety check: Pass | MIT | |
| Duel Agents2aronS/Duel-Agents | 746 | — | ~298 | Automated safety check: Notes | MIT | |
| Add Modelget-convex/convex-evals | 129 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | |
| Phoenix Typescript Package DocsArize-ai/phoenix | 12k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Agent Harness DesignAnastasiyaW/codex-claude-code-config | 154 | — | ~764 | Automated safety check: Pass | MIT |
2aronS/Duel-Agents
Use Duel Agents model routing in Cursor. An agent skill from 2aronS/Duel-Agents.
2aronS/Duel-Agents
Configure and use Duel Agents model routing in Claude Code. An agent skill from 2aronS/Duel-Agents.
get-convex/convex-evals
Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.
Arize-ai/phoenix
Maintain the bundled TypeScript package docs that ship inside Phoenix npm packages.
AnastasiyaW/codex-claude-code-config
Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
diegosouzapw/OmniRoute
Backup and restore OmniRoute data from the CLI. Trigger incremental snapshots, sync to cloud storage, manage backup schedules, and restore from archive files.
diegosouzapw/OmniRoute
Read and update global application settings: system prompts, thinking budget, IP filters, payload rules, combo defaults, and require-login configuration.
diegosouzapw/OmniRoute
Runs a scoped, read-only quality scan on a repository candidate and reports exact evidence, failures and frozen debt, without treating a static scan as release acceptance.
diegosouzapw/OmniRoute
Trigger system backups, restore from backup files, and manage the SQLite database lifecycle. Supports export, import, and incremental snapshot strategies.
diegosouzapw/OmniRoute
Manages AI provider connections, API keys, OAuth flows and connection tests through OmniRoute's REST API across its 327-provider catalog.
diegosouzapw/OmniRoute
Documents OmniRoute's cache endpoints for reading cache statistics and clearing entries, statistics or the reasoning cache, with notes on TTL and similarity settings.
Works with
Categories
Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI. Evals here are automated suites that score model outputs against expected answers or rubrics, stored with their run results in OmniRoute's local database. The command list covers eval suites list and create, eval run with model, combo, concurrency, tag and watch options, eval list with suite, status, since and limit filters, eval get, eval results with a failed-only flag, eval cancel, eval scorecard and a simulate command for a single prompt.
OmniRoute CLI Evals fits situations like: building an eval suite with exact-match or LLM-judge rubrics; running an eval against a model and watching progress live; comparing scorecards across models; reviewing only the failed results of an eval run.
Run `npx skills add diegosouzapw/OmniRoute --skill cli-eval -a claude-code`. Or copy the skill folder (skills/cli-eval in diegosouzapw/OmniRoute) into .claude/skills/cli-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add diegosouzapw/OmniRoute --skill cli-eval -a codex`. Or copy the skill folder (skills/cli-eval in diegosouzapw/OmniRoute) into .agents/skills/cli-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add diegosouzapw/OmniRoute --skill cli-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cli-eval, .gemini/skills/cli-eval, .github/skills/cli-eval and .opencode/skills/cli-eval in your project.
Going by SKILL.md and its folder, OmniRoute CLI Evals needs the command-line tools its instructions call (jq, npm and python3). Our summary lists: The `omniroute` CLI, installed with `npm install -g omniroute` or run through npx; Node.js and npm.
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
OmniRoute CLI Evals is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with OmniRoute CLI Evals: Duel Agents (2aronS/Duel-Agents, 746 stars), Duel Agents (2aronS/Duel-Agents, 746 stars), Add Model (get-convex/convex-evals, 129 stars) and Phoenix Typescript Package Docs (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
diegosouzapw (a GitHub user) maintains it in diegosouzapw/OmniRoute, which has 73,701 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 6, 2026.
Source: diegosouzapw/OmniRoute on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.