Agent skill

OmniRoute Inference Endpoints

by diegosouzapw in diegosouzapw/OmniRoute

Documents OmniRoute's OpenAI-compatible endpoints for chat completions, embeddings, images, speech, transcription, moderation, rerank and the Responses API.

MITAuto-check passedAI & LLM Engineering

Install OmniRoute Inference Endpoints

skills CLI
$ npx skills add diegosouzapw/OmniRoute --skill omni-inference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install diegosouzapw/OmniRoute omni-inference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/diegosouzapw/OmniRoute.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/omni-inference .claude/skills/omni-inference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
omni-inference
GitHub stars
74k
Token cost
~5.5k tokens
SKILL.md length
866 words
Files
2 (incl. references)
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Documents OmniRoute's OpenAI-compatible endpoints for chat completions, embeddings, images, speech, transcription, moderation, rerank and the Responses API.

  • Pointing an agent at OmniRoute's chat completions endpoint
  • SKILL.md covers Overview, Authentication, Endpoints and Payloads, plus 5 more sections
  • Calls curl and jq; reaches anthropic.com; needs OMNIROUTE_KEY and REQUIRE_API_KEY
  • Generating embeddings or images through the gateway

What it does

The main integration surface that OmniRoute offers to AI agents. Its routes follow the OpenAI shape: chat completions, embeddings, image generation, audio speech and transcription, moderations, rerank and the Responses API. Provider-specific variants of the chat, embeddings and image routes sit under a providers path.

The endpoint list also includes web search, session leases, a WebSocket route, a messages route with token counting, multimodal embeddings and a models listing per provider. Full details are in references/endpoints.md. Authentication is a Bearer token or session cookie, and REQUIRE_API_KEY=false is the local development switch.

When your agent uses it

  • Pointing an agent at OmniRoute's chat completions endpoint
  • Generating embeddings or images through the gateway
  • Transcribing audio or synthesizing speech through one API
  • Counting tokens before sending a message

Example prompts

  • “Call OmniRoute's embeddings endpoint on these three paragraphs.”
  • “Send this prompt to the chat completions route of my gateway.”
  • “Transcribe meeting.mp3 through the OmniRoute audio transcription endpoint.”

Requirements

  • A running OmniRoute instance
  • A Bearer token or session cookie

What it can do on your machine

Read from SKILL.md and the folder at commit f8a8131. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • anthropic.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OMNIROUTE_KEY
    • REQUIRE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OmniRoute Inference Endpoints loads about 5.5k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 866 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~16k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from diegosouzapw/OmniRoute at commit f8a8131, republished under its MIT licence (© diegosouzapw). 866 words, ~5,465 tokens.

Download SKILL.mdSave it as .claude/skills/omni-inference/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
omni-inference
description
The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents.
<!-- generated by src/lib/agentSkills/generator.ts; manual edits will be overwritten -->

Overview

The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents.

Authentication

All requests require a valid Bearer token or session cookie. Obtain a token via POST /api/auth/login or configure REQUIRE_API_KEY=false for local development.

Endpoints

Payloads

See the full OpenAPI specification at GET /api/openapi/spec or docs/openapi.yaml for detailed request/response schemas.

<!-- skill:custom-start -->
<!-- Aggregated from: omniroute-chat, omniroute-image, omniroute-tts, omniroute-stt, omniroute-embeddings, omniroute-web-search, omniroute-web-fetch -->

Chat completions

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoints
  • POST $OMNIROUTE_URL/v1/chat/completions — OpenAI format
  • POST $OMNIROUTE_URL/v1/messages — Anthropic Messages format
  • POST $OMNIROUTE_URL/v1/responses — OpenAI Responses API
Discover
bash
curl $OMNIROUTE_URL/v1/models | jq '.data[].id'

Combos (e.g. auto, cost-optimized, subscription) auto-fallback through multiple providers.

OpenAI format example
bash
curl -X POST $OMNIROUTE_URL/v1/chat/completions \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "messages": [{"role": "user", "content": "Refactor this function"}],
    "stream": true
  }'
Anthropic format example
bash
curl -X POST $OMNIROUTE_URL/v1/messages \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "max_tokens": 4096,
    "messages": [{"role": "user", "content": "Hi"}]
  }'
Tool use

Supports OpenAI tools array and Anthropic tools block. Tool results auto-compressed via RTK (47 filters: git-diff, grep, test-jest, terraform-plan, docker-logs, etc.) — 20-40% token savings. Disable per-request with X-Omniroute-Rtk: off header.

Reasoning / thinking

Anthropic extended thinking and OpenAI Responses reasoning blocks are forwarded verbatim. Cached automatically via reasoning cache.

Errors
  • 401 → invalid API key
  • 400 invalid_model → model not in registry; check /v1/models
  • 503 circuit_open → provider circuit breaker tripped; retry later or use combo
  • 429 rate_limited → honor Retry-After; consider using a combo for auto-fallback

Image generation

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoints
  • POST $OMNIROUTE_URL/v1/images/generations — Text-to-image
  • POST $OMNIROUTE_URL/v1/images/edits — Image edit (mask)
  • POST $OMNIROUTE_URL/v1/images/variations — Variations
Discover
bash
curl $OMNIROUTE_URL/v1/models/image | jq '.data[]'

Returns { id, owned_by, sizes:[...], capabilities:[...] } per model.

Generate example
bash
curl -X POST $OMNIROUTE_URL/v1/images/generations \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "dall-e-3",
    "prompt": "a red bicycle on a wet street, photoreal",
    "n": 1,
    "size": "1024x1024",
    "response_format": "b64_json"
  }'

Response: { created, data: [{ url? or b64_json, revised_prompt }] }

Errors
  • 400 invalid_size → not supported by this model; check /v1/models/image
  • 400 content_policy_violation → blocked by provider safety
  • 503 → provider unavailable; try another model in /v1/models/image
Show full SKILL.md (359 more words)Show less

Text-to-speech

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoint
  • POST $OMNIROUTE_URL/v1/audio/speech — returns binary audio (mp3/opus/wav/flac)
Discover
bash
curl $OMNIROUTE_URL/v1/models/tts | jq '.data[]'

Each entry includes voices:[...] for the available voice names per provider.

Example
bash
curl -X POST $OMNIROUTE_URL/v1/audio/speech \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Hello from OmniRoute.",
    "voice": "alloy",
    "response_format": "mp3"
  }' --output speech.mp3
Voices

Voice names vary by provider. Check /v1/models/tts — each entry has voices:[...]. Common OpenAI voices: alloy, echo, fable, onyx, nova, shimmer.

Errors
  • 400 invalid_voice → voice not supported by this model
  • 400 input_too_long → input exceeds model character limit
  • 503 → provider unavailable; try another model in /v1/models/tts

Speech-to-text

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoints
  • POST $OMNIROUTE_URL/v1/audio/transcriptions — multipart upload, returns text
  • POST $OMNIROUTE_URL/v1/audio/translations — transcribe + translate to English
Discover
bash
curl $OMNIROUTE_URL/v1/models/stt | jq '.data[]'
Example
bash
curl -X POST $OMNIROUTE_URL/v1/audio/transcriptions \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -F "file=@audio.mp3" \
  -F "model=whisper-1" \
  -F "response_format=verbose_json"

Response: { text, language, duration, segments?:[{ start, end, text }] }

Supported formats

Audio: mp3, mp4, mpeg, mpga, m4a, wav, webm. Response formats: json, text, srt, verbose_json, vtt.

Errors
  • 400 invalid_file_format → unsupported audio format
  • 400 file_too_large → exceeds provider limit (usually 25MB)
  • 503 → provider unavailable; try another model in /v1/models/stt

Embeddings

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoint
  • POST $OMNIROUTE_URL/v1/embeddings
Discover
bash
curl $OMNIROUTE_URL/v1/models/embedding | jq '.data[]'

Each entry: { id, owned_by, dimensions, max_input_tokens }.

Example
bash
curl -X POST $OMNIROUTE_URL/v1/embeddings \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "text-embedding-3-large",
    "input": ["first text", "second text"],
    "encoding_format": "float"
  }'

Response: { data:[{ embedding:[...], index }], usage:{ prompt_tokens, total_tokens } }

Batch input

input accepts a string or array of strings (up to provider batch limit, typically 2048 items).

Errors
  • 400 input_too_long → input exceeds max_input_tokens for this model
  • 400 invalid_encoding_format → use float or base64
  • 503 → provider unavailable; try another model in /v1/models/embedding

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoint
  • POST $OMNIROUTE_URL/v1/web/search — unified search format
Discover
bash
curl $OMNIROUTE_URL/v1/models/web | jq '.data[] | select(.kind == "webSearch")'
Example
bash
curl -X POST $OMNIROUTE_URL/v1/web/search \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tavily/search",
    "query": "OmniRoute github latest release",
    "max_results": 5,
    "include_answer": true
  }'

Response: { answer?, results:[{ url, title, content, score }] }

Parameters
FieldTypeDescription
modelstringProvider model from /v1/models/web
querystringSearch query
max_resultsnumberMax results (default: 5)
include_answerbooleanInclude AI-synthesized answer
search_depthstringbasic or advanced (Tavily)
Errors
  • 400 query_too_long → shorten the search query
  • 503 → provider unavailable; try another model in /v1/models/web

Web fetch

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoint
  • POST $OMNIROUTE_URL/v1/web/fetch
Discover
bash
curl $OMNIROUTE_URL/v1/models/web | jq '.data[] | select(.kind == "webFetch")'
Example
bash
curl -X POST $OMNIROUTE_URL/v1/web/fetch \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jina/reader",
    "url": "https://anthropic.com",
    "format": "markdown"
  }'

Response: { url, title, markdown, links?:[...], images?:[...] }

Parameters
FieldTypeDescription
modelstringProvider from /v1/models/web (e.g. jina/reader, firecrawl/scrape)
urlstringURL to fetch
formatstringmarkdown (default), html, text
Errors
  • 400 invalid_url → URL must be http/https
  • 403 blocked → provider blocked by target site; try a different model
  • 503 → provider unavailable; try another model in /v1/models/web
<!-- skill:custom-end -->

© diegosouzapw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/omni-inference of diegosouzapw/OmniRoute.

  • SKILL.md
  • references/endpoints.md

Open the folder on GitHubat commit f8a8131

Compare with similar skills

OmniRoute Inference Endpoints next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OmniRoute Inference Endpoints compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OmniRoute Inference Endpoints this skilldiegosouzapw/OmniRoute74k—~5.5kAutomated safety check: PassMIT
Caching Architecturemajiayu000/litellm-rs117—~2kAutomated safety check: PassMIT
Llmobs IntegrationDataDog/dd-trace-js836—~1.4kAutomated safety check: PassCustom licence
Keiroutermydisha/keirouter147—~995Automated safety check: PassMIT
Cognee Integrations Setuptopoteretes/cognee32k—~1kAutomated safety check: NotesApache-2.0
Azure AI Openai Dotnetmicrosoft/skills3.1k6 repos~3.4kAutomated safety check: PassMIT

Similar skills

  • Caching Architecture

    majiayu000/litellm-rs

    LiteLLM-RS response caching architecture. An agent skill from majiayu000/litellm-rs.

    117 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Llmobs Integration

    DataDog/dd-trace-js

    Official

    A skill your agent uses when adding, debugging, or modifying LLMObs plugins for an LLM library in dd-trace-js.

    836 GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Keirouter

    mydisha/keirouter

    Entry point for KeiRouter — local/remote AI gateway with OpenAI-compatible REST for chat, image, TTS, embeddings, web search, web fetch.

    147 GitHub stars~995 tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check passed
  • Cognee Integrations Setup

    topoteretes/cognee

    Switches cognee's LLM, embedding, relational, vector and graph backends through environment variables, with the extras to install and the traps to avoid.

    32k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Azure AI Openai Dotnet

    microsoft/skills

    Official

    Azure OpenAI SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Openai API

    ynulihao/AgentSkillOS

    Build with OpenAI's stateless APIs - Chat Completions (GPT-5, GPT-4o), Embeddings, Images (DALL-E 3), Audio (Whisper + TTS), and Moderation.

    617 GitHub starsUsed in 1 repo~5.9k tokens
    AI & LLM EngineeringAuto-check: notes

More from diegosouzapw/OmniRoute

All 50 skills in this repo
  • OmniRoute Backup and Sync CLI

    diegosouzapw/OmniRoute

    Backup and restore OmniRoute data from the CLI. Trigger incremental snapshots, sync to cloud storage, manage backup schedules, and restore from archive files.

    74k GitHub stars~948 tokensUpdated today
    Auto-check passed
  • OmniRoute Settings API

    diegosouzapw/OmniRoute

    Read and update global application settings: system prompts, thinking budget, IP filters, payload rules, combo defaults, and require-login configuration.

    74k GitHub starsUsed in 1 repo~3.6k tokens
    Auto-check passed
  • Quality Scan

    diegosouzapw/OmniRoute

    Runs a scoped, read-only quality scan on a repository candidate and reports exact evidence, failures and frozen debt, without treating a static scan as release acceptance.

    74k GitHub stars~748 tokensUpdated today
    Auto-check passed
  • OmniRoute Database Backups

    diegosouzapw/OmniRoute

    Trigger system backups, restore from backup files, and manage the SQLite database lifecycle. Supports export, import, and incremental snapshot strategies.

    74k GitHub starsUsed in 1 repo~395 tokens
    Auto-check passed
  • OmniRoute Provider Management

    diegosouzapw/OmniRoute

    Manages AI provider connections, API keys, OAuth flows and connection tests through OmniRoute's REST API across its 327-provider catalog.

    74k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • OmniRoute Prompt Compression CLI

    diegosouzapw/OmniRoute

    Configure and test prompt compression from the CLI. Manage RTK filters, Caveman rules, stacked compression modes, and preview compression output with real…

    74k GitHub starsUsed in 1 repo~521 tokens
    Auto-check passed

Works with

Questions about OmniRoute Inference Endpoints

What does OmniRoute Inference Endpoints do?

Documents OmniRoute's OpenAI-compatible endpoints for chat completions, embeddings, images, speech, transcription, moderation, rerank and the Responses API. The main integration surface that OmniRoute offers to AI agents. Its routes follow the OpenAI shape: chat completions, embeddings, image generation, audio speech and transcription, moderations, rerank and the Responses API.

When should I use OmniRoute Inference Endpoints?

OmniRoute Inference Endpoints fits situations like: pointing an agent at OmniRoute's chat completions endpoint; generating embeddings or images through the gateway; transcribing audio or synthesizing speech through one API; counting tokens before sending a message.

How do I install OmniRoute Inference Endpoints in Claude Code?

Run `npx skills add diegosouzapw/OmniRoute --skill omni-inference -a claude-code`. Or copy the skill folder (skills/omni-inference in diegosouzapw/OmniRoute) into .claude/skills/omni-inference in your project. Claude Code loads it when a task matches its description.

How do I install OmniRoute Inference Endpoints in Codex?

Run `npx skills add diegosouzapw/OmniRoute --skill omni-inference -a codex`. Or copy the skill folder (skills/omni-inference in diegosouzapw/OmniRoute) into .agents/skills/omni-inference in your project. Codex loads it when a task matches its description.

Can I use OmniRoute Inference Endpoints in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add diegosouzapw/OmniRoute --skill omni-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/omni-inference, .gemini/skills/omni-inference, .github/skills/omni-inference and .opencode/skills/omni-inference in your project.

What does OmniRoute Inference Endpoints need to run?

Going by SKILL.md and its folder, OmniRoute Inference Endpoints needs the command-line tools its instructions call (curl and jq) and credentials named OMNIROUTE_KEY and REQUIRE_API_KEY. Our summary lists: A running OmniRoute instance; A Bearer token or session cookie.

Does OmniRoute Inference Endpoints access the network?

SKILL.md names 1 domain. In commands or code: anthropic.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is OmniRoute Inference Endpoints safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does OmniRoute Inference Endpoints use?

OmniRoute Inference Endpoints is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does OmniRoute Inference Endpoints use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to OmniRoute Inference Endpoints?

Skills that share tags, products or a category with OmniRoute Inference Endpoints: Caching Architecture (majiayu000/litellm-rs, 117 stars), Llmobs Integration (DataDog/dd-trace-js, 836 stars), Keirouter (mydisha/keirouter, 147 stars) and Cognee Integrations Setup (topoteretes/cognee, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OmniRoute Inference Endpoints?

diegosouzapw (a GitHub user) maintains it in diegosouzapw/OmniRoute, which has 74,003 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 8, 2026.

Source: diegosouzapw/OmniRoute on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.