OpenCodex Proxy Operations
lidge-jun/opencodex
Operates an opencodex (`ocx`) proxy: finds CLI tasks offline, checks local configuration, and manages accounts, providers, models, routing and usage reports.
Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.
$ npx skills add veniceai/skills --skill venice-text-routing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install veniceai/skills venice-text-routing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-text-routing .claude/skills/venice-text-routing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "venice-text-routing" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-text-routing into .claude/skills/venice-text-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-text-routing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/veniceai/skills/tree/main/skills/venice-text-routingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add veniceai/skills --skill venice-text-routing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install veniceai/skills venice-text-routing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/venice-text-routing .agents/skills/venice-text-routing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "venice-text-routing" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-text-routing into .agents/skills/venice-text-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-text-routing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add veniceai/skills --skill venice-text-routing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install veniceai/skills venice-text-routing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/venice-text-routing .cursor/skills/venice-text-routing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "venice-text-routing" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-text-routing into .cursor/skills/venice-text-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-text-routing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/veniceai/skills.git --path skills/venice-text-routing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add veniceai/skills --skill venice-text-routing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install veniceai/skills venice-text-routing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/venice-text-routing .gemini/skills/venice-text-routing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "venice-text-routing" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-text-routing into .gemini/skills/venice-text-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-text-routing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install veniceai/skills venice-text-routingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add veniceai/skills --skill venice-text-routing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/venice-text-routing .github/skills/venice-text-routing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "venice-text-routing" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-text-routing into .github/skills/venice-text-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-text-routing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add veniceai/skills --skill venice-text-routing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install veniceai/skills venice-text-routing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/venice-text-routing .opencode/skills/venice-text-routing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "venice-text-routing" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-text-routing into .opencode/skills/venice-text-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "venice-text-routing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
venice-text-routingPicks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.
Aimed at a local agent that receives a prompt and must decide whether to answer it locally or escalate to Venice, the skill encodes the routing logic that sits above `venice-chat`, the call surface, and `venice-models`, the discovery API. When escalating, it picks the cheapest model that satisfies the required privacy tier, modality and capabilities.
Decisions rely on a cached `snapshots/text-routing.json`. If that file is missing or its `snapshot_date` is more than 30 days old, the agent runs `python scripts/refresh_routing.py`, which rebuilds the snapshot and `routing-matrix.md` from Venice's public model endpoints. No key is needed, though `VENICE_API_KEY` can tailor the result to one key. Filters such as uncensored models or log-probability support need a direct read of the models endpoint.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
curlpythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.venice.aiAlso links to:
docs.venice.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
VENICE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Venice Text Model Routing loads about 5.2k tokens when it runs. Until then it costs about 149 tokens; SKILL.md has 2,039 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 2,039 words, ~5,224 tokens.
.claude/skills/venice-text-routing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.This skill encodes the decision logic for "which Venice text model do I call?" — the routing layer that sits above venice-chat (the call surface) and consumes venice-models (the discovery API).
Primary use case: a local agent receives a prompt, decides whether the local model can handle it, and — if not — picks the cheapest Venice model that satisfies the privacy / modality / capability requirements.
Before applying the matrix below, read
snapshots/text-routing.json. If the file is missing, orsnapshot_dateis older than 30 days, runpython scripts/refresh_routing.pyto regenerate it fromGET /models?type=text+GET /models/traits?type=text. Otherwise trust the cached file — do not hit/modelson every routing decision.
refresh_routing.py rewrites both snapshots/text-routing.json and routing-matrix.md. Both endpoints are public, so no key is needed; VENICE_API_KEY / --api-key is optional and tailors the result to that key (a modelPrivacy-restricted key sees a filtered catalog). Run it once on first install, then ~monthly (CI nightly is also fine).
In the snapshot, each model's privacy label is e2ee when supportsE2EE is true, tee when only supportsTeeAttestation is true, and otherwise the model's own private / anonymized value (every TEE/E2EE model is private in /models, so treat tee and e2ee as private too); tier is bucketed by input price only; beta_access mirrors model_spec.beta (only listed to beta-access keys, so it is always false in the shipped keyless snapshot) and beta_status mirrors model_spec.betaModel (callable, but may change or disappear).
Per model the snapshot carries only: the 13 capability flags shown in the matrix, max_images, available_context_tokens, max_completion_tokens, input/output price per 1M, beta_access, beta_status, offline, region_restrictions and deprecation_date. Filters that need anything else (uncensored, supportsLogProbs, reasoningEffortOptions, maxVideos, cache_input / extended pricing, replacementModelId) have to read GET /models?type=text.
default_reasoning, most_intelligent, …) and a hand-filtered candidate.For the chat call surface itself, see venice-chat. For raw discovery of every model field, see venice-models.
Pick the least restrictive tier that satisfies the request — restricting tier shrinks the candidate pool and often raises cost.
model_spec.privacy has only two values, private and anonymized. TEE and E2EE are capability flags on top of private.
| Tier | Selector (GET /models) | Guarantee | Use when |
|---|---|---|---|
| Anonymized | privacy: "anonymized" | Venice hides your identity from the upstream provider, but the provider may still see the prompt | Non-sensitive workloads; the only path to several closed frontier models (Claude, GPT, Gemini). |
| Private | privacy: "private" | Zero data retention, contract-enforced: content is processed for inference only and not retained | Default for user data, business logic, anything you wouldn't paste into a public chatbot. |
| TEE | capabilities.supportsTeeAttestation: true | Runs inside a hardware Trusted Execution Environment with remote attestation (GET /api/v1/tee/attestation) | Regulated data, verifiable/signed inference. |
| E2EE | capabilities.supportsE2EE: true | TEE + client-side encryption (ECDH secp256k1 → HKDF-SHA256 → AES-256-GCM); Venice relays ciphertext only | Strongest. Healthcare, legal, secrets. Requires the E2EE flow in venice-chat. |
Today every listed TEE model is also E2EE-capable and uses an e2ee-* ID (e.g. e2ee-glm-5-3-p, e2ee-kimi-k3-p, e2ee-qwen3-8-27b). TEE and E2EE use the same model ID: send E2EE headers for E2EE, omit them (or set venice_parameters.enable_e2ee: false) for TEE-only. Legacy tee-* IDs are unlisted aliases of e2ee-* models — don't route on the prefix; use the capability flags.
E2EE trade-offs: on E2EE requests Venice injects no web search, scraping, character, or Venice system prompt, and file parts return 400; the E2EE guide also requires stream: true and lists function calling as unsupported (supportsFunctionCalling describes the model, not E2EE mode). Several E2EE models also have small limits (e.g. e2ee-qwen-2-5-7b-p: 32K context, 4,096 output tokens). E2EE is not supported on /responses; TEE-only is (with enable_e2ee: false).
API keys can carry a modelPrivacy restriction: PRIVATE_TEXT (text, embedding and decision models must be Private, TEE, or E2EE) or PRIVATE_ONLY (every model). Such a key's /models list is already filtered; a disallowed model returns 403 — filter to privacy: "private" for such keys.
Sources: docs.venice.ai/overview/privacy, docs.venice.ai/guides/features/tee-e2ee-models.
Two endpoints let you check that inference really ran inside an enclave. Both
are GET, unauthenticated on purpose (attestation evidence has to be
verifiable by any party without credentials), and rate limited to 10 requests
per minute per IP (429 beyond that). Neither is a path in the published
OpenAPI spec; they are live, and the supportsTeeAttestation capability
description in GET /models points at them.
| Endpoint | Query | Returns |
|---|---|---|
GET /api/v1/tee/attestation | model (required), nonce (optional, exactly 64 hex chars / 32 bytes, binds the attestation to your challenge) | Attestation report, TEE provider, verification result, and the signing key/address. |
GET /api/v1/tee/signature | model and request_id (required), signing_algo (optional ecdsa | ecdsa-p256 | rsa) | The provider's signature over a specific request, plus request/response hashes. |
Both return 400 if the model exists but is not TEE-attested, and 404 if the
model ID is unknown. The attestation endpoint returns 502 (with
verified: false) when verification fails; both return 502 when the TEE
provider is unavailable. Chat responses from TEE models carry X-Venice-TEE: true and
X-Venice-TEE-Provider.
Verify the chain of trust in this order: fetch the attestation to get the signing public key and hardware type, confirm the recovered signer matches the attestation signing address, then verify the signature over the exact signed text the signature endpoint returned. Treat the request and response hashes as provider-reported values unless you can recompute them yourself from a documented canonical format.
Map prompt requirement → model_spec field (full list in the model_spec.capabilities section of venice-models). Sending image / audio / video parts, tools / tool_choice, a non-text response_format, or logprobs to a model without the matching flag returns 400 before inference; some other features (e.g. enable_x_search) are silently ignored instead.
| Requirement | Filter | Notes |
|---|---|---|
| Vision (single image) | capabilities.supportsVision | Single-image models keep images only from the last image-bearing message. |
| Vision (multiple images) | supportsVision && supportsMultipleImages | Honor capabilities.maxImages. Hard cap: 10 images per message. |
| Audio input | capabilities.supportsAudioInput | Base64 only; URLs are not accepted. |
| Video input | capabilities.supportsVideoInput | data:video/... URLs, direct public URLs (no redirects), or YouTube links on some models. Max 3 videos per request. |
| Documents (PDF, DOCX, …) | — | Document file parts are extracted to text server-side, so any text model works (not on E2EE). Image files sent as file parts become images and need vision. |
| Reasoning | capabilities.supportsReasoning | To dial effort, also require supportsReasoningEffort and pick a value from reasoningEffortOptions (default: defaultReasoningEffort). |
| Tools / function calling | capabilities.supportsFunctionCalling | Required for any agent loop. |
| Code-heavy task | capabilities.optimizedForCode | GET /models?type=code returns only this subset. |
| Web search (Venice) | — | supportsWebSearch is true on every text model. Toggle with venice_parameters.enable_web_search. |
| X / Twitter search | capabilities.supportsXSearch | xAI native (Grok models). ~$0.01 per search. |
| Structured JSON output | capabilities.supportsResponseSchema | response_format: {type: "json_schema", json_schema: {name, schema, strict}}. |
| Large context | availableContextTokens >= N | Pair with prompt_cache_key; prefer models with pricing.cache_input. Watch pricing.extended, which raises prices above context_token_threshold input tokens. |
| Long output | maxCompletionTokens >= N | Requests above it are rejected on models with an enforced cap. |
| Minimal content filtering | model_spec.uncensored: true | Present only on models Venice classifies as uncensored; upstream providers may still filter. |
| Logprobs | capabilities.supportsLogProbs | Niche — eval / sampling debug. |
These buckets are this skill's own convention (the same boundaries refresh_routing.py uses), keyed on model_spec.pricing.input.usd per 1M tokens. Upper bounds are exclusive (a model at exactly $1.00 is in M), and the snapshot's tier_boundaries_usd_per_1m_input gives each bucket's exclusive upper bound. Pick the smallest bucket that hosts a model satisfying your filters.
| Tier | $/1M input | Use when |
|---|---|---|
| XS | < $0.20 | Classification, intent extraction, simple summarization. |
| S | $0.20 – < $1 | General chat, basic agents, light vision. |
| M | $1 – < $4 | Moderate reasoning, strong code, multi-image vision. |
| L | $4 – < $10 | Heavy reasoning, complex tool use. |
| Frontier | ≥ $10 | Most expensive closed models (e.g. claude-fable-5, openai-gpt-6-astra, openai-gpt-54-pro as of 2026-10-02). |
Output prices vary widely within a bucket (from ~1.5× to ~10× input), so tie-break on output price. Authoritative per-model pricing lives in model_spec.pricing (input and output rates are mirrored as pricing_per_1m in the snapshot) — never hard-code dollar figures from this prose.
Walk top-down. Stop at the first rule that applies.
1. Local-first check
- Prompt is ≤ ~500 tokens, no special-capability requirement,
no privacy escalation, no tool calls expected
→ handle on the local model. Do not call Venice.
2. Privacy gate
- User flagged "private" OR prompt contains regulated data (PHI, secrets, legal),
OR the API key has modelPrivacy PRIVATE_TEXT / PRIVATE_ONLY:
require model_spec.privacy === "private" (includes every TEE/E2EE model;
in the snapshot accept privacy "private", "tee" or "e2ee").
- User flagged "E2EE" / "must be encrypted end to end":
require capabilities.supportsE2EE. Use /chat/completions only.
- User flagged "TEE" / "verifiable inference":
require capabilities.supportsTeeAttestation.
- Otherwise: any tier acceptable (still prefer "private" when otherwise tied).
3. Modality gate
- Image input present → require supportsVision (+ supportsMultipleImages if > 1).
- Audio input present → require supportsAudioInput.
- Video input present → require supportsVideoInput.
4. Capability gate
- Tool calls expected → require supportsFunctionCalling.
- Code-heavy task → prefer optimizedForCode (or call /models?type=code).
- Chain-of-thought / planning / hard math
→ require supportsReasoning;
prefer supportsReasoningEffort to dial reasoning effort.
- Structured JSON output required → require supportsResponseSchema.
- X/Twitter content needed → require supportsXSearch.
- Refusal-free / minimal filtering → prefer model_spec.uncensored (or trait most_uncensored).
5. Context size gate
- Estimated prompt tokens > availableContextTokens, or expected output > maxCompletionTokens
→ bump up to a model with sufficient limits. Prefer ones with cache_input pricing.
6. Frontier override
- User asked for "best", "frontier", "most intelligent", "smartest"
→ resolve trait `most_intelligent` from the snapshot. Skip cost-min step.
7. Cost minimization
- From surviving candidates, pick the smallest cost tier (XS → S → M → L → Frontier).
Tie-break by lower output $/1M, then lower input $/1M.
8. Sanity filters (apply throughout)
- Drop model_spec.beta === true (snapshot: beta_access) unless your key has
beta access (such models are only listed to beta-access keys anyway).
betaModel === true (snapshot: beta_status) marks callable beta-status
models that may change or disappear; don't drop them, prefer non-beta when tied.
- Drop model_spec.offline === true.
- Drop candidates whose model_spec.regionRestrictions lists the caller's country (403 otherwise).
- Prefer models without model_spec.deprecation.When the prompt maps cleanly to a named trait, skip the matrix and resolve the trait from the snapshot's traits block (sourced from GET /models/traits?type=text). Trait names also work directly as the model value in a request.
| Trait | Use for | Resolves to (2026-10-02) |
|---|---|---|
default | Generic chat / catch-all. | zai-org-glm-5-2 |
function_calling_default | Agent loops, tool use. | zai-org-glm-5-2 |
default_reasoning | "Think step by step" without specifying a model. | kimi-k3 |
default_code | Code generation / refactor / review. | deepseek-v4-pro-0813 |
default_vision | Vision input, no other special needs. | qwen-3-8-27b |
most_intelligent | "Best available", frontier override. | grok-4-7 |
most_uncensored | Minimal filtering / red-team / creative writing. | venice-uncensored-1-2 |
The mapping changes over time — always read it from the snapshot or the endpoint. Only the keys the endpoint returns exist (there is currently no fastest trait). Cache the resolved trait → ID map at session start (one HTTP call) and reuse.
For a local agent driving Venice as an "escalation backend":
Score the prompt cheaply (locally):
--model frontier, --privacy e2ee).Decide local vs Venice:
Run the decision tree above to pick a Venice model.
Call via venice-chat:
curl https://api.venice.ai/api/v1/chat/completions \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<chosen id>",
"messages": [...],
"venice_parameters": {"include_venice_system_prompt": false}
}'Cap blast radius: log the chosen model + estimated cost before sending; refuse to escalate beyond the user's --max-cost ceiling. When present, the response's cost field reports what was actually charged.
Example 1 — local-first wins
Prompt: "Summarize this email in one sentence: …" (~200 tokens)
Example 2 — privacy + reasoning
Prompt: "Here's our customer churn dataset (PII). Reason about which factors drive churn."
privacy: "private".supportsReasoning: true.e2ee-* model in TEE-only mode if the user also wants attestation.Example 3 — multi-image vision
Prompt: 3 product photos + "Compare these for build quality."
supportsVision && supportsMultipleImages && maxImages >= 3.traits.default_vision.Example 4 — frontier intelligence on a long doc
Prompt: "Give me your absolute best take on this 400-page contract bundle." (~200K tokens)
availableContextTokens below ~200K.most_intelligent); confirm the resolved model still passes step 5.traits.most_intelligent (whatever the snapshot says; grok-4-7 with 500K context as of 2026-10-02).pricing.extended: as of 2026-10-02 grok-4-7 roughly doubles its rates once input exceeds 200,000 tokens, which a ~200K prompt can cross.Example 5 — code agent with tools
Prompt: "Refactor this repo. You have shell + edit tools."
supportsFunctionCalling && optimizedForCode.traits.default_code if it also supports tools, else traits.function_calling_default filtered by optimizedForCode.A scripts/route.py CLI may be added later for runtimes that prefer structured output (route.py --prompt '...' --max-cost 0.001 --need-vision → {"model_id": "...", "estimated_cost": ..., "tier": "..."}). The prose decision tree above remains the source of truth.
Sibling routing skills (venice-image-routing, venice-audio-routing, venice-video-routing) can mirror this layout when needed.
404 on a model ID.privacy, supportsTeeAttestation, and supportsE2EE; tee-* IDs are unlisted legacy aliases.most_uncensored / model_spec.uncensored and privacy: "private" are independent axes.most_intelligent is not the most expensive. It is a curated pick and can sit in a mid cost tier.enable_e2ee defaults to true on E2EE-capable models when E2EE headers are present — see venice-chat. The routing decision selects the model; the chat skill drives the handshake.model_spec.regionRestrictions[] (only present on restricted models) lists blocked countries → 403 for requests from those countries.type. Always pass ?type=text. Don't reuse image traits.venice-models — /models, /models/traits, /models/compatibility_mapping (the discovery API this skill consumes).venice-chat — /chat/completions (the call surface this skill picks a model for).venice-responses — /responses (no E2EE).venice-auth — Bearer vs x402 wallet auth.venice-api-keys — modelPrivacy key restrictions.venice-billing — confirming actual spend matched the routing estimate (ADMIN key).venice-errors — 402 / 403 / 422 / 429 handling on routed requests.snapshots/text-routing.json.routing-matrix.md.© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in skills/venice-text-routing of veniceai/skills.
Open the folder on GitHubat commit 5eaeac5
Venice Text Model Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Venice Text Model Routing this skillveniceai/skills | 144 | — | ~5.2k | Automated safety check: Pass | MIT | |
| OpenCodex Proxy Operationslidge-jun/opencodex | 17k | — | ~3.1k | Automated safety check: Pass | MIT | |
| 9Router AI Gateway Setupdecolua/9router | 31k | — | ~744 | Automated safety check: Pass | MIT | |
| 9Router Chat Completionsdecolua/9router | 31k | — | ~635 | Automated safety check: Pass | MIT | |
| Using Ccproxy Inspectorstarbaser/ccproxy | 350 | — | ~2.7k | Automated safety check: Pass | Custom licence | |
| Pinme LLMglitternetwork/pinme | 3.8k | — | ~2.8k | Automated safety check: Pass | MIT |
lidge-jun/opencodex
Operates an opencodex (`ocx`) proxy: finds CLI tasks offline, checks local configuration, and manages accounts, providers, models, routing and usage reports.
decolua/9router
Sets up access to the 9Router AI gateway, an OpenAI-compatible REST endpoint for chat, images, speech, embeddings, web search and web fetch, and indexes its capability skills.
decolua/9router
Sends chat and code-generation requests through a 9Router gateway using OpenAI or Anthropic message formats, with streaming and auto-fallback combos.
starbaser/ccproxy
Operates the ccproxy inspector MITM system for intercepting, inspecting, and transforming LLM API traffic.
glitternetwork/pinme
A skill your agent uses when a PinMe project (Worker TypeScript) needs to call OpenRouter-backed LLM APIs, including models, chat/completions, streaming, or OpenRouter web search.
starbaser/ccproxy
Guides users through ccproxy as an OpenAI-compatible and Anthropic-compatible LLM API server with SDK integration, OAuth authentication, sentinel key substitution, model routing, and troubleshooting.
veniceai/skills
Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.
veniceai/skills
Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.
veniceai/skills
High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.
veniceai/skills
Async music, sound-effect and long-form voice generation via Venice.
veniceai/skills
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.
veniceai/skills
Transcribe audio files to text via POST /audio/transcriptions.
Works with
Categories
Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent. Aimed at a local agent that receives a prompt and must decide whether to answer it locally or escalate to Venice, the skill encodes the routing logic that sits above `venice-chat`, the call surface, and `venice-models`, the discovery API. When escalating, it picks the cheapest model that satisfies the required privacy tier, modality and capabilities.
Venice Text Model Routing fits situations like: choosing a Venice model for a prompt at runtime; building a local-first agent that escalates hard prompts to Venice; requiring a private, TEE or end-to-end encrypted model; picking a model that handles vision, audio or function calling.
Run `npx skills add veniceai/skills --skill venice-text-routing -a claude-code`. Or copy the skill folder (skills/venice-text-routing in veniceai/skills) into .claude/skills/venice-text-routing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add veniceai/skills --skill venice-text-routing -a codex`. Or copy the skill folder (skills/venice-text-routing in veniceai/skills) into .agents/skills/venice-text-routing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-text-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-text-routing, .gemini/skills/venice-text-routing, .github/skills/venice-text-routing and .opencode/skills/venice-text-routing in your project.
Going by SKILL.md and its folder, Venice Text Model Routing needs Python for the scripts in its folder, the command-line tools its instructions call (curl and python) and credentials named VENICE_API_KEY. Our summary lists: Python to run `scripts/refresh_routing.py`; Network access to the Venice API; Optional `VENICE_API_KEY`.
SKILL.md names 2 domains. In commands or code: api.venice.ai; the agent is likely to contact it when it follows the instructions. As links in the text: docs.venice.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Venice Text Model Routing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Venice Text Model Routing: OpenCodex Proxy Operations (lidge-jun/opencodex, 17k stars), 9Router AI Gateway Setup (decolua/9router, 31k stars), 9Router Chat Completions (decolua/9router, 31k stars) and Using Ccproxy Inspector (starbaser/ccproxy, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.
Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.