Agent skill

Venice Text Model Routing

by veniceai in veniceai/skills

Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.

MITAuto-check passedAI & LLM Engineering

Install Venice Text Model Routing

skills CLI
$ npx skills add veniceai/skills --skill venice-text-routing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install veniceai/skills venice-text-routing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-text-routing .claude/skills/venice-text-routing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
venice-text-routing
GitHub stars
144
Token cost
~5.2k tokens
SKILL.md length
2,039 words
Files
4 (incl. scripts)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.

  • Works in 5 steps: Score the prompt cheaply (locally) → Decide local vs Venice → Run the decision tree above to pick a… → …
  • Choosing a Venice model for a prompt at runtime
  • SKILL.md covers Snapshot freshness, When to load this skill, Privacy tier ladder and Capability filters, plus 8 more sections
  • Runs Python scripts from its folder; calls curl and python; reaches api.venice.ai; needs VENICE_API_KEY

What it does

Aimed at a local agent that receives a prompt and must decide whether to answer it locally or escalate to Venice, the skill encodes the routing logic that sits above `venice-chat`, the call surface, and `venice-models`, the discovery API. When escalating, it picks the cheapest model that satisfies the required privacy tier, modality and capabilities.

Decisions rely on a cached `snapshots/text-routing.json`. If that file is missing or its `snapshot_date` is more than 30 days old, the agent runs `python scripts/refresh_routing.py`, which rebuilds the snapshot and `routing-matrix.md` from Venice's public model endpoints. No key is needed, though `VENICE_API_KEY` can tailor the result to one key. Filters such as uncensored models or log-probability support need a direct read of the models endpoint.

When your agent uses it

  • Choosing a Venice model for a prompt at runtime
  • Building a local-first agent that escalates hard prompts to Venice
  • Requiring a private, TEE or end-to-end encrypted model
  • Picking a model that handles vision, audio or function calling

Example prompts

  • “Which Venice model should handle this long-context code review with reasoning turned on?”
  • “Pick the cheapest private Venice model that accepts image input.”
  • “Refresh the Venice routing snapshot, since it looks out of date.”

Requirements

  • Python to run `scripts/refresh_routing.py`
  • Network access to the Venice API
  • Optional `VENICE_API_KEY`

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Score the prompt cheaply (locally)
  2. Decide local vs Venice
  3. Run the decision tree above to pick a Venice model.
  4. Call via venice-chat
  5. Cap blast radius: log the chosen model + estimated cost before sending; refuse to escalate beyond the user's --max-cost ceiling. When…

What it can do on your machine

Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • curl
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.venice.ai

    Also links to:

    • docs.venice.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VENICE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Venice Text Model Routing loads about 5.2k tokens when it runs. Until then it costs about 149 tokens; SKILL.md has 2,039 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~149
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 2,039 words, ~5,224 tokens.

Download SKILL.mdSave it as .claude/skills/venice-text-routing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
venice-text-routing
description
Route a prompt to the right Venice text model based on privacy tier (anonymized / private / TEE / E2EE), modality (vision / audio / video input), capability (reasoning, reasoning effort, code, function calling, X search, large context, structured output), uncensored needs, and cost. Use when a local agent (Claude Code, Hermes, NanoClaw, Codex CLI, etc.) needs to decide whether to handle a prompt locally or escalate to Venice, and if escalating, which Venice text model to call. Sits one level above venice-models (model discovery) and before venice-chat (the call surface).

Venice Text-Model Routing

This skill encodes the decision logic for "which Venice text model do I call?" — the routing layer that sits above venice-chat (the call surface) and consumes venice-models (the discovery API).

Primary use case: a local agent receives a prompt, decides whether the local model can handle it, and — if not — picks the cheapest Venice model that satisfies the privacy / modality / capability requirements.

Snapshot freshness

Before applying the matrix below, read snapshots/text-routing.json. If the file is missing, or snapshot_date is older than 30 days, run python scripts/refresh_routing.py to regenerate it from GET /models?type=text + GET /models/traits?type=text. Otherwise trust the cached file — do not hit /models on every routing decision.

refresh_routing.py rewrites both snapshots/text-routing.json and routing-matrix.md. Both endpoints are public, so no key is needed; VENICE_API_KEY / --api-key is optional and tailors the result to that key (a modelPrivacy-restricted key sees a filtered catalog). Run it once on first install, then ~monthly (CI nightly is also fine).

In the snapshot, each model's privacy label is e2ee when supportsE2EE is true, tee when only supportsTeeAttestation is true, and otherwise the model's own private / anonymized value (every TEE/E2EE model is private in /models, so treat tee and e2ee as private too); tier is bucketed by input price only; beta_access mirrors model_spec.beta (only listed to beta-access keys, so it is always false in the shipped keyless snapshot) and beta_status mirrors model_spec.betaModel (callable, but may change or disappear).

Per model the snapshot carries only: the 13 capability flags shown in the matrix, max_images, available_context_tokens, max_completion_tokens, input/output price per 1M, beta_access, beta_status, offline, region_restrictions and deprecation_date. Filters that need anything else (uncensored, supportsLogProbs, reasoningEffortOptions, maxVideos, cache_input / extended pricing, replacementModelId) have to read GET /models?type=text.

When to load this skill

  • Picking a Venice text model from a prompt at runtime.
  • Building a local-first agent that escalates to Venice for hard prompts.
  • Deciding privacy tier (anonymized vs private vs TEE vs E2EE).
  • Choosing between a trait shortcut (default_reasoning, most_intelligent, …) and a hand-filtered candidate.

For the chat call surface itself, see venice-chat. For raw discovery of every model field, see venice-models.

Privacy tier ladder

Pick the least restrictive tier that satisfies the request — restricting tier shrinks the candidate pool and often raises cost.

model_spec.privacy has only two values, private and anonymized. TEE and E2EE are capability flags on top of private.

TierSelector (GET /models)GuaranteeUse when
Anonymizedprivacy: "anonymized"Venice hides your identity from the upstream provider, but the provider may still see the promptNon-sensitive workloads; the only path to several closed frontier models (Claude, GPT, Gemini).
Privateprivacy: "private"Zero data retention, contract-enforced: content is processed for inference only and not retainedDefault for user data, business logic, anything you wouldn't paste into a public chatbot.
TEEcapabilities.supportsTeeAttestation: trueRuns inside a hardware Trusted Execution Environment with remote attestation (GET /api/v1/tee/attestation)Regulated data, verifiable/signed inference.
E2EEcapabilities.supportsE2EE: trueTEE + client-side encryption (ECDH secp256k1 → HKDF-SHA256 → AES-256-GCM); Venice relays ciphertext onlyStrongest. Healthcare, legal, secrets. Requires the E2EE flow in venice-chat.

Today every listed TEE model is also E2EE-capable and uses an e2ee-* ID (e.g. e2ee-glm-5-3-p, e2ee-kimi-k3-p, e2ee-qwen3-8-27b). TEE and E2EE use the same model ID: send E2EE headers for E2EE, omit them (or set venice_parameters.enable_e2ee: false) for TEE-only. Legacy tee-* IDs are unlisted aliases of e2ee-* models — don't route on the prefix; use the capability flags.

E2EE trade-offs: on E2EE requests Venice injects no web search, scraping, character, or Venice system prompt, and file parts return 400; the E2EE guide also requires stream: true and lists function calling as unsupported (supportsFunctionCalling describes the model, not E2EE mode). Several E2EE models also have small limits (e.g. e2ee-qwen-2-5-7b-p: 32K context, 4,096 output tokens). E2EE is not supported on /responses; TEE-only is (with enable_e2ee: false).

API keys can carry a modelPrivacy restriction: PRIVATE_TEXT (text, embedding and decision models must be Private, TEE, or E2EE) or PRIVATE_ONLY (every model). Such a key's /models list is already filtered; a disallowed model returns 403 — filter to privacy: "private" for such keys.

Sources: docs.venice.ai/overview/privacy, docs.venice.ai/guides/features/tee-e2ee-models.

Verifying a TEE claim

Two endpoints let you check that inference really ran inside an enclave. Both are GET, unauthenticated on purpose (attestation evidence has to be verifiable by any party without credentials), and rate limited to 10 requests per minute per IP (429 beyond that). Neither is a path in the published OpenAPI spec; they are live, and the supportsTeeAttestation capability description in GET /models points at them.

EndpointQueryReturns
GET /api/v1/tee/attestationmodel (required), nonce (optional, exactly 64 hex chars / 32 bytes, binds the attestation to your challenge)Attestation report, TEE provider, verification result, and the signing key/address.
GET /api/v1/tee/signaturemodel and request_id (required), signing_algo (optional ecdsa | ecdsa-p256 | rsa)The provider's signature over a specific request, plus request/response hashes.

Both return 400 if the model exists but is not TEE-attested, and 404 if the model ID is unknown. The attestation endpoint returns 502 (with verified: false) when verification fails; both return 502 when the TEE provider is unavailable. Chat responses from TEE models carry X-Venice-TEE: true and X-Venice-TEE-Provider.

Verify the chain of trust in this order: fetch the attestation to get the signing public key and hardware type, confirm the recovered signer matches the attestation signing address, then verify the signature over the exact signed text the signature endpoint returned. Treat the request and response hashes as provider-reported values unless you can recompute them yourself from a documented canonical format.

Capability filters

Map prompt requirement → model_spec field (full list in the model_spec.capabilities section of venice-models). Sending image / audio / video parts, tools / tool_choice, a non-text response_format, or logprobs to a model without the matching flag returns 400 before inference; some other features (e.g. enable_x_search) are silently ignored instead.

RequirementFilterNotes
Vision (single image)capabilities.supportsVisionSingle-image models keep images only from the last image-bearing message.
Vision (multiple images)supportsVision && supportsMultipleImagesHonor capabilities.maxImages. Hard cap: 10 images per message.
Audio inputcapabilities.supportsAudioInputBase64 only; URLs are not accepted.
Video inputcapabilities.supportsVideoInputdata:video/... URLs, direct public URLs (no redirects), or YouTube links on some models. Max 3 videos per request.
Documents (PDF, DOCX, …)—Document file parts are extracted to text server-side, so any text model works (not on E2EE). Image files sent as file parts become images and need vision.
Reasoningcapabilities.supportsReasoningTo dial effort, also require supportsReasoningEffort and pick a value from reasoningEffortOptions (default: defaultReasoningEffort).
Tools / function callingcapabilities.supportsFunctionCallingRequired for any agent loop.
Code-heavy taskcapabilities.optimizedForCodeGET /models?type=code returns only this subset.
Web search (Venice)—supportsWebSearch is true on every text model. Toggle with venice_parameters.enable_web_search.
X / Twitter searchcapabilities.supportsXSearchxAI native (Grok models). ~$0.01 per search.
Structured JSON outputcapabilities.supportsResponseSchemaresponse_format: {type: "json_schema", json_schema: {name, schema, strict}}.
Large contextavailableContextTokens >= NPair with prompt_cache_key; prefer models with pricing.cache_input. Watch pricing.extended, which raises prices above context_token_threshold input tokens.
Long outputmaxCompletionTokens >= NRequests above it are rejected on models with an enforced cap.
Minimal content filteringmodel_spec.uncensored: truePresent only on models Venice classifies as uncensored; upstream providers may still filter.
Logprobscapabilities.supportsLogProbsNiche — eval / sampling debug.

Cost tiers

These buckets are this skill's own convention (the same boundaries refresh_routing.py uses), keyed on model_spec.pricing.input.usd per 1M tokens. Upper bounds are exclusive (a model at exactly $1.00 is in M), and the snapshot's tier_boundaries_usd_per_1m_input gives each bucket's exclusive upper bound. Pick the smallest bucket that hosts a model satisfying your filters.

Tier$/1M inputUse when
XS< $0.20Classification, intent extraction, simple summarization.
S$0.20 – < $1General chat, basic agents, light vision.
M$1 – < $4Moderate reasoning, strong code, multi-image vision.
L$4 – < $10Heavy reasoning, complex tool use.
Frontier≥ $10Most expensive closed models (e.g. claude-fable-5, openai-gpt-6-astra, openai-gpt-54-pro as of 2026-10-02).

Output prices vary widely within a bucket (from ~1.5× to ~10× input), so tie-break on output price. Authoritative per-model pricing lives in model_spec.pricing (input and output rates are mirrored as pricing_per_1m in the snapshot) — never hard-code dollar figures from this prose.

Show full SKILL.md (753 more words)Show less

Routing decision tree

Walk top-down. Stop at the first rule that applies.

1. Local-first check
   - Prompt is ≤ ~500 tokens, no special-capability requirement,
     no privacy escalation, no tool calls expected
     → handle on the local model. Do not call Venice.

2. Privacy gate
   - User flagged "private" OR prompt contains regulated data (PHI, secrets, legal),
     OR the API key has modelPrivacy PRIVATE_TEXT / PRIVATE_ONLY:
       require model_spec.privacy === "private"   (includes every TEE/E2EE model;
       in the snapshot accept privacy "private", "tee" or "e2ee").
   - User flagged "E2EE" / "must be encrypted end to end":
       require capabilities.supportsE2EE. Use /chat/completions only.
   - User flagged "TEE" / "verifiable inference":
       require capabilities.supportsTeeAttestation.
   - Otherwise: any tier acceptable (still prefer "private" when otherwise tied).

3. Modality gate
   - Image input present  → require supportsVision (+ supportsMultipleImages if > 1).
   - Audio input present  → require supportsAudioInput.
   - Video input present  → require supportsVideoInput.

4. Capability gate
   - Tool calls expected               → require supportsFunctionCalling.
   - Code-heavy task                   → prefer optimizedForCode (or call /models?type=code).
   - Chain-of-thought / planning / hard math
                                       → require supportsReasoning;
                                         prefer supportsReasoningEffort to dial reasoning effort.
   - Structured JSON output required   → require supportsResponseSchema.
   - X/Twitter content needed          → require supportsXSearch.
   - Refusal-free / minimal filtering  → prefer model_spec.uncensored (or trait most_uncensored).

5. Context size gate
   - Estimated prompt tokens > availableContextTokens, or expected output > maxCompletionTokens
     → bump up to a model with sufficient limits. Prefer ones with cache_input pricing.

6. Frontier override
   - User asked for "best", "frontier", "most intelligent", "smartest"
     → resolve trait `most_intelligent` from the snapshot. Skip cost-min step.

7. Cost minimization
   - From surviving candidates, pick the smallest cost tier (XS → S → M → L → Frontier).
     Tie-break by lower output $/1M, then lower input $/1M.

8. Sanity filters (apply throughout)
   - Drop model_spec.beta === true (snapshot: beta_access) unless your key has
     beta access (such models are only listed to beta-access keys anyway).
     betaModel === true (snapshot: beta_status) marks callable beta-status
     models that may change or disappear; don't drop them, prefer non-beta when tied.
   - Drop model_spec.offline === true.
   - Drop candidates whose model_spec.regionRestrictions lists the caller's country (403 otherwise).
   - Prefer models without model_spec.deprecation.

Trait shortcuts

When the prompt maps cleanly to a named trait, skip the matrix and resolve the trait from the snapshot's traits block (sourced from GET /models/traits?type=text). Trait names also work directly as the model value in a request.

TraitUse forResolves to (2026-10-02)
defaultGeneric chat / catch-all.zai-org-glm-5-2
function_calling_defaultAgent loops, tool use.zai-org-glm-5-2
default_reasoning"Think step by step" without specifying a model.kimi-k3
default_codeCode generation / refactor / review.deepseek-v4-pro-0813
default_visionVision input, no other special needs.qwen-3-8-27b
most_intelligent"Best available", frontier override.grok-4-7
most_uncensoredMinimal filtering / red-team / creative writing.venice-uncensored-1-2

The mapping changes over time — always read it from the snapshot or the endpoint. Only the keys the endpoint returns exist (there is currently no fastest trait). Cache the resolved trait → ID map at session start (one HTTP call) and reuse.

Local-first pattern

For a local agent driving Venice as an "escalation backend":

  1. Score the prompt cheaply (locally):

    • Token count of prompt + expected output.
    • Modality signals (any image / audio / video parts).
    • Keyword signals: "code", "reason", "step by step", "private", "secret", "encrypt", "best", "frontier".
    • User-provided overrides (e.g. --model frontier, --privacy e2ee).
  2. Decide local vs Venice:

    • If estimated tokens ≤ local model's comfort window AND no capability or privacy escalation → stay local.
    • Otherwise → continue to step 3.
  3. Run the decision tree above to pick a Venice model.

  4. Call via venice-chat:

    bash
    curl https://api.venice.ai/api/v1/chat/completions \
      -H "Authorization: Bearer $VENICE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "<chosen id>",
        "messages": [...],
        "venice_parameters": {"include_venice_system_prompt": false}
      }'
  5. Cap blast radius: log the chosen model + estimated cost before sending; refuse to escalate beyond the user's --max-cost ceiling. When present, the response's cost field reports what was actually charged.

Examples

Example 1 — local-first wins

Prompt: "Summarize this email in one sentence: …" (~200 tokens)

  • Step 1 → local handles it. Do not call Venice.

Example 2 — privacy + reasoning

Prompt: "Here's our customer churn dataset (PII). Reason about which factors drive churn."

  • Step 2 → privacy gate requires privacy: "private".
  • Step 4 → reasoning required → supportsReasoning: true.
  • Step 7 → among private candidates with reasoning, pick the cheapest.
  • → a private reasoning model, or an e2ee-* model in TEE-only mode if the user also wants attestation.

Example 3 — multi-image vision

Prompt: 3 product photos + "Compare these for build quality."

  • Step 3 → supportsVision && supportsMultipleImages && maxImages >= 3.
  • Step 7 → cheapest survivor, or traits.default_vision.

Example 4 — frontier intelligence on a long doc

Prompt: "Give me your absolute best take on this 400-page contract bundle." (~200K tokens)

  • Step 5 → context gate drops models with availableContextTokens below ~200K.
  • Step 6 → frontier override fires (most_intelligent); confirm the resolved model still passes step 5.
  • → traits.most_intelligent (whatever the snapshot says; grok-4-7 with 500K context as of 2026-10-02).
  • Estimate cost with pricing.extended: as of 2026-10-02 grok-4-7 roughly doubles its rates once input exceeds 200,000 tokens, which a ~200K prompt can cross.

Example 5 — code agent with tools

Prompt: "Refactor this repo. You have shell + edit tools."

  • Step 4 → supportsFunctionCalling && optimizedForCode.
  • Step 7 → cheapest survivor → traits.default_code if it also supports tools, else traits.function_calling_default filtered by optimizedForCode.

Future extensions

A scripts/route.py CLI may be added later for runtimes that prefer structured output (route.py --prompt '...' --max-cost 0.001 --need-vision → {"model_id": "...", "estimated_cost": ..., "tier": "..."}). The prose decision tree above remains the source of truth.

Sibling routing skills (venice-image-routing, venice-audio-routing, venice-video-routing) can mirror this layout when needed.

Gotchas

  • Don't hard-code model IDs. Venice adds and retires models frequently — resolve via traits or filters against the snapshot.
  • Stale snapshot lies silently. 30 days is the maximum age; refresh sooner if you hit a 404 on a model ID.
  • Don't route on ID prefixes. Use privacy, supportsTeeAttestation, and supportsE2EE; tee-* IDs are unlisted legacy aliases.
  • Privacy ≠ uncensored. most_uncensored / model_spec.uncensored and privacy: "private" are independent axes.
  • most_intelligent is not the most expensive. It is a curated pick and can sit in a mid cost tier.
  • enable_e2ee defaults to true on E2EE-capable models when E2EE headers are present — see venice-chat. The routing decision selects the model; the chat skill drives the handshake.
  • Region restrictions. model_spec.regionRestrictions[] (only present on restricted models) lists blocked countries → 403 for requests from those countries.
  • Trait keys differ by type. Always pass ?type=text. Don't reuse image traits.

See also

© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/venice-text-routing of veniceai/skills.

  • SKILL.md
  • routing-matrix.md
  • scripts/refresh_routing.py
  • snapshots/text-routing.json

Open the folder on GitHubat commit 5eaeac5

Compare with similar skills

Venice Text Model Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Venice Text Model Routing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Venice Text Model Routing this skillveniceai/skills144—~5.2kAutomated safety check: PassMIT
OpenCodex Proxy Operationslidge-jun/opencodex17k—~3.1kAutomated safety check: PassMIT
9Router AI Gateway Setupdecolua/9router31k—~744Automated safety check: PassMIT
9Router Chat Completionsdecolua/9router31k—~635Automated safety check: PassMIT
Using Ccproxy Inspectorstarbaser/ccproxy350—~2.7kAutomated safety check: PassCustom licence
Pinme LLMglitternetwork/pinme3.8k—~2.8kAutomated safety check: PassMIT

Similar skills

  • OpenCodex Proxy Operations

    lidge-jun/opencodex

    Operates an opencodex (`ocx`) proxy: finds CLI tasks offline, checks local configuration, and manages accounts, providers, models, routing and usage reports.

    17k GitHub stars~3.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Sets up access to the 9Router AI gateway, an OpenAI-compatible REST endpoint for chat, images, speech, embeddings, web search and web fetch, and indexes its capability skills.

    31k GitHub stars~744 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Sends chat and code-generation requests through a 9Router gateway using OpenAI or Anthropic message formats, with streaming and auto-fallback combos.

    31k GitHub stars~635 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy Inspector

    starbaser/ccproxy

    Operates the ccproxy inspector MITM system for intercepting, inspecting, and transforming LLM API traffic.

    350 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Pinme LLM

    glitternetwork/pinme

    A skill your agent uses when a PinMe project (Worker TypeScript) needs to call OpenRouter-backed LLM APIs, including models, chat/completions, streaming, or OpenRouter web search.

    3.8k GitHub stars~2.8k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy API

    starbaser/ccproxy

    Guides users through ccproxy as an OpenAI-compatible and Anthropic-compatible LLM API server with SDK integration, OAuth authentication, sentinel key substitution, model routing, and troubleshooting.

    350 GitHub stars~4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from veniceai/skills

All 22 skills in this repo
  • Venice Models API

    veniceai/skills

    Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.

    144 GitHub stars~4.3k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Keys

    veniceai/skills

    Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.

    144 GitHub stars~3.8k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Overview

    veniceai/skills

    High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.

    144 GitHub stars~3.5k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Audio Music

    veniceai/skills

    Async music, sound-effect and long-form voice generation via Venice.

    144 GitHub stars~3.1k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Audio Speech

    veniceai/skills

    Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.

    144 GitHub stars~3.6k tokensUpdated 5 days ago
    Auto-check passed
  • Transcribe audio files to text via POST /audio/transcriptions.

    144 GitHub stars~1.7k tokensUpdated 5 days ago
    Auto-check passed

Works with

Questions about Venice Text Model Routing

What does Venice Text Model Routing do?

Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent. Aimed at a local agent that receives a prompt and must decide whether to answer it locally or escalate to Venice, the skill encodes the routing logic that sits above `venice-chat`, the call surface, and `venice-models`, the discovery API. When escalating, it picks the cheapest model that satisfies the required privacy tier, modality and capabilities.

When should I use Venice Text Model Routing?

Venice Text Model Routing fits situations like: choosing a Venice model for a prompt at runtime; building a local-first agent that escalates hard prompts to Venice; requiring a private, TEE or end-to-end encrypted model; picking a model that handles vision, audio or function calling.

How do I install Venice Text Model Routing in Claude Code?

Run `npx skills add veniceai/skills --skill venice-text-routing -a claude-code`. Or copy the skill folder (skills/venice-text-routing in veniceai/skills) into .claude/skills/venice-text-routing in your project. Claude Code loads it when a task matches its description.

How do I install Venice Text Model Routing in Codex?

Run `npx skills add veniceai/skills --skill venice-text-routing -a codex`. Or copy the skill folder (skills/venice-text-routing in veniceai/skills) into .agents/skills/venice-text-routing in your project. Codex loads it when a task matches its description.

Can I use Venice Text Model Routing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-text-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-text-routing, .gemini/skills/venice-text-routing, .github/skills/venice-text-routing and .opencode/skills/venice-text-routing in your project.

What does Venice Text Model Routing need to run?

Going by SKILL.md and its folder, Venice Text Model Routing needs Python for the scripts in its folder, the command-line tools its instructions call (curl and python) and credentials named VENICE_API_KEY. Our summary lists: Python to run `scripts/refresh_routing.py`; Network access to the Venice API; Optional `VENICE_API_KEY`.

Does Venice Text Model Routing access the network?

SKILL.md names 2 domains. In commands or code: api.venice.ai; the agent is likely to contact it when it follows the instructions. As links in the text: docs.venice.ai. This is read from the text; nothing was executed.

Is Venice Text Model Routing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Venice Text Model Routing use?

Venice Text Model Routing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Venice Text Model Routing use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Venice Text Model Routing?

Skills that share tags, products or a category with Venice Text Model Routing: OpenCodex Proxy Operations (lidge-jun/opencodex, 17k stars), 9Router AI Gateway Setup (decolua/9router, 31k stars), 9Router Chat Completions (decolua/9router, 31k stars) and Using Ccproxy Inspector (starbaser/ccproxy, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Venice Text Model Routing?

veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.

Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.