Agent skill

Venice Embeddings

by veniceai in veniceai/skills

Call POST /embeddings on Venice. An agent skill from veniceai/skills.

MITAuto-check passedAI & LLM Engineering

Install Venice Embeddings

skills CLI
$ npx skills add veniceai/skills --skill venice-embeddings -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install veniceai/skills venice-embeddings --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-embeddings .claude/skills/venice-embeddings && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
venice-embeddings
GitHub stars
144
Token cost
~2.4k tokens
SKILL.md length
892 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Call POST /embeddings on Venice. An agent skill from veniceai/skills.

  • Tasks that involve Embeddings
  • SKILL.md covers Use when, Minimal request, Request schema and Response headers & compression, plus 5 more sections
  • Calls curl; reaches api.venice.ai and api.openai.com; needs VENICE_API_KEY

What it does

Venice Embeddings is an agent skill from veniceai/skills. Call POST /embeddings on Venice. Covers request shape (input, model, encodingformat, dimensions, user), text-only input (token arrays rejected), per-input and batch limits, per-model dimensions/privacy, the text-embedding-ada-002 alias, API-key privacy gating, OpenAI/LangChain compatibility, response compression (gzip/br), and practical usage for retrieval, clustering, and RAG.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Embeddings. It works with OpenAI and LangChain. The repository describes itself as: Agent Skills for the Venice.ai API. One folder per surface area, each with a SKILL.md for agent runtimes (Cursor, Claude, Codex, etc.). The licence is MIT.

When your agent uses it

  • Tasks that involve Embeddings

Example prompts

  • “/venice-embeddings”

Requirements

  • Python 3
  • A credential in VENICE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.venice.ai
    • api.openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VENICE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Venice Embeddings loads about 2.4k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 892 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 892 words, ~2,352 tokens.

Download SKILL.mdSave it as .claude/skills/venice-embeddings/SKILL.md (or your agent's skills folder).
name
venice-embeddings
description
Call POST /embeddings on Venice. Covers request shape (input, model, encoding_format, dimensions, user), text-only input (token arrays rejected), per-input and batch limits, per-model dimensions/privacy, the text-embedding-ada-002 alias, API-key privacy gating, OpenAI/LangChain compatibility, response compression (gzip/br), and practical usage for retrieval, clustering, and RAG.

Venice Embeddings

POST /api/v1/embeddings returns vector embeddings for strings. It's OpenAI-compatible: request and response match https://api.openai.com/v1/embeddings closely enough that the OpenAI SDK works with baseURL: "https://api.venice.ai/api/v1", with one exception: token-ID arrays are not accepted (see below).

Auth: Bearer API key or x402 wallet (SIGN-IN-WITH-X) — see venice-auth.

Use when

  • You're building retrieval / RAG / similarity search.
  • You need text clustering, classification, deduplication, or reranking features.
  • You want embeddings from a model Venice runs as Private (most of the catalog) rather than an anonymized third-party one — check model_spec.privacy on each model.

Text-only: input must be a string or an array of strings. For images, run them through a vision chat model and embed the description.

Minimal request

bash
curl https://api.venice.ai/api/v1/embeddings \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  --compressed \
  -d '{
    "model": "text-embedding-bge-m3",
    "input": "Why is the sky blue?"
  }'
json
{
  "object": "list",
  "model": "text-embedding-bge-m3",
  "data": [
    { "object": "embedding", "index": 0, "embedding": [0.0023, -0.0093, 0.0158, ...] }
  ],
  "usage": { "prompt_tokens": 8, "total_tokens": 8 }
}

Request schema

The body is strict — unknown top-level fields (e.g. input_type, task, truncate) are rejected with 400.

FieldTypeNotes
modelstringRequired. Model ID from GET /models?type=embedding. text-embedding-ada-002 is accepted as an alias for text-embedding-bge-m3 (the response model then reads text-embedding-bge-m3).
inputstring | string[]Required. A non-empty string, or an array of 1–2048 strings. Token-ID arrays (number[] / number[][]) are rejected with 400 — see Gotchas.
encoding_format"float" | "base64"Default "float". "base64" returns each vector as a base64-encoded string — a much smaller payload; decode client-side.
dimensionsinteger ≥ 1Optional. Requested output size. Only honoured by models whose model_spec.supportsCustomDimensions is true; it is forwarded as-is otherwise, and the provider may ignore or reject it.
userstringAccepted for OpenAI compatibility; not used for inference, but it does split the error budget per value (see venice-errors).
Input limits
  • Per item: Venice estimates tokens as ceil(chars / 3.2 × 0.95) and rejects any string over 8192 estimated tokens (≈ 27,500 characters) with 400 "Input text exceeds the maximum token limit of 8192 tokens". This cap is the same for every model, including those whose maxInputTokens is 32768.
  • Model limit: models with a smaller model_spec.maxInputTokens (e.g. 512 for text-embedding-multilingual-e5-large-instruct, 2048 for gemini-embedding-2-preview) enforce that limit themselves; Venice's pre-check does not. Chunk to the model's maxInputTokens.
  • Batch: at most 2048 strings per request. Venice returns one embedding per element, in order, with matching index.

Response headers & compression

Send Accept-Encoding: gzip, br (curl: --compressed; most HTTP clients decode automatically); the response comes back with Content-Encoding set. For large batches this matters — float vectors in JSON are big.

Also returned:

  • x-ratelimit-limit-* / x-ratelimit-remaining-* / x-ratelimit-reset-* (requests, tokens) — see venice-errors.
  • x-venice-balance-usd / x-venice-balance-diem — current balance, when non-zero.
  • X-Balance-Remaining — listed in the spec for x402 callers but not currently set by the server; poll GET /x402/balance/{walletAddress} instead.

Using the OpenAI SDK

ts
import OpenAI from 'openai'

const client = new OpenAI({
  apiKey: process.env.VENICE_API_KEY,
  baseURL: 'https://api.venice.ai/api/v1',
})

const res = await client.embeddings.create({
  model: 'text-embedding-bge-m3',
  input: ['first doc', 'second doc'],
})

const vec0 = res.data[0].embedding

The OpenAI Node SDK asks for base64 when you omit encoding_format and decodes it client-side. Venice forwards encoding_format to the model; if the decoded vectors look wrong, pass encoding_format: 'float' explicitly.

LangChain

LangChain's OpenAIEmbeddings tokenizes input and sends token arrays by default, which Venice rejects. Turn that off:

python
import os
from langchain_openai import OpenAIEmbeddings

emb = OpenAIEmbeddings(
    model="text-embedding-bge-m3",
    base_url="https://api.venice.ai/api/v1",
    api_key=os.environ["VENICE_API_KEY"],
    check_embedding_ctx_length=False,
)

Wrappers that build OpenAIEmbeddings internally without that flag (e.g. gpt-researcher's openai provider) hit the same 400.

Batch-embedding pattern

ts
async function embedBatch(texts: string[], batchSize = 64) {
  const out: number[][] = []
  for (let i = 0; i < texts.length; i += batchSize) {
    const slice = texts.slice(i, i + batchSize)
    const res = await client.embeddings.create({
      model: 'text-embedding-bge-m3',
      input: slice,
      encoding_format: 'float',
    })
    for (const row of res.data) out[i + row.index] = row.embedding
  }
  return out
}
  • Keep each string under the per-item limits above; batch size is capped at 2048.
  • On 429, back off exponentially and halve the batch — see venice-errors.
Show full SKILL.md (392 more words)Show less

Choosing a model

Query GET /models?type=embedding for the current catalog. Each entry's model_spec exposes:

  • embeddingDimensions — native output dimension (e.g. 1024 for text-embedding-bge-m3, 4096 for text-embedding-qwen3-8b).
  • maxInputTokens — the model's per-input token limit.
  • supportsCustomDimensions — present and true only on models that honour dimensions (absent otherwise).
  • privacy — "private" (no retention) or "anonymized" (a third-party model; the request is sent without your identity).
  • pricing.input / pricing.output — { usd, diem } per million tokens.

Representative IDs: text-embedding-bge-m3 (private, 1024-d), text-embedding-qwen3-8b (private, 4096-d, custom dimensions), text-embedding-multilingual-e5-large-instruct (private, 512-token inputs), text-embedding-3-small / text-embedding-3-large (anonymized, OpenAI, custom dimensions), gemini-embedding-2-preview (anonymized, custom dimensions). The list changes — always read it from /models.

Always pin the model ID — cosine distances are not comparable across different embedding models.

Error handling

CodeMeaning
400Missing model, or a validation error (details names the field): token-array input, empty string/array, > 2048 items, item over the 8192-token estimate, unknown field. Non-JSON Content-Type → "'Content-Type' must be 'application/json'". Also model-side rejections (e.g. input over the model's own limit), returned with the model's message when one can be extracted, otherwise a generic "Invalid request parameters…".
401Invalid API key or SIWX signature.
402Insufficient balance or the key's USD/DIEM spend limit reached. Bearer → "Insufficient USD or Diem balance…"; x402 → payment-required body + PAYMENT-REQUIRED header. A request with no credentials at all also gets 402 (x402 discovery challenge), not 401.
403The API key's modelPrivacy is PRIVATE_TEXT or PRIVATE_ONLY and the model is anonymized. Also region / provider restrictions, or API access disabled for the account.
404Unknown model (the message may suggest a close match).
429Rate limited.
500Inference failed; retry with jitter.
503Model temporarily offline; retry later.

Gotchas

  • Token arrays are rejected. input: [101, 2023, ...] or [[101, ...]] returns 400 "Token array inputs are not supported. Pass a string or an array of strings." Send text.
  • API-key privacy applies to embeddings. A key with modelPrivacy: PRIVATE_TEXT (or PRIVATE_ONLY) can only call private embedding models; text-embedding-3-* and gemini-embedding-2-preview return 403. See venice-api-keys.
  • A top-level empty string is rejected with 400, but empty strings inside an array are not pre-checked — filter them out yourself.
  • The request must be Content-Type: application/json; anything else is rejected with 400 "'Content-Type' must be 'application/json'" before auth or validation runs.
  • Whether returned vectors are L2-normalized depends on the model — verify with Math.hypot(...v) ≈ 1 before assuming.
  • For RAG, store model (and dimensions, if set) alongside each vector so you can re-embed on upgrade.

© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/venice-embeddings of veniceai/skills.

Open the folder on GitHubat commit 5eaeac5

Compare with similar skills

Venice Embeddings next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Venice Embeddings compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Venice Embeddings this skillveniceai/skills144—~2.4kAutomated safety check: PassMIT
Llmobs IntegrationDataDog/dd-trace-js837—~1.4kAutomated safety check: PassCustom licence
RAG ArchitectJeffallan/claude-skills12k—~2kAutomated safety check: PassMIT
Langchain RAGlangchain-ai/langchain-skills1.3k—~3.9kAutomated safety check: PassMIT
Sap Cloud SDK AIsecondsky/sap-skills462—~3.2kAutomated safety check: PassGPL-3.0
Sap Cloud SDK AI Pythonsecondsky/sap-skills462—~3.8kAutomated safety check: PassGPL-3.0

Similar skills

  • Llmobs Integration

    DataDog/dd-trace-js

    Official

    A skill your agent uses when adding, debugging, or modifying LLMObs plugins for an LLM library in dd-trace-js.

    837 GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub stars~2k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub stars~3.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Sap Cloud SDK AI

    secondsky/sap-skills

    Integrates SAP Cloud SDK for AI into JavaScript/TypeScript and Java applications.

    462 GitHub stars~3.2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Sap Cloud SDK AI Python

    secondsky/sap-skills

    Integrates the SAP Cloud SDK for AI for Python (sap-ai-sdk-gen, formerly generative-ai-hub-sdk) into Python applications.

    462 GitHub stars~3.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Langchain Embeddings Search

    jeremylongshore/tons-of-skills-marketplace

    Build and query vector stores with LangChain 1.0 without getting burned by flipped score semantics, embedding-dim mismatches, reranker quirks, and chunk-splitter bugs.

    2.8k GitHub stars~2.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from veniceai/skills

All 22 skills in this repo
  • Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.

    144 GitHub stars~5.2k tokensUpdated 4 days ago
    Auto-check passed
  • Venice Models API

    veniceai/skills

    Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.

    144 GitHub stars~4.3k tokensUpdated 4 days ago
    Auto-check passed
  • Venice API Keys

    veniceai/skills

    Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.

    144 GitHub stars~3.8k tokensUpdated 4 days ago
    Auto-check passed
  • Venice API Overview

    veniceai/skills

    High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.

    144 GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check passed
  • Venice Audio Music

    veniceai/skills

    Async music, sound-effect and long-form voice generation via Venice.

    144 GitHub stars~3.1k tokensUpdated 4 days ago
    Auto-check passed
  • Venice Audio Speech

    veniceai/skills

    Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.

    144 GitHub stars~3.6k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Venice Embeddings

What does Venice Embeddings do?

Call POST /embeddings on Venice. An agent skill from veniceai/skills. Venice Embeddings is an agent skill from veniceai/skills. Call POST /embeddings on Venice.

When should I use Venice Embeddings?

Venice Embeddings fits situations like: tasks that involve Embeddings.

How do I install Venice Embeddings in Claude Code?

Run `npx skills add veniceai/skills --skill venice-embeddings -a claude-code`. Or copy the skill folder (skills/venice-embeddings in veniceai/skills) into .claude/skills/venice-embeddings in your project. Claude Code loads it when a task matches its description.

How do I install Venice Embeddings in Codex?

Run `npx skills add veniceai/skills --skill venice-embeddings -a codex`. Or copy the skill folder (skills/venice-embeddings in veniceai/skills) into .agents/skills/venice-embeddings in your project. Codex loads it when a task matches its description.

Can I use Venice Embeddings in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-embeddings -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-embeddings, .gemini/skills/venice-embeddings, .github/skills/venice-embeddings and .opencode/skills/venice-embeddings in your project.

What does Venice Embeddings need to run?

Going by SKILL.md and its folder, Venice Embeddings needs the command-line tools its instructions call (curl) and credentials named VENICE_API_KEY. Our summary lists: Python 3; A credential in VENICE_API_KEY.

Does Venice Embeddings access the network?

SKILL.md names 2 domains. In commands or code: api.venice.ai and api.openai.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Venice Embeddings safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Venice Embeddings use?

Venice Embeddings is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Venice Embeddings use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Venice Embeddings?

Skills that share tags, products or a category with Venice Embeddings: Llmobs Integration (DataDog/dd-trace-js, 837 stars), RAG Architect (Jeffallan/claude-skills, 12k stars), Langchain RAG (langchain-ai/langchain-skills, 1.3k stars) and Sap Cloud SDK AI (secondsky/sap-skills, 462 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Venice Embeddings?

veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.

Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.