Agent skill

Venice Responses

by veniceai in veniceai/skills

Use Venice's Alpha POST /responses endpoint - an OpenAI-compatible, stateless Responses API with typed output blocks (reasoning, message, functioncall, websearchcall).

MITAuto-check passedAI & LLM Engineering

Install Venice Responses

skills CLI
$ npx skills add veniceai/skills --skill venice-responses -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install veniceai/skills venice-responses --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/veniceai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/venice-responses .claude/skills/venice-responses && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
venice-responses
GitHub stars
144
Token cost
~3.7k tokens
SKILL.md length
1,310 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Use Venice's Alpha POST /responses endpoint - an OpenAI-compatible, stateless Responses API with typed output blocks (reasoning, message, functioncall, websearchcall).

  • Tasks that involve LLM API integration
  • SKILL.md covers Use when, Limitations vs /chat/completions, Authentication and Minimal request, plus 7 more sections
  • Calls curl; reaches api.venice.ai; needs VENICE_API_KEY
  • Tasks that involve Web search

What it does

Venice Responses is an agent skill from veniceai/skills. Use Venice's Alpha POST /responses endpoint - an OpenAI-compatible, stateless Responses API with typed output blocks (reasoning, message, functioncall, websearchcall). Covers request shape, input items, tools (function, websearch, xsearch), reasoning controls, incomplete responses, streaming events, differences from /chat/completions, the supported veniceparameters subset, and E2EE behavior.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM API integration and Web search. It works with OpenAI. The repository describes itself as: Agent Skills for the Venice.ai API. One folder per surface area, each with a SKILL.md for agent runtimes (Cursor, Claude, Codex, etc.). The licence is MIT.

When your agent uses it

  • Tasks that involve LLM API integration
  • Tasks that involve Web search

Example prompts

  • “/venice-responses”

Requirements

  • Python 3
  • A credential in VENICE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 5eaeac5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.venice.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VENICE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Venice Responses loads about 3.7k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 1,310 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from veniceai/skills at commit 5eaeac5, republished under its MIT licence (© veniceai). 1,310 words, ~3,688 tokens.

Download SKILL.mdSave it as .claude/skills/venice-responses/SKILL.md (or your agent's skills folder).
name
venice-responses
description
Use Venice's Alpha POST /responses endpoint - an OpenAI-compatible, stateless Responses API with typed output blocks (reasoning, message, function_call, web_search_call). Covers request shape, input items, tools (function, web_search, x_search), reasoning controls, incomplete responses, streaming events, differences from /chat/completions, the supported venice_parameters subset, and E2EE behavior.

Venice Responses API (Alpha)

POST /api/v1/responses is Venice's OpenAI-compatible Responses endpoint. It returns a typed output array instead of a single message.content string — useful for agents that need to separate reasoning, messages, tool calls, and web-search events. Internally the request is translated to a chat completion, so model support matches venice-chat.

Alpha. The spec labels it Alpha (and its description still says "Alpha testers only"), but access is no longer restricted: any Bearer API key or x402 wallet can call it. Schemas may still change.

Use when

  • A client library expects the OpenAI Responses shape (output[] with type: "reasoning" | "message" | "function_call" | "web_search_call").
  • You want reasoning, message, and tool-call output cleanly separated.
  • You want SSE streaming with typed events.

Otherwise use venice-chat — it has structured output, audio/video/file inputs, E2EE, sampling controls, and every venice_parameters field.

Limitations vs /chat/completions

LimitationDetail
StatelessNothing is stored. Send the full history each call. previous_response_id, store, background are ignored.
No E2EEE2EE-capable models return 400 unless venice_parameters.enable_e2ee: false (TEE-only mode). For encrypted inference use /chat/completions.
Text + image input onlyinput_text / input_image (and text / image_url parts). No audio, video, or file parts.
No structured outputtext.format / response_format are dropped. Use /chat/completions.
Subset of venice_parameterscharacter_slug, enable_e2ee, enable_web_search, enable_web_scraping, enable_web_citations, include_venice_system_prompt, include_search_results_in_stream. Other keys (strip_thinking_response, disable_thinking, enable_x_search, return_search_results_as_documents) are silently dropped.
No model feature suffixesmodel: "zai-org-glm-5-1:enable_web_search=on" resolves the model but ignores the suffix.
Few generation controlsOnly temperature, top_p, max_output_tokens.

Authentication

Same as the rest of the API — Authorization: Bearer <key> or SIGN-IN-WITH-X: <SIWX> for x402 wallets. See venice-auth.

Minimal request

bash
curl https://api.venice.ai/api/v1/responses \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org-glm-5-1",
    "input": "Explain why the sky is blue in one paragraph."
  }'

OpenAI SDK (Python):

python
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["VENICE_API_KEY"], base_url="https://api.venice.ai/api/v1")
resp = client.responses.create(model="zai-org-glm-5-1", input="Explain why the sky is blue.")
print(resp.output_text)

Request fields

FieldNotes
modelRequired. Model ID, trait, or compatibility mapping.
inputRequired. A string, or an array of input items (below).
max_output_tokensPositive integer. Mapped to max_tokens; above the model's model_spec.maxCompletionTokens → 400 on models with an enforced cap.
temperature (0–2), top_p (0–1)Sampling.
reasoning.effortnone | minimal | low | medium | high | xhigh | max (per-model support: model_spec.capabilities.reasoningEffortOptions). reasoning may be null.
reasoning.enabledfalse disables reasoning on supported models and suppresses reasoning blocks. Ignored when an effort is set.
reasoning.summaryauto | concise | detailed. Accepted but not forwarded.
toolsSee Tools.
tool_choice"auto" | "none" | "required" | {"type":"function","function":{"name":"..."}}. Dropped when no function tools remain.
web_searchBoolean. true forces web search on (same as a {"type":"web_search"} tool).
includeOnly "reasoning.encrypted_content" has an effect (adds encrypted_content to reasoning blocks when the provider returns it).
streamBoolean. SSE with typed events.
anon_user_idOptional end-user id: 1–128 printable ASCII characters, no `
fallbacksUp to 10 {model} entries. Anthropic beta refusal fallback for Claude Fable 5; forwarded only on direct Anthropic routes.
venice_parametersSubset listed above. Example: {"character_slug":"alan-watts","enable_web_search":"auto"}.

The body is permissive: other fields (instructions, metadata, parallel_tool_calls, n, stop, seed, prompt_cache_key, store, previous_response_id, background, text, user) are accepted without error but never reach inference (user still splits the error budget per value). Put system instructions in the input array instead of instructions.

Input items
ItemShapeHandling
Message{role, content} or {type:"message", role, content}role: user / assistant / system / developer (developer becomes system). content is a string or an array of parts.
Function call{type:"function_call", call_id, name, arguments}Replayed as an assistant tool call.
Function output{type:"function_call_output", call_id, output}output may be a string, array, object, number, boolean, or null. input_image parts inside an array output are forwarded to the model as images.
Reasoning{type:"reasoning", ...}Accepted but discarded — reasoning is not carried between turns.
Item reference{type:"item_reference", id}Accepted but discarded (nothing is stored to reference).

Content parts: input_text, output_text (to replay assistant output), and input_image. input_image.image_url may be a URL string (OpenAI Responses style) or {url, detail}; detail (auto / low / high) may also sit on the part. Messages without type additionally accept Chat-style text and image_url parts. Image URLs get the same validation as on /chat/completions (public, no redirects, ≥ 64 px); failures → 400. Use a vision model: message images are not capability-checked on this endpoint (images inside a function_call_output on a non-vision model do return 400).

Tools
ToolEffect
{"type":"function","function":{name, description, parameters, strict}}Function calling. The flat OpenAI form {"type":"function","name":...,"parameters":...} is also accepted. Use a model with supportsFunctionCalling (not pre-checked on this endpoint, unlike chat).
{"type":"web_search"}Forces Venice web search on (not auto). search_context_size / user_location are accepted but ignored.
{"type":"x_search", ...}xAI native web + X search on models with supportsXSearch (Grok); ignored on other models. Optional filters: allowed_x_handles / excluded_x_handles (≤ 10 each), from_date, to_date, enable_image_understanding, enable_video_understanding.
code_interpreter, file_search, computer_use_preview, othersAccepted and dropped. Unknown tool types that carry a name are treated as function tools.

Response shape

json
{
  "id": "resp_chatcmpl-abc123",
  "object": "response",
  "created_at": 1735689600,
  "model": "zai-org-glm-5-1",
  "status": "completed",
  "output": [
    {"type": "reasoning", "id": "rs_1", "summary": ["I considered Rayleigh scattering..."]},
    {"type": "web_search_call", "id": "ws_1", "status": "completed"},
    {"type": "function_call", "id": "fc_1", "call_id": "call_abc", "name": "get_weather",
     "arguments": "{\"city\":\"Paris\"}", "status": "completed"},
    {"type": "message", "id": "msg_1", "status": "completed", "role": "assistant",
     "content": [{"type": "output_text", "text": "The sky is blue because... ^1^",
       "annotations": [{"type": "url_citation", "url": "https://example.com/rayleigh",
         "title": "Rayleigh scattering", "start_index": 27, "end_index": 30}]}]}
  ],
  "usage": {
    "input_tokens": 20,
    "input_tokens_details": {"cached_tokens": 8},
    "output_tokens": 80,
    "output_tokens_details": {"reasoning_tokens": 40},
    "total_tokens": 100
  }
}
  • Non-streamed output order: reasoning → web_search_call → function_call(s) → message. The message block is omitted when the model returned only tool calls with no text.
  • status is completed or incomplete. Errors before or during inference come back as HTTP errors (streaming uses response.failed).
  • Incomplete responses keep their output and usage. When generation stops on max_output_tokens or a content filter, status: "incomplete", incomplete_details: {"reason": "max_output_tokens" | "content_filter"}, and message / function_call blocks carry status: "incomplete".
  • input_tokens_details appears only when cached tokens are non-zero; output_tokens_details only when the provider reports reasoning tokens. There is no cost field (unlike /chat/completions).
Show full SKILL.md (482 more words)Show less

Output block types

typePurpose
reasoningReasoning from thinking models. summary[] holds text; encrypted_content appears only if you sent include: ["reasoning.encrypted_content"] and the provider returned encrypted reasoning. Sending it back in input has no effect.
messageMain text. content[].type === "output_text" with annotations[].
function_callTool call: name, JSON-string arguments, call_id. Answer with a function_call_output item with the same call_id.
web_search_callMarker that Venice web search ran.

url_citation annotations are built only when the text contains single-index ^n^ markers — set venice_parameters.enable_web_citations: true to get them. Each annotation spans the marker itself; multi-index markers such as ^1,3^ are not annotated.

Streaming

With stream: true, events are event: <type> + data: {...} pairs; payloads carry type (equal to the event name) and an increasing sequence_number — except response.web_search.done, whose payload has type: "web_search_call", id, status, results and no sequence_number. Typical flow:

event: response.created                      # status: in_progress
event: response.web_search.done              # Venice-specific; only when search ran, carries results[{index,url,title,snippet}]
event: response.output_item.added            # item.type = reasoning
event: response.reasoning.delta
event: response.output_item.added            # item.type = message
event: response.content_part.added
event: response.output_text.delta            # repeated
event: response.output_item.added            # item.type = function_call
event: response.function_call_arguments.delta
event: response.output_item.done             # reasoning, then message (after content_part.done), then each function_call
event: response.completed                    # or response.incomplete, with the full response
data: [DONE]
  • On an upstream failure you get response.failed (response.status: "failed", response.error: {code, message}) followed by data: [DONE].
  • The final response.completed / response.incomplete payload differs slightly from a non-streamed response: annotations are always empty and function calls come after the message.
  • include_search_results_in_stream has no effect here; search results always arrive in response.web_search.done.

Errors

StatusWhen
400Invalid body, E2EE-capable model without enable_e2ee: false, invalid image, unsupported reasoning.effort for the model, context too long, max_output_tokens over the cap
401Invalid API key or SIWX sign-in; also a model that requires a paid subscription
402No credentials at all (x402 discovery body — not 401), insufficient balance, or API-key spend limit. x402: PAYMENT_REQUIRED body with topUpInstructions + siwxChallenge and a PAYMENT-REQUIRED header (see venice-x402)
403Model blocked by the key's modelPrivacy, region, or provider restriction
404Unknown model
422Content-policy violation on an input image
429Rate limited
500 / 503Inference failed (upstream overloads and timeouts also surface as 500 here, not 429 / 504; retry with backoff) / model offline

The spec lists an X-Balance-Remaining header on x402 200 responses, but the server does not currently set it — poll GET /x402/balance/{walletAddress} instead. See venice-errors.

Migration notes (from /chat/completions)

  • messages → input (the same role/content objects work; system prompts go in as role: "system" or "developer" items).
  • max_tokens → max_output_tokens; reasoning_effort → reasoning.effort.
  • Tool results → function_call_output items keyed by call_id.
  • venice_parameters.character_slug, enable_web_search, enable_web_citations, enable_web_scraping, include_venice_system_prompt → pass inside venice_parameters (not as model suffixes).
  • enable_x_search → add an {"type":"x_search"} tool instead.
  • strip_thinking_response / disable_thinking → use reasoning.enabled: false.
  • Structured output, audio / video / file inputs, seed / stop / n, logprobs, prompt-cache routing, and full E2EE → stay on /chat/completions.

Gotchas

  • Unknown character_slug is not rejected here (chat returns 404). The request runs without the character, without your system messages, and without the Venice system prompt or web search. Validate slugs first with GET /characters/{slug} (venice-characters; Bearer key only — wallet callers can't, and should use /chat/completions, which returns 404 for unknown slugs).
  • Reasoning items in input are discarded; there is no cross-turn reasoning carry-over on this endpoint.
  • tool_choice objects must be {"type":"function","function":{"name":...}}; the flat {"type":"function","name":...} form fails validation.
  • Stateless: previous_response_id is silently ignored, so omitting history silently loses context.

© veniceai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/venice-responses of veniceai/skills.

Open the folder on GitHubat commit 5eaeac5

Compare with similar skills

Venice Responses next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Venice Responses compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Venice Responses this skillveniceai/skills144—~3.7kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~833Automated safety check: PassApache-2.0
Brave Answers APIbrave/brave-search-skills183—~2.3kAutomated safety check: PassMIT
Openai Docstheowenyoung/home1151 repos~1.4kAutomated safety check: PassApache-2.0
Chatgpt Web Researchbear2u/my-skills932—~3.3kAutomated safety check: PassNone
Yichen Chatgpt Web Researchmcncarl/yichen-skills4.4k—~3.4kAutomated safety check: PassCustom licence

Similar skills

  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~833 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Brave Answers API

    brave/brave-search-skills

    Calls the Brave Search Answers endpoint for AI-grounded, cited answers, either a fast single-search reply or a slower multi-search deep research run.

    183 GitHub stars~2.3k tokensUpdated 18 days ago
    AI & LLM EngineeringAuto-check passed
  • Openai Docs

    theowenyoung/home

    A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or…

    115 GitHub starsUsed in 1 repo~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Chatgpt Web Research

    bear2u/my-skills

    Use the user's already signed-in official ChatGPT website account, especially GPT-5.5 Pro / ChatGPT Pro, to perform product research, market research, competitor research, second-opinion research…

    932 GitHub stars~3.3k tokensUpdated 3 mo ago
    Productivity & AutomationAuto-check passed
  • Yichen Chatgpt Web Research

    mcncarl/yichen-skills

    Use the user's already signed-in official ChatGPT website account, especially GPT-5.5 Pro / ChatGPT Pro, to perform product research, market research, competitor research, second-opinion analysis…

    4.4k GitHub stars~3.4k tokensUpdated 6 days ago
    Productivity & AutomationAuto-check passed
  • Agnes Free Text

    kangarooking/agnes-free-model-skills

    Call the free Agnes text model API for chat completions, streaming answers, coding help, tool-calling experiments, and OpenAI-compatible text generation.

    199 GitHub stars~630 tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from veniceai/skills

All 22 skills in this repo
  • Picks which Venice text model to call for a prompt based on privacy tier, input modality, capabilities and cost, and decides when to escalate from a local agent.

    144 GitHub stars~5.2k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Models API

    veniceai/skills

    Documents Venice's model discovery endpoints, GET /models, /models/traits and /models/compatibility_mapping, so an agent can pick a model by capability, constraint or price.

    144 GitHub stars~4.3k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Keys

    veniceai/skills

    Manages Venice API keys through the /api_keys endpoints: create, list, update and revoke keys, set spending limits, and read rate limits.

    144 GitHub stars~3.8k tokensUpdated 5 days ago
    Auto-check passed
  • Venice API Overview

    veniceai/skills

    High-level map of the Venice.ai API: base URL, auth modes per endpoint, endpoint categories, response headers, pricing model, error shape and versioning.

    144 GitHub stars~3.5k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Audio Music

    veniceai/skills

    Async music, sound-effect and long-form voice generation via Venice.

    144 GitHub stars~3.1k tokensUpdated 5 days ago
    Auto-check passed
  • Venice Audio Speech

    veniceai/skills

    Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices.

    144 GitHub stars~3.6k tokensUpdated 5 days ago
    Auto-check passed

Works with

Questions about Venice Responses

What does Venice Responses do?

Use Venice's Alpha POST /responses endpoint - an OpenAI-compatible, stateless Responses API with typed output blocks (reasoning, message, functioncall, websearchcall). Venice Responses is an agent skill from veniceai/skills. Use Venice's Alpha POST /responses endpoint - an OpenAI-compatible, stateless Responses API with typed output blocks (reasoning, message, functioncall, websearchcall).

When should I use Venice Responses?

Venice Responses fits situations like: tasks that involve LLM API integration; tasks that involve Web search.

How do I install Venice Responses in Claude Code?

Run `npx skills add veniceai/skills --skill venice-responses -a claude-code`. Or copy the skill folder (skills/venice-responses in veniceai/skills) into .claude/skills/venice-responses in your project. Claude Code loads it when a task matches its description.

How do I install Venice Responses in Codex?

Run `npx skills add veniceai/skills --skill venice-responses -a codex`. Or copy the skill folder (skills/venice-responses in veniceai/skills) into .agents/skills/venice-responses in your project. Codex loads it when a task matches its description.

Can I use Venice Responses in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add veniceai/skills --skill venice-responses -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/venice-responses, .gemini/skills/venice-responses, .github/skills/venice-responses and .opencode/skills/venice-responses in your project.

What does Venice Responses need to run?

Going by SKILL.md and its folder, Venice Responses needs the command-line tools its instructions call (curl) and credentials named VENICE_API_KEY. Our summary lists: Python 3; A credential in VENICE_API_KEY.

Does Venice Responses access the network?

SKILL.md names 1 domain. In commands or code: api.venice.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Venice Responses safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Venice Responses use?

Venice Responses is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Venice Responses use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Venice Responses?

Skills that share tags, products or a category with Venice Responses: Dingo Verify (MigoXLab/dingo, 757 stars), Brave Answers API (brave/brave-search-skills, 183 stars), Openai Docs (theowenyoung/home, 115 stars) and Chatgpt Web Research (bear2u/my-skills, 932 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Venice Responses?

veniceai (a GitHub organization) maintains it in veniceai/skills, which has 144 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.

Source: veniceai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.