Dingo Verify
MigoXLab/dingo
A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.).
$ npx skills add google/skills --skill agent-platform-inference -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills agent-platform-inference --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/agent-platform-inference .claude/skills/agent-platform-inference && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-platform-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-inference into .claude/skills/agent-platform-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-inference", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/agent-platform-inferenceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill agent-platform-inference -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills agent-platform-inference --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/agent-platform-inference .agents/skills/agent-platform-inference && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-platform-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-inference into .agents/skills/agent-platform-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-inference", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill agent-platform-inference -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills agent-platform-inference --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/agent-platform-inference .cursor/skills/agent-platform-inference && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-platform-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-inference into .cursor/skills/agent-platform-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-inference", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/agent-platform-inference--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill agent-platform-inference -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills agent-platform-inference --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/agent-platform-inference .gemini/skills/agent-platform-inference && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-platform-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-inference into .gemini/skills/agent-platform-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-inference", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills agent-platform-inferenceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill agent-platform-inference -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/agent-platform-inference .github/skills/agent-platform-inference && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-platform-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-inference into .github/skills/agent-platform-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-inference", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill agent-platform-inference -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills agent-platform-inference --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/agent-platform-inference .opencode/skills/agent-platform-inference && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-platform-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/agent-platform-inference into .opencode/skills/agent-platform-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-platform-inference", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-platform-inferenceConnects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.).
Agent Platform Inference is an agent skill from google/skills, published by the product's own GitHub organization. Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate code for calling Gemini or OpenMaaS models, authenticate with GenAI SDK, OpenAI SDK, or legacy Agent Platform SDK, configure base URLs and global/regional endpoints, or troubleshoot 429 Resource Exhausted (DSQ), 400…
Its SKILL.md is about 9.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts (for example `scripts/gemini_genai_sdk.py`, `scripts/gemini_openai_sdk.py` and `scripts/gemini_vertexai_sdk.py`).
It sits in AI & LLM Engineering, covering LLM API integration and Machine learning. It works with DeepSeek, Google Cloud, OpenAI and Qwen. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 8 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
gcloudcurlpippython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
aiplatform.googleapis.comAlso links to:
docs.cloud.google.comgithub.comconsole.cloud.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Platform Inference loads about 9.5k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 3,651 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 3,651 words, ~9,463 tokens.
.claude/skills/agent-platform-inference/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.This skill provides instructions for authenticating and connecting to Google Cloud Agent Platform to use Generative AI models. It covers:
projects/.../endpoints/<id>
resource — tuned Gemini models, OSS LLMs self-deployed from Model Garden
via the agent-platform-deploy skill, and legacy custom models) —
section 4.Before executing any commands or scripts on behalf of the user, you must adhere to the following safety tiers based on the action requested. (The skill is read-only; other safety tiers are omitted):
client.models.generate_content,
client.chat.completions.create, client.completions.create,
client.embeddings.create)123456789012, my-project).us-central1,
global).gemini-2.5-flash,
deepseek-ai/deepseek-v3.2-maas).Google GenAI SDK (google-genai),
OpenAI SDK).max_output_tokens,
response_schema) if specified.
Natural-language paraphrases without explicitly listing these
parameters are NOT sufficient.I will perform model inference with the following parameters. Please confirm this information before I proceed:
- Project ID:
my-project- Region:
us-central1- Model ID:
gemini-2.5-pro- SDK: Google GenAI SDK (
google-genai)- Input Prompt: "Summarize the plot of Hamlet in 3 sentences"
Do you confirm? [Yes/No]
gemini-2.5-pro
or <MODEL_ID>).Google GenAI SDK (google-genai)
or OpenAI SDK).global
or us-central1).CRITICAL: Before running any of the Python sample scripts in the scripts/
directory (e.g., scripts/openmaas_openai_sdk.py), you MUST ensure the
environment is correctly initialized by following these steps:
Google Cloud Authentication: Authenticate with your Google Cloud credentials and configure active Application Default Credentials (ADC) for Agent Platform access:
gcloud auth login
gcloud auth application-default loginEnable API (if not already enabled):
gcloud services enable aiplatform.googleapis.comPython Dependencies: The scripts import vertexai (from
google-cloud-aiplatform), google-genai, and openai. Do not create
a virtual environment — it starts empty and hides packages the environment
already provides, forcing a redundant install. Probe, and install only what
is missing:
python3 -c "import vertexai, google.genai, openai" \
|| pip install -r scripts/requirements.txtscripts/requirements.txt is a fallback for an environment that does not
already provide these SDKs; do not install it on top of a working
environment.
Verify Setup (Optional): Run all sample scripts at once to verify the environment is working end-to-end:
./scripts/verify_all.shExecution: Run the scripts with a plain python3 scripts/.... There is
no environment to activate first.
[!IMPORTANT] CRITICAL: Model IDs & Availability * Gemini Models: See [Gemini Models][gemini-models-docs] for valid Model IDs and Regions. * OpenMaaS Models: See Use Open Models on Agent Platform for Llama, DeepSeek, Qwen, etc. * Incomplete Lists: The Model IDs listed in this skill are examples only and may be incomplete or outdated. * Action: Always verify the Model ID and Region using the links above before generating code.
[gemini-models-docs]: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/migrate
Before preparing code or presenting a Tier R confirmation card, you MUST ensure all necessary parameters are grounded:
Missing Model ID, Model Family, or SDK (CRITICAL):
gemini-2.5-flash, gemini-2.5-pro, or deepseek-v3.2-maas).
Proposing a defaulted model in a confirmation card without asking
violates parameter grounding.Missing Project ID or Region:
If the user's project ID or region is not specified in the prompt or conversation context, ASK the user for the project ID and region (e.g. "Which project ID and region would you like to use?"). Do not silently assume a project or region.
OpenMaaS Locations: OpenMaaS publisher models are hosted on global
(e.g. deepseek-ai/deepseek-v3.2-maas,
meta/llama-3.3-70b-instruct-maas) or regional endpoints such as
us-central1 (e.g. deepseek-ai/deepseek-r1-0528-maas).
When configuring inference for OpenMaaS models, use the appropriate
endpoint:
https://aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/global/endpoints/openapihttps://{REGION}-aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/{REGION}/endpoints/openapiand explicitly reflect the region in the confirmation card and final response.
SDK Choice:
google-genai for
Gemini, OpenAI SDK openai for OpenMaaS).Sandbox Execution via Python (CRITICAL):
run_command,
ALWAYS run Python code using the official SDKs (e.g., writing and
running a Python script with google-genai, openai, or vertexai).
Do not use raw curl commands for final inference execution.Model Specified?
deepseek-ai/deepseek-r1-0528-maas,
deepseek-ai/deepseek-v3.2-maas, meta/llama-3.3-70b-instruct-maas).Model Family & SDK Selection:
gemini-2.5-pro, gemini-2.5-flash) -> Preferred:
GenAI SDK (google-genai). Proceed to [1. Gemini Models].deepseek-ai/*, meta/llama-*, qwen/*) ->
Preferred: OpenAI SDK (openai). Proceed to [2. OpenMaaS Models].projects/.../endpoints/<id>) -> Proceed to [4. Custom Endpoints].Troubleshooting: Is the user reporting an error (429 Resource Exhausted, 400 User Validation, 404 Not Found, empty response due to token limits, etc.)?
[!NOTE] Skip this section if either of these applies:
- The user is calling a custom endpoint (§4) — a tuned Gemini model served on a numeric
projects/.../endpoints/<id>, a self-deployed OSS LLM (Llama, DeepSeek, Qwen, Gemma, etc.), or a legacy custom model. Those requests hit a specific endpoint resource whose region is fixed at deploy time; if the caller-side region doesn't match, the endpoint lookup returns a clean 404 without incurring inference cost. Go to §4.- The user is calling an OpenMaaS publisher model (§2) — Llama, DeepSeek, Qwen, etc. served via the global
openapibase URL. These don't have per-region availability restrictions in the same way first-party Gemini does. Go to §2.Apply this section only if the user is calling a first-party managed Gemini model (
gemini-*, via §1), including fine-tuned LoRA adapters on top of Gemini — these route through a publisher endpoint whose regional availability actually varies.
Before responding to any inference request that names a specific region for a
first-party managed Gemini model (gemini-*) or a fine-tuned Gemini LoRA
adapter (identified by numeric endpoint ID + user-stated base model), you
MUST verify the model is actually available in that region by making a
live API call. Do not rely on Google Search, training-corpus knowledge, or
publisher documentation for availability claims — regional availability
changes frequently and grounded text can be stale or wrong.
Probe only the exact model and region the user asked about. Do not probe other models as a "control" — you cannot infer anything about model A's availability from model B's status, because a different model may itself be unavailable in the reference region for unrelated reasons.
For first-party Gemini models, probe with a real :generateContent call using
a minimal valid payload:
curl -sS -o /dev/null -w "%{http_code}\n" \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://${LOCATION_ID}-aiplatform.googleapis.com/v1/projects/${PROJECT_ID}/locations/${LOCATION_ID}/publishers/google/${MODEL_ID}:generateContent" \
-d "{\"contents\":{\"role\":\"user\",\"parts\":{\"text\":\"${PROBE_TEXT:-hi}\"}}}"For inference against a fine-tuned Gemini LoRA adapter, probe the base
model in the target region using the same :generateContent call above with
${MODEL_ID} set to the base (e.g. gemini-2.5-flash if the adapter was
tuned on gemini-2.5-flash). The LoRA adapter cannot serve in a region where
its base model isn't available.
Interpret the probe result and act:
gcloud ai model-garden models list --filter="name~$MODEL_NAME" without --region). Do not silently switch
regions. Do not proceed to write inference code or SDK initialization for
the unsupported region. Do not run additional "control" probes to
double-check the 404 — the target-region probe is authoritative.For Gemini models (e.g., gemini-2.5-pro, gemini-3-flash-preview), the
GenAI SDK (google-genai) is the PREFERRED method. The legacy
vertexai SDK is still supported but GenAI SDK is recommended for new projects.
[!IMPORTANT] Preview Models (including Gemini 3.1) are often ONLY available in the
globalregion. Stable models are available inus-central1and other regions.
google-genai) is PREFERRED. Use
OpenAI SDK for compatibility, or Legacy SDK (vertexai) if needed.pip install google-genaiSee scripts/gemini_genai_sdk.py for the
complete code.
Use the standard OpenAI SDK with the Agent Platform endpoint. This is great for cross-compatibility.
See scripts/gemini_openai_sdk.py for the
complete code.
The legacy vertexai SDK is still widely used but google-genai is preferred
for new Gemini projects.
See scripts/gemini_vertexai_sdk.py for the
complete code.
Documentation: Google GenAI SDK
Documentation: Agent Platform Gemini Models
For OpenMaaS (Model-as-a-Service) models, the HIGHLY RECOMMENDED approach is to use the standard OpenAI SDK with a specific Vertex AI endpoint.
[!WARNING] While
GenerativeModelcan support some OpenMaaS models, it is discouraged. Use the OpenAI SDK for best compatibility (especially for Chat Completions).
pip install openai google-authYou MUST use a Google Cloud OAuth access token as the API key for the OpenAI SDK.
import subprocess
def get_gcp_access_token():
return subprocess.check_output(
["gcloud", "auth", "print-access-token"]
).decode("utf-8").strip()[!NOTE] Google Cloud access tokens typically expire after 1 hour. The
get_gcp_access_token()function above retrieves a fresh token at the time it is called. For long-running applications, you implement a refresh mechanism. See Refresh the access token for details.
https://aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/global/endpoints/openapihttps://{REGION}-aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/{REGION}/endpoints/openapiSee scripts/openmaas_openai_sdk.py for the
complete code.
[!TIP] Alternative: Environment Variables You can set environment variables in your shell instead of updating the code.
Alternative: Environment Variables You can set environment variables in your shell instead of updating the code.
export OPENAI_BASE_URL="https://aiplatform.googleapis.com/v1/projects/YOUR_PROJECT_ID/locations/global/endpoints/openapi"
export OPENAI_API_KEY="$(gcloud auth application-default print-access-token)"Then initialize the client without arguments:
client = OpenAI()
The following models support the legacy Completions API: zai-org/glm-5-maas,
moonshotai/kimi-k2-thinking-maas, minimaxai/minimax-m2-maas,
deepseek-ai/deepseek-v3.1-maas, and deepseek-ai/deepseek-v3.2-maas.
response = client.completions.create(
model="deepseek-ai/deepseek-v3.2-maas",
prompt="Once upon a time",
max_tokens=100
)
print(response.choices[0].text)# Verify specific Embedding Model ID on Model Garden (e.g., intfloat/multilingual-e5-small)
response = client.embeddings.create(
model="intfloat/multilingual-e5-large-maas",
input="The quick brown fox jumps over the lazy dog",
)
print(response.data[0].embedding)The google-genai SDK can also access OpenMaaS models via the vertexai
backend.
See scripts/openmaas_genai_sdk.py for the
complete code.
[!IMPORTANT] Model ID Format: For GenAI SDK with OpenMaaS, you MUST use the full path:
publishers/PUBLISHER/models/MODEL(e.g.,publishers/zai-org/models/glm-5-maas).
For OpenMaaS, you can also use GenerativeModel (if supported).
See scripts/openmaas_vertexai_sdk.py for
the complete code.
[!IMPORTANT] Model ID Format: For Agent Platform SDK with OpenMaaS, you MUST use the full path:
publishers/PUBLISHER/models/MODEL.
Documentation: Use Open Models on Agent Platform
[!TIP] Self-Deployment for Control: If you need dedicated hardware (GPUs/TPUs), guaranteed capacity, or specific regional placement not offered by MaaS, you can Self-Deploy these models to Agent Platform Endpoints. Search for the model in Model Garden and click "Deploy" to select your machine type. See the
agent-platform-deployskill for the deployment workflow, and section 4 of this skill for how to invoke the resulting self-deployed endpoint (use/chat/completionson the dedicated endpoint DNS, NOT the OpenMaaS publisher URL above).
[!IMPORTANT] Finding Inference Examples: The list above is a starting point. For the definitive inference snippets (especially for Chat Completions payload structure): 1. Consult the Use Open Models on Agent Platform list. 2. Click the link for your specific model (e.g., "DeepSeek-V3") to visit its Model Garden page. 3. Look for the "Sample Code" or "Use this model" button on the Model Garden page to get the exact
curlor Python code for that specific model version.
[!NOTE] This list is INCOMPLETE. See Use Open Models on Agent Platform for the full list of supported models.
| Model Family | Model ID Examples | Location | Notes |
|---|---|---|---|
| Llama 4 | meta/llama-4-maverick-17b-128e-instruct-maas | us-east5 | |
| Llama 4 | meta/llama-4-scout-17b-16e-instruct-maas | us-east5 | |
| Llama 3.3 | meta/llama-3.3-70b-instruct-maas | us-central1 | |
| DeepSeek | deepseek-ai/deepseek-v3.2-maas | global | Global ONLY |
| DeepSeek | deepseek-ai/deepseek-v3.1-maas | us-west2 | US-West2 ONLY |
| DeepSeek | deepseek-ai/deepseek-r1-0528-maas | us-central1 | |
| Qwen 3 | qwen/qwen3-coder-480b-a35b-instruct-maas | global | |
| Qwen 3 | qwen/qwen3-next-80b-a3b-instruct-maas | global | |
| Kimi | moonshotai/kimi-k2-thinking-maas | global | |
| MiniMax | minimaxai/minimax-m2-maas | global | |
| GLM | zai-org/glm-4.7-maas, zai-org/glm-5-maas | global |
This section covers how to invoke a model on an Agent Platform
Endpoint that belongs to your project — i.e., something with a
numeric resource name like
projects/.../endpoints/5875254126916403200. This is distinct from
calling the publisher MaaS surfaces in sections 2 and 3 (which hit
publishers/.../models/... or endpoints/openapi, not your endpoint
ID).
[!IMPORTANT]
Publisher MaaS vs your endpoint (don't confuse them). Section 3's OpenMaaS examples (e.g.
meta/llama-3.3-70b-instruct-maas) hit a shared publisher URL at/v1/projects/.../locations/.../endpoints/openapi. This section's recipes hit YOUR endpoint at/v1/projects/.../endpoints/<id>. If you have a Llama / Gemma / etc. model deployed via Model Garden "Deploy" (NOT the MaaS publisher product), follow this section — not section 3.
[!IMPORTANT]
Active Endpoint Discovery & Single Source of Truth:
To check if a model or tuned Gemini adapter is deployed and ready for inference, run:
gcloud ai endpoints list --project=<PROJECT_ID> --region=<REGION> --format=json.
gcloud ai endpoints listis the ONLY authoritative source of active serving endpoints.Do NOT rely on historical
job.tunedModel.endpointidentifiers from pastgcloud ai tuning-jobs listrecords — those record where an endpoint was initially created during tuning, but if the endpoint was subsequently deleted, undeployed, or expired, it is no longer active.If
gcloud ai endpoints listreturns empty[]or the requested tuned model is not deployed on any listed endpoint, report directly to the user that no active endpoint was found in that region and halt. NEVER hallucinate that historical/deleted endpoints are ready, and NEVER silently substitute the base model without explicit user instruction and a fresh Tier R confirmation prompt.
[!IMPORTANT]
Two orthogonal axes determine the call shape:
Axis 1 — model family drives the RPC method and payload:
Endpoint serves Method Payload A tuned Gemini model (output of Gemini tuning — the endpoint is already deployed for you) :generateContentcontents/generationConfigA self-deployed OSS LLM (Llama, DeepSeek, Qwen, Gemma, Mistral, etc., deployed via Model Garden) /chat/completionsOpenAI-compatible messagesA legacy custom model (classification, regression, custom-trained, embedding) :predictinstances/parametersRun
gcloud ai endpoints describe <ENDPOINT_ID> --region=<REGION> --format=jsonand inspectdeployedModels[].modelto decide: containsgemini→ tuned Gemini; matches an OSS publisher (meta/,google/gemma-,deepseek-ai/,qwen/, ...) → OSS LLM; otherwise → likely legacy custom.Axis 2 — endpoint type (shared vs dedicated) drives the URL host:
dedicatedEndpointEnabledHost false(default — shared endpoint)<REGION>-aiplatform.googleapis.comtrue(dedicated endpoint, has its own DNS)the value of dedicatedEndpointDns(format:<ENDPOINT_ID>.<REGION>-<PROJECT_NUM>.prediction.vertexai.goog)A dedicated endpoint cannot be reached via the shared
<REGION>-aiplatform.googleapis.comhost (per theEndpoint.dedicated_endpoint_enabledproto: "Once you enabled dedicated endpoint, you won't be able to send request to the shared DNS"). Always checkdedicatedEndpointDnsin the describe output: if it's set, use it as the host; otherwise use the shared host.Path is always
/v1/projects/.../locations/.../endpoints/<id>/...on both hosts. Both/v1/(GA) and/v1beta1/(beta) route to the same backend; the recipes in this skill use/v1/. The public Gemma deployment notebook still uses/v1beta1/, which also works.
Output of Gemini tuning is always an endpoint that's already deployed for
you, reachable on the shared host with :generateContent.
PROJECT_ID=my-project
ENDPOINT_ID=5875254126916403200
REGION=us-central1
TOKEN=$(gcloud auth application-default print-access-token)
curl -sS -X POST \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
"https://${REGION}-aiplatform.googleapis.com/v1/projects/${PROJECT_ID}/locations/${REGION}/endpoints/${ENDPOINT_ID}:generateContent" \
-d '{
"contents": [
{"role": "user", "parts": [{"text": "Hello! Introduce yourself briefly."}]}
],
"generationConfig": {
"temperature": 0.2
}
}'[!WARNING]
If you set
maxOutputTokens, be generous for thinking models. Gemini 2.5 Pro (and other thinking-enabled models) emit "thoughts" tokens that count againstmaxOutputTokensBEFORE any user-visible text. With a small cap (e.g. 100), the entire budget is consumed by thoughts and the response has emptytextparts but a non-zerousageMetadata.candidatesTokenCount.If you don't need to constrain output length, omit
maxOutputTokensentirely and let the model emit as much as it wants. If you do set it:>= 512for any chat-like use,>= 1024for a paragraph of output. If you seefinishReason: "MAX_TOKENS"and notextcontent in the response, your cap is too low.
Self-deployed OSS LLMs may be on a shared or dedicated endpoint
depending on how the deploy was configured (dedicated_endpoint_enabled
at create time). The recipe below handles both cases by checking
dedicatedEndpointDns in the describe output.
PROJECT_ID=my-project
ENDPOINT_ID=5875254126916403200
REGION=us-central1
TOKEN=$(gcloud auth application-default print-access-token)
# Step 1: discover host. dedicatedEndpointDns is empty for shared endpoints.
DEDICATED_DNS=$(gcloud ai endpoints describe "$ENDPOINT_ID" \
--project="$PROJECT_ID" --region="$REGION" \
--format="value(dedicatedEndpointDns)")
if [ -n "$DEDICATED_DNS" ]; then
HOST="$DEDICATED_DNS"
else
HOST="${REGION}-aiplatform.googleapis.com"
fi
# Step 2: call /chat/completions. Path is identical for both hosts.
curl -sS -X POST \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
"https://${HOST}/v1/projects/${PROJECT_ID}/locations/${REGION}/endpoints/${ENDPOINT_ID}/chat/completions" \
-d '{
"messages": [
{"role": "user", "content": "Hello! Introduce yourself briefly."}
]
}'[!NOTE]
max_tokens(NOTmaxOutputTokens) — this is OpenAI-compatible vocabulary, not Vertex. Omit it entirely to let the model emit as much as it wants; set it explicitly only if you need to cap output.- The
"model"field in the OpenAI-style payload can be omitted (or set to"") for endpoint deployments — the endpoint already determines which model serves the request.- Same endpoint also exposes
/completions(legacy text completion) and/embeddingsfor embedding models.- Reasoning models (DeepSeek-R1, Kimi-K2-Thinking, GLM-5 variants, etc.) emit thinking tokens that count against
max_tokensBEFORE the final answer — same pathology as Gemini 2.5 Pro in section 4a. If you DO setmax_tokensand get an emptychoices[0].message.contentorfinish_reason: "length", raise it (>= 1024 for chat, >= 2048 for longer thinking chains) or omit it.
Python equivalent (OpenAI SDK) — mirrors the public Gemma deployment notebook:
import google.auth
from google.auth.transport.requests import Request
import openai
from google.cloud import aiplatform
PROJECT_ID = "my-project"
ENDPOINT_ID = "5875254126916403200"
REGION = "us-central1"
aiplatform.init(project=PROJECT_ID, location=REGION)
endpoint = aiplatform.Endpoint(
f"projects/{PROJECT_ID}/locations/{REGION}/endpoints/{ENDPOINT_ID}"
)
endpoint_resource_name = endpoint.resource_name # full projects/.../endpoints/<id>
dedicated_dns = endpoint.gca_resource.dedicated_endpoint_dns # empty if shared
host = dedicated_dns if dedicated_dns else f"{REGION}-aiplatform.googleapis.com"
base_url = f"https://{host}/v1/{endpoint_resource_name}"
import subprocess
token = subprocess.check_output(
["gcloud", "auth", "print-access-token"]
).decode("utf-8").strip()
client = openai.OpenAI(base_url=base_url, api_key=token)
response = client.chat.completions.create(
model="", # endpoint determines the served model
messages=[{"role": "user", "content": "Hello! Introduce yourself briefly."}],
# Omit max_tokens to let the model emit as much as it wants. Set it
# only if you need to cap output length (see notes above).
)
print(response.choices[0].message.content)See also: agent-platform-deploy skill section 4 "Verifying Deployment",
which uses the same pattern post-deploy.
:predict (custom / classification / embedding)Same host-discovery logic as 4b (shared or dedicated based on
dedicatedEndpointDns):
DEDICATED_DNS=$(gcloud ai endpoints describe "$ENDPOINT_ID" \
--project="$PROJECT_ID" --region="$REGION" \
--format="value(dedicatedEndpointDns)")
HOST=${DEDICATED_DNS:-${REGION}-aiplatform.googleapis.com}
curl -sS -X POST \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
"https://${HOST}/v1/projects/${PROJECT_ID}/locations/${REGION}/endpoints/${ENDPOINT_ID}:predict" \
-d '{
"instances": [{"key": "value"}],
"parameters": {}
}'The exact instances shape is model-specific; consult the deployed
model's documentation or the Model Garden card it was deployed from.
from google import genai
import google.auth
_, project_id = google.auth.default()
client = genai.Client(vertexai=True, project=project_id, location="us-central1")
ENDPOINT_ID = "5875254126916403200"
response = client.models.generate_content(
model=f"projects/{project_id}/locations/us-central1/endpoints/{ENDPOINT_ID}",
contents="Hello! Introduce yourself briefly.",
config={"temperature": 0.2}, # add max_output_tokens only if you need a cap
)
print(response.text):predict → switch to :generateContent
(section 4a). Error mentions "Required instances format mismatch".:generateContent or :predict → switch to /chat/completions
(section 4b). Error may be 404, 405, or "method not allowed".:generateContent or
/chat/completions → switch to :predict (section 4c).Could not resolve host) or 404.<REGION>-aiplatform.googleapis.com, and the dedicated DNS
(*.prediction.vertexai.goog) only exists when
dedicatedEndpointEnabled is true.gcloud ai endpoints describe ... --format=json
for the dedicatedEndpointDns field; use it iff non-empty (per
the host-discovery snippets in section 4b/4c).maxOutputTokens is set too low. Gemini 2.5 Pro and other
thinking models emit "thoughts" tokens that count against the budget
BEFORE any user-visible text. With a small cap (e.g. 100), the entire
budget is consumed by thoughts and the response has empty text parts
but a non-zero usageMetadata.candidatesTokenCount and
finishReason: "MAX_TOKENS".maxOutputTokens entirely (let the model emit as much
as it wants), or raise it to >= 512 for chat-like use, >= 1024 for
longer output. See section 4 "Custom Endpoints" for details.us-central1 or global regions.us-central1, europe-west4, and many other regions.us-central1 or global.© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts) in skills/cloud/agent-platform-inference of google/skills.
Open the folder on GitHubat commit 8a1ac05
Agent Platform Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Platform Inference this skillgoogle/skills | 21k | — | ~9.5k | Automated safety check: Pass | Apache-2.0 | |
| Dingo VerifyMigoXLab/dingo | 757 | — | ~833 | Automated safety check: Pass | Apache-2.0 | |
| Claude Maintain ModelsKiln-AI/Kiln | 5.2k | — | ~15k | Automated safety check: Notes | Custom licence | |
| Awesome Free LLM APIsLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.1k | Automated safety check: Pass | MIT | |
| ModLens Image Vision Bridgeliustack/modlens | 4.1k | — | ~1.3k | Automated safety check: Notes | MIT | |
| Pocketmen With Yousix-nut/PocketMen-with-you | 310 | — | ~2.6k | Automated safety check: Pass | MIT |
MigoXLab/dingo
A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.
Kiln-AI/Kiln
Add new AI models to Kiln's mlmodellist.py and produce a Discord announcement.
LeoYeAI/openclaw-master-skills
Reference guide for permanent free-tier LLM APIs with rate limits, model lists, and OpenAI-compatible integration patterns.
liustack/modlens
Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.
six-nut/PocketMen-with-you
Turn 2+ user reference images into a high-fidelity animated Codex companion using PocketMen's own local stack.
xiincs/claude-code-vision-skill
Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
google/skills
Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.
Works with
Categories
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Agent Platform Inference is an agent skill from google/skills, published by the product's own GitHub organization.).
Agent Platform Inference fits situations like: asked to perform inference; ask a model a question; run a test prompt; execute chat completions.
Run `npx skills add google/skills --skill agent-platform-inference -a claude-code`. Or copy the skill folder (skills/cloud/agent-platform-inference in google/skills) into .claude/skills/agent-platform-inference in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill agent-platform-inference -a codex`. Or copy the skill folder (skills/cloud/agent-platform-inference in google/skills) into .agents/skills/agent-platform-inference in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill agent-platform-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-platform-inference, .gemini/skills/agent-platform-inference, .github/skills/agent-platform-inference and .opencode/skills/agent-platform-inference in your project.
Going by SKILL.md and its folder, Agent Platform Inference needs Python and a shell for the scripts in its folder, the command-line tools its instructions call (gcloud, curl, pip and python3) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A Bash shell.
SKILL.md names 4 domains. In commands or code: aiplatform.googleapis.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.cloud.google.com, github.com and console.cloud.google.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Agent Platform Inference is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.5k tokens (SKILL.md is roughly 38k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agent Platform Inference: Dingo Verify (MigoXLab/dingo, 757 stars), Claude Maintain Models (Kiln-AI/Kiln, 5.2k stars), Awesome Free LLM APIs (LeoYeAI/openclaw-master-skills, 2.2k stars) and ModLens Image Vision Bridge (liustack/modlens, 4.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.