Official agent skill

Model Garden Deployment

by google in google/skills

Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Model Garden Deployment

skills CLI
$ npx skills add google/skills --skill agent-platform-deploy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills agent-platform-deploy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/agent-platform-deploy .claude/skills/agent-platform-deploy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-platform-deploy
GitHub stars
21k
Token cost
~5.1k tokens
SKILL.md length
2,045 words
Files
8 (incl. scripts, references)
Skills in repo
145
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

  • Works in 6 steps: Prerequisites → Discovering Deployable Models → Deploying a Model → …
  • Deploying an open model from Model Garden to a new endpoint
  • SKILL.md covers 1P Tuned Model Copy & Deployment, Safety & Confirmation Tiers…, 1. Prerequisites and 2. Discovering Deployable Models, plus 5 more sections
  • Runs Python and Shell scripts from its folder; calls gcloud and curl

What it does

The skill covers deploying a model, listing the Model Garden catalog, checking whether a model is deployable with gcloud ai model-garden models list-deployment-config, estimating cost with scripts/calculate_cost.py, troubleshooting failures such as quota limits, and undeploying models or deleting endpoints. Separate guides handle copying and deploying a first-party tuned model, regional availability, troubleshooting and undeploy steps. Pure discovery questions such as which endpoints exist go to agent-platform-endpoint-management, and model evaluation goes to agent-platform-eval-flywheel.

Actions are sorted into three safety tiers. Read-only commands (list, describe, list-deployment-config) run without asking. Deploy and undeploy-model require a dry-run card showing the exact request or command, model, project and region, machine and accelerator, estimated hourly cost and endpoint name, and the agent must end its turn and wait for your approval. Before undeploying it verifies the endpoint exists. Delete is irreversible and needs a typed confirmation after the agent explains the consequences.

When your agent uses it

  • Deploying an open model from Model Garden to a new endpoint
  • Checking whether a model is deployable in a region and what it would cost per hour
  • Diagnosing a deployment that failed on quota or availability
  • Undeploying a model and deleting an endpoint to stop charges

Example prompts

  • “Check whether this Model Garden model is deployable in us-central1 and what it costs per hour.”
  • “Deploy the model to a new endpoint, and show me the confirmation card before running anything.”
  • “My deployment failed with a quota error; work out what to change.”
  • “Undeploy the model from the staging endpoint and then delete the endpoint.”

Requirements

  • The gcloud CLI with access to a Google Cloud project
  • Quota and permissions to create endpoints in the target region

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Prerequisites
  2. Discovering Deployable Models
  3. Deploying a Model
  4. Checking Deployment Status
  5. Undeploying and Cleaning Up
  6. Troubleshooting

What it can do on your machine

Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • gcloud
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • cloud.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Garden Deployment loads about 5.1k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 215 tokens; SKILL.md has 2,045 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~215
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 2,045 words, ~5,100 tokens.

Download SKILL.mdSave it as .claude/skills/agent-platform-deploy/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
agent-platform-deploy
description
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available models, check if a specific model is deployable (`gcloud ai model-garden models list-deployment-config`), query deployment cost, troubleshoot deployment errors (like quota limits), or undeploy/clean up endpoints. Also use when copying and deploying a 1P Tuned Model. Don't use for pure listing/discovery questions of the form "is X deployed?", "list my endpoints", or "which regions have models running?" — for those use `agent-platform-endpoint-management`. Don't use for running model evaluations (use `agent-platform-eval-flywheel` skill).
metadata.version
1.0.3
metadata.category
AiAndMachineLearning

Agent Platform Model Garden Deploy Skill

This skill provides instructions for deploying Open Models from Agent Platform Model Garden to endpoints, and subsequently undeploying them to clean up resources.

1P Tuned Model Copy & Deployment

If you need to copy a 1P (First-Party) Tuned Model from a source project to a destination region or project and deploy it to a newly created endpoint, refer to the 1P Tuned Model Copy & Deployment Guide.

Safety & Confirmation Tiers (CRITICAL)

Before executing any commands on behalf of the user, you MUST adhere to the following safety tiers based on the action requested:

  1. Tier R: Read-only (list, describe, list-deployment-config)
    • Rule: No confirmation needed. You may execute these commands immediately to gather information for the user.
  2. Tier M: Mutating & Reversible (deploy, undeploy-model)
    • Rule: This requires explicit user confirmation. You MUST present a clear dry-run confirmation card containing:

      1. Exact proposed request: the :deploy request body for a deployment (§3), or the gcloud command code block for undeploy-model.
      2. Model identifier, destination project ID, and target region.
      3. Machine type and accelerator configuration.
      4. Estimated hourly cost ($/hr).
      5. Endpoint display name.
      6. Explicit confirmation prompt asking the user to approve before execution. You MUST wait for their explicit confirmation before executing. For undeploy-model, you MUST first verify that the endpoint and deployed model exist; if describe or list returns a 404 or empty result, you MUST halt and inform the user rather than attempting undeployment.
    • Same-turn restriction: Do not run the command in the same turn as presenting the confirmation prompt. End your turn after asking and wait for the user's reply; only execute after explicit approval. Printing a preview and then calling the tool before the user can answer does not count as obtaining confirmation.

  3. Tier D: Destructive & Irreversible (delete)
    • Rule: This requires explicit typed confirmation. You MUST output a text message explaining the irreversible nature of endpoint or model deletion and asking the user to type "I confirm" or "Yes, delete it" before executing the deletion command.

[!IMPORTANT]

Always Output Complete Text Response (NEVER Emit Empty Text): After executing any tool call (such as the :deploy API call, gcloud ai endpoints delete, gcloud ai endpoints list, or status checks), you MUST formulate and return a complete, informative textual response to the user. Explicitly report the operation ID, endpoint name/ID, error message, or list of resources. NEVER finish a turn with empty text or silence.

1. Prerequisites

Before deploying, ensure you have the correct project and region set. The commands below use placeholder variables PROJECT_ID and LOCATION_ID.

Ensure you are authenticated:

bash
gcloud auth login
gcloud auth application-default login
gcloud config set project $PROJECT_ID

2. Discovering Deployable Models

You can list models available in Model Garden and check if they can be self-deployed.

bash
gcloud ai model-garden models list

To see what machine types and accelerators are supported for a specific model, pass a MODEL_ID you obtained from the models list output above. Substitute <PUBLISHER>/<FAMILY>@<VERSION-ID> below with the exact string from the catalog output — the placeholder is deliberately not a real model ID:

bash
gcloud ai model-garden models list-deployment-config \
    --model="<PUBLISHER>/<FAMILY>@<VERSION-ID>"

[!NOTE] Some models, especially Hugging Face models, might require a Hugging Face Access Token for deployment.

[!TIP] Model Recommendation Instructions: Whenever you are about to name a specific model version in a response, do NOT recommend from memory. This applies in all of the following situations — not just direct deploy requests:

  • The user asks to deploy a model without naming one.
  • You are volunteering a next-step suggestion after a list, describe, or undeploy operation (e.g. "Would you like me to deploy <model> to this endpoint?").
  • The user asks a general "what should I use?" / "what's a good model for X?" question.
  • You are filling in a MODEL_ID value in an example command you are showing the user (as opposed to a placeholder like <PUBLISHER>/<FAMILY>@<VERSION-ID>).

New model versions ship frequently and older ones may be deprecated, so training-corpus knowledge of which models exist is unreliable. Follow this procedure:

  1. Clarify the use case if it isn't already clear from context (task type, quality vs. latency vs. cost priorities, hardware/quota constraints, license constraints). Skip if the user has already given enough signal.
  2. Query the live catalog with gcloud ai model-garden models list. Narrow with --filter when appropriate (e.g. --filter="name~gemma", --filter="name~llama", --filter="name~qwen", --filter="name~deepseek"). Never name a specific model version to the user until you have seen it in the catalog output for this project.
  3. Pick the latest generally-available version in the family that fits the use case. When multiple size variants exist, pick the one that matches the user's hardware/cost tolerance. Prefer a newer major version over an older one unless it is marked preview/experimental and the user explicitly asked for a stable option.
  4. Verify the exact model ID is deployable with gcloud ai model-garden models list-deployment-config --model="<publisher>/<family>@<version>" before naming it in your response.
  5. Cite the model ID verbatim in your recommendation, exactly as it appears in the catalog. Do not paraphrase to a family label ("Gemma", "Llama").

The MODEL_ID values in the §3 examples below are intentionally non-substantive placeholders (<PUBLISHER>/<FAMILY>@<VERSION-ID>). Do NOT replace them with a remembered model name for a user-facing recommendation — always re-run steps 2-4 first, then cite the exact string from the catalog.

2.1 Region Availability Check (Gemini + LoRA only)

For first-party Gemini or LoRA deploys, you must verify region availability before proceeding. Load the full instructions with load_skill_resource(skill_name='agent-platform-deploy', file_path='references/region_availability.md').

Skip this for open-weights models (Gemma, Llama, DeepSeek, Qwen) and for Gemini-tuned models — they have no per-region publisher endpoint restriction. Go straight to §3.

3. Deploying a Model

[!WARNING] Deploying models, especially large ones, consumes significant compute resources and incurs costs.

  1. You MUST compute an hourly $ estimate for the requested --machine-type before proposing a deploy. Try each source below in order, falling through to the next on any failure:

    • Run scripts/calculate_cost.py. The accelerator type and count are fixed per machine type in Model Garden and derived automatically. Example:

      bash
      python3 scripts/calculate_cost.py \
          --machine-type=g2-standard-48

      If the script exits non-zero (unknown --machine-type — a routine state for machines in the Model Garden catalog but not yet in the price snapshot, e.g. A4/B200 today), fall through to the next source. Do NOT invent a number.

    • Fall back to Agent Platform prediction pricing

if no source above produced a number. Read the accelerator + hourly rate directly off that page and cite the URL in the estimate you present to the user.

  1. You MUST present this cost estimation to the user and warn them that this is the list price, which may differ from their actual bill due to potential discounts, reservations, or non-us-central1 regions.
  2. You MUST ALWAYS request explicit confirmation from the user agreeing to the estimated cost before executing any deploy command.

To deploy an open-weights Model Garden model, call the :deploy API directly with curl.

If a deployment is rejected for quota, report the API's error verbatim.

[!IMPORTANT]

  • Cost Pushback & Hardware Renegotiation: If the user pushes back on cost (e.g., "That is too expensive, can you try a smaller configuration?"), or requests an invalid or unsupported hardware combination (e.g. g2-standard-48g with 4x H100 GPUs), explain the constraint or invalidity clearly, check list-deployment-config to identify the supported alternative (e.g., g2-standard-24 with 2x L4 or g2-standard-12 with 1x L4), compute its cost estimate with a single query, and immediately render a complete Tier M dry-run confirmation card for that recommended configuration in the same response.
  • Region Failover & Quota Exhaustion: When a deployment fails due to quota or capacity in the requested region (e.g. QUOTA_EXCEEDED or RESOURCE_EXHAUSTED), identify an alternative supported region (e.g. us-east4 or us-east1), compute its cost estimate with a single query, and immediately render a complete Tier M dry-run confirmation card with the new --region and exact command in the same response. State the alternative region directly without making unverified capacity claims.
  • Efficient Tool Execution (No Redundant Calls): Do NOT execute redundant models list, list-deployment-config, or --help commands if the model ID, region, or hardware configuration are already known or resolved. Run each discovery command strictly once.
  • Single Status Check & Response Formatting (CRITICAL):
    • When initiating a deployment (the :deploy call in §3), the response immediately returns a long-running operation. Formulate and return your textual confirmation response with the operation ID and endpoint display name immediately. Do NOT call operations describe in the same turn as deployment initiation.
    • When the user explicitly asks to check deployment status (e.g., "Please check to see the status of the deployment" or "Can you check if the deployment has finished?"):
      • NEVER run sleep commands, while loops, or repeated polling calls.
      • Execute gcloud ai operations describe <OP_ID> --region=<REGION> strictly ONCE.
      • ALWAYS output a full textual response reporting the operation status (e.g. "The deployment operation projects/.../operations/... is currently in progress / running (created at ...). Asynchronous model deployment typically takes 10–15 minutes to complete").
      • Only proceed with sending a test prediction if the status check confirms the endpoint is already serving and ready.
  • Alphanumeric Project ID: Always specify the alphanumeric Project ID (e.g. my-gcp-project) for --project, NOT the numeric project number (e.g. 123456789012). If given a numeric project number and its Project ID is not available, pass the number inside a fully-qualified resource name, e.g. gcloud ai endpoints list --region=projects/123456789012/locations/us-central1, or as a positional resource, gcloud ai endpoints describe projects/123456789012/locations/us-central1/endpoints/<ENDPOINT_ID> --region=us-central1. If a command is still refused because core/project is set to a project number, the sandbox's own project was seeded as a number: every gcloud ai call then needs an explicit --project=<PROJECT_ID>, which a resource name cannot substitute for. Report that instead of retrying.
  • Valid User-Specified Hardware Priority: When the user specifies an explicit, valid hardware configuration (e.g. g2-standard-96 with 8 NVIDIA_L4 GPUs, or g2-standard-12 with 1 NVIDIA_L4 GPU), honor that requested configuration for the dry-run preview and cost estimation rather than overriding it with default recommendations. However, if the requested configuration is invalid or unsupported (e.g. mismatched GPU count such as g2-standard-12 with 2 L4 GPUs, or non-existent machine shapes), follow the Cost Pushback & Hardware Renegotiation rule above: explain the invalidity clearly, identify the supported alternative (e.g. g2-standard-24 with 2 L4 GPUs), calculate its cost, and immediately present the confirmation card for the valid alternative.
  • Endpoint Display Name: If the user specifies or requests an endpoint name or display name (e.g. 'usersim-gemma-eval-...'), you MUST always include --endpoint-display-name="<NAME>" in the deploy command.
Show full SKILL.md (356 more words)Show less
Example: Deploying an open-weights model from Model Garden

Here is a typical bash script to deploy a model. You can run this block directly.

bash
#!/bin/bash
# Example script to deploy an open-weights model from Model Garden.
#
# NOTE: MODEL_ID below is a PLACEHOLDER, not a real model ID. Substitute it
# with a value from a live `gcloud ai model-garden models list` (see §2)
# before running this script, and do NOT quote the placeholder back to the
# user as a recommended model.

PROJECT_ID=$(gcloud config get-value project)
LOCATION_ID="us-central1" # Recommended default region
# Replace placeholder with exact ID from `gcloud ai model-garden models list`:
MODEL_ID="<PUBLISHER>/<FAMILY>@<VERSION-ID>"

echo "Deploying model $MODEL_ID to project $PROJECT_ID in $LOCATION_ID..."

# The API takes the model as a resource name, while the catalog ID is
# "<PUBLISHER>/<FAMILY>@<VERSION-ID>". Split on the first "/" to convert.
PUBLISHER_MODEL="publishers/${MODEL_ID%%/*}/models/${MODEL_ID#*/}"

# Omit deployConfig entirely to select the recommended default config.
# Comprehensive request with supported fields. Returns a long-running
# operation, so this is inherently asynchronous.
curl -sS -X POST \
    "https://${LOCATION_ID}-aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/${LOCATION_ID}:deploy" \
    -H "Authorization: Bearer $(gcloud auth print-access-token)" \
    -H "Content-Type: application/json" \
    -d "{
      \"publisherModelName\": \"${PUBLISHER_MODEL}\",
      \"modelConfig\": {
        \"acceptEula\": true
      },
      \"endpointConfig\": {
        \"endpointDisplayName\": \"my-open-model-deployment\"
      },
      \"deployConfig\": {
        \"dedicatedResources\": {
          \"machineSpec\": {
            \"machineType\": \"g2-standard-12\",
            \"acceleratorType\": \"NVIDIA_L4\",
            \"acceleratorCount\": 1
          },
          \"minReplicaCount\": 1
        }
      }
    }"

echo "Deployment initiated asynchronously."

The response is a GoogleLongrunningOperation. Its name field is the full operation path, projects/<PROJECT>/locations/<REGION>/operations/<OP_ID>; §4 accepts either that or the bare <OP_ID>.

  • Set modelConfig.huggingFaceAccessToken when deploying gated Hugging Face models that require authentication.
  • Set deployConfig.dedicatedResources.machineSpec.reservationAffinity if using reserved compute.
1P Tuned Model Cross-Region Copy and Deployment

For the detailed tuned model copy and deployment workflow, load load_skill_resource(skill_name='agent-platform-deploy', file_path='references/copy_deploy_guide.md'). That guide covers the execution sequence, tier assignments for copy/deploy/delete commands, hardware renegotiation, test prediction verification, and the in-progress operation lock.

4. Checking Deployment Status

The :deploy call in §3 is asynchronous in itself -- there is no flag to pass -- and returns a long-running operation whose name is the operation ID. You can use that ID to check the ongoing status of the deployment.

bash
gcloud ai operations describe YOUR_OPERATION_ID \
    --region=$LOCATION_ID

[!IMPORTANT]

Single Status Check Only (No Sleep / Polling Loops): Model deployment operations take 10–30 minutes. NEVER run sleep commands (e.g. sleep 45 && ...) or loop operations describe repeatedly in a turn. Run gcloud ai operations describe strictly ONCE. If done is not true, immediately return the operation ID and in-progress status to the user and explain that deployment takes 10–15 minutes.

Note: Large models (roughly 20B+ parameters) may take 15-20 minutes to fully deploy and start serving.

Verifying Deployment

If the model is successfully deployed, verify by making a prediction call to test. Because Model Garden models are often deployed to Dedicated Endpoints, you shouldn't use gcloud ai endpoints predict. Instead, you must fetch the endpoint's dedicated DNS name and send a curl request.

[!TIP] Ask the user to try using their own prompt to see the results. Otherwise use the default.

Use the following script:

bash
#!/bin/bash
PROJECT_ID=$(gcloud config get-value project)
LOCATION_ID="us-central1"
ENDPOINT_ID="YOUR_ENDPOINT_ID"
PROMPT=${1:-"Explain quantum computing in simple terms."}

echo "Fetching dedicated Endpoint DNS..."
ENDPOINT_URL=$(gcloud ai endpoints describe $ENDPOINT_ID \
    --project=$PROJECT_ID \
    --region=$LOCATION_ID \
    --format="value(dedicatedEndpointDns)")

if [ -z "$ENDPOINT_URL" ]; then
    echo "Error: Could not retrieve dedicated endpoint URL for $ENDPOINT_ID."
    exit 1
fi

echo "Sending prediction request to $ENDPOINT_URL..."

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  "https://${ENDPOINT_URL}/v1beta1/projects/${PROJECT_ID}/locations/${LOCATION_ID}/endpoints/${ENDPOINT_ID}/chat/completions" \
  -d '{
    "model": "'"$ENDPOINT_ID"'",
    "messages": [
      {
        "role": "user",
        "content": "'"$PROMPT"'"
      }
    ]
  }'

5. Undeploying and Cleaning Up

For the full undeploy and cleanup procedure (find endpoint, undeploy model, delete endpoint, delete model), load load_skill_resource(skill_name='agent-platform-deploy', file_path='references/undeploy_guide.md').

[!WARNING] Failing to undeploy a model will result in continuous charges for the allocated compute resources, even if you are not sending prediction requests. Always clean up after testing.

6. Troubleshooting

For troubleshooting quota/resource exhausted errors and hardware fallback, load load_skill_resource(skill_name='agent-platform-deploy', file_path='references/troubleshooting.md').

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in skills/cloud/agent-platform-deploy of google/skills.

  • SKILL.md
  • references/copy_deploy_guide.md
  • references/region_availability.md
  • references/troubleshooting.md
  • references/undeploy_guide.md
  • references/usage.md
  • scripts/calculate_cost.py
  • scripts/config_gcloud_cli.sh

Open the folder on GitHubat commit 8a1ac05

Compare with similar skills

Model Garden Deployment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Garden Deployment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Garden Deployment this skillgoogle/skills21k—~5.1kAutomated safety check: PassApache-2.0
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0
AWS AI MLaws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0
Vertex AI Geminimajiayu000/claude-skill-registry6661 repos~2.1kAutomated safety check: PassApache-2.0
SkyPilot Multi-Cloud OrchestrationOrchestra-Research/AI-Research-SKILLs13k4 repos~2.4kAutomated safety check: PassMIT
Model Serving Kubernetessickn33/agentic-awesome-skills47k1 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vertex AI Gemini

    majiayu000/claude-skill-registry

    Google Cloud Vertex AI for enterprise Gemini deployments — production scaling, fine-tuning, and MLOps.

    666 GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • SkyPilot Multi-Cloud Orchestration

    Orchestra-Research/AI-Research-SKILLs

    Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

    13k GitHub starsUsed in 4 repos~2.4k tokens
    DevOps & CloudAuto-check passed
  • Model Serving Kubernetes

    sickn33/agentic-awesome-skills

    Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

    47k GitHub starsUsed in 1 repo~2.3k tokens
    DevOps & CloudAuto-check passed
  • ML System Design Interviewer

    PrepLabsAI/InterviewMentor

    A Principal ML Engineer interviewer that simulates a FAANG-style ML system design interview covering the full lifecycle from data to production.

    112 GitHub stars~4.2k tokensUpdated today
    DevOps & CloudAuto-check passed

More from google/skills

All 145 skills in this repo
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated today
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Official

    Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.

    21k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Official

    Interviews you about data model, workload and scale, then recommends one Google Cloud database from a decision matrix and drafts starter provisioning code for review.

    21k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Works with

Questions about Model Garden Deployment

What does Model Garden Deployment do?

Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change. py, troubleshooting failures such as quota limits, and undeploying models or deleting endpoints. Separate guides handle copying and deploying a first-party tuned model, regional availability, troubleshooting and undeploy steps.

When should I use Model Garden Deployment?

Model Garden Deployment fits situations like: deploying an open model from Model Garden to a new endpoint; checking whether a model is deployable in a region and what it would cost per hour; diagnosing a deployment that failed on quota or availability; undeploying a model and deleting an endpoint to stop charges.

How do I install Model Garden Deployment in Claude Code?

Run `npx skills add google/skills --skill agent-platform-deploy -a claude-code`. Or copy the skill folder (skills/cloud/agent-platform-deploy in google/skills) into .claude/skills/agent-platform-deploy in your project. Claude Code loads it when a task matches its description.

How do I install Model Garden Deployment in Codex?

Run `npx skills add google/skills --skill agent-platform-deploy -a codex`. Or copy the skill folder (skills/cloud/agent-platform-deploy in google/skills) into .agents/skills/agent-platform-deploy in your project. Codex loads it when a task matches its description.

Can I use Model Garden Deployment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill agent-platform-deploy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-platform-deploy, .gemini/skills/agent-platform-deploy, .github/skills/agent-platform-deploy and .opencode/skills/agent-platform-deploy in your project.

What does Model Garden Deployment need to run?

Going by SKILL.md and its folder, Model Garden Deployment needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (gcloud and curl). Our summary lists: The gcloud CLI with access to a Google Cloud project; Quota and permissions to create endpoints in the target region.

Does Model Garden Deployment access the network?

SKILL.md names 1 domain. As links in the text: cloud.google.com. This is read from the text; nothing was executed.

Is Model Garden Deployment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Model Garden Deployment use?

Model Garden Deployment is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Garden Deployment use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.6k tokens, read only when the agent opens those files.

What are the alternatives to Model Garden Deployment?

Skills that share tags, products or a category with Model Garden Deployment: SageMaker Production Defaults (huggingface/skills, 11k stars), AWS AI ML (aws/agent-toolkit-for-aws, 2.8k stars), Vertex AI Gemini (majiayu000/claude-skill-registry, 666 stars) and SkyPilot Multi-Cloud Orchestration (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Garden Deployment?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.