All skills
google avatar

/agent-platform-deploy

@99c871e
by googlegoogle/skills21k stars
1,698

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available models, check if a specific model is deployable (`gcloud ai model-garden models list-deployment-config`), query deployment cost, troubleshoot deployment errors (like quota limits), or undeploy/clean up endpoints. Also use when copying and deploying a 1P Tuned Model. Don't use for pure listing/discovery questions of the form "is X deployed?", "list my endpoints", or "which regions have models running?" — for those use `agent-platform-endpoint-management`. Don't use for running model evaluations (use `agent-platform-eval-flywheel` skill).

Use this Skill: https://skilld.dev/gh/google/skills/agent-platform-deploy

This session only. Nothing lands on disk.

referencesregion_availability.md

≈901 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Region Availability Check for Publisher Endpoints

[!NOTE] Skip this section if the user is deploying an open-weights model from Model Garden (Gemma, Llama, DeepSeek, Qwen, or any user-supplied weights) or a Gemini-tuned model — i.e. anything served via the §3 :deploy API onto a dedicated endpoint. These have no per-region publisher endpoint restriction. The real failure modes for an unusual region are (a) the requested accelerator/machine type isn't offered in that region, or (b) the project has no quota — both surface as a clean error at deploy time before any resources are provisioned (§3's cost-confirm gate catches them). Go straight to §3.

Apply this section only if the user is asking to serve a first-party managed Gemini model (google/gemini-*) via a publisher endpoint whose regional availability varies.

Before responding to any deploy request that names a specific region for a first-party managed Gemini model (google/gemini-*), you MUST verify the model is actually available in that region by making a live API call. Do not rely on Google Search, training-corpus knowledge, or publisher documentation for availability claims — regional availability changes frequently and grounded text can be stale or wrong.

Probe only the exact model and region the user asked about. Do not probe other models as a "control" — you cannot infer anything about model A's availability from model B's status, because a different model may itself be unavailable in the reference region for unrelated reasons.

For first-party publisher models (google/*), probe with a real :generateContent call using a minimal valid payload. This probe is agent-executed in the sandbox — do not ask the user to run it manually:


curl -sS -o /dev/null -w "%{http_code}\n" \
    -H "Authorization: Bearer $(gcloud auth print-access-token)" \
    -H "Content-Type: application/json" \
    "https://${LOCATION_ID}-aiplatform.googleapis.com/v1/projects/${PROJECT_ID}/locations/${LOCATION_ID}/publishers/google/${MODEL_ID}:generateContent" \
    -d "{\"contents\":{\"role\":\"user\",\"parts\":{\"text\":\"${PROBE_TEXT:-hi}\"}}}"

For LoRA adapters deployed on top of a first-party Gemini base model, probe the base model in the target region using the same :generateContent call above with ${MODEL_ID} set to the base (e.g. gemini-2.5-flash). The adapter cannot serve in a region where its base model isn't available. This does not apply to Gemini Tuning models, which are deployed via a different path and have no publisher-endpoint region restriction.

Interpret the probe result and act:

  • 200 — model is available in that region. Proceed with the deploy.
  • 404 — model is not available in that region. STOP. Tell the user plainly that the model isn't offered in that region and list the regions where it is available (from gcloud ai model-garden models list --filter="name~${MODEL_NAME}" without --region). Do not silently switch regions. Do not proceed to write deploy code or SDK initialization for the unsupported region. Do not run additional "control" probes to double-check the 404 — the target-region probe is authoritative.
  • Any other outcome (permission denied, quota, transient failure, etc.) — do not conclude the model is available or unavailable. Explain the underlying cause in plain language (e.g. "your account doesn't have access to this project's Vertex AI API — enable it in the console or switch projects") and the concrete next action.

Source: SKILL.md on GitHub

No alerts1d3 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    This skill facilitates model deployment on Google Cloud's Agent Platform with built-in safety confirmations for mutating tasks. There are security considerations related to potential command injection in helper scripts and the use of administrative IAM commands, which should be reviewed to ensure proper sanitization and least-privilege access.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: LOW · No issues

Signed by skilld at 99c871e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 days ago
metadata
{
  "version": "1.0.3",
  "category": "AiAndMachineLearning"
}

README badge

README badge for google/skills/agent-platform-deploy