All skills
google avatar

/gke-manifest-generation

@becc4b8
by googlegoogle/skills21k stars
1,698

Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters. Use when creating or modifying GKE deployment manifests, configuring container security contexts, setting CPU/memory resource limits, defining readiness/liveness/startup probes, mounting secrets and volumes, configuring GKE Gateway API routes, targeting Spot VMs, or deploying AI model inference workloads (vLLM, TGI, Gemma). Don't use for live cluster operations, pod troubleshooting (use gke-workload-troubleshooting), or cluster infrastructure provisioning (use gke-cluster-creation).

Use this Skill: https://skilld.dev/gh/google/skills/gke-manifest-generation

This session only. Nothing lands on disk.

referencesai-inference.md

≈767 tokens on demand. Your agent reads this file only when SKILL.md points to it.

AI/LLM Inference Workload Example (vLLM & GCS FUSE)

This reference example demonstrates deploying an AI/LLM model serving workload (such as Gemma 2 27B) on GKE.

It includes:

  • Workload Identity annotation for GCP authentication
  • GPU resource allocation (nvidia.com/gpu) and nodeSelector targeting accelerator types
  • GCS FUSE CSI driver volume mount (csi.storage.gke.io) for read-only model weight loading
  • Shared memory volume mount at /dev/shm (emptyDir with medium: Memory)
  • Extended startupProbe failure threshold to accommodate slow model initialization
apiVersion: v1
kind: Namespace
metadata:
  name: gemma-ns
---
apiVersion: v1
kind: ServiceAccount
metadata:
  name: gemma-sa
  namespace: gemma-ns
  annotations:
    iam.gke.io/gcp-service-account: {gcp_service_account_email}
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: gemma-27b-deployment
  namespace: gemma-ns
  labels:
    app.kubernetes.io/name: gemma-27b
spec:
  replicas: 1
  selector:
    matchLabels:
      app.kubernetes.io/name: gemma-27b
  template:
    metadata:
      labels:
        app.kubernetes.io/name: gemma-27b
      annotations:
        gke-gcsfuse/volumes: "true"
    spec:
      serviceAccountName: gemma-sa
      securityContext:
        runAsNonRoot: true
        runAsUser: 10000
        runAsGroup: 10000
        seccompProfile:
          type: RuntimeDefault
      containers:
        - name: gemma-server
          image: vllm/vllm-openai:gemma2 # Example optimized image
          args: ["--model", "/models", "--tensor-parallel-size", "4"]
          ports:
            - name: http-api
              containerPort: 8000
          securityContext:
            allowPrivilegeEscalation: false
            capabilities:
              drop:
                - ALL
          resources:
            requests:
              cpu: "32"
              memory: "128Gi"
              nvidia.com/gpu: 4
            limits:
              cpu: "32"
              memory: "128Gi"
              nvidia.com/gpu: 4
          livenessProbe:
            httpGet:
              path: /healthz
              port: http-api
            periodSeconds: 30
          readinessProbe:
            httpGet:
              path: /healthz
              port: http-api
            periodSeconds: 10
          startupProbe:
            httpGet:
              path: /healthz
              port: http-api
            failureThreshold: 60
            periodSeconds: 10
          volumeMounts:
            - name: model-weights
              mountPath: /models
              readOnly: true
            - name: dshm
              mountPath: /dev/shm
      nodeSelector:
        cloud.google.com/gke-accelerator: "nvidia-l4"
      volumes:
        - name: model-weights
          csi:
            driver: gcsfuse.csi.storage.gke.io
            readOnly: true
            volumeAttributes:
              bucketName: {gcs_bucket_name}
              mountOptions: "implicit-dirs"
        - name: dshm
          emptyDir:
            medium: Memory

Source: SKILL.md on GitHub

1 warning9d3 checks · Risk SAFE
  • Gen Agent Trust Hub9d

    This skill includes security considerations such as a potential surface for indirect prompt injection when processing user-provided application code or natural language descriptions. These elements are part of the skill's core manifest-generation functionality. See the detailed analysis for additional context on tool usage and data handling.

  • Socket9d

    No alerts

  • Snyk9d

    Risk: MEDIUM · 1 issue

Signed by skilld at becc4b8. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 weeks ago
metadata
{
  "version": "1.0.0",
  "category": "Containers"
}

README badge

README badge for google/skills/gke-manifest-generation