All skills
google avatar

/google-cloud-solution-agentic-ai-data-science-workflow

@becc4b8
by googlegoogle/skills21k stars
1,698

Designs a tailored multi-product agentic data science architecture on Google Cloud that incorporates opinionated best practices. Use when architecting multi-product solutions for agent-based data analytics or ML workloads. Don't use for simple queries, non-agentic pipelines, general cloud reviews, or writing agent code.

Use this Skill: https://skilld.dev/gh/google/skills/google-cloud-solution-agentic-ai-data-science-workflow

This session only. Nothing lands on disk.

referencesdesign-recommendations.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Design Recommendations Guidance

Generate design guidance based on the products used in the solution architecture. Use the latest guidelines from https://docs.cloud.google.com/architecture/multiagent-ai-system#design_considerations and Google Cloud best practices to ground the guidance that you generate.

  • Security, privacy, and compliance:
    • Enforce least-privilege Identity and Access Management (IAM) controls. Grant each agent only the permissions that it needs to perform its tasks and to communicate with tools and with other agents.
    • To discover and de-identify sensitive data in the prompts and responses and in log data, use the Cloud Data Loss Prevention API.
    • Secure frontend Cloud Run endpoints by disabling default run.app URLs and use a regional external Application Load Balancer and Google Cloud Armor security policies for added protection.
    • To authenticate internal user access to the frontend Cloud Run service, use Identity-Aware Proxy (IAP).
    • To authenticate external user access to the frontend service, use Identity Platform or Firebase Authentication.
    • When you configure your agents to use MCP, ensure that you authorize access to external data and tools, implement privacy controls like encryption, apply filters to protect sensitive data, and monitor agent interactions.
    • Configure Direct VPC egress on Cloud Run to ensure agent runtimes communicate with databases and internal services over private VPC IPs without exposing traffic to the public internet.
    • Securely manage database credentials, API keys, and sensitive configuration using Secret Manager.
    • Use Model Armor to inspect and sanitize inference prompts and responses for prompt injection, threat defense, and sensitive data protection.
    • For more information, see https://docs.cloud.google.com/architecture/framework/perspectives/ai-ml/security.md.txt
  • Reliability:
    • Design the agentic system to tolerate or handle agent-level failures. Where feasible, use a decentralized approach where agents can operate independently.
    • Implement coordinator agent fallback logic and graceful error handling for specialized agent, SQL generation, or code execution failures.
    • Handle API rate limits and 429 resource exhaustion using exponential backoff with retries or provisioned throughput for business-critical workloads.
    • Set strict query timeouts, execution time limits, and memory/CPU caps on generated code and database queries.
    • Deploy Cloud Run frontend and agent runtimes across multiple regional zones for automatic failover during single-zone outages.
    • For more information, see https://docs.cloud.google.com/architecture/framework/perspectives/ai-ml/reliability.md.txt
  • Operational excellence:
    • To isolate generated code execution and prevent unauthorized filesystem or network access, use sandbox code execution.
    • Route structured agent logs to Cloud Logging to capture prompt interpretation, generated SQL or Python code, tool invocations, and execution outcomes.
    • Trace end-to-end multi-agent request flows and inter-agent communication loops using Cloud Trace / OpenTelemetry.
    • For more information, see https://docs.cloud.google.com/architecture/framework/perspectives/ai-ml/operational-excellence.md.txt
  • Cost optimization:
    • Enable Cloud Run autoscaling to zero for idle frontend and agent runtimes in non-production or low-traffic environments.
    • To reduce the cost of requests that contain repeated content with high input token counts, use context caching.
    • Use a tiered model strategy: route routine intent classification and SQL/code generation to Gemini Flash, reserving Gemini Pro for complex data science reasoning and multi-agent coordination.
    • For more information, see https://docs.cloud.google.com/architecture/framework/perspectives/ai-ml/cost-optimization.md.txt
  • Performance efficiency:
  • Sustainability:

Source: SKILL.md on GitHub

No alerts9d3 checks · Risk SAFE
  • Gen Agent Trust Hub9d

    This skill provides a structured framework for designing and implementing agentic AI workflows on Google Cloud. It includes security considerations such as the generation of infrastructure-as-code and validation scripts based on user requirements. These features are fundamental to the skill's utility and are supported by references to official documentation and security best practices.

  • Socket9d

    No alerts

  • Snyk9d

    Risk: LOW · No issues

Signed by skilld at becc4b8. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 weeks ago
metadata
{
  "version": "1.0.0",
  "category": "MultiProductSolutions"
}

README badge

README badge for google/skills/google-cloud-solution-agentic-ai-data-science-workflow