All skills
google avatar

/google-cloud-solution-agentic-ai-borderless-data-lakehouse

@becc4b8
by googlegoogle/skills21k stars
1,698

Discovers requirements and designs a borderless open data lakehouse using Lakehouse for Apache Iceberg and BigQuery data agents. Use when architecting multi-cloud storage infrastructure (Cloud Storage, AWS S3, Azure Blob), establishing ingestion and AI serving subsystems, configuring Cross-Cloud Interconnect, or deploying Gemini Enterprise Agent Platform and BigQuery data agents. Don't use for single-cloud data warehouses, or when the focus is on Knowledge Catalog metadata governance and Spark-driven IDE analytics workflows (use google-cloud-solution-agentic-analytics-spark-knowledge-catalog instead).

Use this Skill: https://skilld.dev/gh/google/skills/google-cloud-solution-agentic-ai-borderless-data-lakehouse

This session only. Nothing lands on disk.

referencesproduct_mapping.md

≈959 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Product Mapping Guidance

Explain to the user that the solution consists of two subsystems:

  • The data ingestion subsystem ingests data from external sources and uses a central lakehouse to unify and process fragmented databases into a unified data profile in Google Cloud.
  • The serving subsystem lets users query an AI assistant and a data analysis agent to analyze the consolidated data.

For each component in the confirmed technical decomposition, identify the appropriate Google Cloud products and features, based on the following guidance:

  • Data ingestion subsystem components:
    • Central metadata and governance:
      • Recommended primary product: Lakehouse for Apache Iceberg
      • Alternative product 1: Dataproc Metastore
        • Pros: Better for legacy open-source heavy pipelines.
        • Cons: Can have lower performance for borderless federation and Apache Iceberg.
    • Processing engine:
      • Recommended primary product: Managed Service for Apache Spark with Lightning Engine
      • Alternative product 1: BigQuery
        • Pros: Allows querying data in place on external clouds using standard SQL, reducing data movement.
        • Cons: Less flexible than Spark for highly complex, programmatic transformations or custom code.
      • Alternative product 2: Dataflow
        • Pros: Powerful for complex, unified batch or stream ETL.
        • Cons: Requires learning the Apache Beam programming model and managing a Cloud Storage bucket for job error logs.
    • Internal data storage:
      • Recommended primary product: Cloud Storage using Apache Iceberg format or Apache Parquet format.
      • Alternative product 1: BigQuery storage
        • Pros: Provides high performance for native BigQuery queries.
        • Cons: Less portable for other open-source processing engines compared to open formats on Cloud Storage.
      • Alternative product 2: Cloud Storage
        • Pros: Eliminates borderless egress fees and latency if you choose to consolidate your workload to a single cloud.
    • Security:
      • Recommended primary product: Use Secret Manager to securely hold authentication credentials for federated REST catalogs. Manage direct storage object access using BigQuery Cloud Resource connections and runtime credential vending.
      • Alternative product 1: Cloud Key Management Service (KMS)
        • Pros: Provides hardware-backed key management for encryption.
        • Cons: Not designed to store plain-text secrets like API tokens.
    • Borderless networking:
      • Recommended primary product: Cross-Cloud Interconnect
      • Alternative product 1: Cloud VPN (HA VPN)
        • Pros: Offers lower costs during low-traffic periods.
        • Cons: Can have higher latency and lower bandwidth compared to dedicated Cross-Cloud Interconnect.
  • Serving subsystem components:
    • AI serving and agentic workflows:
      • Recommended primary product: BigQuery data agent with Antigravity CLI
      • Alternative product 1: Gemini Enterprise Agent Platform
        • Pros: Provides built-in orchestration, native enterprise grounding, and managed chat UIs.
        • Cons: Offers less granular control over the prompt loop, and can be more expensive than a lightweight MCP server.
      • Alternative product 2: Google Cloud Data Agent Kit
        • Pros: Optimized for data practitioners, data engineers, and data scientists to manage the data lifecycle and perform interactive analysis directly within their IDE.
        • Cons: Designed for developer-centric workflows rather than serving production end-to-end business applications.

Source: SKILL.md on GitHub

No alerts9d3 checks · Risk SAFE
  • Gen Agent Trust Hub9d

    This skill provides a structured framework for architecting and implementing borderless data lakehouse solutions on Google Cloud. It leverages official documentation and standard automation practices to assist users in creating secure and reliable multi-cloud infrastructures.

  • Socket9d

    No alerts

  • Snyk9d

    Risk: LOW · No issues

Signed by skilld at becc4b8. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 weeks ago
metadata
{
  "version": "1.0.0",
  "category": "MultiProductSolutions"
}

README badge

README badge for google/skills/google-cloud-solution-agentic-ai-borderless-data-lakehouse