All skills
google avatar

/google-cloud-solution-agentic-analytics-spark-knowledge-catalog

@8f9a457
by googlegoogle/skills21k stars
1,698

Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine). Use when designing data science and analytics workflows across structured and unstructured distributed data (including in S3, Azure Blob, AlloyDB, and Iceberg), establishing metadata governance with Knowledge Catalog aspect types, or grounding agentic IDEs (VS Code, Antigravity) by using the Google Cloud Data Agent Kit. Don't use for provisioning borderless data lakehouse infrastructure (use google-cloud-solution-agentic-ai-borderless-data-lakehouse instead).

Use this Skill: https://skilld.dev/gh/google/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog

This session only. Nothing lands on disk.

referencesproduct-selection-guidance.md

≈849 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Product selection guide

Use the guidance in this file to recommend products and features for a governed, secure pipeline for agentic analytics solution across structured and unstructured data that's distributed across Google Cloud, on-premises systems, and other cloud providers.

Note: When generating solution designs, architecture diagrams, and documentation, use the latest Google Cloud product names, as listed in the following table:

Old name New name
BigLake Lakehouse for Apache Iceberg
Dataproc Serverless Managed Service for Apache Spark
Dataplex Knowledge Catalog

The underlying APIs, gcloud commands, Terraform resources, and IAM roles often retain the old names.

Important:

  • Don't recommend any products that are deprecated, retired, decommissioned, or unsupported. Verify the status of the products by using the resources that are listed in the "Ground all generated content" section in ../SKILL.md.
  • Don't recommend any features that are deprecated, retired, decommissioned, or unsupported. Verify the status of the features by using the resources that are listed in the "Ground all generated content" section in ../SKILL.md.
  • If multiple products or features can be used for a component of the workload, then do the following:
    • Recommend the most appropriate product or feature. When alternative products exist, the relevant product documentation might provide guidance on when to recommend each product. Follow that guidance.
    • Mention the available alternative products or features.
    • Explain the pros and cons of each alternative product or feature.

Product recommendations

  • Central metadata and data governance
    • Recommended product: Knowledge Catalog (to capture business glossaries, columns descriptions, and custom metadata aspect types) integrated with Lakehouse for Apache Iceberg
  • Processing engine
    • Recommended product: Managed Service for Apache Spark with Lightning Engine (configured with Iceberg REST Catalog)
    • Alternative product: BigQuery
      • Pros: Easy to write standard SQL queries directly over external BigLake tables.
  • Raw data storage
    • Recommended product: Cloud Storage
    • Alternative product: BigQuery
      • Cons: Not designed for raw blob storage of unstructured files like PDFs.
  • Connectivity to external data sources
    • Recommended products:
      • For cross-cloud connectivity: Cross-Cloud Interconnect
      • For on-premises to cloud connectivity: Cloud Interconnect
  • Operational data storage
    • Recommended product: AlloyDB for PostgreSQL
      • Pros: Supports column-cache vector acceleration for fast joins, which is ideal for real-time operational query acceleration.
    • Alternative product: Cloud SQL for PostgreSQL
      • Pros: Highly cost-effective for simpler, smaller relational workloads.
      • Cons: Lacks advanced column-cache vector acceleration (AlloyDB columnar engine) for fast joins.
  • User and agentic interactions:
    • Recommended product: Antigravity IDE or VS Code with the Google Cloud Data Agent Kit (provides IDE grounding)

Source: SKILL.md on GitHub

No alerts6d3 checks · Risk SAFE
  • Gen Agent Trust Hub6d

    This skill facilitates the design and implementation of governed agentic analytics solutions on Google Cloud. It includes several security considerations, such as the generation of validation scripts and references to external documentation and starter packs. These activities are conducted with explicit user oversight and are grounded in authoritative technical resources.

  • Socket6d

    No alerts

  • Snyk6d

    Risk: LOW · No issues

Signed by skilld at 8f9a457. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
metadata
{
  "version": "1.0.0",
  "category": "MultiProductSolutions"
}

README badge

README badge for google/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog