All skills
google avatar

/google-cloud-solution-agentic-analytics-spark-knowledge-catalog

@8f9a457
by googlegoogle/skills21k stars
1,698

Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine). Use when designing data science and analytics workflows across structured and unstructured distributed data (including in S3, Azure Blob, AlloyDB, and Iceberg), establishing metadata governance with Knowledge Catalog aspect types, or grounding agentic IDEs (VS Code, Antigravity) by using the Google Cloud Data Agent Kit. Don't use for provisioning borderless data lakehouse infrastructure (use google-cloud-solution-agentic-ai-borderless-data-lakehouse instead).

Use this Skill: https://skilld.dev/gh/google/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog

This session only. Nothing lands on disk.

referencesdesign-recommendations.md

≈601 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Design recommendations

Use the following guidance to generate design recommendations for a governed, secure pipeline for agentic analytics solution across structured and unstructured data that's distributed across Google Cloud, on-premises systems, and other cloud providers.

  • Security, privacy, and compliance:
    • Assign the required roles and permissions.
    • Configure token federation or Secret Manager credentials to allow Spark REST catalogs to authenticate with remote AWS S3 / Azure storage.
    • Establish the required aspect types and business glossaries to classify table sensitivities (e.g., private vs public) and map business metrics.
    • Use Knowledge Catalog lineage tracking to identify PII data leakage paths across database conversions.
  • Reliability:
    • Configure Cross-Cloud Interconnect with redundant links in separate domains to prevent network drops during massive queries.
  • Operational excellence:
    • Organize Spark sessions by deploying reusable session templates (template.yaml) in Managed Service for Apache Spark.
    • Track data lineage in Knowledge Catalog to perform impact analysis on upstream schema breaks.
  • Cost optimization:
    • Run zero-copy Spark REST catalog queries to AWS S3 instead of transferring files over public endpoints, eliminating network egress fees.
    • Use Knowledge Catalog lineage cost-optimization tools to find and prune orphaned tables.
  • Performance efficiency:
    • Enable Lightning Engine with native runtime (spark.dataproc.lightningEngine.runtime=native) and set the compute tier to premium (dataproc.tier=premium) inside the Spark session templates.
    • Configure off-heap memory settings (spark.memory.offHeap.size=1g) to boost vector operations.
    • To minimize geographic latency, select Google Cloud regions that let you co-locate Google Cloud resources alongside AWS/Azure datacenters where S3 buckets reside.
  • Sustainability:
    • Run Managed Service for Apache Spark Serverless with Lightning Engine to perform operations in memory, reducing server compute waste.
    • Use zero-copy architecture (BigQuery Omni/Spark REST catalogs) to prevent storing duplicate multi-petabyte datasets.

Source: SKILL.md on GitHub

No alerts6d3 checks · Risk SAFE
  • Gen Agent Trust Hub6d

    This skill facilitates the design and implementation of governed agentic analytics solutions on Google Cloud. It includes several security considerations, such as the generation of validation scripts and references to external documentation and starter packs. These activities are conducted with explicit user oversight and are grounded in authoritative technical resources.

  • Socket6d

    No alerts

  • Snyk6d

    Risk: LOW · No issues

Signed by skilld at 8f9a457. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
metadata
{
  "version": "1.0.0",
  "category": "MultiProductSolutions"
}

README badge

README badge for google/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog