All skills
aws avatar

/ingesting-into-data-lake

@b33847d

Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL, pull from Snowflake, query BigQuery into S3, export DynamoDB, CTAS, convert to Iceberg. Do NOT use for setting up or troubleshooting Glue connections (use connecting-to-data-source), creating empty tables (use creating-data-lake-table), running queries (use querying-data-lake), finding tables by fuzzy name (use finding-data-lake-assets), catalog audit (use exploring-data-catalog), or SaaS platforms like Salesforce, ServiceNow, SAP, MongoDB, Kafka.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/ingesting-into-data-lake

This session only. Nothing lands on disk.

referencesmigration-troubleshooting.md

≈660 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Migration Troubleshooting

Common issues when migrating Glue Data Catalog tables to S3 Tables.

CTAS Errors

Problem Cause Fix
GENERIC_INTERNAL_ERROR: Invalid table or column names Uppercase in table or column names Lowercase all names in the CTAS SELECT and table name
CTAS times out Table too large for single Athena query Use Glue ETL (Path B) or migrate in partitioned batches
LOCATION is not supported Included LOCATION clause in CTAS Remove LOCATION -- S3 Tables manages storage automatically
SYNTAX_ERROR: line X:Y: mismatched input Malformed partition transform or missing quotes Check partitioning = ARRAY[...] syntax and quote the catalog path
TABLE_NOT_FOUND on source Wrong catalog prefix for source table Use awsdatacatalog as the source catalog for standard Glue tables

Validation Failures

Problem Cause Fix
Row count mismatch WHERE filter excluded rows, or source has duplicates Check filter clause; run dedup analysis on source
Schema mismatch (extra/missing columns) SELECT * picked up partition columns or metadata Explicitly list columns in SELECT
Null count differs Type coercion converted empty strings to nulls Check source data for empty strings vs actual nulls
Boundary values differ Timezone or precision differences Compare with explicit CAST to same type

Visibility Issues

Problem Cause Fix
Target table not visible in Athena Analytics integration not enabled Create s3tablescatalog federated catalog. See ctas-patterns.md for catalog path syntax.
Table visible but returns no data CTAS succeeded but wrote zero rows Check WHERE filter; verify source table has data
Table visible but columns show as _col0, _col1 Used SELECT * with incompatible source format Explicitly name columns with aliases

Partition Issues

Problem Cause Fix
CTAS fails with too many partitions Over 100 target partitions in single CTAS Batch with WHERE filters or use coarser partition transform (e.g., month() instead of day())
Partition column missing in target Iceberg hidden partitions derive from source column The source column must be in the SELECT; the transform is in partitioning
Uneven partition sizes Poor transform choice for data distribution Consider bucket() for high-cardinality columns

Glue ETL Issues

See glue-etl-migration.md for Glue-specific errors.

Source: SKILL.md on GitHub

2 warnings17d3 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    This skill facilitates the ingestion of data from various external sources into an AWS data lake. It contains security considerations related to the processing of untrusted data and the dynamic generation of Spark scripts, which are characteristic of ETL (Extract, Transform, Load) operations. These patterns are consistent with the skill's purpose and are used within a managed cloud environment.

  • Socket17d

    1 alert: gptAnomaly

  • Snyk17d

    Risk: MEDIUM · 1 issue

Signed by skilld at b33847d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 months ago
Other metadata
metadata
{
  "version": "1",
  "argument-hint": "'[source-path|connection-name|table-name] [--target s3-tables|iceberg|parquet]'"
}

README badge

README badge for aws/agent-toolkit-for-aws/ingesting-into-data-lake