All skills
aws avatar

/ingesting-into-data-lake

@b33847d

Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL, pull from Snowflake, query BigQuery into S3, export DynamoDB, CTAS, convert to Iceberg. Do NOT use for setting up or troubleshooting Glue connections (use connecting-to-data-source), creating empty tables (use creating-data-lake-table), running queries (use querying-data-lake), finding tables by fuzzy name (use finding-data-lake-assets), catalog audit (use exploring-data-catalog), or SaaS platforms like Salesforce, ServiceNow, SAP, MongoDB, Kafka.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/ingesting-into-data-lake

This session only. Nothing lands on disk.

referencesctas-patterns.md

≈757 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Athena CTAS Patterns for S3 Tables Migration

Basic Migration (no partitions)

CREATE TABLE "s3tablescatalog/my-bucket"."my_namespace"."customers"
WITH (format = 'PARQUET') AS
SELECT * FROM "awsdatacatalog"."legacy_db"."customers"

Migration with Iceberg Partition Transforms

Convert Hive-style explicit partitions to Iceberg hidden partitions:

-- Source has explicit year/month/day columns from Hive partitioning
-- Target uses Iceberg day() transform on the timestamp column
CREATE TABLE "s3tablescatalog/my-bucket"."analytics"."events"
WITH (
    format = 'PARQUET',
    partitioning = ARRAY['day(event_timestamp)']
) AS
SELECT
    event_id,
    user_id,
    event_type,
    event_timestamp,
    payload
FROM "awsdatacatalog"."raw_db"."events_hive"

Available Partition Transforms

Transform Example Use when
year(col) ARRAY['year(created_at)'] Multi-year data, infrequent queries
month(col) ARRAY['month(created_at)'] Monthly reporting, medium cardinality
day(col) ARRAY['day(event_time)'] Daily data, time-series workloads
hour(col) ARRAY['hour(event_time)'] High-volume streaming data
bucket(col, N) ARRAY['bucket(user_id, 16)'] High-cardinality columns, even distribution
Multiple ARRAY['month(ts)', 'bucket(id, 8)'] Compound partitioning

Batched Migration (over 100 partitions)

Athena CTAS has a 100-partition limit per statement. Migrate in batches:

-- Batch 1: 2023 data
CREATE TABLE "s3tablescatalog/my-bucket"."ns"."orders"
WITH (format = 'PARQUET', partitioning = ARRAY['month(order_date)']) AS
SELECT * FROM "awsdatacatalog"."sales"."orders"
WHERE order_date >= DATE '2023-01-01' AND order_date < DATE '2024-01-01'

-- Batch 2+: INSERT INTO for subsequent years
INSERT INTO "s3tablescatalog/my-bucket"."ns"."orders"
SELECT * FROM "awsdatacatalog"."sales"."orders"
WHERE order_date >= DATE '2024-01-01' AND order_date < DATE '2025-01-01'

Migration with Column Transformations

CREATE TABLE "s3tablescatalog/my-bucket"."clean"."users"
WITH (format = 'PARQUET') AS
SELECT
    user_id,
    LOWER(email) AS email,
    COALESCE(display_name, username) AS name,
    CAST(created_at AS timestamp) AS created_at,
    CASE WHEN status = 'A' THEN 'active' ELSE 'inactive' END AS status
FROM "awsdatacatalog"."legacy"."users_raw"

Cross-Catalog Migration (self-managed Iceberg)

CREATE TABLE "s3tablescatalog/my-bucket"."analytics"."transactions"
WITH (
    format = 'PARQUET',
    partitioning = ARRAY['day(transaction_date)']
) AS
SELECT * FROM "awsdatacatalog"."iceberg_db"."transactions_selfmanaged"

Format Options

Format Best for Notes
PARQUET (default) Most analytical workloads Columnar, good compression, wide tool support
AVRO Write-heavy, schema evolution Row-based, fast writes
ORC Hive ecosystem compatibility Columnar, good for Hive migrations

Source: SKILL.md on GitHub

2 warnings16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill facilitates the ingestion of data from various external sources into an AWS data lake. It contains security considerations related to the processing of untrusted data and the dynamic generation of Spark scripts, which are characteristic of ETL (Extract, Transform, Load) operations. These patterns are consistent with the skill's purpose and are used within a managed cloud environment.

  • Socket16d

    1 alert: gptAnomaly

  • Snyk16d

    Risk: MEDIUM · 1 issue

Signed by skilld at b33847d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 months ago
Other metadata
metadata
{
  "version": "1",
  "argument-hint": "'[source-path|connection-name|table-name] [--target s3-tables|iceberg|parquet]'"
}

README badge

README badge for aws/agent-toolkit-for-aws/ingesting-into-data-lake