All skills
aws avatar

/ingesting-into-data-lake

@b33847d

Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL, pull from Snowflake, query BigQuery into S3, export DynamoDB, CTAS, convert to Iceberg. Do NOT use for setting up or troubleshooting Glue connections (use connecting-to-data-source), creating empty tables (use creating-data-lake-table), running queries (use querying-data-lake), finding tables by fuzzy name (use finding-data-lake-assets), catalog audit (use exploring-data-catalog), or SaaS platforms like Salesforce, ServiceNow, SAP, MongoDB, Kafka.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/ingesting-into-data-lake

This session only. Nothing lands on disk.

referencesmigration-validation.md

≈693 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Migration Validation Checklist

Run all checks after migration. Do not skip any.

1. Row Count Match

SELECT 'source' AS tbl, COUNT(*) AS cnt
FROM "<source_catalog>"."<source_db>"."<source_table>"
UNION ALL
SELECT 'target' AS tbl, COUNT(*) AS cnt
FROM "s3tablescatalog/<bucket>"."<namespace>"."<target_table>"

Counts must match exactly unless a WHERE filter was applied during migration. If filtered, document the expected difference.

2. Schema Comparison

-- Source schema
DESCRIBE "<source_catalog>"."<source_db>"."<source_table>"

-- Target schema
DESCRIBE "s3tablescatalog/<bucket>"."<namespace>"."<target_table>"

Check:

  • All expected columns are present
  • Column order matches (or is acceptable if reordered)
  • Types are compatible (minor promotions like int->bigint are OK)
  • No unexpected columns added or dropped

3. Null Count Comparison

-- Run for each column, or generate dynamically
SELECT
    COUNT(*) - COUNT(col1) AS col1_nulls,
    COUNT(*) - COUNT(col2) AS col2_nulls
FROM "s3tablescatalog/<bucket>"."<namespace>"."<target_table>"

Compare against the same query on the source. Null counts should match.

4. Boundary Value Check

SELECT
    MIN(numeric_col) AS min_val,
    MAX(numeric_col) AS max_val,
    MIN(date_col) AS min_date,
    MAX(date_col) AS max_date
FROM "s3tablescatalog/<bucket>"."<namespace>"."<target_table>"

Compare against source. Min/max values should match (accounting for any WHERE filters).

5. Distinct Count Check

SELECT
    COUNT(DISTINCT key_col) AS distinct_keys
FROM "s3tablescatalog/<bucket>"."<namespace>"."<target_table>"

Compare against source. Mismatches indicate duplicates introduced or rows lost.

6. Partition Verification (if partitioned)

SELECT <partition_expression>, COUNT(*) AS row_count
FROM "s3tablescatalog/<bucket>"."<namespace>"."<target_table>"
GROUP BY 1
ORDER BY 1

Verify partition distribution is reasonable and no partitions are missing.

7. Sample Row Comparison

-- Pick a specific key value and compare full rows
SELECT * FROM "<source_catalog>"."<source_db>"."<source_table>"
WHERE key_col = '<known_value>'

SELECT * FROM "s3tablescatalog/<bucket>"."<namespace>"."<target_table>"
WHERE key_col = '<known_value>'

Spot-check 3-5 specific rows. All column values should match.

Pass Criteria

Check Pass condition
Row count Exact match (or documented delta if filtered)
Schema All columns present with compatible types
Null counts Match within tolerance (0 difference expected)
Boundary values Match exactly
Distinct counts Match exactly
Partitions All expected partitions present
Sample rows All values match

Source: SKILL.md on GitHub

2 warnings17d3 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    This skill facilitates the ingestion of data from various external sources into an AWS data lake. It contains security considerations related to the processing of untrusted data and the dynamic generation of Spark scripts, which are characteristic of ETL (Extract, Transform, Load) operations. These patterns are consistent with the skill's purpose and are used within a managed cloud environment.

  • Socket17d

    1 alert: gptAnomaly

  • Snyk17d

    Risk: MEDIUM · 1 issue

Signed by skilld at b33847d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 months ago
Other metadata
metadata
{
  "version": "1",
  "argument-hint": "'[source-path|connection-name|table-name] [--target s3-tables|iceberg|parquet]'"
}

README badge

README badge for aws/agent-toolkit-for-aws/ingesting-into-data-lake