All skills
aws avatar

/ingesting-into-data-lake

@b33847d

Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL, pull from Snowflake, query BigQuery into S3, export DynamoDB, CTAS, convert to Iceberg. Do NOT use for setting up or troubleshooting Glue connections (use connecting-to-data-source), creating empty tables (use creating-data-lake-table), running queries (use querying-data-lake), finding tables by fuzzy name (use finding-data-lake-assets), catalog audit (use exploring-data-catalog), or SaaS platforms like Salesforce, ServiceNow, SAP, MongoDB, Kafka.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/ingesting-into-data-lake

This session only. Nothing lands on disk.

referencesupload-options.md

≈264 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Upload Options Reference

cp vs sync

Command Use when
aws s3 cp Single file, or directory with --recursive
aws s3 sync Directory upload, skips unchanged files on re-run

sync is idempotent — safe to re-run after interruption. Prefer sync for directories.

Multipart Upload

aws s3 cp automatically uses multipart for files over 8 MB (default threshold). No flags needed. To tune:

aws configure set default.s3.multipart_threshold 64MB
aws configure set default.s3.multipart_chunksize 64MB

Compression Before Upload

Compressing locally saves transfer time and storage cost. Downstream tools (Athena, Glue) read gzip natively.

gzip file.csv
aws s3 cp file.csv.gz s3://<bucket>/<prefix>/

Do NOT compress Parquet, Avro, or ORC — they have built-in compression.

Overwrite Protection

Check if target exists before uploading:

aws s3 ls s3://<bucket>/<prefix>/<filename>

If it exists, warn the user. aws s3 cp overwrites without confirmation.

Source: SKILL.md on GitHub

2 warnings17d3 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    This skill facilitates the ingestion of data from various external sources into an AWS data lake. It contains security considerations related to the processing of untrusted data and the dynamic generation of Spark scripts, which are characteristic of ETL (Extract, Transform, Load) operations. These patterns are consistent with the skill's purpose and are used within a managed cloud environment.

  • Socket17d

    1 alert: gptAnomaly

  • Snyk17d

    Risk: MEDIUM · 1 issue

Signed by skilld at b33847d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 months ago
Other metadata
metadata
{
  "version": "1",
  "argument-hint": "'[source-path|connection-name|table-name] [--target s3-tables|iceberg|parquet]'"
}

README badge

README badge for aws/agent-toolkit-for-aws/ingesting-into-data-lake