Basin Catalog Patterns
Choose the engine based on the project's existing runtime and workload, then retrieve its current connection example.
| Need | Starting point |
|---|---|
| Python catalog operations and ingestion without a Spark deployment | PyIceberg |
| Existing Spark ETL and distributed table processing | PySpark |
| Connect an existing SQL engine | Engine configuration guides |
| Query through Cloudflare's serverless SQL service | Basin SQL |
| Stream events into tables | Basin Pipelines patterns |
Use the discovered Catalog URI and Warehouse name from configuration. Match dependencies to the installed engine and the current guide instead of adopting a universal pinned Spark/Iceberg combination.
Plan ingestion, query, and maintenance responsibilities together. Prefer automatic table maintenance when it meets the workload; align retention with time-travel needs before enabling expiration. For engine-specific partitioning, schema evolution, or manual procedures, consult that engine's linked upstream documentation and verify behavior on representative data.
When multiple writers share a table, design recovery around the actual failed operation and the engine's commit semantics. Reproduce conflicts and ensure retries do not duplicate application work. See API selection and troubleshooting.