All skills
wshobson avatar

/spark-optimization

@be57c0b
by Seth Hobsonwshobson/agents40k stars
4,281

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

Use this Skill: https://skilld.dev/gh/wshobson/agents/spark-optimization

This session only. Nothing lands on disk.

SKILL.md

≈53 tokens always: the name and description. ≈730 when used: this file. ≈2.5k more on demand in 1 file.

Apache Spark Optimization

Production patterns for optimizing Apache Spark jobs including partitioning strategies, memory management, shuffle optimization, and performance tuning.

When to Use This Skill

  • Optimizing slow Spark jobs
  • Tuning memory and executor configuration
  • Implementing efficient partitioning strategies
  • Debugging Spark performance issues
  • Scaling Spark pipelines for large datasets
  • Reducing shuffle and data skew

Core Concepts

1. Spark Execution Model

Driver Program
    ↓
Job (triggered by action)
    ↓
Stages (separated by shuffles)
    ↓
Tasks (one per partition)

2. Key Performance Factors

Factor Impact Solution
Shuffle Network I/O, disk I/O Minimize wide transformations
Data Skew Uneven task duration Salting, broadcast joins
Serialization CPU overhead Use Kryo, columnar formats
Memory GC pressure, spills Tune executor memory
Partitions Parallelism Right-size partitions

Quick Start

from pyspark.sql import SparkSession
from pyspark.sql import functions as F

# Create optimized Spark session
spark = (SparkSession.builder
    .appName("OptimizedJob")
    .config("spark.sql.adaptive.enabled", "true")
    .config("spark.sql.adaptive.coalescePartitions.enabled", "true")
    .config("spark.sql.adaptive.skewJoin.enabled", "true")
    .config("spark.serializer", "org.apache.spark.serializer.KryoSerializer")
    .config("spark.sql.shuffle.partitions", "200")
    .getOrCreate())

# Read with optimized settings
df = (spark.read
    .format("parquet")
    .option("mergeSchema", "false")
    .load("s3://bucket/data/"))

# Efficient transformations
result = (df
    .filter(F.col("date") >= "2024-01-01")
    .select("id", "amount", "category")
    .groupBy("category")
    .agg(F.sum("amount").alias("total")))

result.write.mode("overwrite").parquet("s3://bucket/output/")

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Best Practices

Do's

  • Enable AQE - Adaptive query execution handles many issues
  • Use Parquet/Delta - Columnar formats with compression
  • Broadcast small tables - Avoid shuffle for small joins
  • Monitor Spark UI - Check for skew, spills, GC
  • Right-size partitions - 128MB - 256MB per partition

Don'ts

  • Don't collect large data - Keep data distributed
  • Don't use UDFs unnecessarily - Use built-in functions
  • Don't over-cache - Memory is limited
  • Don't ignore data skew - It dominates job time
  • Don't use .count() for existence - Use .take(1) or .isEmpty()

Source: SKILL.md on GitHub

No alerts16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides legitimate technical documentation and code patterns for optimizing Apache Spark jobs. It covers standard performance tuning topics such as partitioning, memory management, and shuffle optimization without any malicious code or security risks.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    1/1 file flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at be57c0b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 4 months ago
  • apache-spark
  • performance-tuning
  • partitioning
  • shuffle-optimization
  • memory-management
  • data-skew
  • pyspark
  • distributed-computing

README badge

README badge for wshobson/agents/spark-optimization

Optimizes Apache Spark jobs through partitioning strategies, memory tuning, shuffle reduction, and data skew mitigation. Targets production Spark pipelines and includes configuration patterns for adaptive query execution, serialization, and broadcast joins.

Generated from the current SKILL.md.

Does this skill cover Spark on Kubernetes or just standalone/YARN clusters?
The skill focuses on general Spark optimization patterns (partitioning, caching, shuffle, memory tuning) that apply across all cluster managers. Kubernetes-specific deployment or resource allocation is not covered.
Does this include Spark Streaming or structured streaming optimization?
No. The skill targets batch Spark jobs and SQL workloads. Streaming-specific patterns like micro-batch tuning or backpressure handling are not included.
What Python and Spark versions does this assume?
The skill uses PySpark and assumes Spark 3.x with Adaptive Query Execution (AQE) available. Specific version pinning is not mentioned in the documentation.
Does this cover Delta Lake optimization or just Parquet?
The skill recommends Delta Lake and Parquet as columnar formats but does not include Delta-specific optimization patterns like Z-order clustering or merge/upsert tuning.
Are GPU or machine learning workloads covered?
No. The skill is limited to CPU-based distributed data processing optimization. GPU acceleration and ML-specific tuning are out of scope.

Generated from the current SKILL.md. These answers refresh after source changes.