Apache Spark is an open-source engine for processing large datasets in parallel across a cluster of machines.
Why Spark Exists
When data is too large for one machine's memory, work must be split across many. Spark distributes data and computation, handling failures and data movement for you.
Key Ideas
- DataFrames: distributed tables, similar in spirit to pandas DataFrames, that you query with SQL or a Python, Scala or R API.
- Lazy evaluation: transformations build a plan; nothing runs until an action (such as writing results) requires it, letting Spark optimise the whole plan.
- Partitions: data is split into chunks processed in parallel.
- Shuffles: operations such as joins and group-bys move data between machines and are the main performance cost.
What It's Used For
Large-scale ETL, joining and aggregating big datasets, feature engineering for machine learning, streaming (Structured Streaming) and distributed machine learning.
Example
from pyspark.sql import SparkSession, functions as F
spark = SparkSession.builder.getOrCreate()
trips = spark.read.parquet("s3://bucket/trips/")
daily = trips.groupBy(F.to_date("pickup_time").alias("day")).agg(F.sum("fare").alias("revenue"))
daily.write.parquet("s3://bucket/daily_revenue/")
Do You Need It?
Often not. Tools such as DuckDB and Polars process tens or hundreds of gigabytes on a single machine quickly, and cloud warehouses handle large SQL workloads. Choose Spark when data volume, existing platforms or processing needs genuinely require a cluster.
Performance Tips
Filter early, select only needed columns, use columnar formats, and watch for skewed keys that overload single partitions.