Data pipelines extract data from sources, transform it and load it into a destination. The order matters.
ETL: Extract, Transform, Load
Data is transformed before loading into the warehouse, typically in a separate processing tool.
Pros: only clean, shaped data reaches the warehouse; sensitive fields can be removed early; suited to limited warehouse capacity. Cons: transformations are harder to change; raw data may be lost; separate infrastructure to run.
ELT: Extract, Load, Transform
Raw data is loaded first, then transformed inside the warehouse using SQL.
Pros: raw data is preserved, so transformations can be changed and re-run; uses the warehouse's scalable computing; transformations written in SQL are accessible to analysts. Cons: raw data, including sensitive data, lands in the warehouse and must be governed; warehouse compute costs can grow.
Why ELT Became Popular
Cloud warehouses made storage cheap and computing elastic, and tools like dbt made in-warehouse transformation manageable with version control and tests.
Choosing
- Use ELT as a default with modern cloud warehouses and analytics workloads.
- Use ETL (or transform-before-load for specific fields) when data must be cleaned, masked or reduced before it lands — for privacy, regulation or very large raw volumes.
In Practice
Many pipelines mix both: light cleansing and masking on the way in, heavier modelling in the warehouse.