Data Engineering - Intermediate

Apache Spark

Process large datasets with Spark SQL, DataFrames, partitioning, and performance-aware transformations.

Instructor: OSEKOO ACADEMY - 3 units

This module can be taken independently, in any order.

Course overview

Move from single-machine processing to distributed computation. Understand Spark execution, design efficient transformations, and ship a maintainable analytics workload.

Standalone unit guides

Every unit is designed to be followed independently: review the guide, complete the lessons in order, finish the practical lab, pass the checkpoint, and deliver the unit project.

1. Distributed Processing3 lessons

Lesson sequence

  1. Spark architecture and execution model - Free preview
  2. DataFrames, schemas, and types
  3. Transformations and actions
2. Spark SQL and Storage3 lessons

Lesson sequence

  1. Spark SQL workflows
  2. Joins and window functions
  3. Parquet and partitioning
3. Performance and Delivery3 lessons

Lesson sequence

  1. Query plans and optimization
  2. Testing Spark applications
  3. Capstone: distributed ETL
Verified learner reviews

Learner feedback

No published reviews yet. Enrolled learners can submit the first review.