Data Engineering Learning Path

From Python scripts to trusted real-time platforms.

A focused, evolving curriculum for engineers learning reliable batch, distributed, streaming, and event-driven data systems.

Your learning modules

7 modules
Python + Scala
Typed foundations
Kafka + Spark
Production systems
7technical courses
30structured units
178structured lessons, including premium lessons in Python Basics
1flexible curriculum
The curriculum

Choose the module that matches your goals

Beginner

Python Basics

Learn Python 3 through a beginner-first progression of guided practice, mini-projects, debugging, testing, and a final command-line application.

12 units
View course →
Intermediate

Scala

Build type-safe data applications with Scala 3, immutable domain models, functional transformations, concurrency, and production-grade testing.

3 units
View course →
Intermediate

Apache Spark

Process large datasets with Spark SQL, DataFrames, partitioning, and performance-aware transformations.

3 units
View course →
Intermediate

Data Quality

Build trustworthy data products with profiling, contracts, automated validation, and quality observability.

3 units
View course →
Advanced

Apache Kafka

Build event-driven systems with Kafka topics, partitions, consumer groups, delivery semantics, and schemas.

3 units
View course →
Engineering-first learning

Understand, practise, progress

01

Learn the model

Study execution, quality, state, delivery guarantees, and failure modes.

02

Apply the concepts

Work through technical examples, checkpoints, and the growing set of guided labs.

03

Build sound judgement

Learn to reason about testing, monitoring, optimization, and operational trade-offs.

Start with the module that fits your goals.

Choose any technical module and progress at your own pace.

Create your account