Disclosure: I wrote this handbook. It is free and MIT licensed.
Tutorials rot quietly. A dependency changes, an option is renamed, a default flips, and the instructions stop working, usually for the first reader who follows them on a clean machine.
I maintain an open-source data engineering handbook with eight hands-on labs and two capstones: SQL on DuckDB, dbt, Spark with Delta Lake, Kafka, Airflow, data quality gates with Great Expectations, Iceberg, and change data capture with Debezium. They all run locally on one e-commerce dataset with planted defects. To stop them rotting, every lab runs in CI. Five of them (Kafka, Airflow, data quality, Iceberg and Debezium) also have a script that replays the README exercises against the real service and asserts the numbers the README quotes.
That work found more mistakes in my own instructions than I expected. Here are six.
1. The Airflow lab failed for every Linux user on a fresh clone
The first CI run of the Airflow lab failed on every task:
PermissionError: [Errno 13] Permission denied: '/opt/lab/warehouse/landing'
The lab bind-mounts ./warehouse into the container. That folder is git-ignored, so a fresh clone does not have it, and Docker creates a missing bind-mount source owned by root. The Airflow process runs as another user and cannot write to it. I had no way to run Docker where I wrote the lab, so CI was the first real run, and it found what a reader on Linux would have hit. The fix is one line, mkdir -p warehouse, before docker compose up.
2. An Iceberg tutorial option no longer works
The Iceberg guide in my own handbook read an old snapshot with spark.read.option("snapshot-id", ...). On Spark 4.1 with Iceberg 1.11 it raises:
Time travel option `snapshot-id` is no longer supported, use Spark built-in `versionAsOf` instead
as-of-timestamp fails the same way; the replacements are versionAsOf and timestampAsOf. Pinning also mattered. Iceberg publishes its Spark runtime only for some Spark versions, and there is still no runtime for Spark 4.2 (checked again on Maven Central the day this went up), while the neighbouring Delta lab runs 4.2. Pin the engine and the table format together, and say why.
3. In Debezium, "before" is usually empty
Every Debezium explanation shows an update event with a before and an after. On Postgres, with the default replica identity, an update's before is null, and a delete's before holds only the primary key, with placeholder values (empty strings, 1970 timestamps) in the other columns. Reading a non-key column from a delete's before gives you plausible-looking garbage. ALTER TABLE ... REPLICA IDENTITY FULL makes the old row available, at the cost of more data in the write-ahead log.
4. A MERGE that changed nothing still rewrote the table
In the Iceberg lab, a MERGE that inserted 200 rows and updated 1 removed 6,000 records and added 6,200. Iceberg's default is copy-on-write: any data file with a matched row is rewritten in full, even when the row does not change. My first pipeline merged the whole source on every run, so a rerun with no new data rewrote every file. The fix is to send only new or changed rows. Merge-on-read writes a small file and a delete file instead (2 rows written for a 2-row update, versus 6,200 rewritten), at a cost paid on every read until compaction.
5. In Great Expectations, a warning still fails the run
Great Expectations lets an expectation carry severity="warning". But result.success is False as soon as any expectation fails, warnings included. A gate that blocks only on critical failures has to read the severity of each result itself. If it checks success, every warning stops the pipeline.
6. The flakiest thing in the repo was my own test
The most recent failure was in the test I wrote to catch these problems. The Iceberg smoke test passed on the pull request and failed after the merge, with row counts that ignored a commit made moments earlier. The cause: Iceberg's catalog caches each table's metadata in memory for 30 seconds, so a session can miss a commit made by another process inside that window. Whether the test passed depended on timing. I reproduced it on demand (6,000 rows read instead of 6,200 with the cache on) and turned the cache off in the lab, which is also the right setting for a lab that writes in one process and reads in another.
How the tests are checked
Two habits made the tests worth having.
Assert the numbers the README quotes. If the README says a delete produces 10 events plus 10 tombstones, the test counts them in Kafka.
Break the reference solution on purpose. A test that cannot fail proves nothing. For each lab I mutate the reference answer (use max instead of min for a Kafka watermark, delete the DELETE from an idempotent load, keep the oldest change per key instead of the newest) and confirm the test fails. This finds gaps in tests. While testing a SQL guard for an AI assistant, two of my mutations survived: removing a rule that rejects schema-qualified table names, and removing the one that rejects table functions, and every test still passed. I added a test for each. It is the only way I know to find a gap in a test.
What I could not check
Some things I could not run where I wrote them. There is no Docker in my development environment, so the Compose-based labs were checked by running the real services natively (Postgres, Kafka, Kafka Connect and the Debezium connector as plain processes) and then by CI on GitHub. The pull request says what was run locally and what was left to CI. Every guide in the handbook now carries a review date I can stand behind — the last batch, covering the remaining non-AI guides, wrapped up the same week this article was checked. Before that, most pages said Not yet individually reviewed rather than claiming a check that never happened, because their date was a shared starting point, not a review.
Try it, and tell me what is missing
The labs are in the repository. Each starts with pip install and a dataset generator; nothing to sign up for. I would like to know which lab is missing, and where an exercise is too easy or too guided.
Top comments (0)