DEV Community

Cover image for How I Built a Two-Tier Data Platform in 9 Phases (Spark, Kafka, dbt, Kubernetes)
Adnan
Adnan

Posted on

How I Built a Two-Tier Data Platform in 9 Phases (Spark, Kafka, dbt, Kubernetes)

TL;DR — I built OpenBI, an end-to-end data platform that runs at two scales: a single-node Postgres warehouse on 10K rows, and a distributed Spark + Delta Lake pipeline on 1M+ rows. It also has a streaming tier (Kafka + Structured Streaming) and a Kubernetes deployment (Kind + Terraform). 100+ tests, 2 green CI workflows. This post covers the architecture, the specific bugs I hit, and what I'd do differently.


Why I built this

Most data engineering portfolio projects pick a scale: either a small Postgres demo or a full Spark pipeline. I wanted both — running side-by-side, serving the same dashboards, with zero code changes when switching between them.

That constraint turned out to be interesting because it forced me to think about:

  • What changes when you go from 10K to 1M rows
  • Where the BI layer should be decoupled from compute
  • What the "serving layer" actually is

The architecture

The platform has three tiers.

Tier 1 — v1: Postgres + pandas (10K rows)

The original version. Python ETL loads a CSV into a staging table, transforms it into a star schema (5 dimensions + 1 fact), and computes 5 KPI views. A scikit-learn layer does RFM + KMeans segmentation and ETS forecasting. Results are served via Superset and FastAPI.

CSV → staging → star schema → KPI views → Superset / FastAPI
                                      ├── RFM + KMeans
                                      └── ETS forecasting
Enter fullscreen mode Exit fullscreen mode

Revenue processed: $2,297,200.86 (verified against source data).

Executive dashboard

Tier 2 — v2: Spark + Delta Lake (1M rows)

Same domain, but distributed. A synthetic 1M-row Parquet dataset is ingested by PySpark into a Delta Bronze table, transformed into a Delta Silver star schema (partitioned by year), and aggregated into Delta Gold tables. Those are published to Postgres via JDBC.

Parquet → Delta Bronze → Delta Silver → Delta Gold → Postgres warehouse_big
                                      ├── Spark MLlib (RFM + KMeans)
                                      └── Spark MLlib forecasting
Enter fullscreen mode Exit fullscreen mode

Revenue processed: $287,833,061.24 (100× the v1 scale).

Tier 3 — Streaming: Kafka + Structured Streaming

Real-time orders flow through Kafka, get consumed by a Spark Structured Streaming job, and land in Delta Bronze. A second streaming job publishes micro-batches to a Postgres table.

producer.py → Kafka → Structured Streaming → Delta Bronze → Postgres
                              │
                              ▼
                    Superset real-time dashboard
Enter fullscreen mode Exit fullscreen mode

Events processed during the demo run: 12,000+.

Streaming dashboard

Serving layer

The BI layer is decoupled from compute:

  • Superset dashboards read from warehouse (v1) or warehouse_big (v2). Same column names, same charts.
  • FastAPI exposes /kpis, /customers, /forecasts — same endpoints regardless of which schema backs them.

Zero BI code changes when switching between v1 and v2. That's the whole point.


The stack

Layer Tools
Ingestion pandas, PySpark, Kafka
Storage PostgreSQL, Delta Lake (Parquet)
Compute Apache Spark (batch + streaming), Spark SQL, Spark MLlib
Orchestration Apache Airflow, Make
Transformation dbt
BI Apache Superset
API FastAPI
Monitoring Prometheus, Grafana, statsd-exporter
Testing pytest (100+ tests)
CI GitHub Actions (2 workflows)
Deployment Docker Compose, Kubernetes (Kind), Terraform, Kustomize

All 100% open source. Total cost to run: $0.


The parts that were hard

1. Kind's DNS doesn't behave like Docker Compose's

Kind runs each cluster node as a Docker container. On Linux with systemd-resolved, the resolver inside a Kind container can't reach the host's DNS. Image pulls fail with:

dial tcp: lookup registry-1.docker.io on 172.19.0.1:53:
server misbehaving
Enter fullscreen mode Exit fullscreen mode

The fix is to pre-load images into the cluster from the host:

docker pull postgres:16-alpine
kind load docker-image postgres:16-alpine --name openbi
Enter fullscreen mode Exit fullscreen mode

For images you build yourself (like the FastAPI service), you build on the host and load them the same way:

docker build -t openbi-fastapi:latest fastapi-app/
kind load docker-image openbi-fastapi:latest --name openbi
Enter fullscreen mode Exit fullscreen mode

Then the manifest uses imagePullPolicy: IfNotPresent, which tells Kubernetes to use the loaded image instead of pulling.

Why this matters: cloud emulators are not the cloud. Kind is close, but its networking is different enough to break things that work in a real cluster.


2. Terraform doesn't expand ~ in path variables

I set:

variable "kubeconfig_path" {
  default = "~/.kube/config"
}
Enter fullscreen mode Exit fullscreen mode

Terraform interpreted ~ as a literal directory name. It created:

infrastructure/terraform/~/.kube/config
Enter fullscreen mode Exit fullscreen mode

inside my project — and worse, I committed it to git.

That kubeconfig contains client certificates. Anyone with access to the public repo could have connected to my cluster.

The fix:

  • Removed the file from git history (git rm --cached, then rewrote the commit)
  • Added infrastructure/terraform/~/ and **/kubeconfig to .gitignore
  • Changed the variable default to an absolute path:
/home/adnan/.kube/config
Enter fullscreen mode Exit fullscreen mode

The lesson: never trust ~ in any IaC tool. Always use absolute paths. And audit git status carefully — I would have missed this if I hadn't grepped the commit output.


3. Kubernetes probes need longer initial delays than you think

Superset takes 30–60 seconds to warm up on first start. My original readiness probe had initialDelaySeconds: 5.

Kubernetes killed the pod before it finished initializing, restarted it, killed it again — a crash loop that looked like Superset was broken.

The actual fix was:

readinessProbe:
  httpGet:
    path: /health
    port: http
  initialDelaySeconds: 30
  periodSeconds: 10
  failureThreshold: 10
  timeoutSeconds: 5
Enter fullscreen mode Exit fullscreen mode

Superset's first start now succeeds. Subsequent starts are faster because the metadata DB is warm.

Kubernetes pods


4. Superset's CLI has a bug

superset run in Superset 3.1.3 fails with:

Error: 'tcp' is not a valid port number.
Enter fullscreen mode Exit fullscreen mode

The fix is to call the image's own run-server.sh directly, which wraps Gunicorn:

command:
  - /usr/bin/run-server.sh

env:
  - name: SUPERSET_BIND_ADDRESS
    value: "0.0.0.0"
  - name: SUPERSET_PORT
    value: "8088"
  - name: FLASK_APP
    value: "superset.app:create_app()"
Enter fullscreen mode Exit fullscreen mode

The superset run wrapper is what breaks — the underlying Gunicorn server is fine. This is documented in several Superset GitHub issues.


5. Kustomize doesn't always detect file changes

kubectl apply -k . sometimes uses a cached version of the manifest, even after you edit the file.

If the deployed resource doesn't match what's on disk, force a clean apply:

kubectl delete -k .
kubectl apply -k .
Enter fullscreen mode Exit fullscreen mode

Or, for a single resource:

kubectl delete deployment superset -n openbi
kubectl apply -k .
Enter fullscreen mode Exit fullscreen mode

Why this happens: Kustomize hashes each resource. If two applies happen close together, the second can use the first's cached state.


What I'd do differently

Set up CI on day one

I hit the same class of bug — missing dependency — three times in different layers.

Each time, the fix was one line in a requirements.txt or a Dockerfile. But I only caught them because I ran CI after pushing.

If I'd set up GitHub Actions in Phase 1, I would have caught all three in the first hour instead of the third day.

Always set up CI before writing production code.


Add a backup target earlier

I lost a Postgres volume mid-project because I ran:

docker compose down -v
Enter fullscreen mode Exit fullscreen mode

instead of:

docker compose down
Enter fullscreen mode Exit fullscreen mode

The -v option wipes volumes. Rebuilding from scratch took about 15 minutes.

The fix is trivial:

backup:
    @mkdir -p backups
    docker compose exec -T postgres pg_dump -U openbi openbi > \
        backups/openbi_$$(date +%Y%m%d_%H%M%S).sql
Enter fullscreen mode Exit fullscreen mode

Run:

make backup
Enter fullscreen mode Exit fullscreen mode

before any risky operation.


Write the tests as I went, not after

I wrote tests in batches after each phase.

If I'd written them incrementally, I would have caught bugs earlier — and I wouldn't have had to reverse-engineer test cases from working code.


What's next

The platform is functional: 100+ tests passing, 2 green CI workflows.

The next steps I'd consider:

  • Ingress controller — expose services on real hostnames
  • Helm chart — package OpenBI for one-command installation
  • Real cloud deployment — GKE free tier
  • Companion ML paper — I already published one on a related experiment

Links

If you're building something similar — or hitting any of the same bugs — I'd love to hear about it in the comments.


Top comments (0)