<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Adnan</title>
    <description>The latest articles on DEV Community by Adnan (@adnanphp).</description>
    <link>https://dev.to/adnanphp</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4162425%2F061bb1b7-7794-42b5-b6c3-a328df74495c.png</url>
      <title>DEV Community: Adnan</title>
      <link>https://dev.to/adnanphp</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/adnanphp"/>
    <language>en</language>
    <item>
      <title>How I Built a Two-Tier Data Platform in 9 Phases (Spark, Kafka, dbt, Kubernetes)</title>
      <dc:creator>Adnan</dc:creator>
      <pubDate>Sun, 04 Oct 2026 20:37:06 +0000</pubDate>
      <link>https://dev.to/adnanphp/how-i-built-a-two-tier-data-platform-in-9-phases-spark-kafka-dbt-kubernetes-1ebe</link>
      <guid>https://dev.to/adnanphp/how-i-built-a-two-tier-data-platform-in-9-phases-spark-kafka-dbt-kubernetes-1ebe</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — I built &lt;a href="https://github.com/adnanphp/openbi" rel="noopener noreferrer"&gt;OpenBI&lt;/a&gt;, an end-to-end data platform that runs at two scales: a single-node Postgres warehouse on 10K rows, and a distributed Spark + Delta Lake pipeline on 1M+ rows. It also has a streaming tier (Kafka + Structured Streaming) and a Kubernetes deployment (Kind + Terraform). 100+ tests, 2 green CI workflows. This post covers the architecture, the specific bugs I hit, and what I'd do differently.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Why I built this
&lt;/h2&gt;

&lt;p&gt;Most data engineering portfolio projects pick a scale: either a small Postgres demo or a full Spark pipeline. I wanted &lt;strong&gt;both&lt;/strong&gt; — running side-by-side, serving the same dashboards, with &lt;strong&gt;zero code changes&lt;/strong&gt; when switching between them.&lt;/p&gt;

&lt;p&gt;That constraint turned out to be interesting because it forced me to think about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What changes when you go from 10K to 1M rows&lt;/li&gt;
&lt;li&gt;Where the BI layer &lt;em&gt;should&lt;/em&gt; be decoupled from compute&lt;/li&gt;
&lt;li&gt;What the "serving layer" actually is&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The architecture
&lt;/h2&gt;

&lt;p&gt;The platform has three tiers.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tier 1 — v1: Postgres + pandas (10K rows)
&lt;/h3&gt;

&lt;p&gt;The original version. Python ETL loads a CSV into a staging table, transforms it into a star schema (5 dimensions + 1 fact), and computes 5 KPI views. A scikit-learn layer does RFM + KMeans segmentation and ETS forecasting. Results are served via Superset and FastAPI.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CSV → staging → star schema → KPI views → Superset / FastAPI
                                      ├── RFM + KMeans
                                      └── ETS forecasting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Revenue processed:&lt;/strong&gt; $2,297,200.86 (verified against source data).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0inlxvz5mk63qqe66m6p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0inlxvz5mk63qqe66m6p.png" alt="Executive dashboard" width="800" height="356"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Tier 2 — v2: Spark + Delta Lake (1M rows)
&lt;/h3&gt;

&lt;p&gt;Same domain, but distributed. A synthetic 1M-row Parquet dataset is ingested by PySpark into a Delta Bronze table, transformed into a Delta Silver star schema (partitioned by year), and aggregated into Delta Gold tables. Those are published to Postgres via JDBC.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Parquet → Delta Bronze → Delta Silver → Delta Gold → Postgres warehouse_big
                                      ├── Spark MLlib (RFM + KMeans)
                                      └── Spark MLlib forecasting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Revenue processed:&lt;/strong&gt; $287,833,061.24 (100× the v1 scale).&lt;/p&gt;

&lt;h3&gt;
  
  
  Tier 3 — Streaming: Kafka + Structured Streaming
&lt;/h3&gt;

&lt;p&gt;Real-time orders flow through Kafka, get consumed by a Spark Structured Streaming job, and land in Delta Bronze. A second streaming job publishes micro-batches to a Postgres table.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;producer.py → Kafka → Structured Streaming → Delta Bronze → Postgres
                              │
                              ▼
                    Superset real-time dashboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Events processed during the demo run:&lt;/strong&gt; 12,000+.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgknxibp6vii34794v5y2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgknxibp6vii34794v5y2.png" alt="Streaming dashboard" width="800" height="251"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Serving layer
&lt;/h3&gt;

&lt;p&gt;The BI layer is decoupled from compute:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Superset&lt;/strong&gt; dashboards read from &lt;code&gt;warehouse&lt;/code&gt; (v1) or &lt;code&gt;warehouse_big&lt;/code&gt; (v2). Same column names, same charts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FastAPI&lt;/strong&gt; exposes &lt;code&gt;/kpis&lt;/code&gt;, &lt;code&gt;/customers&lt;/code&gt;, &lt;code&gt;/forecasts&lt;/code&gt; — same endpoints regardless of which schema backs them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Zero BI code changes&lt;/strong&gt; when switching between v1 and v2. That's the whole point.&lt;/p&gt;




&lt;h2&gt;
  
  
  The stack
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Tools&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ingestion&lt;/td&gt;
&lt;td&gt;pandas, PySpark, Kafka&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;PostgreSQL, Delta Lake (Parquet)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compute&lt;/td&gt;
&lt;td&gt;Apache Spark (batch + streaming), Spark SQL, Spark MLlib&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Orchestration&lt;/td&gt;
&lt;td&gt;Apache Airflow, Make&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transformation&lt;/td&gt;
&lt;td&gt;dbt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BI&lt;/td&gt;
&lt;td&gt;Apache Superset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API&lt;/td&gt;
&lt;td&gt;FastAPI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monitoring&lt;/td&gt;
&lt;td&gt;Prometheus, Grafana, statsd-exporter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Testing&lt;/td&gt;
&lt;td&gt;pytest (100+ tests)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CI&lt;/td&gt;
&lt;td&gt;GitHub Actions (2 workflows)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deployment&lt;/td&gt;
&lt;td&gt;Docker Compose, Kubernetes (Kind), Terraform, Kustomize&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All 100% open source. Total cost to run: &lt;strong&gt;$0&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The parts that were hard
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Kind's DNS doesn't behave like Docker Compose's
&lt;/h3&gt;

&lt;p&gt;Kind runs each cluster node as a Docker container. On Linux with &lt;code&gt;systemd-resolved&lt;/code&gt;, the resolver inside a Kind container can't reach the host's DNS. Image pulls fail with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;dial tcp: lookup registry-1.docker.io on 172.19.0.1:53:
server misbehaving
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fix is to pre-load images into the cluster from the host:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker pull postgres:16-alpine
kind load docker-image postgres:16-alpine &lt;span class="nt"&gt;--name&lt;/span&gt; openbi
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For images you build yourself (like the FastAPI service), you build on the host and load them the same way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker build &lt;span class="nt"&gt;-t&lt;/span&gt; openbi-fastapi:latest fastapi-app/
kind load docker-image openbi-fastapi:latest &lt;span class="nt"&gt;--name&lt;/span&gt; openbi
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the manifest uses &lt;code&gt;imagePullPolicy: IfNotPresent&lt;/code&gt;, which tells Kubernetes to use the loaded image instead of pulling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why this matters:&lt;/strong&gt; cloud emulators are not the cloud. Kind is close, but its networking is different enough to break things that work in a real cluster.&lt;/p&gt;




&lt;h3&gt;
  
  
  2. Terraform doesn't expand &lt;code&gt;~&lt;/code&gt; in path variables
&lt;/h3&gt;

&lt;p&gt;I set:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;variable&lt;/span&gt; &lt;span class="s2"&gt;"kubeconfig_path"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;default&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"~/.kube/config"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Terraform interpreted &lt;code&gt;~&lt;/code&gt; as a literal directory name. It created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;infrastructure/terraform/~/.kube/config
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;inside my project — and worse, I committed it to git.&lt;/p&gt;

&lt;p&gt;That kubeconfig contains client certificates. Anyone with access to the public repo could have connected to my cluster.&lt;/p&gt;

&lt;p&gt;The fix:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Removed the file from git history (&lt;code&gt;git rm --cached&lt;/code&gt;, then rewrote the commit)&lt;/li&gt;
&lt;li&gt;Added &lt;code&gt;infrastructure/terraform/~/&lt;/code&gt; and &lt;code&gt;**/kubeconfig&lt;/code&gt; to &lt;code&gt;.gitignore&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Changed the variable default to an absolute path:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/home/adnan/.kube/config
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The lesson:&lt;/strong&gt; never trust &lt;code&gt;~&lt;/code&gt; in any IaC tool. Always use absolute paths. And audit &lt;code&gt;git status&lt;/code&gt; carefully — I would have missed this if I hadn't grepped the commit output.&lt;/p&gt;




&lt;h3&gt;
  
  
  3. Kubernetes probes need longer initial delays than you think
&lt;/h3&gt;

&lt;p&gt;Superset takes 30–60 seconds to warm up on first start. My original readiness probe had &lt;code&gt;initialDelaySeconds: 5&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Kubernetes killed the pod before it finished initializing, restarted it, killed it again — a crash loop that looked like Superset was broken.&lt;/p&gt;

&lt;p&gt;The actual fix was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;readinessProbe&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;httpGet&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/health&lt;/span&gt;
    &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http&lt;/span&gt;
  &lt;span class="na"&gt;initialDelaySeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;30&lt;/span&gt;
  &lt;span class="na"&gt;periodSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;
  &lt;span class="na"&gt;failureThreshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;
  &lt;span class="na"&gt;timeoutSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Superset's first start now succeeds. Subsequent starts are faster because the metadata DB is warm.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzku7fb1sz08xhedgvw5n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzku7fb1sz08xhedgvw5n.png" alt="Kubernetes pods" width="799" height="546"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  4. Superset's CLI has a bug
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;superset run&lt;/code&gt; in Superset 3.1.3 fails with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: 'tcp' is not a valid port number.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fix is to call the image's own &lt;code&gt;run-server.sh&lt;/code&gt; directly, which wraps Gunicorn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/usr/bin/run-server.sh&lt;/span&gt;

&lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;SUPERSET_BIND_ADDRESS&lt;/span&gt;
    &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0.0.0.0"&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;SUPERSET_PORT&lt;/span&gt;
    &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8088"&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;FLASK_APP&lt;/span&gt;
    &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;superset.app:create_app()"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;superset run&lt;/code&gt; wrapper is what breaks — the underlying Gunicorn server is fine. This is documented in several Superset GitHub issues.&lt;/p&gt;




&lt;h3&gt;
  
  
  5. Kustomize doesn't always detect file changes
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;kubectl apply -k .&lt;/code&gt; sometimes uses a cached version of the manifest, even after you edit the file.&lt;/p&gt;

&lt;p&gt;If the deployed resource doesn't match what's on disk, force a clean apply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl delete &lt;span class="nt"&gt;-k&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
kubectl apply &lt;span class="nt"&gt;-k&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or, for a single resource:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl delete deployment superset &lt;span class="nt"&gt;-n&lt;/span&gt; openbi
kubectl apply &lt;span class="nt"&gt;-k&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why this happens: Kustomize hashes each resource. If two applies happen close together, the second can use the first's cached state.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Set up CI on day one
&lt;/h3&gt;

&lt;p&gt;I hit the same class of bug — missing dependency — three times in different layers.&lt;/p&gt;

&lt;p&gt;Each time, the fix was one line in a &lt;code&gt;requirements.txt&lt;/code&gt; or a Dockerfile. But I only caught them because I ran CI after pushing.&lt;/p&gt;

&lt;p&gt;If I'd set up GitHub Actions in Phase 1, I would have caught all three in the first hour instead of the third day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Always set up CI before writing production code.&lt;/strong&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Add a backup target earlier
&lt;/h3&gt;

&lt;p&gt;I lost a Postgres volume mid-project because I ran:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker compose down &lt;span class="nt"&gt;-v&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker compose down
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;-v&lt;/code&gt; option wipes volumes. Rebuilding from scratch took about 15 minutes.&lt;/p&gt;

&lt;p&gt;The fix is trivial:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight make"&gt;&lt;code&gt;&lt;span class="nl"&gt;backup&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
    &lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; backups
    docker compose &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-T&lt;/span&gt; postgres pg_dump &lt;span class="nt"&gt;-U&lt;/span&gt; openbi openbi &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
        backups/openbi_&lt;span class="p"&gt;$$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%Y%m%d_%H%M%S&lt;span class="p"&gt;)&lt;/span&gt;.sql
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make backup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;before any risky operation.&lt;/p&gt;




&lt;h3&gt;
  
  
  Write the tests as I went, not after
&lt;/h3&gt;

&lt;p&gt;I wrote tests in batches after each phase.&lt;/p&gt;

&lt;p&gt;If I'd written them incrementally, I would have caught bugs earlier — and I wouldn't have had to reverse-engineer test cases from working code.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The platform is functional: &lt;strong&gt;100+ tests passing, 2 green CI workflows.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The next steps I'd consider:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ingress controller&lt;/strong&gt; — expose services on real hostnames&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Helm chart&lt;/strong&gt; — package OpenBI for one-command installation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real cloud deployment&lt;/strong&gt; — GKE free tier&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Companion ML paper&lt;/strong&gt; — I already published one on a related experiment&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/adnanphp/openbi" rel="noopener noreferrer"&gt;github.com/adnanphp/openbi&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kubernetes docs:&lt;/strong&gt; &lt;a href="https://github.com/adnanphp/openbi/blob/main/docs/kubernetes.md" rel="noopener noreferrer"&gt;docs/kubernetes.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;v2.0.0 release:&lt;/strong&gt; &lt;a href="https://github.com/adnanphp/openbi/releases/tag/v2.0.0" rel="noopener noreferrer"&gt;GitHub Release&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're building something similar — or hitting any of the same bugs — I'd love to hear about it in the comments.&lt;/p&gt;




</description>
      <category>dataengineering</category>
      <category>spark</category>
      <category>kafka</category>
      <category>kubernetes</category>
    </item>
  </channel>
</rss>
