DEV Community

Hariesh Kai
Hariesh Kai

Posted on

Stop Writing the Same Pandas Boilerplate: How We Built a Visual Pipeline Studio for ML Preprocessing

Hey everyone,

Every time I start a new machine learning project or Kaggle competition, I end up spending the first hour doing the exact same chores:

  • handling null values and missing label rows
  • writing manual IQR outlier clipping formulas
  • coercing corrupt object columns to numeric/datetimes
  • one-hot / label encoding categories
  • trimming invisible whitespace

I built DataForger to make preprocessing fast, modular, and visual:

What it does:

  1. Drag-and-Drop Ingestion: Instant health check, missing value diagnostics, and distributions.
  2. Visual Node Pipeline: Drag, connect, and reorder transformation stages on a React Flow graph canvas.
  3. Live Diff Previews: Inspect exactly what changed before and after applying an operation.
  4. Export Clean Artifacts: Download the cleaned .csv + the reproducible JSON pipeline configuration to use in your training code.

Early Access & Feedback:
I'm rolling this out as a Founding Member Early Access Pass for ₹299 INR (~$3.50 USD one-time lifetime) for the first 50 users before we switch to a subscription model.

Try it out here: [https://shirogani-dataforger-beta.vercel.app]

I really want your raw feedback:

  • What's the most annoying preprocessing step in your daily workflow that you wish was automated?
  • Does the node graph feel snappy or would you prefer a linear list view?

I'll be in the comments all day to answer questions and patch any bugs you find!

Top comments (0)