DEV Community

Armen-Aris Shahinyan
Armen-Aris Shahinyan

Posted on

How much of your ML workflow is still held together manually?

I've been working on a DataOps/MLOps product called Datryc, and before we go much further with it, I want to challenge some of the assumptions we're making.

The idea came from a fairly simple observation: getting from raw data to a model running somewhere involves much more than just training the model.

You may need to:

  • connect to different data sources;
  • validate and clean the data;
  • build preprocessing pipelines;
  • monitor data quality;
  • train and evaluate models;
  • version model artifacts;
  • deploy models for inference;
  • monitor what happens afterward.

There are excellent tools solving individual parts of this problem. What I'm interested in is the integration cost between those parts.

With Datryc, we're experimenting with putting these workflows behind a common API/CLI while keeping the underlying components modular. Internally, we're using asynchronous workers and Kafka so that pipeline execution isn't coupled directly to the API layer.

But there's a danger when building infrastructure products: solving the architecture before proving that the workflow is actually painful for users.

So I'd really like to hear from people working with production data/ML systems:

What is the most painful or unnecessarily manual part of your current Data/ML workflow?

I'm particularly curious about:

  • data preparation and quality;
  • moving pipelines between environments;
  • connecting different ML/data tools;
  • model deployment;
  • monitoring;
  • reproducibility;
  • integrating ML workflows into existing CI/CD.

And if your current stack already handles all of this well, I'd love to know what you're using.

I'm building Datryc, so I'm obviously not approaching the problem as a neutral observer. I'm mainly interested in finding out where our assumptions are wrong before we build too much around them.

Top comments (0)