DEV Community

Cover image for From State Machines to Configurable Workflows
Mina Golzari Dalir
Mina Golzari Dalir

Posted on

From State Machines to Configurable Workflows

There is a stage in software development where the hardest problem is no longer writing the code.

It is deciding whether the architecture you chose six months ago is still the right architecture today.

Recently, I found myself dealing with exactly this situation while designing a workflow-heavy .NET system.

The initial problem looked like a perfect use case for a state machine.

An entity had a lifecycle.

It moved from one state to another.

Some transitions were allowed, some were not.

The rules could be expressed clearly in code.

So we implemented a state-machine-oriented design.

It worked.

And that is important.

The state machine was not a bad decision.

In fact, it was the right decision for the problem we had at that time.

The problem appeared later, when the business scenarios started changing.

New scenarios appeared.

Different customers needed different processes.

Some steps became optional.

Some steps were conditional.

Some processes needed additional approval.

Existing running processes had to continue using their original rules.
At that point, the question changed from:

"What state is this entity in?"

to:

"Which business process should this entity follow?"

That distinction completely changed my view of the architecture.

This article explains that evolution.

1. The Original Problem

Let's use a fictional example.

Imagine we are building an Order Fulfillment Platform for an e-commerce company.

An order goes through a process like:

Created
   ↓
Payment Verified
   ↓
Inventory Reserved
   ↓
Packed
   ↓
Shipped
   ↓
Delivered
Enter fullscreen mode Exit fullscreen mode

Initially, this looks like a textbook state machine.

We can define:

public enum OrderState
{
    Created,
    PaymentVerified,
    InventoryReserved,
    Packed,
    Shipped,
    Delivered,
    Cancelled
}

Enter fullscreen mode Exit fullscreen mode

And events:

public enum OrderEvent
{
    VerifyPayment,
    ReserveInventory,
    Pack,
    Ship,
    Deliver,
    Cancel
}
Enter fullscreen mode Exit fullscreen mode

Then transitions:

Created
   ├── VerifyPayment → PaymentVerified
   └── Cancel → Cancelled

PaymentVerified
   ├── ReserveInventory → InventoryReserved
   └── Cancel → Cancelled

InventoryReserved
   └── Pack → Packed

Packed
   └── Ship → Shipped

Shipped
   └── Deliver → Delivered
Enter fullscreen mode Exit fullscreen mode

This is clean.

It is easy to understand.

It is easy to test.

And most importantly, it protects the business rules.

An order cannot simply jump from:

Created → Delivered

because that transition doesn't exist.

For a stable process, this is exactly what we want.

2. Why We Initially Chose a State Machine

A state machine gave us several useful properties.

Explicit states

The possible lifecycle states were clearly defined.

Explicit transitions

We could see exactly which transitions were legal.

Centralized rules

Instead of scattering conditions throughout controllers and services, transition rules were concentrated in one place.

Type safety

The compiler knew about states and events.

Testability

We could easily test:

Given Created
When VerifyPayment
Then PaymentVerified
Enter fullscreen mode Exit fullscreen mode

and:

Given Created
When Deliver
Then Reject
Enter fullscreen mode Exit fullscreen mode

This is much safer than having arbitrary status updates throughout the application.

At this point, choosing a state machine was a reasonable engineering decision.

3. Then the Business Changed

The first new requirement was simple:

Premium customers should have priority packing.

We added a condition.

Then another requirement appeared:

Orders above $1,000 require manual fraud review.

So the process became:

Created
   ↓
Payment Verified
   ↓
Fraud Review
   ↓
Inventory Reserved
   ↓
Packed
   ↓
Shipped
Enter fullscreen mode Exit fullscreen mode

But only for some orders.

Then:

Digital products don't need shipping.

Now we had:

Physical Order:

Payment
 ↓
Inventory
 ↓
Packing
 ↓
Shipping
 ↓
Delivery
Enter fullscreen mode Exit fullscreen mode

while:

Digital Order:

Payment
 ↓
Fulfillment
 ↓
Completed
Enter fullscreen mode Exit fullscreen mode

Then another customer introduced:

Orders containing regulated products require compliance verification.

And another:

Enterprise customers require account-manager approval before fulfillment.

The process was no longer one fixed sequence.

It had become a collection of possible processes.

4. The State Machine Started Growing

The state machine itself could still handle these requirements.

That is important.

A state machine is not incapable of handling conditional logic.

We could add guards:

if (order.IsDigital)
{
    // Skip physical fulfillment
}

if (order.TotalAmount > 1000)
{
    // Require fraud review
}

if (order.ContainsRegulatedProducts)
{
    // Require compliance
}

if (order.CustomerType == CustomerType.Enterprise)
{
    // Require account manager approval
}
Enter fullscreen mode Exit fullscreen mode

But something started bothering me.

The state machine was increasingly responsible for understanding:

customer type
product type
order value
risk level
geography
compliance requirements
approval policies
fulfillment strategy
Enter fullscreen mode Exit fullscreen mode

The question was no longer simply:

"What state is the order in?"

The state machine was starting to answer:

"Which business process should this order follow?"

That is a different responsibility.

5. The Real Problem: State vs Process

This became the key architectural insight.

A state describes the current condition of something.

For example:

PaymentVerified

A workflow describes how something should progress through a business process.

For example:

Payment
 → Fraud Review
 → Inventory
 → Packing
 → Shipping
Enter fullscreen mode Exit fullscreen mode

These are related, but they are not the same concept.

A state machine is excellent at answering:

"Can this entity move from state A to state B?"

A workflow engine needs to answer:

"Which steps should this business process execute, in what order, under which conditions?"

Once we recognized this distinction, the architecture became much clearer.

6. What We Actually Needed

We needed the process itself to become configurable.

Instead of hard-coding:

Payment
 → Fraud Review
 → Inventory
 → Packing
 → Shipping
Enter fullscreen mode Exit fullscreen mode

we wanted to define:

Workflow
    ↓
Version
    ↓
Steps
    ↓
Conditions
    ↓
Handlers

Enter fullscreen mode Exit fullscreen mode

For example:

OrderFulfillment
Version 1

1. VerifyPayment
2. ReserveInventory
3. Pack
4. Ship
5. Deliver
Enter fullscreen mode Exit fullscreen mode

Another version could be:

OrderFulfillment
Version 2

1. VerifyPayment
2. FraudReview
3. ReserveInventory
4. Pack
5. Ship
6. Deliver
Enter fullscreen mode Exit fullscreen mode

And a digital-product workflow could be:

DigitalOrderFulfillment
Version 1

1. VerifyPayment
2. PrepareDownload
3. CompleteOrder
Enter fullscreen mode Exit fullscreen mode

The workflow itself is now data.

7. But We Didn't Want a "Everything Is JSON" System

This was an important design constraint.

A configurable workflow can easily become dangerous.

It is tempting to put everything into JSON:

{
  "condition": "order.amount > 1000 && customer.type == 'Enterprise'"
}
Enter fullscreen mode Exit fullscreen mode

Then create a generic interpreter.

Eventually, you have effectively invented your own programming language.

That creates a different class of problems:

difficult debugging
weak type safety
difficult refactoring
runtime failures
complicated validation
security concerns
difficult IDE support

We wanted configurability without abandoning normal .NET engineering practices.

That led us to the handler-based model.

8. The Handler-Based Workflow

The workflow defines what should happen.

The handler defines how it happens.

For example:

Workflow Step
      ↓
Handler Key
      ↓
Handler
Enter fullscreen mode Exit fullscreen mode

A workflow might contain:

VerifyPayment
ReserveInventory
FraudReview
PackOrder
ShipOrder
Enter fullscreen mode Exit fullscreen mode

These map to handlers:

VerifyPayment
      ↓
VerifyPaymentHandler

ReserveInventory
      ↓
ReserveInventoryHandler

FraudReview
      ↓
FraudReviewHandler

PackOrder
      ↓
PackOrderHandler

ShipOrder
      ↓
ShipOrderHandler
Enter fullscreen mode Exit fullscreen mode

A simple interface could look like:

public interface IWorkflowHandler
{
    string Key { get; }

    Task<WorkflowStepResult> ExecuteAsync(
        WorkflowContext context,
        CancellationToken cancellationToken);
}
Enter fullscreen mode Exit fullscreen mode

An implementation:

public sealed class VerifyPaymentHandler
    : IWorkflowHandler
{
    public string Key => "VerifyPayment";

    public async Task<WorkflowStepResult> ExecuteAsync(
        WorkflowContext context,
        CancellationToken cancellationToken)
    {
        // Payment verification logic

        return WorkflowStepResult.Success();
    }
}

Enter fullscreen mode Exit fullscreen mode

The workflow engine doesn't need to know how payment verification works.

It only knows:

Execute "VerifyPayment"

The handler contains the actual business/application logic.

9. Handler Resolution

The application can register handlers using normal .NET dependency injection:

services.AddScoped<IWorkflowHandler, VerifyPaymentHandler>();
services.AddScoped<IWorkflowHandler, ReserveInventoryHandler>();
services.AddScoped<IWorkflowHandler, FraudReviewHandler>();
services.AddScoped<IWorkflowHandler, PackOrderHandler>();
services.AddScoped<IWorkflowHandler, ShipOrderHandler>();
Enter fullscreen mode Exit fullscreen mode

A resolver can then find a handler by its key:

public sealed class WorkflowHandlerResolver
{
    private readonly IEnumerable<IWorkflowHandler> _handlers;

    public WorkflowHandlerResolver(
        IEnumerable<IWorkflowHandler> handlers)
    {
        _handlers = handlers;
    }

    public IWorkflowHandler Resolve(string key)
    {
        return _handlers.FirstOrDefault(x => x.Key == key)
            ?? throw new InvalidOperationException(
                $"Handler '{key}' was not found.");
    }
}
Enter fullscreen mode Exit fullscreen mode

This gives us a useful boundary.

The workflow is configurable.

The executable business capabilities remain strongly typed.

10. Workflow Configuration

A workflow definition might conceptually look like:

WorkflowDefinition
-------------------------
Key
Name
Version
Status

Enter fullscreen mode Exit fullscreen mode

And its steps:

WorkflowStep
-------------------------
WorkflowVersionId
Key
Order
Configuration
Enter fullscreen mode Exit fullscreen mode

Transitions:

WorkflowTransition
-------------------------
FromStep
ToStep
Condition
Enter fullscreen mode Exit fullscreen mode

For example:

OrderFulfillment v3

VerifyPayment
      ↓
FraudReview
      ↓
ReserveInventory
      ↓
PackOrder
      ↓
ShipOrder
      ↓
DeliverOrder
Enter fullscreen mode Exit fullscreen mode

The important part is that the workflow engine doesn't need to be changed when we rearrange these existing capabilities.

11. Versioning Changes Everything

This was one of the strongest reasons to move toward a versioned workflow.

Imagine that version 1 is currently being used by thousands of orders.

Then the business introduces mandatory fraud screening.

We don't want to modify the meaning of an existing order's process.

Instead:

OrderFulfillment v1

remains unchanged.

We publish:

OrderFulfillment v2

with:

VerifyPayment
    ↓
FraudReview
    ↓
ReserveInventory
    ↓
PackOrder
    ↓
ShipOrder
Enter fullscreen mode Exit fullscreen mode

New orders use v2.

Existing orders continue using v1.

This gives us historical reproducibility.

We can answer:

"Which process governed this order?"

with:

Order #12345
Workflow: OrderFulfillment
Version: 1

rather than trying to reconstruct the rules from the current application code.

12. The Most Important Rule: A Workflow Instance Has a Version

I would strongly recommend that a running workflow instance stores its workflow version.

Conceptually:

WorkflowInstance
-------------------------
Id
WorkflowDefinitionId
WorkflowVersionId
BusinessEntityId
CurrentStep
Status
StartedAt
CompletedAt
Enter fullscreen mode Exit fullscreen mode

This means:

Order #1001 → Fulfillment v1
Order #1002 → Fulfillment v1
Order #1003 → Fulfillment v2
Order #1004 → Fulfillment v2
Enter fullscreen mode Exit fullscreen mode

The active workflow can evolve without rewriting history.

This becomes especially important for systems involving:

regulations
financial processes
approvals
contracts
compliance
long-running operations
Enter fullscreen mode Exit fullscreen mode

13. What Happens When the Business Changes Again?

Suppose we currently have:

v2:

Payment
 ↓
Fraud Review
 ↓
Inventory
 ↓
Packing
 ↓
Shipping
Enter fullscreen mode Exit fullscreen mode

Now the business says:

Orders above $5,000 require manager approval.

We can create:

v3:

Payment
 ↓
Fraud Review
 ↓
Inventory
 ↓
Manager Approval
 ↓
Packing
 ↓
Shipping
Enter fullscreen mode Exit fullscreen mode

The workflow engine remains unchanged.

If ManagerApprovalHandler already exists, this can potentially be configuration-only.

That is the key advantage.

14. Three Levels of Change

This led me to a useful way of thinking about the architecture.

Level 1 — Change the process

Existing handlers are sufficient.

A → B → C
Enter fullscreen mode Exit fullscreen mode

becomes:

A → C → B
Enter fullscreen mode Exit fullscreen mode

No application code needs to change.

Only the workflow definition changes.

Level 2 — Change handler configuration

The capability already exists, but its behavior needs parameters.

For example:

FraudReviewHandler

could receive configuration:

{
    "riskThreshold": 700,
    "manualReviewRequired": true
}
Enter fullscreen mode Exit fullscreen mode

Different workflow versions can configure the same handler differently.

Level 3 — Introduce a new capability

The business introduces something that doesn't exist.

For example:

GovernmentComplianceCheck

There is no existing handler.

We implement:

GovernmentComplianceCheckHandler

After that, the handler becomes another reusable building block for future workflows.

This is the boundary I find most useful.

Configuration should compose capabilities. Code should create new capabilities.

15. Conditional Branching

The workflow can also represent conditional paths.

For example:

                 Payment
                    ↓
               Fraud Check
                    ↓
              Risk Level?
              /          \
           Low            High
            ↓               ↓
       Inventory        Manual Review
            │               │
            └───────┬───────┘
                    ↓
                  Packing
Enter fullscreen mode Exit fullscreen mode

The workflow definition can reference a rule:

HighRiskOrder

rather than embedding arbitrary code into configuration.

The rule itself can remain strongly typed:

public interface IWorkflowRule
{
    string Key { get; }

    Task<bool> EvaluateAsync(
        WorkflowContext context,
        CancellationToken cancellationToken);
}

Enter fullscreen mode Exit fullscreen mode

This keeps the dynamic part controlled.

16. Parallel Steps

Another situation where the workflow abstraction becomes attractive is parallel execution.

Suppose after payment verification we need:

              Payment
                ↓
        ┌───────┼────────┐
        ↓       ↓        ↓
      Fraud   Inventory  Compliance
        │       │        │
        └───────┼────────┘
                ↓
              Packing
Enter fullscreen mode Exit fullscreen mode

These checks are conceptually independent.

A workflow engine can represent:

Parallel
   ├── Fraud
   ├── Inventory
   └── Compliance
        ↓
       Join
        ↓
      Packing
Enter fullscreen mode Exit fullscreen mode

A state machine can also model this.

But as the number of parallel activities increases, a state machine often starts requiring composite states, flags, guards, and synchronization logic.

At that point, we're gradually building a workflow engine inside the state machine.

17. Human Tasks and Long-Running Processes

Another difference appears when the process needs human interaction.

For example:

Fraud Review
      ↓
Manual Approval
      ↓
Wait for Manager
      ↓
Continue
Enter fullscreen mode Exit fullscreen mode

The manager may approve it:

in 30 seconds

or:

in 3 days

A workflow instance can persist:

Status = Waiting
CurrentStep = ManagerApproval
AssignedTo = Manager
DueAt = ...

and resume later.

Again, a state machine can support this, but now we need to build persistence, resumption, scheduling, timeout handling, and related infrastructure around it.

The workflow abstraction naturally starts to fit the problem better.

18. So Is a Configurable Workflow Always Better?

No.

This is probably the most important conclusion of this whole exercise.

A configurable workflow introduces significant complexity.

You now need to think about:

workflow definitions
versions
instance persistence
step execution
retries
idempotency
concurrency
conditions
failure handling
timeouts
audit history
authorization
migrations
debugging
configuration validation
Enter fullscreen mode Exit fullscreen mode

A simple state machine does not need all of this.

If the process is:

Created
↓
Approved
↓
Completed

building a workflow engine would be overengineering.

A state machine is probably the better solution.

19. State Machine vs Configurable Workflow

The comparison I now use is not:

"Which one is more powerful?"

It is:

"Which type of change does the system need to optimize for?"

20. The Architectural Lesson

The biggest lesson for me was not:

"State machines are bad."

They aren't.

The lesson was:

We had initially modeled the problem as state transitions, but the problem eventually became process orchestration.

That is a very different problem.

A state machine asks:

What state am I in?
What transitions are legal?

A configurable workflow asks:

Which process applies?
Which version applies?
Which step comes next?
Which conditions apply?
Which handler executes it?
Can this step run in parallel?
Should we wait?
Can we retry?
How do we resume?
Enter fullscreen mode Exit fullscreen mode

Once the business starts asking those questions, a workflow abstraction becomes much more appropriate.

21. The Architecture I Would Choose

The resulting architecture looks roughly like this:

                 ┌───────────────────────┐
                 │ Workflow Definition   │
                 │       Version         │
                 └───────────┬───────────┘
                             │
                             ↓
                 ┌───────────────────────┐
                 │    Workflow Engine    │
                 └───────────┬───────────┘
                             │
              ┌──────────────┼──────────────┐
              ↓              ↓              ↓
           Steps        Conditions       Rules
              │              │              │
              ↓              ↓              ↓
       Handler Keys      Rule Keys       Rule Handlers
              │
              ↓
       Handler Resolver
              │
      ┌───────┼────────┐
      ↓       ↓        ↓
   Payment  Fraud   Inventory
   Handler  Handler  Handler
      │       │        │
      └───────┼────────┘
              ↓
        Domain/Application
            Services
Enter fullscreen mode Exit fullscreen mode

The responsibilities are separated:

Workflow

Defines:

What, when, order, branching, version.

Handler

Defines:

How a specific capability is executed.

Domain/Application services

Define:

Actual business behavior and invariants.

Database

Stores:

Definitions, versions, instances, history, and execution state.

22. The Most Important Design Principle

If I had to summarize the entire architectural decision in one sentence:

Configuration should compose existing capabilities; code should create new capabilities.

That gives us a useful balance.

We don't need to deploy the application every time the business rearranges an existing process.

But we also don't pretend that every possible future behavior can magically be represented by configuration.

When a genuinely new capability appears, we write code.

Then that capability becomes available to future workflow versions.

23. Final Takeaway

I would still happily use a state machine in a new .NET project.

If I know the lifecycle is stable, I would probably prefer it.

It is simple, explicit, strongly typed, testable, and easy to reason about.

But if I discover that:

the sequence of steps changes frequently,
different customers need different processes,
business rules evolve independently of deployments,
old processes must remain reproducible,
steps can be conditional or parallel,
human approval is involved,
processes can run for days or weeks,

then I would stop trying to make the state machine handle everything.

At that point, I would consider a versioned configurable workflow composed of strongly typed handlers.

Not because it is more sophisticated.

Because it models the problem more accurately.

And that is perhaps the most important lesson in system design:

The best architecture is not the one with the most features. It is the one whose abstractions match the kind of change the system is expected to absorb.

Top comments (0)