<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yu Ting Chen</title>
    <description>The latest articles on DEV Community by Yu Ting Chen (@yu_ting_chen).</description>
    <link>https://dev.to/yu_ting_chen</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3964318%2F37e4045a-ed50-4d8f-9673-88ab576b15f8.jpg</url>
      <title>DEV Community: Yu Ting Chen</title>
      <link>https://dev.to/yu_ting_chen</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yu_ting_chen"/>
    <language>en</language>
    <item>
      <title>A Team Asked for One More Service, and My Kubernetes Platform Quietly Refused</title>
      <dc:creator>Yu Ting Chen</dc:creator>
      <pubDate>Thu, 01 Oct 2026 04:03:32 +0000</pubDate>
      <link>https://dev.to/yu_ting_chen/a-team-asked-for-one-more-service-and-my-kubernetes-platform-quietly-refused-okm</link>
      <guid>https://dev.to/yu_ting_chen/a-team-asked-for-one-more-service-and-my-kubernetes-platform-quietly-refused-okm</guid>
      <description>&lt;p&gt;I'm building a small internal developer platform on Kubernetes. The idea is simple. A team writes a short request that says "this is my team, this is my service, run it". The platform does the rest: it gives the team its own space in the cluster, the right access, a resource budget, and it deploys the service.&lt;/p&gt;

&lt;p&gt;It worked for the first service. Then I added a second service for the same team. Kubernetes accepted the request. No error came back. The new service just never became ready.&lt;/p&gt;

&lt;p&gt;This post is about what was going on, how I thought it through, and what I changed. If you build things on Kubernetes that create other things, you may run into the same wall.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the platform does
&lt;/h2&gt;

&lt;p&gt;Kubernetes lets you add your own kinds of objects, called custom resources. You then write a controller: a program that watches those objects and makes the cluster match what they describe.&lt;/p&gt;

&lt;p&gt;My platform had one custom resource, a &lt;code&gt;ServiceClaim&lt;/code&gt;. A team's request looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ServiceClaim&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payments&lt;/span&gt;          &lt;span class="c1"&gt;# the service name&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;team&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payments&lt;/span&gt;
  &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx:1.27-alpine&lt;/span&gt;   &lt;span class="c1"&gt;# a stand-in for the real service; any image works&lt;/span&gt;
  &lt;span class="na"&gt;replicas&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
  &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;              &lt;span class="c1"&gt;# the team's resource budget&lt;/span&gt;
    &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2"&lt;/span&gt;
    &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;4Gi&lt;/span&gt;
    &lt;span class="na"&gt;pods&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For each claim, the controller created four things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;a namespace for the team, &lt;code&gt;team-payments&lt;/code&gt; (a namespace is a separate space inside the cluster)&lt;/li&gt;
&lt;li&gt;a RoleBinding that gives the team edit access in that namespace&lt;/li&gt;
&lt;li&gt;a ResourceQuota that caps what the team can use&lt;/li&gt;
&lt;li&gt;an ArgoCD &lt;code&gt;Application&lt;/code&gt; that deploys the service (ArgoCD is a tool that keeps the cluster in sync with manifests in Git)&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Why the second service got stuck
&lt;/h2&gt;

&lt;p&gt;In Kubernetes, an object can record who owns it. One of those owners can be marked as the controller: the one object responsible for it. Kubernetes allows only one controller owner per object. The library my controller uses, controller-runtime, checks this rule before it even sends an update, and returns an &lt;code&gt;AlreadyOwnedError&lt;/code&gt; if the object already has one.&lt;/p&gt;

&lt;p&gt;My controller made the &lt;code&gt;ServiceClaim&lt;/code&gt; the controller owner of all four objects. That's fine for one service. But look at the list again. Only the last one belongs to a service. The namespace, the access and the budget belong to the team.&lt;/p&gt;

&lt;p&gt;So the second claim for the same team asked to own a namespace that the first claim already owned. Kubernetes' rule said no. The controller stopped at step one and retried, over and over. The claim's status showed it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NamespaceReady=False (NamespaceError) Object /team-payments is already owned by another ServiceClaim controller payments
Ready=False (ResourcesNotReady) one of the claim's resources is not ready
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing crashed. The request was accepted, the first service kept running, and the second one waited forever. That kind of failure is easy to miss, because nothing looks broken until someone asks where their service is.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqf1xo197fotx763f9mol.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqf1xo197fotx763f9mol.png" alt="Before: one ServiceClaim owns the namespace, RoleBinding, quota and its Application, so a second claim for the same team hits AlreadyOwnedError. After: a Tenant owns the three team-level objects and each ServiceClaim owns only its own Application." width="800" height="387"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I thought about it
&lt;/h2&gt;

&lt;p&gt;The error message was clear. The harder question was why my design asked for this at all. So before fixing anything, I listed the obvious fixes and asked what each one would lock in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Block the second claim.&lt;/strong&gt; A validation check could reject any second claim for the same team. It's the simplest fix, and it's the wrong one. It turns a bug into a rule. Teams still can't run two services, they just get told so earlier.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let claims share the team's objects.&lt;/strong&gt; Each claim could add itself to the team's objects and keep a count. The namespace would go away only when the last claim did. But the controller would then have to add up every claim's budget into one quota, and track who is still around. That's a lot of code to work around the shape of one object.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make one object per team, with a list of services inside.&lt;/strong&gt; No conflict, since there is only one owner. But changing one service would mean editing the whole team's list. That's not how teams think. They want to "deploy a service", not "re-declare my team".&lt;/p&gt;

&lt;p&gt;There was also a second problem hiding behind the first. Even if the ownership conflict were solved another way, each claim carried its own &lt;code&gt;resources&lt;/code&gt; budget. Two claims would keep writing two different budgets into the same quota. Nobody had hit that yet, because the second claim never got past the namespace. But on paper, the design was already wrong.&lt;/p&gt;

&lt;p&gt;Looking at the three options side by side made the real issue easier to see. One object was carrying two different sizes of thing: team things and service things. The fix was to split it along that line.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: one object for the team, one for each service
&lt;/h2&gt;

&lt;p&gt;I added a second custom resource, the &lt;code&gt;Tenant&lt;/code&gt;, for the team:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Tenant&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payments&lt;/span&gt;          &lt;span class="c1"&gt;# the team name&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2"&lt;/span&gt;
    &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;4Gi&lt;/span&gt;
    &lt;span class="na"&gt;pods&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;Tenant&lt;/code&gt; owns the namespace, the access and the budget. Its name is the team name, so "one Tenant per team" is guaranteed by Kubernetes itself. A &lt;code&gt;Tenant&lt;/code&gt; is cluster-wide, and two cluster-wide objects of the same kind can't share a name. I didn't have to write any code for it.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;ServiceClaim&lt;/code&gt; became smaller. It keeps only what a service needs, &lt;code&gt;team&lt;/code&gt;, &lt;code&gt;image&lt;/code&gt; and &lt;code&gt;replicas&lt;/code&gt;, and it owns only its own ArgoCD &lt;code&gt;Application&lt;/code&gt;. A team can now have as many claims as it wants, all pointing at the same &lt;code&gt;Tenant&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The split raised new questions
&lt;/h2&gt;

&lt;p&gt;Two objects that depend on each other bring questions that one object never had. I answered two of them before calling the fix done.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if the service arrives before the team?&lt;/strong&gt; Teams often apply a folder with a &lt;code&gt;Tenant&lt;/code&gt; and several &lt;code&gt;ServiceClaim&lt;/code&gt;s in one go, and Kubernetes doesn't promise any order. If a claim were rejected because its &lt;code&gt;Tenant&lt;/code&gt; wasn't there yet, a valid folder would fail or succeed by luck. So a claim without a ready &lt;code&gt;Tenant&lt;/code&gt; doesn't fail. It reports &lt;code&gt;TenantReady=False&lt;/code&gt; and waits, and it continues on its own when the &lt;code&gt;Tenant&lt;/code&gt; is ready. "Not yet" and "wrong" are different states, and only "wrong" deserves an error.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens on delete?&lt;/strong&gt; Kubernetes cleans up owned objects automatically when their owner is deleted, but it doesn't do it in any order. Deleting a &lt;code&gt;Tenant&lt;/code&gt; could remove the namespace while services were still running in it. To control the order, I used finalizers. A finalizer is a marker on an object that makes Kubernetes wait, before deleting it, until a controller has done its cleanup. Three of them work together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;code&gt;ServiceClaim&lt;/code&gt; finalizer deletes the claim's ArgoCD &lt;code&gt;Application&lt;/code&gt;, and waits until it's really gone.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;Application&lt;/code&gt; carries ArgoCD's own finalizer. That's what makes ArgoCD remove the running service before the &lt;code&gt;Application&lt;/code&gt; itself disappears. I found this out while building the delete path: without it, deleting the &lt;code&gt;Application&lt;/code&gt; only makes ArgoCD forget the service, and the service keeps running with nothing managing it.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;Tenant&lt;/code&gt; finalizer refuses to delete the team while any claim still points at it. So the namespace stays alive long enough for every service to be removed first.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One gap is known and written down rather than fixed: deleting a &lt;code&gt;Tenant&lt;/code&gt; with &lt;code&gt;kubectl delete --cascade=foreground&lt;/code&gt; removes the namespace before the finalizer runs. The default delete is the supported path.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cost
&lt;/h2&gt;

&lt;p&gt;The split isn't free. One simple object became two objects with a relationship: a reference, a waiting state, and finalizers that have to run in the right order. That's more to reason about and more to test. I think it's the right trade here, but it's a trade, not a pure win.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you run into something similar
&lt;/h2&gt;

&lt;p&gt;This is the approach I'd take again:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Treat the error as a question about the design.&lt;/strong&gt; &lt;code&gt;AlreadyOwnedError&lt;/code&gt; wasn't the bug. It was Kubernetes pointing at an object that owned things at two different levels.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For each quick fix, ask what it would lock in.&lt;/strong&gt; The simplest fix here would have made the limitation permanent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split objects by what they really belong to.&lt;/strong&gt; Team things and service things have different lifetimes. Putting them in one object is where the conflict came from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer the questions the fix creates.&lt;/strong&gt; Order of arrival and order of deletion were new problems that only existed after the split.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write down the questions you're not solving yet.&lt;/strong&gt; A month before this bug, I had noted that "who owns what" between my custom resources would need real design once there was more than one. The note didn't predict this bug, but when it came, the question was already on paper.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;The original code that failed isn't in the repository's history, because the history was squashed. So I rebuilt that earlier version to show the failure. It runs in its own small local cluster, and it prints the same status you saw above.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The failing version: &lt;a href="https://github.com/ChenYuTingJerry/idp-platform-lab/tree/main/scenarios/ownership-conflict" rel="noopener noreferrer"&gt;scenarios/ownership-conflict&lt;/a&gt;. The README there has the steps. You need Docker, k3d, kubectl, the &lt;code&gt;task&lt;/code&gt; runner and Go.&lt;/li&gt;
&lt;li&gt;The current design: &lt;a href="https://github.com/ChenYuTingJerry/idp-platform-lab" rel="noopener noreferrer"&gt;idp-platform-lab&lt;/a&gt;. The README's quick start brings up the whole platform, and &lt;code&gt;docs/verification.md&lt;/code&gt; walks from a request to a running service.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;I used AI to help draft and edit the English. I checked the technical claims against the code and a live reproduction.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>platformengineering</category>
      <category>architecture</category>
      <category>devops</category>
    </item>
    <item>
      <title>When Celery Chains Get Hard to Maintain: Introducing CeleryFlow</title>
      <dc:creator>Yu Ting Chen</dc:creator>
      <pubDate>Thu, 06 Aug 2026 07:48:35 +0000</pubDate>
      <link>https://dev.to/yu_ting_chen/when-celery-chains-get-hard-to-maintain-introducing-celeryflow-1idi</link>
      <guid>https://dev.to/yu_ting_chen/when-celery-chains-get-hard-to-maintain-introducing-celeryflow-1idi</guid>
      <description>&lt;p&gt;If you run Celery in production and your multi-step jobs have turned&lt;br&gt;
into a pile of &lt;code&gt;chain()&lt;/code&gt; calls nobody wants to touch, this post is&lt;br&gt;
about that specific problem.&lt;/p&gt;

&lt;p&gt;I ran into this problem at a previous job and built a small internal library to solve it. Earlier this year, I released an open-source version as &lt;a href="https://pypi.org/project/celeryflow/" rel="noopener noreferrer"&gt;CeleryFlow&lt;/a&gt;. It lets you describe Celery workflows in YAML instead of assembling them in Python.&lt;/p&gt;

&lt;p&gt;The honest disclaimer first: it's not the right tool for most new projects, and it is not a durable execution engine. If you're starting fresh in 2026, there are better places to look, and I'll point at them near the end. What follows is written for people already running&lt;br&gt;
Celery.&lt;/p&gt;
&lt;h2&gt;
  
  
  When you outgrow &lt;code&gt;chain()&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;At that job, we had dozens of multi-step workflows in an e-commerce platform: order placement, refund, plan upgrade, subscription renewal, etc. Each was 5-10 steps. They shared common sub-flows ("validate order", "charge card", "send receipt"). Conditions like "only run the loyalty-points step if buyer is a returning customer" appeared all over.&lt;/p&gt;

&lt;p&gt;Celery's &lt;code&gt;chain()&lt;/code&gt;, &lt;code&gt;group()&lt;/code&gt;, and &lt;code&gt;chord()&lt;/code&gt; primitives are powerful,&lt;br&gt;
and none of this was impossible. But this approach had two specific pain points:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Composing chains was painful.&lt;/strong&gt; When &lt;code&gt;place_order_flow&lt;/code&gt; and
&lt;code&gt;refund_flow&lt;/code&gt; needed the same opening sequence, you'd either
duplicate the chain construction or wrap it in helper functions
that ballooned over time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conditionals were inline.&lt;/strong&gt; "Only run step X when payload field
Y looks like Z" became &lt;code&gt;if&lt;/code&gt; statements inside the task body, mixing
business logic with flow control.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvqhyoqxqjg2uaom3d346.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvqhyoqxqjg2uaom3d346.png" alt="Two-column comparison. Left, raw chain(): PlaceOrder and Refund each repeat the validate and charge steps (highlighted as duplicated), and the loyalty step carries an inline 'if returning' check. Right, CeleryFlow: a single shared checkout sub-flow (validate then charge) that both PlaceOrder and Refund point to, and the loyalty step carries a declarative condition instead of an inline if." width="799" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Same two flows, two approaches. On the left, the steps are duplicated across flows, and the condition is implemented as an inline &lt;code&gt;if&lt;/code&gt; statement. On the right, the sub-flow is defined once, and the condition is defined in config.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The natural response was to pull the workflow structure out of the code and into config files. We did, and it stuck. CeleryFlow brings that design to open source.&lt;/p&gt;
&lt;h2&gt;
  
  
  What it looks like
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;work-flows&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
    &lt;span class="na"&gt;tasks&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;order.validate&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;order.charge&lt;/span&gt;

&lt;span class="na"&gt;main-flows&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;PlaceOrder&lt;/span&gt;
    &lt;span class="na"&gt;flows&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;flow&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;task&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;order.send_receipt&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;task&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;order.award_loyalty_points&lt;/span&gt;
        &lt;span class="na"&gt;condition&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;customer_type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;$eq&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;returning"&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That's the whole workflow definition, but not the whole setup. To wire it in, use the &lt;code&gt;CeleryFlow&lt;/code&gt; app class instead of &lt;code&gt;Celery&lt;/code&gt;, give participating tasks &lt;code&gt;base=EventTask&lt;/code&gt;, and register the file with &lt;code&gt;FlowBuilder.from_yaml()&lt;/code&gt;. Once registered, the flow runs on your existing Celery workers and broker. There is no new worker model or orchestration server, and your existing retry behavior, routing, Flower setup, and monitoring continue to work.&lt;/p&gt;

&lt;p&gt;One important caveat: in v0.2.0, a condition that doesn't match raises &lt;code&gt;ConditionFailed&lt;/code&gt; before the task is queued. There is no built-in quiet skip; unless the exception is handled explicitly, the chain fails. A missing field counts as passing, so conditions only gate fields that are present in the payload.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"celeryflow[yaml]"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;yaml&lt;/code&gt; extra pulls in PyYAML, which &lt;code&gt;FlowBuilder.from_yaml()&lt;/code&gt;&lt;br&gt;
needs. Plain &lt;code&gt;pip install celeryflow&lt;/code&gt; works too if you'd rather keep&lt;br&gt;
your flow definitions as Python dicts or JSON.&lt;/p&gt;

&lt;p&gt;Quickstart and full docs: &lt;a href="https://github.com/ChenYuTingJerry/CeleryFlow" rel="noopener noreferrer"&gt;https://github.com/ChenYuTingJerry/CeleryFlow&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When you don't need it
&lt;/h2&gt;

&lt;p&gt;You don't need any workflow framework if:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You only have single tasks or short chains with one or two follow-ups, and you're happy using &lt;code&gt;Task.s() | Task.s()&lt;/code&gt; for the chains.&lt;/li&gt;
&lt;li&gt;You don't have many of them.&lt;/li&gt;
&lt;li&gt;The chains rarely change shape.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a handful of simple workflows, just use the Celery primitives. Adding a layer on top is overhead you don't need.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it isn't
&lt;/h2&gt;

&lt;p&gt;The most important question before choosing an approach is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If a worker process dies in the middle of a multi-step workflow,&lt;br&gt;
what happens?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;With Celery, and therefore with CeleryFlow, individual tasks may be redelivered after certain failures, depending on &lt;code&gt;acks_late&lt;/code&gt;, the broker, and how the worker failed. A task may also run more than once. What Celery doesn't provide is durable workflow state from which the whole flow can resume.&lt;/p&gt;

&lt;p&gt;If that's not acceptable, look at &lt;em&gt;durable execution&lt;/em&gt; platforms:&lt;br&gt;
Temporal, Hatchet, Restate, AWS Step Functions. They persist workflow&lt;br&gt;
progress, so execution can recover after a failure and wait between&lt;br&gt;
steps. Adopting one means moving to a different execution and operational model, whether self-hosted or managed. CeleryFlow has no equivalent persistence or durable timers, so it is a poor fit for flows that wait for hours or days between steps. It is also Python-only, while most of those platforms support several languages.&lt;/p&gt;

&lt;p&gt;Durable workflow recovery does not mean that every side effect happens exactly once. Individual operations may still be retried, so anything involving money or an external API needs to be idempotent. The exact guarantees vary by platform.&lt;/p&gt;

&lt;p&gt;And if your "workflow" is really a &lt;strong&gt;data pipeline&lt;/strong&gt; (ETL, batch ML&lt;br&gt;
training, scheduled report generation), Prefect and Dagster are built&lt;br&gt;
for that, with first-class scheduling, observability UIs, and&lt;br&gt;
integrations with warehouses and dbt. CeleryFlow can run scheduled&lt;br&gt;
tasks via Celery Beat, but that's not what it's for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it sits
&lt;/h2&gt;

&lt;p&gt;My first-hand experience here is with Celery and CeleryFlow. I've also operated Airflow at the infrastructure level. For the other tools, I'm relying on their intended design rather than direct use.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;When to use it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Celery&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Task queue&lt;/td&gt;
&lt;td&gt;You need workers that run individual tasks. Mature, battle-tested.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Celery + &lt;code&gt;chain()&lt;/code&gt;/&lt;code&gt;group()&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Manual workflow&lt;/td&gt;
&lt;td&gt;You have Celery and a handful of multi-step jobs. Works, gets messy at scale.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CeleryFlow&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Config-driven workflow on Celery&lt;/td&gt;
&lt;td&gt;You already use Celery, want declarative workflows, don't want a separate service.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Temporal&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Durable execution platform&lt;/td&gt;
&lt;td&gt;You need workflows that survive worker crashes, run for days, and resume from where they stopped.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Prefect&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Modern data orchestration&lt;/td&gt;
&lt;td&gt;Data pipelines, ML workflows, want a UI, want hybrid cloud / on-prem.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Dagster&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Asset-oriented orchestration&lt;/td&gt;
&lt;td&gt;Data engineering teams who think in terms of "data assets" not "tasks."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Airflow&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Classic DAG scheduler&lt;/td&gt;
&lt;td&gt;You're at a company that already runs Airflow, or your team is comfortable with it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;arq / TaskIQ / Procrastinate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Async-native task queues&lt;/td&gt;
&lt;td&gt;Brand-new project, async-first, don't need Celery's surface area.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm0i5wp5fkgpvajbv07k7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm0i5wp5fkgpvajbv07k7.png" alt="Quadrant chart, y-axis durable execution low to high, x-axis separate service / infra weight low to high. Celery + chain(), CeleryFlow and arq/TaskIQ sit bottom-left; Prefect, Dagster and Airflow sit mid-right; Temporal and Step Functions sit top-right." width="800" height="508"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Two questions capture most of the trade-off: do you need durable execution, and how much additional infrastructure are you willing to run?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  So is it for you?
&lt;/h2&gt;

&lt;p&gt;I'll be specific. CeleryFlow is a good fit if:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You have a Python backend that already runs Celery in production.&lt;/li&gt;
&lt;li&gt;You have between &lt;strong&gt;5 and 50&lt;/strong&gt; multi-step workflows.&lt;/li&gt;
&lt;li&gt;Each workflow is &lt;strong&gt;seconds to minutes&lt;/strong&gt; long, not days.&lt;/li&gt;
&lt;li&gt;Workflows have &lt;strong&gt;conditions&lt;/strong&gt; that gate a step on the payload, or
share &lt;strong&gt;common sub-flows&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;You &lt;strong&gt;don't want a separate service&lt;/strong&gt; for workflow orchestration.&lt;/li&gt;
&lt;li&gt;Re-triggering the whole flow from the upstream event is an acceptable recovery strategy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqhq53v26olf4ix6euaxb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqhq53v26olf4ix6euaxb.png" alt="Top-down decision tree. Multi-step background jobs? If the workflow must resume automatically after a crash, use durable execution (Temporal / Hatchet / Restate). Else if it is really a data pipeline, use Prefect / Dagster. Else if you are not already running Celery, use a plain task queue (Celery / arq / TaskIQ). Else if you have 5-50 multi-step flows with conditions or shared sub-flows, use CeleryFlow; otherwise plain Celery chain() / group()." width="800" height="638"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The whole article as one picture. CeleryFlow is the leaf at the end of a fairly narrow path, and that's on purpose.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd like feedback on
&lt;/h2&gt;

&lt;p&gt;I'm trying to understand whether this design still fits real-world needs in 2026 or whether teams have moved on. If you've used Celery at any scale, I'd love to hear your perspective:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Have you ever written something like CeleryFlow? Why or why not?&lt;/li&gt;
&lt;li&gt;Did you switch from Celery to Temporal / Prefect / something else?
What pushed you?&lt;/li&gt;
&lt;li&gt;If you tried CeleryFlow and didn't keep it, what made you bounce?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;GitHub issues and comments here on DEV are both welcome.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I used AI to help draft and edit the English. I checked the technical claims against the CeleryFlow code.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>celery</category>
      <category>workflow</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
