DEV Community

Django Mastery
Django Mastery

Posted on AI-assisted

Your Celery task will run twice. Here's how to make that safe

Celery can't promise a task runs exactly once. It can run twice, and one day it will. If that task sends an email or charges a card, your user gets two.

The fix isn't to stop duplicates. It's to make the second run do nothing.

How a task ends up running twice

With acks_late=True, the message stays unacknowledged while the task runs. Now picture this:

  1. A worker runs send_order_confirmation. The email goes out.
  2. The worker is killed (out of memory, a pod eviction) before it acks.
  3. The broker redelivers: RabbitMQ when the connection drops, Redis after the visibility timeout.
  4. Another worker runs the task again. Second email.

And it's not only crashes:

  • an autoretry after a timeout where the other side actually succeeded,
  • a user double-clicking a button,
  • a scheduled task overlapping its previous run,
  • a task that runs longer than the visibility timeout.

Pick your failure mode

Setting Crash during the task Guarantee
Default (early ack) The task is lost At-most-once
acks_late=True The task runs again At-least-once

For anything important, pick acks_late plus idempotency.

One catch: with task_reject_on_worker_lost, a task that always runs out of memory gets redelivered forever. Cap it with a delivery counter or time limits.

So exactly-once execution isn't available. Aim for an exactly-once effect instead.

The strongest guard: a unique constraint

class EmailLog(models.Model):
    kind = models.CharField(max_length=50)
    object_id = models.BigIntegerField()
    sent_at = models.DateTimeField(auto_now_add=True)

    class Meta:
        constraints = [
            models.UniqueConstraint(
                fields=["kind", "object_id"],
                name="uniq_email_per_object",
            ),
        ]


@shared_task
def send_order_confirmation(order_id: int) -> None:
    try:
        with transaction.atomic():
            EmailLog.objects.create(kind="order_confirmation", object_id=order_id)
            send_confirmation_email(order_id)
    except IntegrityError:
        return
Enter fullscreen mode Exit fullscreen mode

How it works:

  • Claim the work first by inserting the log row. A duplicate run raises IntegrityError right there, so it exits quietly.
  • If sending fails, the exception rolls back the log row too, so a retry can still run.
  • The database enforces it, even when two workers hit the same order at the same moment.

Use this for anything that should happen once per object: emails, notifications, certificates, payouts.

Bonus: task states and their gotchas

State Meaning Gotcha
PENDING Waiting, or unknown id A typo'd id is PENDING forever
STARTED A worker began Only with task_track_started = True
RETRY Failed, retry scheduled A new message, same task id
SUCCESS Returned normally Needs a result backend to be visible
FAILURE Raised, no retries left Traceback stored if enabled
REVOKED Cancelled with revoke() A running task goes on unless terminate=True

Quick check

acks_late=True is on. A worker sends an email, then is killed before acking. What happens?

It runs again and sends a second email, unless the task is idempotent. Celery has no idea the email already went out.

Remember

  • Exactly-once execution doesn't exist; aim for an exactly-once effect.
  • acks_late plus idempotency for important work.
  • A unique constraint is the strongest duplicate guard.

This article and the video are a free sample from Django Mastery, a video course with 306 short animated lessons that takes you from your first model to production deployment (Django 6.1 and 5.2 LTS), including a full level on Celery and background jobs.

👉 Full course: https://djangomastery.gumroad.com/l/django-mastery
📺 More free lessons: https://www.youtube.com/@Django-Mastery

The course and this article were produced with AI assistance, and the video uses AI narration.

Top comments (1)

Collapse
 
beusebiu profile image
Eusebiu Balan •

Laravel has the same trap under different names. The queue's retry_after is the visibility timeout, and if it's shorter than the worker's --timeout, a slow job gets handed to a second worker while the first one is still running.

It's one line in the docs, so it keeps catching people. Your log-row constraint catches that one as well.