DEV Community

Cover image for Out-of-Order Events Break Your AI Agent on AWS

Out-of-Order Events Break Your AI Agent on AWS

Ciao 👋

Make some chai, this is a long one, and it has a plot twist I did not see coming.

Here is the bug we are chasing today. I built a small AI agent that reacts to account events, and it billed the wrong outcome to a customer who paid. Not because it crashed. Not because of a logic error I could grep for. It billed wrong because two events reached it in the wrong order, and by the time it looked at the account, it had convinced itself the customer was suspended.

No exception. No error log. Just a quietly incorrect decision written to a database and, in the real world, an angry email to someone who did nothing wrong.

That is the kind of bug that does not show up on your dashboards. It shows up in your churn numbers three months later. And it is exactly the kind of bug that gets more common the moment you put an AI agent on the consuming end of an event stream, which, in 2026, a lot of us are doing.

So I built the whole thing on real AWS, broke it on purpose, and then the pipeline taught me something I genuinely did not expect. Let's go.

Why I am writing this now

Amazon EventBridge just shipped enhanced custom event buses, and the headline feature is native event ordering. You tag related events with an EventGroupId, and events in the same group are delivered in order. If that sounds familiar, it is the same idea SQS FIFO has had for years with MessageGroupId, except now it lives on the event bus that sits at the center of most serverless architectures, and it can span an entire AWS organization. That is a genuinely big deal, and half of AWS Twitter said so the week it dropped.

Now the honest part, up front, because this is the first roadblock and it shaped everything.

I sat down to build directly on the enhanced bus and hit a wall in the first five minutes: the ordering APIs are not in the AWS SDKs yet. I checked boto3 by hand. The latest version has no EventGroupId field on PutEvents, no Subscriber resource, none of the new operations. The feature is real in the console and the announcement, but you cannot write code against it today. This happens with brand-new AWS features, the blog post ships before the SDK catches up.

So I made a call: build the whole thing on the ordering primitive that IS in the SDK, SQS FIFO, which exposes the identical group-id ordering model, and frame the enhanced EventBridge bus as the future-work sequel. Everything in this article is deployed and run on real AWS. When boto3 gains the enhanced-bus operations, the same design points at EventBridge with almost no change. The lesson, ordering keyed by a group id, is the same either way.

That decision is itself the first thing worth taking away: check the SDK before you design around a launch. A console demo is not an API.

What I built: a real pipeline, not a script

I did not want a Python script that pretends to be an architecture. I wanted the real thing, deployed, the way you would actually wire it.

ordered-pipeline

One SAM template, one sam deploy, eleven resources: an EventBridge custom bus and a rule, two SQS FIFO queues (one per ordering path, I will get to that), a FIFO dead-letter queue for poison events, the Lambda that runs the agent, its IAM role, the event source mappings, and the DynamoDB table that holds each account's state. sam delete tears the whole thing back down in one command, which, as you will see, I ended up being very grateful for.

The agent itself is a deterministic mock. That is deliberate: it keeps the demo free, keeps it reproducible, and keeps the ordering lesson clean instead of muddied by model randomness. Swap in a Bedrock call and nothing about the lesson changes, an out-of-order event still hands the model a false picture of the account's state, and a confident model acting on a false picture is exactly the nightmare we are trying to prevent.

The agent, and where order quietly becomes money

The agent tracks one thing, the account's state, and reacts to three event types:

def handle(self, event):
    etype = event["type"]
    if etype == "account.suspended":
        self.state = "suspended"
        self.actions.append("send_dunning_email")
    elif etype == "account.reactivated":
        self.state = "active"
        self.actions.append("restore_access")
    elif etype == "usage.reported":
        if self.state == "active":
            self.actions.append(f"bill_usage:{event['units']}")
        else:
            self.actions.append(f"hold_usage_suspended:{event['units']}")
Enter fullscreen mode Exit fullscreen mode

Look hard at that last branch. The agent only bills usage if it believes the account is active, and its belief about "active" comes entirely from the order it saw the suspend and reactivate events. Order is not metadata here. Order is the input that decides whether a paying customer gets billed correctly or gets a nastygram.

First, prove the concept offline

Before spending a cent on AWS, I built a local FIFO simulator so the divergence could be run and asserted offline. It mirrors the two guarantees that matter, per-group ordering and content-based dedup, and it let me watch the bug happen in a plain terminal.

Screenshot_2026-09-28_105820

In order, the account ends active and gets billed. Out of order, reactivation before suspension, it ends suspended and holds the usage. Same three events, same agent, opposite outcome. And because I do not trust a bug I cannot reproduce, I wired it into the test suite so the divergence is an assertion, not an anecdote.

Screenshot_2026-09-28_153114

Six passing tests: the ordering divergence, content-based dedup, and the pipeline handler logic checked against a mocked DynamoDB. Green locally. Now for the real thing.

Roadblock parade: getting it onto real AWS

This is the part every clean tutorial hides. Here is everything that went wrong between "it works locally" and "it runs on AWS," in order, because you will hit some of these too.

Roadblock one: the demo IAM user could not create a queue. My day-to-day IAM user is scoped tight, so the first create-queue came back with a flat AccessDenied on sqs:CreateQueue. Fixable with a scoped inline policy, but a reminder that a purpose-scoped user is not a deploy user.

Roadblock two: a shell quoting gremlin. My setup script printed two lines, a friendly "Queue created" and an export QUEUE_URL=.... I ran it with source <(...) to capture the export, and bash promptly tried to execute the word "Queue" as a command. bash: Queue: command not found. Harmless, but it sent me chasing a ghost for a minute before I realized the queue was fine and only the friendly line had confused the shell.

Roadblock three: SAM was installed but invisible. I installed the SAM CLI with winget, and Git Bash refused to see it. sam: command not found, even after restarting. Turns out winget dropped a sam.cmd wrapper, and Git Bash only auto-resolves .exe, not .cmd. An alias pointing at the full sam.cmd path fixed it. Ten minutes I will never get back.

Roadblock four, the big one: my user could not deploy CloudFormation. sam deploy needs to create a stack, IAM roles, Lambda, DynamoDB, SQS, EventBridge, and an S3 bucket for artifacts. Granting all of that to my scoped user piecemeal would have turned it into an admin in everything but name. The right move, and the one I took, was to create a dedicated admin-deploy IAM user with proper permissions and deploy under that profile. Deploying infrastructure from a narrowly-scoped tool user is precisely the thing you are not supposed to do. Do not paper over it, use the right identity.

With the admin profile, the stack finally went up clean, all eleven resources.

Screenshot_2026-09-28_160052

And SAM handed back the outputs I needed to drive it, the direct queue URL, the bus name, and the state table.

Screenshot_2026-09-28_155906

The betrayal, on real AWS

Now the fun part. Two runs against the live pipeline, each on a fresh account id.

Why fresh account ids? Because of roadblock five, and it is a good one. My first out-of-order run came back CORRECT, which made no sense. The culprit was leftover state in DynamoDB: the account was already active from a previous run, so re-sending events left it looking correct. Content-based deduplication on the FIFO queue may also have dropped some re-sent messages within the five-minute window, though the stale state alone was enough to poison the result. The pipeline was not lying, I was measuring a contaminated run. The fix was a small script that uses a brand-new random account id every time and a unique dedup id per message, so nothing carries over and nothing gets deduped. Clean inputs, honest outputs.

With that, in order:

Screenshot_2026-09-28_203030

state: active, bill_usage:512, verdict CORRECT. Now the same three events, sent out of order:

Screenshot_2026-09-28_203047

state: suspended, hold_usage_suspended:512, verdict WRONG ACTION ON A PAID ACCOUNT. That state was written to a real DynamoDB table by a real Lambda, triggered by a real FIFO queue. No exception anywhere in the pipeline. The system did exactly what it was told, and what it was told was wrong.

ordered-divergence

The plot twist: FIFO did not actually save me

Here is where the real pipeline taught me something the local simulator never could, and honestly it is the most interesting thing in this whole build.

I sent three events. I assumed the Lambda would receive all three in a single, tidy, ordered batch. It did not. Go back and look closely at that in-order screenshot: the final DynamoDB write recorded only two events in its events seen list, reactivated and usage. The suspend event was processed in a completely separate Lambda invocation.

SQS-to-Lambda delivers in batches, and a single message group can be split across multiple invocations. FIFO guarantees order within delivery, but the event source can hand your consumer two events now and the third a moment later, in a separate invocation that only sees its own batch. Whether that invocation reuses a warm container or not, it cannot safely rely on in-memory state from the batch before it.

Sit with what that means. If my agent had held its state only in memory, that split would have shredded it. Invocation A sees "suspended" and ends. Invocation B sees only "reactivated, usage" in its batch, and if it trusted in-memory state alone it would have no idea the account was ever suspended, and would do whatever that partial view dictates. The ordering guarantee at the queue would have been perfectly intact, and the agent would still have been wrong.

ordered-fragmentation

The only reason the in-order run stayed correct is that the consumer is stateful the right way. Before processing a batch, it loads the account's current state from DynamoDB, applies the new events, and writes back:

def process_batch(records, table):
    by_account = defaultdict(list)
    for r in records:
        body = json.loads(r["body"])
        by_account[body["account"]].append(body)

    for account, events in by_account.items():
        agent = AccountAgent(state=_load_state(table, account))  # resume, not restart
        for e in events:
            agent.handle(e)
        table.put_item(Item={
            "account": account,
            "state": agent.state,
            "last_actions": agent.actions,
            "events_seen": [e["type"] for e in events],
        })
Enter fullscreen mode Exit fullscreen mode

That one line, _load_state before processing, is the least glamorous line in the entire repository and by far the most important. State lives in the database, not in the Lambda. A fragmented group still converges, as long as the pieces arrive in order and every write is a safe continuation of the last. Take that line out, and a shuffled batch quietly corrupts an account.

The lesson most "just use FIFO" advice skips

If you have ever asked "how do I get ordering in my event system" and been told "use FIFO," you got a third of the answer. Here is the whole thing, and it is three parts, not one.

ordered-contract

The producer must assign a correct sequence. FIFO faithfully delivers whatever order it was handed, including a wrong one. In my out-of-order run, FIFO did its job flawlessly, and the outcome was still wrong, because the producer sent nonsense in a tidy line.

The queue keys ordering to a group ID. Put the account ID in the group so one account's events never overtake each other, while different accounts stay parallel and you do not pay a global-ordering tax you do not need.

The consumer must be stateful and idempotent, because batching can fragment a group across invocations. State belongs in a store, not in memory, and every write has to be a safe continuation of the last.

Miss any one of the three and the guarantee leaks, silently, into a wrong decision. That is the honest shape of ordered event processing, and an agent on the consuming end raises the stakes because a stale state does not throw. It just acts.

The EventBridge constraint I ran into

The pipeline has two ordering paths on purpose, and building the second one surfaced a real constraint worth knowing before you design around it.

Path A sends to the FIFO queue directly, setting MessageGroupId per event to the account id. Clean per-account ordering.

Path B routes through the EventBridge bus to a FIFO target. And here is the catch: when EventBridge targets a FIFO queue, it sets the MessageGroupId statically, in the target configuration, one fixed value for every event through that target. It cannot derive the group ID from the event payload. I confirmed this in the API shape itself: the target's SQS parameters have exactly one field for it, and it is a constant.

So a single EventBridge-to-FIFO target orders the entire stream globally, not per account. Fine for one account, far too coarse at scale, because now every account shares one ordering lane and they all queue behind each other. To keep per-account ordering through EventBridge today, you send to FIFO directly or fan out to per-group targets.

Which is precisely the gap the enhanced bus is built to close.

Future work: the enhanced EventBridge bus sequel

This is why the launch I opened with actually matters, and it is the natural next chapter.

EventGroupId on the enhanced custom bus is the per-event group key that the classic bus could never give a FIFO target. It brings per-entity ordering to the organisation-scale event bus most teams already run on, without the send-directly-to-FIFO workaround and without the static-group-id ceiling. When its APIs land in the SDKs, the next version of this pipeline drops the SQS FIFO hop entirely and orders on EventBridge itself, with the same account-as-group-id pattern.

And here is the part I find reassuring: almost everything in this article carries straight over. The stateful consumer that survives fragmentation, the producer-sequencing discipline, the three-part contract, the DynamoDB-as-source-of-truth pattern - none of that changes. Only the ordering primitive underneath does. So this build is not throwaway scaffolding waiting for a better API. It is the durable half of the design, and the enhanced bus just makes the ordering layer prettier.

I will build that sequel the day boto3 ships the operations. Consider it teased.

What I am taking away

  • For an agent that reacts to a stream, event order is not metadata; it is an input that changes the action. Out-of-order events do not crash the agent. They make it confidently wrong, and you find that in your churn numbers, not your logs.
  • FIFO is necessary, not sufficient. Ordered delivery still gets fragmented across batched Lambda invocations, so the consumer must carry state in a store and treat every write as a continuation. That one _load_state line is the whole game.
  • Ordering is a three-part contract: a producer that sequences correctly, a queue keyed by a group id, and a stateful idempotent consumer. Any missing part leaks the guarantee.
  • EventBridge sets a FIFO target's group id statically. Per-entity ordering through EventBridge needed a workaround, and the enhanced bus with EventGroupId exists to remove it.
  • Check the SDK before you design around a launch. The enhanced bus is real in the console and absent from boto3, and knowing that up front saved me from building against an API that does not exist yet.
  • Deploy infrastructure from a real deploy identity, not your scoped tool user. And put the whole thing in one SAM stack, because being able to sam delete an eleven-resource pipeline in one command is what makes experiments like this cheap to run and safe to clean up.

The whole thing is on GitHub: the SAM template, the Lambda consumer, the producer for both ordering paths, the local FIFO simulator so you can see the concept offline, and the tests that assert the divergence. MIT licensed. Clone it, deploy it, and watch a paid account get suspended by a shuffled stream, then watch one line of state-loading save it.

Repo: https://github.com/mursalfk/ordered-agents

Happy Coding! 👋

Top comments (0)