DEV Community

Alejandro Hernández
Alejandro Hernández

Posted on

Your Prompt Is Not an Architecture: Event-Driven Serverless Agents

In the previous article,
I proposed a simple idea: agents don't start with prompts; they start
with events.

Most current conversations about agents begin in the wrong place. They
begin with a prompt, an LLM, a list of tools, and a final response. That
model can be useful, but it doesn't fully describe how many real systems
behave in production.

Most agent frameworks start with the prompt. This architecture starts
with the event.

While many frameworks model agents as conversations, this architecture
models them as event-driven control systems.

A business process doesn't begin because someone wrote a prompt. It
begins because something happened: a candidate applied for a job, a
customer sent an email, a document was uploaded, a payment was approved,
or an order changed status. The event is the trigger. The prompt, when
one exists, comes later.

In this article, I want to turn that idea into a concrete architecture.
If we define an agent as a system that observes events, evaluates
information, and executes actions, then we can build it as an
event-driven control loop.

And that loop fits serverless extremely well.

We don't need to imagine the agent as a living process thinking all day.
We can build it as a set of small Lambdas that wake up when something
happens, consult durable memory, make decisions, and emit new events.

Not:

prompt -> LLM -> tool -> response

Instead:

event -> collect -> correlate -> evaluate -> emit -> action -> event
Enter fullscreen mode Exit fullscreen mode

The LLM can participate. In fact, it will often be a very powerful tool
inside the process. But it isn't the center of the architecture.

The control loop is.

Control services as an implementation pattern

Let's use the example from the previous article: recruiting.

When someone applies for a job, the system doesn't need to "wait for a
prompt." A relevant fact has already occurred:

CandidateApplied
Enter fullscreen mode Exit fullscreen mode

From there, many other events may happen:

CandidateApplied
        ↓
ResumeAnalyzed
        ↓
InterviewScheduled
        ↓
InterviewCompleted
        ↓
FeedbackReceived
        ↓
OfferSent
Enter fullscreen mode Exit fullscreen mode

The interesting question is: where is the agent?

It isn't only in the resume analysis. It also isn't only an LLM
producing a recommendation. The agent is the complete system that
observes the process, accumulates evidence, evaluates the current state,
and generates new facts to move things forward.

In a serverless architecture, control-service patterns fit this way of
building event-driven agents very well.

A control service observes lower-level events, maintains operational
memory, and produces higher-level events. That doesn't mean every agent
must be a control service. It means this pattern gives us a natural way
to implement event-driven agents in production.

Lower-level events describe things that happened:

CandidateApplied
ResumeUploaded
ResumeAnalyzed
InterviewCompleted
FeedbackReceived
Enter fullscreen mode Exit fullscreen mode

Higher-level events represent derived conclusions or intentions:

CandidateQualified
InterviewRequested
HumanReviewRequired
OfferRecommended
CandidateRejected
Enter fullscreen mode Exit fullscreen mode

The agent lives in the transition between the two.

The architecture

To make this concrete, imagine a recruiting agent that answers a limited
question:

When a candidate applies for a job, do we have enough information to
recommend the next step?

We're not building a complete ATS (Applicant Tracking System). We're
also not automating the final hiring decision. That's not the point.

The point is to show how an agent can:

  • observe an application;
  • store the event;
  • correlate it with the candidate, job, and application;
  • request resume analysis;
  • use an LLM for a specific task;
  • evaluate the result;
  • emit an intention for the next step.

This architecture uses managed serverless services for what they do
best:

  • EventBridge routes business events.
  • SQS or Kinesis decouple producers and consumers.
  • Lambda executes small pieces of logic.
  • DynamoDB stores events and correlations.
  • DynamoDB Streams wakes up the next control-loop iteration.

The agent doesn't need to run all the time. It waits for events.

For the examples, I'll use
modmex-lambda, a library
for building 100% serverless cloud-native applications on AWS Lambda
with Python. It supports synchronous services such as API Gateway APIs
as well as asynchronous services for event-driven architectures.

The part we're interested in here is modmex_lambda.stream: the module
for building event-driven pipelines with sources, rules, and flavors
such as Collect, Correlate, Evaluate, Expired, and Task.

It isn't a prompt-based agent framework. It's a small layer for
implementing control patterns over events: normalizing AWS events,
executing rules, persisting or querying DynamoDB, and publishing new
events to EventBridge.

The minimum event contract

Before talking about prompts or models, we need to talk about events.

An application event might look like this:

{
  "id": "candidate-applied-001",
  "type": "candidate.applied",
  "timestamp": 1718050000000,
  "partition_key": "application-789",
  "application": {
    "id": "application-789",
    "status": "submitted",
    "resume_s3_key": "resumes/johndoe.pdf",
    "candidate": {
      "id": "candidate-123",
      "name": "John Doe",
      "email": "johndoe@example.com"
    },
    "job": {
      "id": "job-456",
      "title": "Backend Engineer",
      "required_skills": ["python", "aws", "serverless"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

A few fields matter:

  • id: unique event identity.
  • type: what happened.
  • timestamp: when it happened.
  • partition_key: the natural key for distribution and correlation.
  • application: the domain entity associated with the event.

This convention matters. Events are normally emitted around a domain
entity. If the event is candidate.applied, the primary entity may be
application; if the event is resume.analyzed, the primary entity may
still be application, now enriched with the analysis.

An event shouldn't be only a minimal notification. Downstream services
often need enough information to react without coupling themselves back
to the producer. That's why the canonical entity can contain snapshots
of related data such as the candidate and job.

The control service can also enrich that entity when it emits
higher-level events. For example, it can take an application received
from a core service and emit another application containing
analysis, recommendation, and next_step.

This contract matters more than it may seem. If events are poor, the
agent has nothing meaningful to observe. If events are ambiguous, the
agent can't explain why it acted. If events don't have correlation keys,
the agent can't build context.

The agent's memory doesn't have to live in an LLM context window. It
can live in durable events.

Collect: observing facts

The first step in the loop is Collect.

Collect doesn't decide anything. It doesn't call the LLM. It doesn't
send emails. It doesn't schedule interviews.

Its responsibility is simpler and more important: take an observed event
and store it as a durable fact.

With modmex-lambda, a Lambda can consume messages from SQS and
register the events in DynamoDB:

from modmex_lambda.stream.flavors.collect import Collect
from modmex_lambda.stream.rules_registry import RulesRegistry
from modmex_lambda.stream.sources.sqs import sqs_source

listener_registry = RulesRegistry().registry(
    Collect({
        "id": "collect-candidate-events",
        "event_type": [
            "candidate.applied",
            "resume.analyzed",
            "interview.completed",
            "feedback.received",
        ],
        "correlation_key": "application.id",
        "table_name": "RecruitingEvents",
        "include_raw": True,
    })
)

@sqs_source(listener_registry, concurrency=True)
def handler(event, context):
    return {"statusCode": 200}
Enter fullscreen mode Exit fullscreen mode

sqs_source adapts the AWS event into a standard unit of work.
RulesRegistry registers the pipelines that should execute. Collect
filters by event_type, calculates a correlation key, and writes to
DynamoDB.

Conceptually, the persisted item looks like this:

{
  "pk": "candidate-applied-001",
  "sk": "EVENT",
  "discriminator": "EVENT",
  "timestamp": 1718050000000,
  "data": "application-789",
  "event": {
    "id": "candidate-applied-001",
    "type": "candidate.applied",
    "partition_key": "application-789"
  }
}
Enter fullscreen mode Exit fullscreen mode

This turns a transient event into operational memory.

From here, DynamoDB Streams can wake up the next step of the loop.

Correlate: building context

An isolated event is rarely enough to make a decision.

A candidate can have multiple applications. A job can have many
candidates. An application can accumulate analysis, interviews,
feedback, and approvals. That's why we need correlation.

Correlate takes events that have already been stored and rewrites
references under alternative keys. It doesn't duplicate business
reality; it creates efficient ways to query context.

Original event:
pk=candidate-applied-001, sk=EVENT

Correlations:
pk=application-789.application, sk=candidate-applied-001
pk=candidate-123.candidate, sk=candidate-applied-001
pk=job-456.job, sk=candidate-applied-001
Enter fullscreen mode Exit fullscreen mode

With modmex-lambda, we can declare those correlations like this:

from modmex_lambda.stream.flavors.correlate import Correlate
from modmex_lambda.stream.rules_registry import RulesRegistry
from modmex_lambda.stream.sources.dynamodb import dynamodb_source

trigger_registry = RulesRegistry().registry(
    Correlate({
        "id": "correlate-by-application",
        "event_type": ["candidate.applied", "resume.analyzed", "interview.completed", "feedback.received"],
        "correlation_key": "application.id",
        "correlation_key_suffix": "application",
        "table_name": "RecruitingEvents",
    }),
    Correlate({
        "id": "correlate-by-candidate",
        "event_type": ["candidate.applied", "resume.analyzed", "interview.completed", "feedback.received"],
        "correlation_key": "application.candidate.id",
        "correlation_key_suffix": "candidate",
        "table_name": "RecruitingEvents",
    }),
    Correlate({
        "id": "correlate-by-job",
        "event_type": ["candidate.applied", "resume.analyzed", "interview.completed", "feedback.received"],
        "correlation_key": "application.job.id",
        "correlation_key_suffix": "job",
        "table_name": "RecruitingEvents",
    }),
)

@dynamodb_source(trigger_registry, concurrency=True)
def handler(event, context):
    return {"statusCode": 200}
Enter fullscreen mode Exit fullscreen mode

This trigger isn't listening directly to the ATS. It listens to the
DynamoDB stream.

candidate.applied
        ↓
Collect → pk=candidate-applied-001, sk=EVENT
        ↓
DynamoDB Streams
        ↓
Correlate
        ↳ pk=application-789.application, sk=candidate-applied-001
        ↳ pk=candidate-123.candidate, sk=candidate-applied-001
        ↳ pk=job-456.job, sk=candidate-applied-001
Enter fullscreen mode Exit fullscreen mode

Now the agent can answer different questions:

  • What do we know about this application?
  • What do we know about this candidate?
  • What's happening with this job?

Without correlation, the agent only reacts. With correlation, it can
evaluate with context.

Evaluate: making decisions

Evaluate is where the decision lives. It can emit a higher-level event
directly, or query correlated events before deciding.

If an application arrives with a resume, we want to request an analysis:

from pydash import get
from modmex_lambda.stream.flavors.evaluate import Evaluate

trigger_registry.registry(
    Evaluate({
        "id": "request-resume-analysis",
        "event_type": "candidate.applied",
        "correlation_key_suffix": "application",
        "expression": lambda uow: bool(get(uow, "event.application.resume_s3_key")),
        "emit": lambda uow, rule, template: {
            **template,
            "type": "resume.analysis.requested",
            "application": get(uow, "event.application"),
        },
    })
)
Enter fullscreen mode Exit fullscreen mode

This rule doesn't analyze the resume. It only decides that the analysis
should happen. The result isn't a textual response. It's a new event:

{
  "id": "0.request-resume-analysis",
  "type": "resume.analysis.requested",
  "timestamp": 1718050000000,
  "partition_key": "application-789",
  "application": {
    "id": "application-789",
    "status": "submitted",
    "resume_s3_key": "resumes/johndoe.pdf",
    "candidate": {"id": "candidate-123"},
    "job": {"id": "job-456"}
  },
  "triggers": [{
    "id": "candidate-applied-001",
    "type": "candidate.applied",
    "timestamp": 1718050000000
  }]
}
Enter fullscreen mode Exit fullscreen mode

That event can be published to EventBridge and consumed by another
Lambda.

Up to this point, we haven't used an LLM at all. And that's deliberate.

Many agent decisions don't require generative AI. They require clear
rules, correct data, and well-modeled events.

The LLM as a Task

Now the LLM appears---but it appears late.

Before the prompt, we already had:

  • a business event;
  • persistence in the micro event store;
  • correlation;
  • an evaluation rule;
  • an event requesting analysis.

The LLM enters as a bounded task: analyzing the resume against the job
opening. For practical purposes, here's an implementation using OpenAI
Agents SDK:

from agents import Agent, Runner
from modmex import BaseModel
from modmex_lambda.stream.sources.sqs import sqs_source
from modmex_lambda.stream.flavors.task import Task

class ResumeAnalysis(BaseModel):
    matched_skills: list[str]
    missing_skills: list[str]
    confidence: float
    summary: str

resume_analyzer_agent = Agent(
    name="Resume analyzer",
    instructions=(
        "Analyze a resume against a job opening. "
        "Return a structured, auditable, and conservative evaluation. "
        "Do not invent skills that are not supported by the available evidence. "
        "confidence must be a number between 0 and 1."
    ),
    output_type=ResumeAnalysis,
)

class ResumeTextExtractor:
    def extract(self, resume_s3_key: str) -> str:
        # Real implementation:
        # 1. download the file from S3;
        # 2. extract text from the PDF/DOCX;
        # 3. clean or truncate the content if necessary.
        return "Extracted resume text..."

def analyze_resume_with_llm(uow, task):
    application = get(uow, "event.application")
    candidate = get(application, "candidate")
    job = get(application, "job")
    extractor = task.resolve(ResumeTextExtractor)
    resume_text = extractor.extract(application["resume_s3_key"])

    result = Runner.run_sync(
        resume_analyzer_agent,
        f"""
        Evaluate whether the candidate meets the job requirements.

        Candidate:
        {candidate}

        Job:
        {job}

        Resume:
        {resume_text}
        """,
    )

    analysis = result.final_output.model_dump()
    return {
        "candidate_id": candidate["id"],
        "job_id": job["id"],
        "application_id": application["id"],
        **analysis,
    }

listener_registry = RulesRegistry().registry(
    Task({
        "id": "analyze-resume",
        "event_type": "resume.analysis.requested",
        "execute": analyze_resume_with_llm,
        "emit": lambda uow, task, template: {
            **template,
            "type": "resume.analyzed",
            "application": {
                **get(uow, "event.application"),
                "analysis": uow["result"],
            },
        },
    })
)

@sqs_source(listener_registry, concurrency=True)
def handler(event, context):
    return {"statusCode": 200}
Enter fullscreen mode Exit fullscreen mode

The LLM didn't decide to hire anyone. It produced a structured signal:

{
  "type": "resume.analyzed",
  "application": {
    "id": "application-789",
    "analysis": {
      "matched_skills": ["python", "aws", "serverless"],
      "missing_skills": [],
      "confidence": 0.87,
      "summary": "The candidate meets the primary requirements."
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

That signal returns to the flow as an event.

The important output isn't a polished response for a user. It's a new
fact that other components can evaluate, audit, and process.

The LLM produces signals. The control service decides what to do with
them.

Evaluate with context

When resume.analyzed arrives, the system returns to the control loop.

def qualifies_for_screening(uow):
    analysis = get(uow, "event.application.analysis", {})
    correlated = uow["correlated"]

    has_application = any(
        event["type"] == "candidate.applied"
        for event in correlated
    )

    return (
        has_application
        and analysis.get("confidence", 0) >= 0.80
        and len(analysis.get("missing_skills", [])) == 0
    )

trigger_registry.registry(
    Evaluate({
        "id": "qualify-candidate",
        "event_type": "resume.analyzed",
        "correlation_key_suffix": "application",
        "expression": qualifies_for_screening,
        "emit": lambda uow, rule, template: {
            **template,
            "type": "candidate.qualified",
            "application": {
                **get(uow, "event.application"),
                "status": "qualified",
                "recommendation": "screening",
            },
        },
    })
)
Enter fullscreen mode Exit fullscreen mode

Here, Evaluate queries DynamoDB using the correlation key and receives
related events in uow["correlated"].

The agent's memory isn't a long conversation. It's a chain of durable
events.

If the rule passes, another event is emitted:

candidate.qualified
Enter fullscreen mode Exit fullscreen mode

And that event can activate another decision:

trigger_registry.registry(
    Evaluate({
        "id": "request-interview",
        "event_type": "candidate.qualified",
        "emit": lambda uow, rule, template: {
            **template,
            "type": "interview.requested",
            "application": {
                **get(uow, "event.application"),
                "next_step": "interview",
            },
        },
    })
)
Enter fullscreen mode Exit fullscreen mode

The complete flow becomes:

candidate.applied
        ↓
resume.analysis.requested
        ↓
resume.analyzed
        ↓
candidate.qualified
        ↓
interview.requested
Enter fullscreen mode Exit fullscreen mode

This looks much more like a real agent---not because it has a longer
prompt, but because it observes, remembers, evaluates, and acts.

Where exactly does the LLM fit?

It's worth repeating because this is the most important point:

  • The LLM doesn't listen to events.
  • The LLM doesn't coordinate the process.
  • The LLM isn't the memory system.
  • The LLM doesn't decide everything.
  • The LLM participates when a rule needs probabilistic evaluation.

The right mental model is:

Event:
  candidate.applied

Control service:
  collect
  correlate
  evaluate

Decision:
  we need to analyze the resume

Event:
  resume.analysis.requested

Task:
  call LLM

Event:
  resume.analyzed

Control service:
  evaluate with context

Event:
  candidate.qualified
Enter fullscreen mode Exit fullscreen mode

The prompt appears late. An event happened first.

Expire: time without live processes

There's another important detail in real systems: time.

Many decisions depend not only on something happening, but on something
not happening within a certain period.

For example:

InterviewRequested
        ↓ 48 hours without a response
InterviewRequestExpired
        ↓
HumanReviewRequired
Enter fullscreen mode Exit fullscreen mode

In a traditional design, we might create a permanent worker, a
scheduler, or a table of pending jobs. In this serverless
implementation, we can use DynamoDB TTL and turn expirations into events
through Expired.

This lets us express temporal logic without keeping a live process
constantly checking state.

Again, the agent doesn't need to be "thinking." It needs to react when
the system observes a new fact---even when that fact is that a time
window expired.

Why not let the LLM orchestrate everything?

At this point, a reasonable question appears: if today's models can call
tools natively, why build all this infrastructure with Lambdas, queues,
and databases? Why not connect an LLM to a set of tools, give it
context, and let it orchestrate the flow autonomously?

It's a valid question.

But the answer is fundamental to designing production agents:

An LLM is an excellent probabilistic reasoning engine, but a poor
infrastructure controller.

Delegating the primary orchestrator to an LLM creates serious problems:

  • Determinism vs. probability. Critical business routing cannot depend on a model's temperature. If a strict rule says the flow must stop when a required document is missing, you need deterministic code. Rules evaluate precisely; LLMs predict the next token.
  • Durable memory vs. context window. Traditional agents often try to preserve memory by injecting execution history into the prompt. If a process lasts days or accumulates dozens of events, context grows, latency increases, costs rise, and attention degrades. Operational memory belongs in a database, not a token window.
  • Fault tolerance. What happens when an external API times out? In an event-driven architecture, a failed task can be retried from the queue or moved to a Dead Letter Queue (DLQ). The whole system doesn't need to collapse.
  • Real asynchrony. Business decisions take time. If we need to wait 48 hours for a candidate's response, an LLM cannot remain in an infinite loop waiting. An asynchronous design with expirable events handles time naturally without idle processes.

The LLM doesn't disappear from this architecture. It simply occupies the
place where it provides the most value.

  • The control service coordinates the flow.
  • Events preserve state.
  • Infrastructure provides resilience.
  • The LLM participates when we need interpretation, classification, summarization, or probabilistic judgment.

Each component does what it was designed to do.

Why does serverless fit so well?

Serverless fits event-driven agents for a simple reason:

An event-driven agent spends most of its time waiting.

You don't need to pay for an idle process just so the agent can exist.
You need the system to wake up correctly when something relevant
happens.

Pay per use

If there are no applications, analyses, interviews, or feedback, there
are no invocations.

no events = no invocations = near-zero compute cost
Enter fullscreen mode Exit fullscreen mode

This matters even more when LLMs are involved because we don't only want
to control compute cost. We also want to control when and why inference
cost is incurred.

Automatic scalability

If 10 applications arrive, you process 10.

If 10,000 arrive, Lambda scales with the event volume while SQS,
Kinesis, or DynamoDB Streams absorb the spike.

The architecture doesn't need to assume a fixed number of running
agents. Throughput follows the event flow.

Failure isolation

Separating actions through events gives each capability its own
boundary:

  • resume analysis;
  • correlation;
  • decision;
  • notification;
  • interview scheduling;
  • human review.

If the LLM fails, the whole agent doesn't have to fail. The failure can
be recorded as an event. It can be retried. It can be escalated. It can
be sent to human review.

Auditability

Every important decision produces an event:

resume.analysis.requested
resume.analyzed
candidate.qualified
interview.requested
Enter fullscreen mode Exit fullscreen mode

This lets us reconstruct why the system acted.

In systems that use LLMs, that traceability matters. It isn't enough to
know that "the model said something." We need to know which event
triggered it, what data was available, which rule ran, and which action
was emitted.

Evolution

Adding a new capability doesn't require rewriting the entire agent.

You can add another rule:

RulesRegistry().registry(
    Evaluate({...}),
    Task({...}),
)
Enter fullscreen mode Exit fullscreen mode

Or add a new EventBridge consumer.

This lets the system evolve by adding capabilities rather than through
massive rewrites.

The complete agent

When we put the pieces together, the resulting system doesn't look like
a chatbot.

It also doesn't look like a linear workflow of prompts.

It is implemented using control-system patterns.

It observes lower-level events. It stores them. It correlates them. It
evaluates conditions. It emits higher-level events. It executes actions.
And those actions produce new events.

That loop is a practical way to implement the agent.

Conclusion

I think much of the current confusion around agents comes from starting
the conversation with the prompt. The important shift isn't only
technical. It's a change in mental model.

If we start with the prompt, we end up designing conversations.

If we start with the event, we can design control systems.

Prompts matter. LLMs matter too. But before a prompt exists, something
has usually already happened.

  • A candidate applied.
  • A customer wrote.
  • A document was uploaded.
  • A payment was approved.

A production agent needs to react to those facts. It needs durable
memory. It needs correlation. It needs rules. It needs auditability. And
at some points, it may need an LLM.

But the LLM doesn't define the agent.

In this architecture, the agent is implemented as a control loop:

event
  ↓
collect
  ↓
correlate
  ↓
evaluate
  ↓
emit
  ↓
action
  ↓
event
Enter fullscreen mode Exit fullscreen mode

Agents don't start with prompts. In production, they start with
events.

Links

Top comments (0)