<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ram Ji  Tripathi</title>
    <description>The latest articles on DEV Community by Ram Ji  Tripathi (@ramji_tripathi_095c7f4810).</description>
    <link>https://dev.to/ramji_tripathi_095c7f4810</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2066272%2F32e36d61-8064-47fa-9de4-c8b24132fa4b.png</url>
      <title>DEV Community: Ram Ji  Tripathi</title>
      <link>https://dev.to/ramji_tripathi_095c7f4810</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ramji_tripathi_095c7f4810"/>
    <language>en</language>
    <item>
      <title>Redis + BullMQ for AI Workloads: Retries Are Not Enough — You Also Need Idempotency and Ordering</title>
      <dc:creator>Ram Ji  Tripathi</dc:creator>
      <pubDate>Sun, 04 Oct 2026 12:43:57 +0000</pubDate>
      <link>https://dev.to/ramji_tripathi_095c7f4810/redis-bullmq-for-ai-workloads-retries-are-not-enough-you-also-need-idempotency-and-ordering-51hj</link>
      <guid>https://dev.to/ramji_tripathi_095c7f4810/redis-bullmq-for-ai-workloads-retries-are-not-enough-you-also-need-idempotency-and-ordering-51hj</guid>
      <description>&lt;p&gt;Moving AI analysis out of an HTTP request improves the request path. Queueing that work does not make it correct.&lt;/p&gt;

&lt;p&gt;In my AI Support Assistant, a new message triggers BullMQ analysis of the ticket conversation. The request waits for database writes and queue publication, then returns without waiting for the AI results.&lt;/p&gt;

&lt;p&gt;But a retryable job can still run more than once, reload different context, race with another job, repeat provider work or overwrite newer results. The message can also be saved without its analysis job reaching Redis.&lt;/p&gt;

&lt;p&gt;Examining the current implementation makes those boundaries concrete. The scenarios below are correctness risks exposed by the code, not production incidents; the proposed fixes are not implemented features.&lt;/p&gt;

&lt;h2&gt;
  
  
  The System This Came From
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://blog.ramjwork.in/ai-engineering/ai-customer-support-architecture" rel="noopener noreferrer"&gt;The first article&lt;/a&gt; described the broader React, Node.js, PostgreSQL, Redis and OpenAI architecture.&lt;/p&gt;

&lt;p&gt;This article focuses on one workflow: a message is added to a support ticket, and background processing generates a summary, sentiment and priority. Those values are written back to the ticket. The worker also records an &lt;code&gt;AIInteraction&lt;/code&gt; containing the conversation, summary response and summary token usage.&lt;/p&gt;

&lt;p&gt;The queue is named &lt;code&gt;ai-processing&lt;/code&gt;. Its job type is &lt;code&gt;generate-ticket-summary&lt;/code&gt;, although the processor does more than summarize. It analyzes the whole ticket conversation, rather than just the message that triggered publication.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the AI Work Went to BullMQ
&lt;/h2&gt;

&lt;p&gt;Saving a message and enriching a ticket have different completion requirements. The message is the user-authored record. AI metadata can arrive later.&lt;/p&gt;

&lt;p&gt;The current boundary looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;React client → Node.js API → PostgreSQL: Message + Ticket timestamp
                         → Redis/BullMQ: publish analysis job
                         → HTTP 201 after publication

Redis/BullMQ → separate worker → load ticket messages
                              → summary → sentiment → priority
                              → update Ticket → insert AIInteraction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Publication makes the job available to the worker. The worker may start before HTTP 201 is sent; the diagram shows two execution paths, not a guarantee that background processing begins after the response.&lt;/p&gt;

&lt;p&gt;The important boundary is what the request awaits: persistence and publication, followed by a return to the caller. Summary, sentiment and priority belong to the worker's lifecycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Actual Producer Flow
&lt;/h2&gt;

&lt;p&gt;From &lt;code&gt;backend/src/modules/messages/message.service.ts&lt;/code&gt;, lines 26–58. Authorization precedes this excerpt. The excerpt is otherwise complete; blank lines are condensed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;createMessageRepository&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;updateTicketRepository&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;updatedAt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;aiQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;GENERATE_TICKET_SUMMARY_JOB&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;messageId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;backoff&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;exponential&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;removeOnComplete&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;removeOnFail&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The repository helpers execute &lt;code&gt;prisma.message.create&lt;/code&gt; and &lt;code&gt;prisma.ticket.update&lt;/code&gt;. These are separate writes, followed by publication to Redis. There is no encompassing transaction or outbox in this flow.&lt;/p&gt;

&lt;p&gt;The controller awaits this service before returning HTTP 201. Consequently, a message can already exist even if the service does not reach its successful return. Queue publication is part of request completion; AI processing is not.&lt;/p&gt;

&lt;p&gt;The payload includes &lt;code&gt;messageId&lt;/code&gt;, but the worker's processing logic only consumes &lt;code&gt;ticketId&lt;/code&gt;. The message ID currently provides neither a processing boundary nor an idempotency key.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Worker Actually Does
&lt;/h2&gt;

&lt;p&gt;The worker dispatches &lt;code&gt;generate-ticket-summary&lt;/code&gt; to &lt;code&gt;processTicketSummaryJob&lt;/code&gt; using only the ticket ID. Its configuration in &lt;code&gt;backend/src/queues/ai.worker.ts&lt;/code&gt;, lines 71–75, is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;connection&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;concurrency&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the complete options object from the worker constructor, with blank lines condensed.&lt;/p&gt;

&lt;p&gt;Five jobs can run concurrently within this worker instance, including jobs for the same ticket. This is a local worker limit, not a global or per-ticket limit. See BullMQ’s &lt;a href="https://docs.bullmq.io/guide/workers/concurrency" rel="noopener noreferrer"&gt;worker concurrency guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The processor loads messages with &lt;code&gt;where: { ticketId }&lt;/code&gt;, includes sender role and name, and sorts by &lt;code&gt;createdAt: "asc"&lt;/code&gt;. It returns early when there are no messages. Otherwise it formats the conversation and executes this sequence from lines 113–127:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;aiResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generateTicketSummary&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="nx"&gt;aiContext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sentimentResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generateSentiment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="nx"&gt;aiContext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;priorityResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;detectTicketPriority&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;aiContext&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These operations are sequential within one job. Different jobs can still overlap. The service routes generation through an OpenAI-primary orchestrator with optional Gemini fallback for eligible failures, so three logical operations do not necessarily mean exactly three external HTTP requests.&lt;/p&gt;

&lt;p&gt;After sentiment and priority validation, lines 133–155 perform two separate writes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;prisma&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;where&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;aiSummary&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;aiResult&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;sentiment&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;createAIInteraction&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;aiResult&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;tokensUsed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;aiResult&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tokensUsed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Blank lines are condensed. &lt;code&gt;createAIInteraction&lt;/code&gt; is a plain Prisma insert. There is no job identity or conversation version attached to that record, and no conditional version check on the ticket update.&lt;/p&gt;

&lt;p&gt;The worker runs through a separate entry point in &lt;code&gt;backend/src/worker.ts&lt;/code&gt;. That separates its execution lifecycle from HTTP handling; it does not add coordination between jobs for one ticket.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retries Solve Failure Recovery, Not Correctness
&lt;/h2&gt;

&lt;p&gt;The producer configures three total attempts, including the initial execution, with exponential backoff and a five-second base. For ordinary thrown-error retries, that gives nominal delays of five and ten seconds before the remaining attempts; scheduling can add waiting time. See BullMQ's &lt;a href="https://docs.bullmq.io/guide/retrying-failing-jobs" rel="noopener noreferrer"&gt;retry documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Completed jobs are removed. Failed jobs are retained. Those choices affect recovery and diagnosis, but they do not make processing idempotent.&lt;/p&gt;

&lt;p&gt;Retry asks, “Should I try this work again after failure?” Idempotency asks, “If I perform it again, will the resulting state remain correct?” Deduplication asks, “Should equivalent work be admitted or executed again at all?”&lt;/p&gt;

&lt;p&gt;Here, retries restart a processor that calls providers and performs database writes. The retry configuration supplies no rule for deciding whether those effects have already happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure Mode 1: The Database Commits but the Queue Does Not
&lt;/h2&gt;

&lt;p&gt;Consider the producer's actual order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The message insert succeeds.&lt;/li&gt;
&lt;li&gt;The ticket timestamp update succeeds.&lt;/li&gt;
&lt;li&gt;Publication rejects, or the process exits before publication completes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The database now contains the conversation change without assured corresponding queue work. Retrying a worker cannot recover a job that was never published.&lt;/p&gt;

&lt;p&gt;A publication error can also leave an ambiguous outcome: Redis might have accepted the job before the caller lost its acknowledgement. Recovery must handle both missing work and repeated publication.&lt;/p&gt;

&lt;p&gt;The earlier boundary also matters: if the ticket update fails after message creation, the message remains persisted and publication is never reached. Wrapping those two database writes in a transaction would address that database-only partial state. It would not atomically commit a Redis job with PostgreSQL data.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;proposed transactional outbox&lt;/strong&gt; changes where publication intent becomes durable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BEFORE — CURRENT
Message insert → Ticket update → Redis publication
     database writes complete       can fail separately

IMPROVED — PROPOSED
PostgreSQL transaction:
  Message insert + Ticket update + Outbox event
                         ↓ committed publication intent
Outbox dispatcher → Redis/BullMQ → worker
       ↓
mark event published after successful queue publication
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dispatcher retries unpublished events. If it publishes and crashes before marking the event, it may publish again. An outbox therefore needs stable event identity and duplicate-safe consumption; it does not confer exactly-once execution.&lt;/p&gt;

&lt;p&gt;This adds a table, dispatcher lifecycle, backlog monitoring and retention policy. For optional enrichment in an early application, explicit reconciliation may be an interim choice. If every accepted message must eventually be analyzed, I would prioritize an outbox regardless of traffic volume. Reliability requirements, rather than scale alone, determine when it is needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure Mode 2: Retrying AI Work Is Not Automatically Idempotent
&lt;/h2&gt;

&lt;p&gt;Suppose summary generation succeeds and sentiment generation fails with an error that escapes the provider layer. BullMQ can retry the processor. It starts again by loading messages and generating the summary. There is no checkpoint that reuses the previous summary.&lt;/p&gt;

&lt;p&gt;Now move the failure later. Summary, sentiment and priority succeed; the ticket update commits; the interaction insert fails. The ticket already contains AI metadata, but the job has failed. A retry repeats the AI sequence and attempts both database writes again.&lt;/p&gt;

&lt;p&gt;The resulting values need not match the previous output. More messages may also have arrived before the retry reloads the conversation. The same job payload therefore does not necessarily identify the same input snapshot.&lt;/p&gt;

&lt;p&gt;A duplicate execution that reaches the insert can append another &lt;code&gt;AIInteraction&lt;/code&gt;: its schema uses a generated UUID and has no unique analysis key. Interruption after application writes but before queue completion is recorded creates another window for repeated effects. The relevant effects here are provider requests, ticket metadata writes and interaction records.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deterministic job identity helps at publication
&lt;/h3&gt;

&lt;p&gt;A proposed identity such as &lt;code&gt;ticket-analysis-&amp;lt;messageId&amp;gt;&lt;/code&gt; could suppress repeated publication for the same saved message while that queue job exists. It would not suppress jobs triggered by two different messages, nor deduplicate message creation if a repeated HTTP request creates a new message ID.&lt;/p&gt;

&lt;p&gt;BullMQ's &lt;a href="https://docs.bullmq.io/guide/jobs/job-ids" rel="noopener noreferrer"&gt;job ID documentation&lt;/a&gt; explains that an existing custom ID prevents another job with that ID from being added. Once the job is removed, the ID no longer protects against duplicates. That is directly relevant to this application's &lt;code&gt;removeOnComplete: true&lt;/code&gt; setting.&lt;/p&gt;

&lt;h3&gt;
  
  
  Durable idempotency belongs at the application effect
&lt;/h3&gt;

&lt;p&gt;For this workflow, I would define a logical analysis key from ticket identity, conversation revision and analysis/prompt version. A durable record with a unique constraint could track whether that analysis has already been applied.&lt;/p&gt;

&lt;p&gt;The ticket update, interaction insert and successful analysis marker should commit together, with concurrent duplicate attempts resolved by database constraints and conditional writes. A separate initial “already done?” read is insufficient because two workers can both pass it.&lt;/p&gt;

&lt;p&gt;That would protect database effects. Provider requests made before commit could still repeat. Reusing stored generation results would require checkpointing and invalidation rules; I would start with atomic result persistence and durable identity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure Mode 3: Ordering Is a Business Constraint
&lt;/h2&gt;

&lt;p&gt;Two messages arrive quickly on ticket T:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Message A is saved and job A starts. It reads the conversation containing A.&lt;/li&gt;
&lt;li&gt;Message B is saved and job B starts in another available worker slot. It reads A and B.&lt;/li&gt;
&lt;li&gt;Job B finishes first and writes metadata based on A and B.&lt;/li&gt;
&lt;li&gt;Job A finishes later and overwrites that metadata using its older context.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The current ticket update matches only &lt;code&gt;id: ticketId&lt;/code&gt;. Nothing prevents step four. This is a possible stale, last-writer-wins result visible from the architecture.&lt;/p&gt;

&lt;p&gt;Both jobs could also start after B exists, read overlapping or identical context, and perform redundant analysis. Because &lt;code&gt;messageId&lt;/code&gt; is unused, neither job is restricted to the conversation as it stood when its triggering message was saved.&lt;/p&gt;

&lt;p&gt;Sorting messages inside each job does not order jobs or their writes. The timestamp sort also has no explicit tie-breaker for equal timestamps.&lt;/p&gt;

&lt;p&gt;Global throughput and per-ticket correctness are separate constraints. Lowering one worker's concurrency to one would reduce overlap there, but also serialize unrelated tickets and would not establish a cross-process ordering policy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Idempotency, Deduplication and Ordering Are Different
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concept&lt;/th&gt;
&lt;th&gt;Question it answers&lt;/th&gt;
&lt;th&gt;Example in this system&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Retry&lt;/td&gt;
&lt;td&gt;Should failed work be attempted again?&lt;/td&gt;
&lt;td&gt;Retry a ticket analysis after a transient failure.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deduplication&lt;/td&gt;
&lt;td&gt;Should equivalent work be admitted again?&lt;/td&gt;
&lt;td&gt;Proposed: suppress publication while the same job ID exists.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Idempotency&lt;/td&gt;
&lt;td&gt;Will repeating the work preserve correct state?&lt;/td&gt;
&lt;td&gt;Proposed: commit one analysis record for a ticket revision.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ordering&lt;/td&gt;
&lt;td&gt;In what sequence may work take effect?&lt;/td&gt;
&lt;td&gt;Proposed: apply ticket analyses in conversation-revision order.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concurrency&lt;/td&gt;
&lt;td&gt;How much work may run at the same time?&lt;/td&gt;
&lt;td&gt;Allow five active jobs within this worker instance.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Only retries and worker concurrency are configured in the current path. The other examples describe proposed behavior. A version check can reject stale results without executing every job in strict order; serialization still needs a policy for retries and earlier failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Better Per-Ticket Processing Model
&lt;/h2&gt;

&lt;p&gt;The following is a &lt;strong&gt;proposed model&lt;/strong&gt;, not current code. For summary, sentiment and priority, the business goal is usually a current conversation view. It may not require publishing every intermediate analysis.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Messages A, B, C → atomically advance ticket conversation revision
                                   ↓
                       record latest requested revision
                                   ↓
                 coalesce pending work for this ticket
                                   ↓
                 one active analysis per ticket, if needed
                                   ↓
                    read consistent snapshot at revision R
                                   ↓
                            generate AI results
                                   ↓
                  atomic apply only if revision is still R
                          ↙                      ↘
                stale: discard              current: commit
                    ↓                    results + analysis record
         ensure latest revision remains scheduled
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Version checks are the first correctness protection I would add.&lt;/strong&gt; A dedicated conversation revision should advance atomically with message persistence. The worker must read the revision and messages as a consistent snapshot, then conditionally apply results in a short database transaction. A separate check followed by an unconditional update leaves a race.&lt;/p&gt;

&lt;p&gt;I would avoid using &lt;code&gt;Ticket.updatedAt&lt;/code&gt; as the conversation revision. Prisma marks it &lt;code&gt;@updatedAt&lt;/code&gt;, so AI writes and unrelated ticket edits also change it. A dedicated revision makes the input contract explicit. The tradeoff is schema and transaction work; discarded stale analyses still consume processing. This protection is useful now because the current concurrency already permits stale writes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Coalescing can reduce unnecessary work.&lt;/strong&gt; Several pending triggers can represent one request to analyze the latest conversation. A short debounce period trades freshness for fewer intermediate analyses. BullMQ offers &lt;a href="https://docs.bullmq.io/guide/jobs/deduplication" rel="noopener noreferrer"&gt;deduplication modes&lt;/a&gt;, but their lifetime semantics must match the application.&lt;/p&gt;

&lt;p&gt;In particular, simply suppressing every new job while one analysis is active can miss updates arriving after that analysis loaded its messages. The design must durably track the latest requested revision and arrange another pass when needed. Checking once at the end without an atomic handoff can still lose a concurrent update. Coalescing is a pragmatic next step when message bursts cause redundant work; it is not a substitute for version checks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Per-ticket serialization can limit overlap.&lt;/strong&gt; A coordinated lease or partitioned processing scheme can allow different tickets to progress concurrently while restricting one ticket's active analysis. An in-memory mutex covers only one process. Distributed leases introduce expiry, renewal and crash-recovery concerns; conditional version checks remain valuable when a lease expires while work is still running.&lt;/p&gt;

&lt;p&gt;If a future job must act on every message in business order, it needs explicit sequence tracking and a policy for earlier failures. Coalescing intermediate analyses would no longer satisfy that requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Change First
&lt;/h2&gt;

&lt;p&gt;These are recommendations for this application, not features already implemented.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Define analysis identity and protect freshness together.&lt;/strong&gt; Introduce a conversation revision, conditional result application and an atomic transaction for ticket metadata, interaction and an analysis marker. Then use deterministic publication identity for the chosen unit of work. This addresses correctness exposed by today's concurrency; it costs a schema change and careful transaction design.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coalesce pending ticket analysis where the product permits it.&lt;/strong&gt; Treat the job as “refresh this ticket” rather than “produce every intermediate result,” while preserving a durable follow-up for changes during active processing. This costs scheduling complexity and may delay freshness slightly. It becomes more useful when bursts create redundant work. Add coordinated per-ticket serialization if overlap remains significant; version protection comes first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Close the database-to-queue gap when eventual analysis is required.&lt;/strong&gt; Add an outbox and replay-safe dispatcher. Also review producer connection policy: the shared Redis module sets &lt;code&gt;maxRetriesPerRequest: null&lt;/code&gt;, and the message service has no explicit publication deadline. BullMQ’s &lt;a href="https://docs.bullmq.io/guide/connections" rel="noopener noreferrer"&gt;connection guidance&lt;/a&gt; distinguishes worker recovery from bounded waiting in HTTP producers. The operational cost is real, but the need depends on the reliability contract, not a hypothetical high-traffic milestone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build recovery on the visibility already present.&lt;/strong&gt; The worker logs receipt, completion, failures, ticket/job IDs and duration. Before using the attempt fields for alerts, I would verify their meaning at receipt and failure events. Add exhausted-attempt alerts, queue age/backlog visibility, stale-result counters and revision correlation. These are useful now; a larger telemetry platform can wait.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add a deliberate replay policy.&lt;/strong&gt; Retained failed jobs are a starting point, not an implemented dead-letter workflow. An operator should distinguish transient failures from invalid input, replay after repair, and know whether replay means the original snapshot or latest conversation. Controlled replay tooling and retention add maintenance work. A separate dead-letter queue is optional until volume or operational ownership warrants it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I would evaluate these changes through duplicate work, result freshness and publication failures. Infrastructure changes should follow a demonstrated constraint.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Not Add Yet
&lt;/h2&gt;

&lt;p&gt;I would not introduce Kafka, Kubernetes or decompose the application into microservices solely because it has a queue.&lt;/p&gt;

&lt;p&gt;The concrete concerns are durable publication intent, safe repeated writes and valid application of conversation-derived results. They can be addressed within PostgreSQL, Redis, BullMQ and the existing API/worker boundary.&lt;/p&gt;

&lt;p&gt;More infrastructure may become appropriate under specific throughput, ownership or deployment constraints. It does not remove the need to define what makes an analysis current or safe to repeat.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why AI Workloads Make This More Visible
&lt;/h2&gt;

&lt;p&gt;In this worker, three sequential generation operations separate the initial conversation read from the final ticket update. While those external requests are in flight, another message can arrive and another job can complete. The queue can function normally while the result becomes obsolete.&lt;/p&gt;

&lt;p&gt;Provider calls can be relatively slow, and successful generation consumes usage billed under the provider’s terms; OpenAI documents &lt;a href="https://developers.openai.com/api/docs/pricing" rel="noopener noreferrer"&gt;token-based API pricing&lt;/a&gt;. Repeating a completed summary because sentiment failed can therefore repeat meaningful work, even when no database row has yet changed. This is a reason to control repetition, not a measured cost or latency claim about this application.&lt;/p&gt;

&lt;p&gt;Generated text can also differ between attempts. Overwriting &lt;code&gt;aiSummary&lt;/code&gt; with a second output is not automatically idempotent just because both writes target the same row. Meanwhile, this processor reloads messages on every attempt, so a retry can change both its input and its output.&lt;/p&gt;

&lt;p&gt;For this workload, the useful contract is: which conversation revision does this analysis describe, and when may it take effect? The queue schedules execution. The application must answer that question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Takeaway
&lt;/h2&gt;

&lt;p&gt;BullMQ gives this application a clear boundary between message creation and background AI enrichment. Correctness still needs an application-level design: durable publication intent, duplicate-safe effects and protection against stale conversation results.&lt;/p&gt;

&lt;p&gt;Retries make another attempt possible. Idempotency and ordering rules determine whether that attempt is safe to apply.&lt;/p&gt;

&lt;p&gt;Originally part of the engineering work behind my AI Support Assistant architecture.&lt;/p&gt;

&lt;p&gt;Originally published on my engineering blog:&lt;br&gt;
&lt;a href="https://blog.ramjwork.in/ai-engineering/redis-bullmq-ai-background-jobs" rel="noopener noreferrer"&gt;https://blog.ramjwork.in/ai-engineering/redis-bullmq-ai-background-jobs&lt;/a&gt;&lt;/p&gt;

</description>
      <category>node</category>
      <category>redis</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
    <item>
      <title>How I Architected an AI Customer Support System with React, Node.js, PostgreSQL, Redis &amp; OpenAI</title>
      <dc:creator>Ram Ji  Tripathi</dc:creator>
      <pubDate>Sun, 04 Oct 2026 05:11:40 +0000</pubDate>
      <link>https://dev.to/ramji_tripathi_095c7f4810/how-i-architected-an-ai-customer-support-system-with-react-nodejs-postgresql-redis-openai-2jcf</link>
      <guid>https://dev.to/ramji_tripathi_095c7f4810/how-i-architected-an-ai-customer-support-system-with-react-nodejs-postgresql-redis-openai-2jcf</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Adding AI wasn’t the difficult part. Deciding where AI belongs in the application architecture was.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Adding an LLM API to an application is relatively straightforward.&lt;/p&gt;

&lt;p&gt;The harder questions begin after the first API call works.&lt;/p&gt;

&lt;p&gt;Should AI run inside the user’s HTTP request? What happens when the provider is slow? Which AI operations should be asynchronous? What belongs in React state versus server state? Should an AI-generated support reply be sent automatically? What happens when a background job retries after partially completing?&lt;/p&gt;

&lt;p&gt;I ran into these questions while building an AI-powered customer support application with React, TypeScript, Node.js, PostgreSQL, Redis, BullMQ and OpenAI.&lt;/p&gt;

&lt;p&gt;The interesting part of the project wasn’t the prompt.&lt;/p&gt;

&lt;p&gt;It was deciding &lt;strong&gt;where AI should—and should not—sit inside the system.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This article walks through that architecture, shows selected code from the actual implementation, and discusses several things I would change before scaling the design further.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the Application Does
&lt;/h2&gt;

&lt;p&gt;The application revolves around support tickets and conversations.&lt;/p&gt;

&lt;p&gt;It has three roles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Customers&lt;/strong&gt; create tickets, read their own tickets and add messages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agents&lt;/strong&gt; work with assigned tickets, update their status and request AI-generated reply suggestions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Admins&lt;/strong&gt; have broader ticket access and assignment capabilities.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI assists the workflow in two different ways.&lt;/p&gt;

&lt;p&gt;First, when a new message is added, background processing analyzes the ticket conversation to generate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a ticket summary,&lt;/li&gt;
&lt;li&gt;sentiment,&lt;/li&gt;
&lt;li&gt;priority.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Second, an agent or admin can explicitly request an AI-generated reply suggestion.&lt;/p&gt;

&lt;p&gt;Those two AI workloads look similar at first—they both call an LLM—but they have very different execution requirements.&lt;/p&gt;

&lt;p&gt;That distinction shaped much of the architecture.&lt;/p&gt;




&lt;h2&gt;
  
  
  High-Level Architecture
&lt;/h2&gt;

&lt;p&gt;At a high level, the system looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    ┌─────────────────────────┐
                    │       React Client      │
                    │                         │
                    │ Auth / Tickets /        │
                    │ Messages / AI Replies   │
                    └────────────┬────────────┘
                                 │
                              REST API
                                 │
                    ┌────────────▼────────────┐
                    │     Node.js / Express   │
                    │                         │
                    │ Auth → Controller →     │
                    │ Service → Repository    │
                    └───────┬─────────┬───────┘
                            │         │
                            │         │ enqueue
                            │         ▼
                    ┌───────▼───┐  ┌──────────────┐
                    │PostgreSQL │  │ Redis/BullMQ │
                    └───────────┘  └──────┬───────┘
                                         │
                                         ▼
                                  ┌─────────────┐
                                  │  AI Worker  │
                                  │             │
                                  │ Summary     │
                                  │ Sentiment   │
                                  │ Priority    │
                                  └──────┬──────┘
                                         │
                                    LLM provider
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application is deliberately not split into many network services.&lt;/p&gt;

&lt;p&gt;The API is a &lt;strong&gt;modular monolith&lt;/strong&gt;, while AI background execution runs in a separate worker process.&lt;/p&gt;

&lt;p&gt;That gives me an important boundary without introducing a distributed-service architecture prematurely:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;HTTP request handling and background AI processing have independent execution lifecycles.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Frontend State: Don’t Put Everything in One Store
&lt;/h2&gt;

&lt;p&gt;One frontend decision was to separate state based on &lt;strong&gt;who owns it&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I use three broad categories:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;authentication/session state → React Context,&lt;/li&gt;
&lt;li&gt;remote server state → TanStack Query,&lt;/li&gt;
&lt;li&gt;transient UI state → local component state.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, tickets are server-owned resources, so their fetching and cache lifecycle belong to TanStack Query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;useTickets&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="nx"&gt;TicketListParams&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;useQuery&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;queryKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ticketKeys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;queryFn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;getTickets&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;useTicket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;useQuery&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;queryKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ticketKeys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;queryFn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;getTicketById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Boolean&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;useCreateTicket&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;queryClient&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useQueryClient&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;useMutation&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;mutationFn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="na"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;CreateTicketRequest&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;createTicket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;

    &lt;span class="na"&gt;onSuccess&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;queryClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;invalidateQueries&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;queryKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ticketKeys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lists&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part here isn’t the library.&lt;/p&gt;

&lt;p&gt;It’s the ownership model.&lt;/p&gt;

&lt;p&gt;A ticket fetched from the API is not fundamentally client state. The server remains authoritative.&lt;/p&gt;

&lt;p&gt;TanStack Query handles fetching, caching and invalidation, while component-local state handles things such as editable drafts and temporary UI interactions.&lt;/p&gt;

&lt;p&gt;This prevents a common React architecture problem: turning one global store into a mirror of the backend.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Backend Is a Modular Monolith
&lt;/h2&gt;

&lt;p&gt;The backend follows a layered flow that is roughly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Route
  ↓
Authentication / Authorization / Validation
  ↓
Controller
  ↓
Service
  ↓
Repository
  ↓
Prisma
  ↓
PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives features such as tickets and messages clear boundaries without requiring each feature to become an independent service.&lt;/p&gt;

&lt;p&gt;The service layer owns application workflows.&lt;/p&gt;

&lt;p&gt;Repositories isolate most persistence operations.&lt;/p&gt;

&lt;p&gt;Controllers remain focused primarily on translating HTTP requests and responses.&lt;/p&gt;

&lt;p&gt;AI is a partial exception because some worker-side operations interact more directly with Prisma, but the overall HTTP architecture follows this separation.&lt;/p&gt;

&lt;p&gt;For the current scope, I prefer this to creating multiple microservices.&lt;/p&gt;

&lt;p&gt;A service boundary should solve an actual deployment, ownership or scaling problem—not exist simply because the system has multiple features.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why I Moved AI Analysis Out of the Request Path
&lt;/h2&gt;

&lt;p&gt;Consider what happens when a customer adds a message.&lt;/p&gt;

&lt;p&gt;The important user-facing operation is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Persist my message successfully.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Summary, sentiment and priority enrichment are useful, but the user should not have to wait for several LLM operations to complete before the message request can finish.&lt;/p&gt;

&lt;p&gt;The service therefore persists the message and publishes a BullMQ job:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;createMessageService&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;CreateMessageInput&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getAuthorizedTicketService&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;accessContext&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;createMessageRepository&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;updateTicketRepository&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;updatedAt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;aiQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nx"&gt;GENERATE_TICKET_SUMMARY_JOB&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;messageId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;backoff&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;exponential&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="na"&gt;removeOnComplete&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;removeOnFail&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is an important architectural detail here.&lt;/p&gt;

&lt;p&gt;The HTTP request &lt;strong&gt;does wait for queue publication&lt;/strong&gt;, but it does &lt;strong&gt;not wait for the AI analysis itself&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST message
     │
     ▼
Authorize access
     │
     ▼
Persist message
     │
     ▼
Update ticket timestamp
     │
     ▼
Publish BullMQ job
     │
     ├──────────────► HTTP response
     │
     ▼
Background worker
     │
     ▼
AI processing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That removes LLM execution time from the main request path while still ensuring the API knows whether it successfully handed work to the queue.&lt;/p&gt;

&lt;p&gt;It also allows the worker lifecycle to be operated separately from the HTTP server.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the Background Worker Actually Does
&lt;/h2&gt;

&lt;p&gt;BullMQ uses Redis as the queue infrastructure.&lt;/p&gt;

&lt;p&gt;A separate worker consumes AI-processing jobs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;aiWorker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Worker&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;AI_QUEUE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

  &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;startedAt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="nx"&gt;jobStartTimes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nx"&gt;startedAt&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ticketId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
      &lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="nx"&gt;workerLogger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ai_job_received&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;jobId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;jobName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;attempt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;attemptsMade&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Job received&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;switch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="na"&gt;GENERATE_TICKET_SUMMARY_JOB&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;processTicketSummaryJob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

      &lt;span class="nl"&gt;default&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;workerLogger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
          &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="na"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ai_job_unknown&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="na"&gt;jobId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="na"&gt;jobName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="p"&gt;},&lt;/span&gt;
          &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Unknown job&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;

  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;connection&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;concurrency&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Inside the ticket-processing workflow, messages are loaded in conversation order and transformed into the conversation passed to the AI layer.&lt;/p&gt;

&lt;p&gt;The relevant part of the processing pipeline is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;conversation&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;formatConversation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;aiResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generateTicketSummary&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="nx"&gt;aiContext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sentimentResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generateSentiment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="nx"&gt;aiContext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;priorityResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;detectTicketPriority&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;aiContext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sentiment&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;validateSentiment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;sentimentResult&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sentiment&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;priority&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;validatePriority&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;priorityResult&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;priority&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Summary, sentiment and priority are currently generated sequentially inside a job.&lt;/p&gt;

&lt;p&gt;The normalized results are then used to update ticket-level AI fields, and an AI interaction record is created for the summary workflow.&lt;/p&gt;

&lt;p&gt;This design gives the application retryable asynchronous execution without coupling LLM latency directly to message creation.&lt;/p&gt;

&lt;p&gt;But queues introduce their own problems.&lt;/p&gt;

&lt;p&gt;I’ll return to those shortly.&lt;/p&gt;




&lt;h2&gt;
  
  
  Not Every AI Operation Belongs in a Queue
&lt;/h2&gt;

&lt;p&gt;Once background processing existed, it would have been easy to conclude:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI is slow, therefore every AI operation should be queued.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I don’t think that is the right abstraction.&lt;/p&gt;

&lt;p&gt;The correct execution model depends on what the user is doing.&lt;/p&gt;

&lt;p&gt;Ticket analysis is enrichment. It can happen asynchronously.&lt;/p&gt;

&lt;p&gt;An agent clicking &lt;strong&gt;Generate Reply&lt;/strong&gt;, however, is explicitly waiting for an answer.&lt;/p&gt;

&lt;p&gt;For that workflow, the backend generates the suggestion synchronously:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;conversation&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;formatConversation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;aiResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generateReplySuggestion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;requestId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;res&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ApiResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Reply suggestion generated&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;aiResult&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important distinction is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Background analysis
Message → Queue → Worker → AI → Ticket metadata

Interactive assistance
Agent → Request suggestion → AI → Suggested text → Agent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second flow doesn’t automatically create or send a support message.&lt;/p&gt;

&lt;p&gt;That is intentional.&lt;/p&gt;




&lt;h2&gt;
  
  
  Human-in-the-Loop AI
&lt;/h2&gt;

&lt;p&gt;A generated support reply can be wrong, incomplete or inappropriate for the context.&lt;/p&gt;

&lt;p&gt;So the system treats AI output as a &lt;strong&gt;suggestion&lt;/strong&gt;, not an autonomous action.&lt;/p&gt;

&lt;p&gt;On the frontend:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;ReplySuggestion&lt;/span&gt;
  &lt;span class="na"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;onUseReply&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;setComposerDraft&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nf"&gt;setComposerKey&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;current&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;current&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;

&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;MessageComposer&lt;/span&gt;
  &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;composerKey&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;initialContent&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;composerDraft&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;onMessageSent&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="nf"&gt;setMessageRefreshToken&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;current&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;current&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="si"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the user selects &lt;strong&gt;Use Reply&lt;/strong&gt;, the suggestion is moved into the normal message composer.&lt;/p&gt;

&lt;p&gt;The agent can then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Generate
   ↓
Review
   ↓
Edit
   ↓
Send
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI does not own the final action.&lt;/p&gt;

&lt;p&gt;That boundary matters more to me than simply adding another confirmation dialog.&lt;/p&gt;

&lt;p&gt;The application architecture itself separates &lt;strong&gt;generation&lt;/strong&gt; from &lt;strong&gt;execution&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This is a pattern I would reuse in many AI-assisted workflows where the output affects another person or an important business process.&lt;/p&gt;




&lt;h2&gt;
  
  
  AI Provider Logic Belongs Behind an Application Boundary
&lt;/h2&gt;

&lt;p&gt;The application currently uses an OpenAI-backed adapter for operations such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;summarization,&lt;/li&gt;
&lt;li&gt;sentiment analysis,&lt;/li&gt;
&lt;li&gt;priority detection,&lt;/li&gt;
&lt;li&gt;reply suggestions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is also provider orchestration that can use Gemini as a fallback for selected classes of provider, quota, rate-limit, network or timeout failures.&lt;/p&gt;

&lt;p&gt;I don’t want ticket or message business logic to know the details of every provider.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Business workflow
       │
       ▼
AI operation
       │
       ▼
Provider orchestration
       │
       ├── OpenAI
       │
       └── Gemini fallback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This boundary becomes useful even if the application never changes providers.&lt;/p&gt;

&lt;p&gt;It gives one place to reason about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;prompts,&lt;/li&gt;
&lt;li&gt;provider failures,&lt;/li&gt;
&lt;li&gt;normalization,&lt;/li&gt;
&lt;li&gt;fallback rules,&lt;/li&gt;
&lt;li&gt;model configuration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Provider abstraction does not mean every model is interchangeable.&lt;/p&gt;

&lt;p&gt;Different models behave differently.&lt;/p&gt;

&lt;p&gt;The goal is simply to stop provider-specific concerns from leaking through the rest of the application.&lt;/p&gt;




&lt;h2&gt;
  
  
  Authentication Is Not Authorization
&lt;/h2&gt;

&lt;p&gt;The application uses JWT authentication, but authentication alone does not answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is this user allowed to access this particular ticket?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A valid token identifies the user and role.&lt;/p&gt;

&lt;p&gt;Resource authorization still has to be enforced.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;customers should only access their own tickets,&lt;/li&gt;
&lt;li&gt;agents operate on tickets available to their role and assignment rules,&lt;/li&gt;
&lt;li&gt;admins have broader access.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The frontend also hides or exposes functionality according to role, but that is a UX decision—not a security boundary.&lt;/p&gt;

&lt;p&gt;The backend remains authoritative.&lt;/p&gt;

&lt;p&gt;This distinction is especially important in React applications because hiding a button is easy to mistake for enforcing permission.&lt;/p&gt;

&lt;p&gt;It isn’t.&lt;/p&gt;




&lt;h2&gt;
  
  
  Redis Has More Than One Responsibility
&lt;/h2&gt;

&lt;p&gt;Redis is not used only because BullMQ needs it.&lt;/p&gt;

&lt;p&gt;In this project it also supports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;short-lived ticket caching,&lt;/li&gt;
&lt;li&gt;distributed rate-limit counters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes Redis shared infrastructure serving different concerns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Redis
├── BullMQ queue infrastructure
├── ticket cache
└── rate-limit counters
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those responsibilities should still remain conceptually separate even if they use the same underlying technology.&lt;/p&gt;

&lt;p&gt;A queue is not a cache.&lt;/p&gt;

&lt;p&gt;A cache is not a source of truth.&lt;/p&gt;

&lt;p&gt;The database remains the authoritative store for ticket and conversation data.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the Architecture Gets Right
&lt;/h2&gt;

&lt;p&gt;There are several boundaries in this design that I would keep.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. User-facing writes are separated from AI enrichment
&lt;/h3&gt;

&lt;p&gt;Message creation does not wait for summary, sentiment and priority generation.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Interactive AI and background AI use different execution models
&lt;/h3&gt;

&lt;p&gt;The system chooses between synchronous and asynchronous execution based on the user workflow rather than treating all AI calls identically.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Server state remains server state
&lt;/h3&gt;

&lt;p&gt;TanStack Query manages API-backed ticket state rather than copying everything into a global client store.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. AI suggestions do not automatically become actions
&lt;/h3&gt;

&lt;p&gt;Generated replies remain editable suggestions.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. HTTP and worker lifecycles are separate
&lt;/h3&gt;

&lt;p&gt;The API server and AI worker can fail, restart and evolve independently at the process level.&lt;/p&gt;

&lt;p&gt;But none of those decisions makes the architecture finished.&lt;/p&gt;

&lt;p&gt;Several weaknesses become more important as concurrency and operational requirements increase.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Would Change Before Scaling
&lt;/h2&gt;

&lt;p&gt;This is the part of the architecture I find most useful to examine.&lt;/p&gt;

&lt;p&gt;Queues solve one category of problem while creating others.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Fix the Database-to-Queue Consistency Gap
&lt;/h3&gt;

&lt;p&gt;Look again at message creation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Write message to PostgreSQL
        ↓
Update ticket
        ↓
Publish BullMQ job
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are separate operations.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;PostgreSQL successfully saves the message.&lt;/li&gt;
&lt;li&gt;Redis becomes unavailable before &lt;code&gt;aiQueue.add()&lt;/code&gt; succeeds.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The application now has committed business data without the corresponding background event.&lt;/p&gt;

&lt;p&gt;Retries do not solve this consistency problem.&lt;/p&gt;

&lt;p&gt;A stronger design would introduce a &lt;strong&gt;transactional outbox&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             PostgreSQL transaction
        ┌────────────────────────────┐
        │ Save message               │
        │ Save outbox event          │
        └──────────────┬─────────────┘
                       │ commit
                       ▼
                Outbox publisher
                       │
                       ▼
                   BullMQ
                       │
                       ▼
                    Worker
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The message and the intent to perform AI processing would be committed atomically in PostgreSQL.&lt;/p&gt;

&lt;p&gt;A separate publisher could then reliably move pending outbox events into BullMQ.&lt;/p&gt;

&lt;p&gt;That doesn’t make the entire system exactly-once.&lt;/p&gt;

&lt;p&gt;It closes a specific and important consistency gap.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Retries Need Idempotency
&lt;/h3&gt;

&lt;p&gt;The queue currently configures multiple attempts with exponential backoff.&lt;/p&gt;

&lt;p&gt;That improves recoverability from transient failures.&lt;/p&gt;

&lt;p&gt;But:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;retryable does not automatically mean safe to retry.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Suppose a job:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;calls the AI provider,&lt;/li&gt;
&lt;li&gt;updates the ticket,&lt;/li&gt;
&lt;li&gt;writes an AI interaction,&lt;/li&gt;
&lt;li&gt;fails somewhere around those operations,&lt;/li&gt;
&lt;li&gt;retries.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Without explicit idempotency, repeated attempts can repeat side effects.&lt;/p&gt;

&lt;p&gt;A stronger implementation would give each logical processing operation a stable identity and make persisted effects safe to replay.&lt;/p&gt;

&lt;p&gt;Possible approaches include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;deterministic job IDs,&lt;/li&gt;
&lt;li&gt;processing-version records,&lt;/li&gt;
&lt;li&gt;unique database constraints,&lt;/li&gt;
&lt;li&gt;upsert semantics,&lt;/li&gt;
&lt;li&gt;explicit processed-event records.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The exact mechanism depends on the workflow.&lt;/p&gt;

&lt;p&gt;The architectural requirement is simpler:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A retry should not accidentally turn one logical operation into multiple business effects.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Protect Against Same-Ticket Concurrency
&lt;/h3&gt;

&lt;p&gt;The worker has concurrency greater than one.&lt;/p&gt;

&lt;p&gt;That’s useful for processing unrelated tickets.&lt;/p&gt;

&lt;p&gt;But it also means two jobs for the same ticket can overlap.&lt;/p&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Ticket conversation V1
        │
        └── Job A starts

New message arrives
        │
Ticket conversation V2
        │
        └── Job B starts

Job B finishes first → stores V2 analysis
Job A finishes later → stores older V1 analysis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The final ticket can now contain stale AI metadata.&lt;/p&gt;

&lt;p&gt;This is not solved by generic worker concurrency configuration alone.&lt;/p&gt;

&lt;p&gt;Before scaling the workflow, I would introduce some combination of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;per-ticket serialization,&lt;/li&gt;
&lt;li&gt;conversation version numbers,&lt;/li&gt;
&lt;li&gt;compare-before-write logic,&lt;/li&gt;
&lt;li&gt;coalescing pending analysis jobs,&lt;/li&gt;
&lt;li&gt;stale-result rejection.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important principle is that &lt;strong&gt;global concurrency and entity-level ordering are different problems&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Give the UI an Explicit AI-Result Delivery Mechanism
&lt;/h3&gt;

&lt;p&gt;Background processing creates another question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How does the browser learn that AI analysis has finished?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The current architecture does not implement a dedicated real-time result-delivery channel.&lt;/p&gt;

&lt;p&gt;At larger scale or with a richer UI, I would make that explicit.&lt;/p&gt;

&lt;p&gt;Depending on the product requirements:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Simplest
Polling

        ↓

More event-driven
Server-Sent Events

        ↓

Bidirectional realtime needs
WebSockets
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I would not automatically choose WebSockets.&lt;/p&gt;

&lt;p&gt;If the browser only needs server-to-client status updates, SSE may be enough.&lt;/p&gt;

&lt;p&gt;Architecture should follow the interaction requirement, not the popularity of the technology.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Tighten Cache Freshness Rules
&lt;/h3&gt;

&lt;p&gt;Caching ticket data can reduce repeated reads, but mutable support conversations create invalidation challenges.&lt;/p&gt;

&lt;p&gt;Every path that changes data relevant to a cached ticket needs a clear freshness strategy.&lt;/p&gt;

&lt;p&gt;Before relying more heavily on caching, I would explicitly document:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what is cached,&lt;/li&gt;
&lt;li&gt;the cache key,&lt;/li&gt;
&lt;li&gt;TTL,&lt;/li&gt;
&lt;li&gt;which writes invalidate it,&lt;/li&gt;
&lt;li&gt;whether stale reads are acceptable,&lt;/li&gt;
&lt;li&gt;what happens when invalidation fails.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A short TTL is useful, but TTL alone isn’t a complete consistency model.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Improve the AI Context and Evaluation Layer
&lt;/h3&gt;

&lt;p&gt;The current AI pipeline is intentionally straightforward.&lt;/p&gt;

&lt;p&gt;There are several improvements I would investigate before treating AI behavior as a mature subsystem:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;explicit structured outputs,&lt;/li&gt;
&lt;li&gt;conversation-size/token management,&lt;/li&gt;
&lt;li&gt;better context construction,&lt;/li&gt;
&lt;li&gt;prompt/model version tracking,&lt;/li&gt;
&lt;li&gt;evaluation datasets,&lt;/li&gt;
&lt;li&gt;regression testing for AI behavior,&lt;/li&gt;
&lt;li&gt;provider health telemetry,&lt;/li&gt;
&lt;li&gt;clearer usage accounting,&lt;/li&gt;
&lt;li&gt;privacy and retention policies for model inputs and outputs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I would add these based on actual product requirements rather than building an elaborate AI platform prematurely.&lt;/p&gt;

&lt;p&gt;The important shift is from:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does the model return something?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can I measure whether this AI behavior remains useful and reliable as the application changes?&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What I Would Not Add Yet
&lt;/h2&gt;

&lt;p&gt;Scaling discussions can easily turn into architecture shopping lists.&lt;/p&gt;

&lt;p&gt;I would &lt;strong&gt;not&lt;/strong&gt; automatically add:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Kafka,&lt;/li&gt;
&lt;li&gt;Kubernetes,&lt;/li&gt;
&lt;li&gt;multiple microservices,&lt;/li&gt;
&lt;li&gt;vector databases,&lt;/li&gt;
&lt;li&gt;WebSockets,&lt;/li&gt;
&lt;li&gt;event sourcing,&lt;/li&gt;
&lt;li&gt;multiple caches.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of those technologies is inherently an upgrade.&lt;/p&gt;

&lt;p&gt;For this system, I would first strengthen the guarantees around the architecture that already exists:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB/queue consistency
        ↓
Idempotent processing
        ↓
Per-ticket ordering/versioning
        ↓
Observable job state
        ↓
AI evaluation
        ↓
Scale infrastructure when evidence requires it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Complexity should be purchased with a requirement.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Most Important Lesson: AI Is a Workload, Not the Architecture
&lt;/h2&gt;

&lt;p&gt;It is tempting to describe this application as an OpenAI integration.&lt;/p&gt;

&lt;p&gt;That misses most of the engineering.&lt;/p&gt;

&lt;p&gt;The LLM is one dependency inside a larger system.&lt;/p&gt;

&lt;p&gt;The architecture still has to answer ordinary software-engineering questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who owns state?&lt;/li&gt;
&lt;li&gt;Which operations block the user?&lt;/li&gt;
&lt;li&gt;What should happen asynchronously?&lt;/li&gt;
&lt;li&gt;Where is authorization enforced?&lt;/li&gt;
&lt;li&gt;What happens after partial failure?&lt;/li&gt;
&lt;li&gt;Are retries safe?&lt;/li&gt;
&lt;li&gt;Can concurrent work overwrite newer state?&lt;/li&gt;
&lt;li&gt;How does the UI learn that background work finished?&lt;/li&gt;
&lt;li&gt;Where does human approval belong?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those questions existed before LLMs.&lt;/p&gt;

&lt;p&gt;AI simply makes several of them more visible because model calls are external, comparatively slow, probabilistic and failure-prone.&lt;/p&gt;




&lt;h2&gt;
  
  
  Final Architecture
&lt;/h2&gt;

&lt;p&gt;The current design can be summarized as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                           React + TypeScript
                                  │
                 ┌────────────────┼────────────────┐
                 │                │                │
            Auth Context     TanStack Query    Local UI State
                 │                │                │
                 └────────────────┼────────────────┘
                                  │
                                  ▼
                           Express REST API
                                  │
                    Auth / AuthZ / Validation
                                  │
                                  ▼
                       Application Services
                         │               │
                         │               │
                         ▼               ▼
                    PostgreSQL      Redis / BullMQ
                                         │
                                         ▼
                                     AI Worker
                                         │
                             ┌───────────┼───────────┐
                             │           │           │
                          Summary    Sentiment    Priority
                             │           │           │
                             └───────────┼───────────┘
                                         │
                                    AI Providers

Interactive reply path:

Agent → API → AI reply suggestion → editable composer → human sends
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is intentionally simpler than a large distributed system.&lt;/p&gt;

&lt;p&gt;And it has known limitations.&lt;/p&gt;

&lt;p&gt;I consider both of those facts healthy.&lt;/p&gt;

&lt;p&gt;Good architecture is not about pretending the current system can handle every future requirement.&lt;/p&gt;

&lt;p&gt;It is about knowing &lt;strong&gt;which guarantees the system currently provides, which ones it doesn’t, and where the next architectural pressure points will appear.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Building this project changed how I think about adding AI to existing applications.&lt;/p&gt;

&lt;p&gt;I wouldn’t start by asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Where can I call the LLM?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I would start with:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What role should AI have in this workflow?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then decide:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;synchronous or asynchronous,&lt;/li&gt;
&lt;li&gt;advisory or autonomous,&lt;/li&gt;
&lt;li&gt;user-facing or background,&lt;/li&gt;
&lt;li&gt;retryable or idempotent,&lt;/li&gt;
&lt;li&gt;cached or authoritative,&lt;/li&gt;
&lt;li&gt;human-reviewed or automatically executed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once those boundaries are clear, choosing the API is usually the easier part.&lt;/p&gt;

&lt;p&gt;In the next article in this architecture series, I’ll go deeper into &lt;strong&gt;BullMQ background jobs for AI workloads—especially retries, idempotency and ordering&lt;/strong&gt;, because adding a queue is only the beginning of making background processing reliable.&lt;/p&gt;




&lt;h2&gt;
  
  
  About the author
&lt;/h2&gt;

&lt;p&gt;I’m Ram Ji Tripathi, a senior frontend-focused full-stack engineer working across React, TypeScript, Node.js and AI-enabled applications.&lt;/p&gt;

&lt;p&gt;I write about frontend architecture, backend systems, AI engineering, system design and the engineering decisions behind building real software.&lt;/p&gt;

&lt;p&gt;More engineering notes and projects are available on &lt;a href="https://ramjwork.in" rel="noopener noreferrer"&gt;my portfolio&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Originally published on my engineering blog:&lt;br&gt;
&lt;a href="https://blog.ramjwork.in/ai-engineering/ai-customer-support-architecture" rel="noopener noreferrer"&gt;https://blog.ramjwork.in/ai-engineering/ai-customer-support-architecture&lt;/a&gt;&lt;/p&gt;

</description>
      <category>react</category>
      <category>node</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
