DEV Community

Cover image for Building IntelliDesk AI: How I Architected a Production-Grade Enterprise ITSM Platform with RAG, WebSockets, and Celery
Pruthviraj Janwade
Pruthviraj Janwade

Posted on Originally published at pruthvirajjanwade.hashnode.dev

Building IntelliDesk AI: How I Architected a Production-Grade Enterprise ITSM Platform with RAG, WebSockets, and Celery

By Pruthviraj Janwade


If you have ever worked in an enterprise environment, you know the dread of filing an IT support ticket.

You navigate through a labyrinthine portal, fill out a 12-field form with dropdowns you do not understand, and wait 24 to 48 hours just to receive an email asking: "Have you tried restarting your machine?"

Tools like ServiceNow, Jira Service Management, and Zendesk are enterprise powerhouses, but they were built in an era of manual triage and static forms. While large language models (LLMs) have taken consumer tech by storm, enterprise IT service management (ITSM) has remained largely trapped in old paradigms.

Over the past few months, I set out to bridge this gap by building IntelliDesk AI — a production-grade, AI-powered enterprise ITSM platform featuring a conversational RAG assistant, automated ticket lifecycle management, real-time analytics, and role-based access control.

In this deep dive, I want to share the architectural decisions, design patterns, engineering challenges, and lessons learned while building this platform from scratch.


1. The Core Philosophy: Conversational-First ITSM

The primary goal of IntelliDesk AI was simple: eliminate the friction of IT support for both employees and agents.

Instead of forcing users to fill out static forms, the primary entry point is IntelliBot — an intelligent conversational assistant that:

  1. Understands Natural Language: An employee simply types, "My Wi-Fi keeps disconnecting every 10 minutes on the 3rd floor."
  2. Performs Semantic Document Search (RAG): Retrieves relevant company knowledge base guides with exact source attribution.
  3. Resolves Autonomously: Walks the employee through step-by-step diagnostic and troubleshooting steps.
  4. Auto-Creates & Escalates Tickets: If self-service fails, IntelliBot extracts the key details (category, urgency, root symptoms), provisions a ticket in the database, and assigns it to the on-duty IT team — without the user filling out a single form field.

2. High-Level System Architecture

To ensure modularity, scalability, and maintainability, the system is designed around Clean Architecture principles:

┌────────────────────────────────────────────────────────┐
│                   Browser (React SPA)                  │
└───────────────────────────┬────────────────────────────┘
                            │ (HTTP / WSS)
                            ▼
┌────────────────────────────────────────────────────────┐
│           NGINX (Reverse Proxy & Rate Limiter)         │
└──────────────┬──────────────────────────┬──────────────┘
               │ (Proxy API)              │ (Proxy WebSocket)
               ▼                          ▼
┌────────────────────────────────────────────────────────┐
│         Flask 3 REST API + Socket.IO (Eventlet)        │
│       Controller ──> Service ──> Repository ──> Model  │
└───────┬──────────────┬──────────────┬──────────────┬───┘
        │              │              │              │
        ▼              ▼              ▼              ▼
┌──────────────┐┌──────────────┐┌──────────────┐┌──────────────┐
│  PostgreSQL  ││ Redis Cache  ││   ChromaDB   ││   Groq API   │
│ (Primary DB) ││  & Broker    ││(Vector Store)││(Llama 3.3 70B│
└──────────────┘└──────┬───────┘└──────┬───────┘└──────────────┘
                       │               │
                       ▼               │
               ┌───────────────┐       │
               │ Celery Worker │◄──────┘
               │ (Async Tasks) │ (Local Sentence Transformers)
               └───────────────┘
Enter fullscreen mode Exit fullscreen mode

Architectural Highlights

  • Strict Layer Separation:
    • Controllers: Pure HTTP routing, input validation (Marshmallow), and status codes.
    • Services: Business logic, domain rules, and workflow orchestration.
    • Repositories: Database queries abstracted via SQLAlchemy 2.0 ORM.
    • Models: Data definitions and relationship mappings.
  • Microservices Orchestration: Fully containerized using Docker and orchestrated with Docker Compose (Postgres, Redis, ChromaDB, Flask, Celery Worker, Celery Beat, Flower, NGINX).

3. Designing the Retrieval-Augmented Generation (RAG) Pipeline

A major pitfall of many LLM projects is high hallucination rates and exorbitant API costs. Here is how I addressed both:

Step 1: Document Processing & Chunking

When an IT administrator uploads internal documentation (PDFs, DOCX, TXT):

  1. Text is extracted cleanly using PyPDF2 or python-docx.
  2. A recursive character text splitter splits documents into chunks of 800 characters with an overlap of 150 characters. This overlap preserves semantic context across chunk boundaries.

Step 2: Zero-Cost Local Embeddings

Instead of calling paid embedding APIs (e.g., OpenAI text-embedding-ada-002 or text-embedding-3-small), I integrated sentence-transformers/all-MiniLM-L6-v2 directly into the container.

  • It runs locally on CPU with inference times under 15ms per chunk.
  • Output dimensionality: 384 vectors.
  • Result: $0 embedding cost and zero network overhead.

Step 3: Vector Indexing & Semantic Search

Embeddings are indexed in ChromaDB. When a query comes in:

  1. The user's prompt is embedded using the same MiniLM model.
  2. ChromaDB runs cosine similarity to fetch the Top-K most relevant chunks.
  3. A confidence score threshold filters out low-relevance matches to prevent hallucinations.

Step 4: LLM Generation with Strategy Pattern

To avoid vendor lock-in, I implemented an AI provider abstraction using the Strategy Pattern:

class LLMProvider(ABC):
    @abstractmethod
    def generate_stream(self, prompt: str, system_message: str):
        pass

class GroqProvider(LLMProvider):
    def generate_stream(self, prompt: str, system_message: str):
        # Ultra-fast inference using Groq Llama 3.3 70B
        ...
Enter fullscreen mode Exit fullscreen mode

Groq’s LPU (Language Processing Unit) delivers inference speeds of ~250-300 tokens/sec, making real-time streaming feel instantaneous.


4. Real-Time Streaming: Replacing Polling with WebSockets

A common issue with AI chat interfaces is the delay while waiting for the LLM to complete its full response.

Initially, simple REST polling or Server-Sent Events (SSE) were considered, but because IntelliDesk already needed bidirectional communication for live dashboard updates, I chose Flask-SocketIO with an Eventlet worker.

The Streaming Protocol

  1. Client emits ai:chat with user query and session ID.
  2. Server validates authentication token and emits ai:stream:start.
  3. As Groq yields text tokens, server immediately emits ai:stream:chunk payloads.
  4. Upon completion, server emits ai:stream:done along with formatted source citations and confidence metrics.
Client (React)                             Server (Flask + SocketIO)
      │                                                │
      │ ─── emit('ai:chat', { prompt, sessionId }) ──> │
      │                                                │ ──> Vector Search (ChromaDB)
      │                                                │ ──> Stream from Groq (Llama 3.3)
      │ <── emit('ai:stream:start') ────────────────── │
      │ <── emit('ai:stream:chunk', { token: '1.' }) ──│
      │ <── emit('ai:stream:chunk', { token: ' Turn' })│
      │ <── emit('ai:stream:done', { citations }) ──── │
Enter fullscreen mode Exit fullscreen mode

This reduced perceived latency from 4–6 seconds down to under 200ms.


5. Background Jobs & Asynchronous Workflows (Celery + Redis)

Heavy operations should never block an HTTP request. I set up Celery 5 with Redis 7 using dedicated priority queues:

Queue Tasks Handled
documents PDF parsing, chunking, vector embedding generation
ai Background ticket intent classification & summarization
email SMTP notifications for ticket status and SLA warnings
reports Periodic CSAT, SLA metrics, and analytics compilation

By pairing Celery with Celery Beat and RedBeat, scheduled tasks run continuously in the background (e.g., checking for SLA breach thresholds every 60 seconds). For monitoring, Flower provides a visual dashboard of task throughput and worker health.


6. Frontend Engineering with React 18 & TypeScript

The frontend was built to feel like modern software from Linear or Vercel:

  • State Strategy: Clear separation between server cache and client state:
    • TanStack Query (React Query) handles server synchronization, automatic caching, and background invalidation for tickets and analytics.
    • Redux Toolkit manages local state (active AI chat session, dark/light theme, UI modals).
  • Socket Lifecycle Management: Custom React hooks manage WebSocket connection lifecycles, graceful reconnection, and event buffering to ensure messages aren't lost during page navigation.

7. Zero-Dollar Infrastructure: Running Production for $0/Month

One of my proudest milestones with this project was achieving enterprise-grade capability on a $0/month infrastructure footprint:

  • Compute API: Render Free Tier
  • Frontend SPA: Vercel Global Edge Network
  • Primary Database: Neon Serverless PostgreSQL
  • Embeddings: Local CPU execution (all-MiniLM-L6-v2)
  • LLM Inference: Groq Free Developer Tier (Llama 3.3 70B)
  • Vector DB: Embedded ChromaDB instance

8. What's Next: Enterprise Kubernetes Deployment on AWS EKS

With the application architecture fully validated, I am now moving to the next engineering milestone: deploying IntelliDesk AI onto an enterprise-grade AWS EKS (Elastic Kubernetes Service) cluster.

The AWS Deployment Blueprint:

  1. Infrastructure as Code (IaC): Provisioning AWS VPC, subnets, IAM Roles for Service Accounts (IRSA), and managed node groups using Terraform.
  2. Kubernetes Packaging: Writing modular Helm Charts for the Flask backend, Celery workers, and NGINX Ingress Controller.
  3. Managed Services Integration:
    • Amazon RDS (PostgreSQL Multi-AZ)
    • Amazon ElastiCache (Redis)
    • Amazon S3 for secure document storage
  4. GitOps Continuous Delivery: Automated cluster reconciliation and deployments using ArgoCD.
  5. Observability: Prometheus for cluster metrics and Grafana for executive operational dashboards.

9. Conclusion & Takeaways

Building IntelliDesk AI taught me several core engineering lessons:

  • Design for abstraction early: Isolating the LLM provider behind a clean interface saved hours when switching between models.
  • RAG is only as good as chunking: Tuning chunk overlap and semantic boundaries matters far more than simply picking a larger LLM.
  • User experience is latency-bound: Real-time streaming via WebSockets fundamentally transforms conversational AI from a sluggish utility into a delightful product.

Explore the Code

The entire project is open-source under the MIT license:

If you found this breakdown valuable, feel free to star the repo or connect with me on LinkedIn!

Top comments (1)

Collapse
 
kyisaiah47 profile image
kyisaiah47 •

When an admin replaces a document, does the old embedding get removed from ChromaDB before retrieval can see it?