DEV Community

Praveen Raj Thulasi S
Praveen Raj Thulasi S

Posted on

I Built an Agentic Analytics Platform — Here's What I Learned

I Built an Agentic Analytics Platform — Here's What I Learned

What if you could ask your analytics dashboard:

"Why did sales decrease last month?"

and instead of manually filtering charts and writing database queries, an AI system could investigate the data, generate a query, validate it, analyze the results, and explain what it found?

That's what I wanted to explore with AgentVerse, an agentic analytics platform I built using React, Node.js, MongoDB, Ollama, and MCP.

This project started as an experiment with multi-agent AI. Along the way, I learned that building an agentic application is much less about simply connecting an LLM to a database—and much more about controlling, validating, and observing what the AI does.


🚀 What is AgentVerse?

AgentVerse is a natural-language analytics platform.

Instead of manually writing MongoDB queries, users can ask questions such as:

Show monthly sales for the last 12 months.
Enter fullscreen mode Exit fullscreen mode

or:

Show the top 5 products by revenue.
Enter fullscreen mode Exit fullscreen mode

or:

Why did sales decrease last month?
Enter fullscreen mode Exit fullscreen mode

The platform translates these questions into an analytics workflow.

At a high level:


User
  ↓
Orchestrator Agent
  ↓
Query Generation
  ↓
Query Guardian
  ↓
MCP Server
  ↓
MongoDB
  ↓
Evidence Analysis
  ↓
Insight Agent
  ↓
Visualization
Enter fullscreen mode Exit fullscreen mode

The goal isn't just to generate an answer.

The goal is to make the entire analytical process controllable and observable.


🧠 Why Multi-Agent?

One approach would be:

User → LLM → Database → Answer
Enter fullscreen mode Exit fullscreen mode

But that gives one model too many responsibilities.

It has to understand the question, understand the schema, generate a query, execute it, analyze the result, and produce a visualization.

Instead, I separated the workflow into specialized components.

                    User
                      │
                      ▼
               Orchestrator
                      │
          ┌───────────┼───────────┐
          ▼           ▼           ▼
       Intent       Schema     Session
       Planning    Discovery    State
          │
          ▼
    Query Generation
          │
          ▼
    Query Guardian
          │
          ▼
       MCP Server
          │
          ▼
       MongoDB
          │
          ▼
    Evidence Engine
          │
          ▼
     Insight Agent
          │
          ▼
    Visualization
Enter fullscreen mode Exit fullscreen mode

Each component has a specific responsibility.

This makes the system easier to debug and gives me more control over what each part of the AI workflow is allowed to do.


🔌 MCP as the Data Boundary

One of the most interesting parts of the project was using Model Context Protocol (MCP) as a boundary between the AI agents and the database.

Instead of allowing the agent to directly access MongoDB:

Agent
  ↓
MongoDB
Enter fullscreen mode Exit fullscreen mode

the architecture becomes:

Agent
  ↓
MCP Server
  ↓
MongoDB
Enter fullscreen mode Exit fullscreen mode

The MCP layer exposes controlled tools such as:

get_schema
execute_query
Enter fullscreen mode Exit fullscreen mode

This means the AI can request the tools it needs without having unrestricted access to the underlying database.

For me, MCP became more than just a way to connect an LLM to tools.

It became an architectural boundary between:

AI reasoning → Data access


🛡️ The Problem I Didn't Expect: AI-Generated Queries

Getting an LLM to generate a MongoDB aggregation pipeline isn't particularly difficult.

The difficult part is deciding:

Should I trust the generated query?

The answer is no.

LLMs can generate invalid queries, use incorrect fields, or potentially generate operations that shouldn't be allowed in an analytics application.

That's why I built a Query Guardian.

The flow is:

LLM generates query
        ↓
   Query Guardian
        ↓
     Validation
        ↓
   ┌────┴────┐
   │         │
 Valid     Invalid
   │         │
   ▼         ▼
Execute    Repair
             │
             ▼
        Validate Again
Enter fullscreen mode Exit fullscreen mode

The Guardian checks things such as:

  • Allowed aggregation stages
  • Allowed operators
  • Collection names
  • Field names
  • Pipeline length
  • Result limits
  • Forbidden operations
  • Query structure

The analytics system is designed around read-only operations rather than allowing the AI to modify the database.

This was one of the biggest lessons from the project:

LLM output should be treated as untrusted input, not executable truth.


📊 A Real Example

Let's say a user asks:

"Why did sales decrease last month?"

The system doesn't simply send that sentence to an LLM and return whatever it says.

Instead, the request goes through several stages.

1. Understand the request

The Orchestrator identifies the request as a root-cause analysis task.

2. Discover the schema

The system retrieves the available database structure through MCP.

3. Generate a query

The Analytics Query Agent creates a MongoDB aggregation pipeline.

4. Validate it

Query Guardian checks the generated pipeline.

5. Execute it

The validated query is sent through the MCP server to MongoDB.

6. Analyze the evidence

The Evidence Engine calculates measurable changes in the returned data.

For example:

Overall Revenue:    -18.2%

South Region:       -31.4%
Electronics:        -24.7%
Product A:          -28.1%
Enter fullscreen mode Exit fullscreen mode

7. Generate the insight

The Insight Agent receives the structured evidence and creates the explanation.

Rather than blindly claiming:

"The South region caused the decline."

the system can use more careful language such as:

"The South region recorded the largest observed regional decline and may represent a contributing factor."

This distinction matters.

A correlation in the data isn't automatically proof of causation.


📈 Dynamic Visualizations

The result isn't just text.

AgentVerse can determine an appropriate visualization based on the analytical result.

For example:

Trend over time
       ↓
   Line Chart

Category comparison
       ↓
    Bar Chart

Distribution
       ↓
    Pie Chart

Single metric
       ↓
    KPI Card
Enter fullscreen mode Exit fullscreen mode

The React frontend uses Recharts to render these visualizations.

The workflow becomes:

Question
   ↓
Intent
   ↓
Query
   ↓
Data
   ↓
Evidence
   ↓
Insight
   ↓
Visualization
Enter fullscreen mode Exit fullscreen mode

🔍 Making the AI Workflow Observable

One thing I didn't want was a black box that simply says:

"Here's your answer."

AgentVerse includes an execution trace showing the different stages of the workflow.

For example:

✓ Orchestrator Agent
  Intent Classification

✓ MCP Schema Discovery
  Database Schema

✓ Analytics Query Agent
  MQL Generation

✓ Query Guardian
  Security Validation

✓ MCP Execution
  MongoDB Query

✓ Insight Agent
  Evidence Synthesis

✓ Visualization
  Chart Selection
Enter fullscreen mode Exit fullscreen mode

This makes the system easier to understand and debug.

If something goes wrong, I can investigate the intermediate steps instead of only looking at the final response.


📝 Audit Logging

Agentic applications can be difficult to debug because there are multiple intermediate operations.

So I also added an audit layer that can track information such as:

Request ID
User Question
Generated Pipeline
Validation Status
Rows Returned
Execution Time
Timestamp
Enter fullscreen mode Exit fullscreen mode

This gives me a history of what the system actually did.

For example, if an insight looks incorrect, I can trace:

User Question
      ↓
Generated Query
      ↓
Guardian Validation
      ↓
Database Result
      ↓
Evidence
      ↓
Final Insight
Enter fullscreen mode Exit fullscreen mode

That is much more useful than debugging only the final LLM response.


🧰 Tech Stack

Frontend

  • React
  • TypeScript
  • Vite
  • Tailwind CSS
  • Recharts
  • Axios

Backend

  • Node.js
  • Express
  • TypeScript
  • MongoDB
  • Mongoose

AI / Agent Layer

  • Ollama
  • LLM-based agents
  • MCP
  • Multi-agent orchestration
  • Structured JSON responses

💡 What I Learned

1. LLMs should not be trusted blindly

The model can generate something that looks valid but isn't.

Validation needs to happen outside the model.

2. Not everything needs AI

Things like query limits, security checks, percentage calculations, and schema validation are better handled deterministically.

Use AI where reasoning is useful.

Use traditional code where deterministic correctness matters.

3. More agents don't automatically mean a better system

Adding agents increases complexity.

The reason for separating components should be clear responsibilities—not simply having "more AI."

4. Observability matters

With a traditional API, debugging can be relatively straightforward.

With an agentic system, there may be several intermediate decisions.

Execution traces and audit logs become extremely valuable.

5. AI should not become a single point of failure

If the LLM is unavailable, the application shouldn't necessarily become completely unusable.

Fallback strategies and deterministic logic can make the system more resilient.


🚧 What's Next?

AgentVerse is still evolving.

Some areas I want to explore next are:

  • More MCP tools
  • Multiple data sources
  • Better agent evaluation
  • Streaming agent responses
  • Improved query validation
  • More advanced anomaly detection
  • Long-term agent memory
  • Cloud deployment
  • Production-grade observability
  • Automated evaluation of analytical accuracy

The biggest question I want to explore is:

How do we reliably evaluate an agentic analytics system?

A fluent AI response isn't necessarily a correct one.

That's where I think a lot of interesting engineering work remains.


🎯 Final Thoughts

Building AgentVerse changed how I think about AI applications.

Earlier, my main question was:

"Which LLM should I use?"

Now I think more about:

"What should the LLM be allowed to do?"

That shift completely changed how I approached the architecture.

The LLM is only one component.

The interesting engineering happens around it:

LLM
 ↓
Tools
 ↓
Validation
 ↓
Security
 ↓
Evidence
 ↓
Observability
 ↓
Reliable Application
Enter fullscreen mode Exit fullscreen mode

That's what I've learned while building AgentVerse.

And I'm still learning.


What's your experience with Agentic AI?

If you're building an agentic application, what has been the hardest part for you?

Orchestration? Tool calling? Security? Reliability? Evaluation?

I'd love to hear your experience.


🛠️ Built With

React TypeScript Node.js Express MongoDB Ollama MCP Multi-Agent AI Tailwind CSS Recharts

AI #AgenticAI #MCP #LLM #MongoDB #React #NodeJS #TypeScript #WebDevelopment

Top comments (3)

Collapse
 
praveen007 profile image
Praveen Raj Thulasi S •

Share your thoughts, will learn from your valuable comments.

Collapse
 
gopika_9725 profile image
Gopika •

Great work!🚀Really interesting to see your learning journey in Agentic AI and MCP. Keep exploring!👋

Collapse
 
praveen007 profile image
Praveen Raj Thulasi S •

Thanks Gopika.