DEV Community

Kokku Sri Madhavi
Kokku Sri Madhavi

Posted on

DataCenter Guardian AI.

๐Ÿ›ก๏ธ DataCenter Guardian AI: Building an AI Incident Response Agent That Learns from Past Incidents

An AI-powered data center intelligence platform that monitors infrastructure, investigates incidents, recalls historical operational experience, and helps engineers respond more effectively.

Modern data centers generate huge amounts of infrastructure data every second. CPU usage, RAM, temperature, network traffic, GPU utilization, power consumption, cooling systems, and other resources constantly change.

When something goes wrong, traditional monitoring systems can tell us:

"There is an incident."

But an infrastructure engineer needs something more useful:

"Have we seen a similar incident before, what caused it, and what worked last time?"

That is the problem we wanted to address with DataCenter Guardian AI.

Our project was developed for HackWithHyderabad 3.0 under the theme:

๐Ÿค– AI Agents That Learn Using Hindsight
๐Ÿšจ The Problem

Data center incidents are often repetitive.

Some common examples include:

CPU saturation
Memory pressure
Cooling anomalies
Database connection failures
Network problems
Power-related issues
Runaway processes

A traditional monitoring system may detect high CPU usage and generate an alert such as:

CPU Usage: 96%
Status: Critical
Action: Investigate Server

But the alert does not necessarily tell the engineer:

Have we experienced this before?
What was the root cause?
What action solved the previous incident?
Did the previous solution actually work?
Can that experience help with the current incident?

This creates a gap between monitoring and operational intelligence.

๐Ÿ’ก Our Solution

We built DataCenter Guardian AI, an AI-powered infrastructure monitoring and incident-response prototype.

The platform brings together:

Infrastructure monitoring
Incident management
Historical incident memory
AI-assisted investigation
Risk assessment
Server diagnostics
What-if simulation
Azure integration readiness

The core workflow is:

DETECT
โ†“
INVESTIGATE
โ†“
RECALL
โ†“
RECOMMEND
โ†“
RESOLVE
โ†“
RETAIN
โ†“
USE EXPERIENCE IN FUTURE INVESTIGATIONS

Instead of treating every incident as a completely new problem, the system can use previously retained operational experience as context.

๐Ÿง  Why Hindsight Memory Matters

The central idea of our project is persistent operational memory.

When an incident is resolved, important information can be retained:

Incident
Affected server
Category
Action taken
Outcome
Lesson learned

When a similar incident appears again, the system can search its stored operational memories and identify relevant previous cases.

This changes the investigation process from:

Current Incident
โ†“
Generic Recommendation

to:

Current Incident
โ†“
Recall Similar Historical Incidents
โ†“
Compare Previous Experience
โ†“
Generate Context-Aware Recommendation

The goal is to make incident investigation more useful as operational experience accumulates.

๐Ÿ—๏ธ How DataCenter Guardian AI Works

  1. Live Infrastructure Monitoring

The platform provides a monitoring interface for simulated data center infrastructure.

It tracks metrics such as:

CPU
RAM
Temperature
Network
GPU
Power
Cooling
Water usage

The monitoring dashboard gives operators a quick view of infrastructure health.

For example:

Server: DC-SRV-024

CPU: 98.5%
RAM: 96%
GPU: 94.8%
Temperature: 84.9ยฐC
Cooling: 54.6%
Power: 766.9 W

Status: CRITICAL

When abnormal conditions are detected, the engineer can move from monitoring to incident investigation.

๐Ÿšจ 2. Incident Center

The Incident Center acts as the operational workspace for infrastructure incidents.

It provides information such as:

Server ID
Severity
Incident type
Root cause
Recommended action
Status
Resolution information

Incidents can be viewed based on severity and status.

The engineer can then investigate an incident and use the available historical operational experience.

๐Ÿค– 3. AI Guardian

The AI Guardian provides an interface for investigating infrastructure incidents.

An engineer can ask questions such as:

Have we seen a similar CPU incident before?

or:

What happened the last time we had a database connection timeout?

or:

What action should we take for this cooling anomaly?

The investigation engine evaluates the current query, identifies the relevant server when available, checks telemetry information, searches historical Guardian Memory records, and generates an investigation result.

The response can include:

Incident summary
Telemetry evidence
Historical similar cases
Risk assessment
Recommended action
Confidence information

This allows the agent to provide more context than a simple monitoring alert.

๐Ÿง  4. Guardian Memory

One of the most important parts of DataCenter Guardian AI is Guardian Memory.

Guardian Memory stores previous operational incidents together with their actions, outcomes, and lessons.

For example:

Memory โ€” Cooling Anomaly

Problem: Cooling anomaly

Server: DC-SRV-024

Action: Moved workload and increased cooling capacity

Outcome: RESOLVED

Lesson: Workload redistribution prevented thermal shutdown.

Memory โ€” Runaway Process

Problem: Runaway process detected

Server: DC-SRV-009

Action: Restarted affected service

Outcome: RESOLVED

Lesson: Restarting the affected service resolved the recurring CPU spike.

Memory โ€” Database Connection Timeout

Problem: Database connection timeout

Server: DC-SRV-017

Action: Increased connection pool and restarted database service

Outcome: RESOLVED

Lesson: Restarting alone was insufficient; connection-pool tuning resolved the incident.

๐Ÿ” 5. How Memory Improves Investigation

Consider a new CPU saturation incident.

Without historical experience, an agent might provide a generic recommendation:

CPU utilization is high. Investigate the running processes and consider restarting the server.

With historical operational memory, the investigation can identify a previous CPU-related incident involving a runaway background process.

The recommendation can then become more specific:

A similar CPU saturation incident was previously associated with a runaway background process. Restarting the affected service resolved the previous incident without requiring a complete server restart.

The important difference is context from previous operational experience.

๐Ÿ”„ 6. The Learning Loop

The project is designed around a continuous operational learning loop:

Incident Detected
โ†“
AI Investigation
โ†“
Historical Memory Recall
โ†“
Compare With Current Incident
โ†“
Generate Recommendation
โ†“
Engineer Takes Action
โ†“
Incident Resolved
โ†“
Store Outcome
โ†“
Future Investigations Can Use This Experience

The prototype demonstrates how retaining incident outcomes can help provide context during future investigations.

๐Ÿ–ฅ๏ธ 7. Data Center Overview

The Overview Dashboard provides a high-level view of the infrastructure.

It includes:

Infrastructure health
Active incidents
Predicted failure information
Energy efficiency
CPU utilization
RAM utilization
GPU utilization
Temperature
Network throughput
Power consumption
Cooling
Water usage
Infrastructure activity charts

This gives an operator a quick understanding of the overall environment before investigating individual servers.

๐Ÿ“ก 8. Live Monitoring

The Live Monitoring page provides infrastructure telemetry in a more detailed operational view.

It displays metrics such as:

CPU utilization
RAM utilization
Temperature
Network throughput
GPU utilization
Power consumption
Cooling
Water usage

The dashboard also provides infrastructure activity and temperature charts.

This helps an operator identify changing conditions across the simulated infrastructure.

๐Ÿ–ฅ๏ธ 9. Server Details

The Server Details module allows engineers to inspect an individual server.

For example, a server diagnostic page can display:

CPU utilization
RAM memory
GPU load
Core temperature
Network bandwidth
Power draw
Cooling loop flow
Uptime
Risk score
Previous incident information

The AI Guardian can then evaluate the server condition and provide:

Current condition
Predicted issue
Possible root cause
Recommended action
Confidence information

This allows the investigation to move from the overall data center to a specific infrastructure node.

๐Ÿงช 10. What-If Simulator

The What-If Simulator allows engineers to explore possible infrastructure changes before applying them.

Parameters include:

GPU workload
CPU workload
Ambient temperature
Cooling capacity
Server rack count
Network traffic

For example, an engineer can increase GPU workload and observe how the simulated infrastructure responds.

The simulator can show changes in:

Power consumption
Temperature
Cooling requirements
GPU load
Failure risk

This provides a way to explore infrastructure scenarios rather than only observing the current state.

โ˜๏ธ 11. Azure Integration Layer

The project also includes an Azure integration abstraction layer.

The architecture is prepared for integration with services such as:

Azure Monitor
Azure IoT Hub / Digital Twins
Azure SQL
Azure Machine Learning
Azure OpenAI
Power BI

In the current prototype, these integrations are represented through an abstraction/readiness layer rather than claiming that the application is connected to live production Azure infrastructure.

The purpose of this architecture is to make it easier to connect the prototype to real infrastructure telemetry and cloud services in future versions.

๐Ÿ—๏ธ System Architecture

The application follows a frontend-backend architecture:

             DATA CENTER GUARDIAN AI
                      โ”‚
                      โ–ผ
            React + TypeScript UI
                      โ”‚
                      โ–ผ
                   FastAPI
                      โ”‚
      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
      โ–ผ               โ–ผ                โ–ผ
 Monitoring       Incidents       AI Guardian
      โ”‚               โ”‚                โ”‚
      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                      โ–ผ
               Guardian Memory
                      โ”‚
                      โ–ผ
                   SQLite
                      โ”‚
                      โ–ผ
            Historical Experience
Enter fullscreen mode Exit fullscreen mode

The frontend provides the operational interface.

The backend handles:

Monitoring
Server information
Incident management
AI investigation
Guardian Memory
Simulation
API operations

SQLite is used for storing the prototype's operational data and historical memory.

โš™๏ธ Technology Stack
Frontend
React 19
TypeScript
Vite
Tailwind CSS
Recharts
Lucide Icons
Backend
Python
FastAPI
SQLAlchemy
Uvicorn
Database
SQLite
AI and Operational Intelligence
AI-assisted incident investigation
Historical incident memory
Rule and keyword-based memory matching
Risk assessment
Context-aware recommendations
What-if infrastructure simulation
๐Ÿ”— API Layer

The application separates its major capabilities through API endpoints.

Examples include:

GET /api/health
GET /api/dashboard
GET /api/monitoring
GET /api/servers
GET /api/servers/{id}
GET /api/incidents
GET /api/incidents/{id}
POST /api/incidents/{id}/resolve
GET /api/memory
GET /api/memory/{id}
POST /api/ai/investigate
POST /api/simulation
GET /api/search
GET /api/azure/status

This separation makes the prototype easier to extend and connect with future infrastructure services.

๐Ÿ” Example Incident Investigation

Imagine that the monitoring system detects:

Server: DC-SRV-009

CPU: 96%
RAM: 82%
Temperature: 77ยฐC

Status: HIGH

An engineer asks the AI Guardian:

Have we seen a similar CPU incident before?

The investigation engine evaluates the query and searches the historical Guardian Memory records.

It can identify a previous CPU-related incident associated with a runaway process.

The previous incident was resolved by restarting the affected service.

Instead of starting from zero, the current investigation can use that previous experience as context.

If the new incident is resolved, its operational outcome can also become part of the system's retained experience.

๐ŸŽฏ Why This Is Different from a Normal Monitoring Dashboard

A traditional monitoring dashboard mainly answers:

What is happening?

DataCenter Guardian AI attempts to answer additional operational questions:

What is happening?

Why might it be happening?

Have we seen something similar before?

What happened previously?

What action worked?

What should the engineer investigate next?

This is the core idea behind combining infrastructure monitoring with operational memory.

๐Ÿš€ Future Improvements

There are several ways we can extend the prototype:

Connect to real infrastructure telemetry
Integrate production cloud monitoring
Add advanced anomaly detection
Integrate a dedicated Hindsight memory service
Improve semantic memory retrieval
Improve root-cause analysis
Add automated runbook execution with human approval
Generate automated incident post-mortems
Add real-time alert integrations
Improve predictive maintenance
Support multi-agent infrastructure operations

A dedicated Hindsight memory integration would be an important next step for making the prototype more closely aligned with the hackathon's AI Agents That Learn Using Hindsight theme.

๐Ÿ’ญ What We Learned

The biggest lesson from building DataCenter Guardian AI is that monitoring alone is not enough.

A monitoring system can tell an engineer:

Something is wrong.

An intelligent incident-response system should help answer:

What happened?

Have we seen this before?

What worked previously?

What should we investigate next?

What did we learn from the resolution?

That is why persistent operational memory is an important part of our architecture.

๐Ÿ† Conclusion

DataCenter Guardian AI is our attempt to move beyond traditional infrastructure monitoring toward an AI-assisted operational intelligence platform.

The core idea is simple:

Detect the incident.
โ†“
Investigate the incident.
โ†“
Recall previous experience.
โ†“
Use that experience as context.
โ†“
Recommend an action.
โ†“
Resolve the incident.
โ†“
Retain the outcome.
โ†“
Use it during future investigations.

We built this prototype for HackWithHyderabad 3.0 under the theme:

๐Ÿง  AI Agents That Learn Using Hindsight

Our goal is to make infrastructure operations more context-aware, explainable, and experience-driven.

๐Ÿ”— Project Links

GitHub Repository:
https://github.com/SRIMADHAVI99/Datacenter-Guardian-AI

Live Prototype:
https://datacenter-guardian-ai1.vercel.app/

The source code and working prototype are available through the links above.






Top comments (0)