Azure Operations Copilot for Microsoft Teams
Overview
Azure administrators currently use multiple Azure services and portal pages to monitor resources, investigate alerts, check application performance, review security issues, and understand cloud costs.
This proposal is for an Azure Operations Copilot for Microsoft Teams that brings these operations into one conversational experience.
The goal is not to replace the Azure portal. Instead, Teams would become a convenient starting point for investigating Azure issues.
An engineer could receive an alert in Teams and immediately ask questions about it without having to open several different Azure tools.
The Problem
When an incident happens, engineers often need to move between different Azure services to understand what is going on.
For example, an alert might come from Azure Monitor, but investigating it may require checking:
Log Analytics for logs
Application Insights for application errors and dependencies
Azure Resource Graph for resource information
Azure Service Health for platform issues
Microsoft Defender for Cloud for security findings
Azure Cost Management for cost information
The information is available, but it is spread across different services.
The proposed Copilot would bring these sources together through a single conversational interface in Teams.
Proposed Experience
Users could ask questions such as:
“What are the critical issues in production?”
“Why did this VM trigger an alert?”
“Show unhealthy resources.”
“What changed in the last 24 hours?”
“What is our Azure cost this month?”
“Which resources caused the cost increase?”
“Are there any critical security findings?”
“Is Azure Service Health reporting an issue affecting us?”
The assistant would understand the question, identify the relevant Azure services, retrieve the available information, and provide a short explanation.
Example: Investigating an Alert
An Azure Monitor alert could still be delivered to Teams as it is today.
For example:
Production API Alert
Response time increased to 4.8 seconds.
Instead of simply acknowledging the alert, the engineer could ask:
Engineer:
Why did this happen?
Azure Operations Copilot:
Application Insights shows that SQL dependency latency increased around 14:22. Database calls account for most of the increase in response time. Three timeout exceptions were also recorded during the same period.
Engineer:
Is there an infrastructure problem?
Copilot:
VM CPU, memory, and network metrics are within their normal ranges. No related infrastructure issue was identified. The available data points mainly to the database dependency.
Engineer:
Show related errors from the last 30 minutes.
Copilot:
Here are the errors related to the API and database dependency during that period...
The engineer can continue the investigation without leaving Teams.
Azure Services
The initial solution could integrate with:
Azure Monitor
Log Analytics
Application Insights
Azure Resource Graph
Azure Cost Management
Microsoft Defender for Cloud
Azure Service Health
The architecture could allow additional Azure services to be added later.
Security and Access Control
Security should be a core part of the design.
Authentication would use Microsoft Entra ID, and the assistant would respect existing Azure RBAC permissions.
For example, if an engineer does not have access to a particular subscription or resource, the Copilot should not expose information from that resource.
The same principle should apply to:
Resource information
Logs
Application data
Cost information
Security findings
Subscription information
The Copilot should work within the user's existing permissions rather than creating a separate access model.
Initial Scope
The first version should focus on read-only operations.
The main capabilities could include:
Alert investigation
Resource health checks
Log and error investigation
Application performance analysis
Recent configuration/change investigation
Azure Service Health checks
Security finding summaries
Cost and usage analysis
This keeps the initial implementation focused on visibility and investigation.
Future Scope
Once the read-only experience is mature, controlled remediation could be considered.
For example:
Restart a VM
Scale a resource
Stop a non-production resource
Acknowledge or update an incident
Trigger an approved runbook
These actions should require appropriate permissions and, where necessary, explicit user confirmation or approval.
What Makes This Different
The existing model is mainly:
Azure Alert → Teams Notification
The proposed experience is:
Azure Alert → Teams → Ask Questions → Query Azure Data → Correlate Information → Explain the Issue → Continue Investigation
The important difference is that the alert becomes the starting point of an investigation, rather than the end of the workflow.
Proposed Architecture
A simplified architecture could look like this:
Azure Services
Azure Monitor
Log Analytics
Application Insights
Resource Graph
Cost Management
Defender for Cloud
Service Health
↓
Azure Operations AI Layer
Query understanding
Data retrieval
Cross-service correlation
Incident context
Response generation
Permission enforcement
↓
Microsoft Teams
Alerts
Conversational investigation
Incident context
Follow-up questions
Investigation summaries
Expected Benefit
This could reduce the amount of context switching required during Azure operations.
Instead of asking an engineer to manually check several Azure portal pages, Teams could provide a single place to start the investigation and bring the relevant information together.
The objective is simple:
Turn Microsoft Teams from an Azure alert destination into an interactive Azure operations workspace.
The Azure portal would remain the primary place for detailed management and configuration, while Teams would provide a fast conversational interface for day-to-day monitoring, investigation, and troubleshooting.
Thank You
Shivanna Gundanavar
Top comments (0)