A few months back, I was setting up a Central Monitoring system for all my projects where I could do log monitoring, cost monitoring, compute utilisation by each project, and so on. For that, I had already set up an ML-based algorithm to detect anomalies in cost and RAM, which then sent me alerts through my Telegram bot. You can read about it in my blog.
But once I set my project in production, I was receiving tons of logs, whether they were Django-specific, Docker-specific, Nginx-specific, or Database-specific. I sorted them in my Grafana Dashboard, but again, the problem was the amount of logs for each project, and not all logs necessarily needed my attention. However, error ones or critical ones did need attention for continuous rectification of the project.

I was having two ideas to resolve this issue. The first was either to go with a Small LLM and build a RAG application for retrieval of related logs and summarise them for me on a daily basis. The second was to build an MCP server on this monitoring server at a different port, where I could have different projects listed as options, different technology stacks listed as options, and then connect that to any of my coding agents, where I could ask it whatever information I wanted, whether it be a daily report or a critical error.
The first solution seems fine, but the roadblock comes in front of me while training a RAG application: model selection, weight training, compute resources, etc. But still, it would be only a half addition to my project. It’s like a doctor giving me only the diagnosis after analysing the report without knowing my medical history.

So basically, it was missing the context of the code on which that application was built and the solution to rectify the error that arises.
That’s where I came up with the solution where I can build the MCP server with various options of projects and the tech stack used in them. Using the skills of that project specifically, I can make sure that in the .md, every time I open up the project, it will give me the report at certain intervals.
As that coding agent will be having the code as context on which the application or project is built, the same coding agent can help me to rectify the error. And as the MCP server will be there, basically, you have created a controlled, monitored, self-healing system.
Here, my problem of separate project views and summaries based on each tech stack, and only needed attention being given to any project, saved me a lot of things.
This project was particularly interesting because I was new to SRE, as I had created just the Dashboard for my ease and kept on adding things to it.
Comment me for the project link; I will provide the GitHub link to my repo.
Top comments (0)