<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mansur Fattakhov</title>
    <description>The latest articles on DEV Community by Mansur Fattakhov (@fattakhov).</description>
    <link>https://dev.to/fattakhov</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122926%2Fd92a7478-2662-4406-9360-c4c7b3797890.jpg</url>
      <title>DEV Community: Mansur Fattakhov</title>
      <link>https://dev.to/fattakhov</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/fattakhov"/>
    <language>en</language>
    <item>
      <title>Yes, I Vibecode</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Sun, 27 Sep 2026 15:40:42 +0000</pubDate>
      <link>https://dev.to/fattakhov/yes-i-vibecode-dmc</link>
      <guid>https://dev.to/fattakhov/yes-i-vibecode-dmc</guid>
      <description>&lt;p&gt;Only a year ago, putting together a working application from a single description in plain language was a party trick people bragged about on Twitter. Today that is how internal tools get built in teams that used to wait months for a slot in a developer's queue, and nobody finds it remarkable any more. Vibecoding has stopped being a sideshow and become something anyone can simply do — much the way building a website without knowing HTML once stopped being a miracle.&lt;/p&gt;

&lt;p&gt;Along with that, a stock reply has appeared in developers' conversations, and I hear it constantly now: "vibecoding is not coding." It is delivered in every possible tone, from the condescending to the openly angry, but the meaning is always the same: what those people are doing does not count. I thought that way myself for quite a while, and that is exactly why I want to work out why it is a bad reaction — not because it is factually wrong, but because it misses the target and spoils the person who says it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the irritation comes from
&lt;/h2&gt;

&lt;p&gt;It is worth starting with the fact that the irritation is entirely understandable. Behind the person feeling it lie years spent on things nobody sees from the outside: algorithms that had to be learned twice, nights spent debugging someone else's code without a single line of documentation, postmortems in which it turned out that you were the cause yourself, and the slowly collected bruises out of which professional instinct is eventually made. And right next to that, someone describes a task in the words they would use with their grandmother, and gets the result in one evening.&lt;/p&gt;

&lt;p&gt;It feels as though the part that was hard is precisely the part being devalued. As though years of experience had collapsed into an ability to phrase things neatly in plain language. This is not greed and not snobbery — it is an ordinary human jealousy over one's own labour, and it would be strange not to feel it. The problem is not the feeling but the conclusion drawn from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honestly: I thought the same
&lt;/h2&gt;

&lt;p&gt;At first I caught myself on exactly that thought — nonsense, not serious, toys for people who do not want to learn properly. I watched a couple of demos, looked inside the resulting code, saw things there I would not let into prod even at the point of a deadline, and calmed down: not a competitor.&lt;/p&gt;

&lt;p&gt;The mistake was not in my assessment of the code — the code really was mediocre. The mistake was that I was comparing the wrong things with each other: someone else's result after one evening against my own result after months of work. The comparison that would have made sense sounds different — what I myself can do with this tool, and how that will differ from what I do without it. The moment I put the question in that form, there was no one left to argue with.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tool does not level, it amplifies
&lt;/h2&gt;

&lt;p&gt;It turned out that the better a person understands how systems are built, the more they get out of code generation. That contradicts the first impression — you would think a tool that writes code for you ought to level out the strong and the weak. It does not level anything out; it multiplies what is already there.&lt;/p&gt;

&lt;p&gt;The reason is simple: an engineer's work never did come down to typing text. It begins long before the first line — with framing the problem, where you have to formulate not only what the result should be, but also what happens when everything goes wrong: what we do on a timeout, what we do when two requests race each other, where the load limit lies beyond which the solution stops working. Half the problems of the future task are visible right here, and they are visible to the person who has already stepped on them.&lt;/p&gt;

&lt;p&gt;Next you need context the model cannot know in principle: how this particular service is put together, why there is a queue here rather than a direct call, which decisions were already taken and written down a year ago and why they should not be reopened along the way. Then the resulting code has to be read with the eyes of the person who will answer for it in prod — which means looking not for typos but for the quiet places: an unhandled timeout, idempotency lost somewhere along the road, a query that works beautifully on a hundred test rows and falls over on a real million. And finally all of it has to be rolled out, the metrics watched, rolled back if it went wrong, and answered for — that part is delegated to absolutely no one and is covered by no tool at all.&lt;/p&gt;

&lt;p&gt;Generation covers the middle: the actual writing of the code, the very part that was already the fastest for an experienced person. The edges stay where they were, and it is in them that everything an engineer is paid for lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two different occupations under one name
&lt;/h2&gt;

&lt;p&gt;This, it seems to me, is where the whole argument grows from: one word is used for two different things, and then people are surprised that they cannot come to an agreement.&lt;/p&gt;

&lt;p&gt;The first is vibecoding as entertainment and as a way into the profession. Someone puts together something of their own, takes pleasure in the fact that what they had in mind works, and sees a result they would never in their life have seen without this tool. For most of them it will stay an evening pastime, and there is absolutely nothing wrong with that. More than that, it is through such trifles that new people come into the industry — and they come several years earlier than they would have come through a textbook on algorithms, if they ever arrived at all.&lt;/p&gt;

&lt;p&gt;The second is an engineer who uses generation professionally. The tool is exactly the same, but the tasks are of a different scale and carry a different cost of error: data migrations that cannot be replayed, integrations with someone else's payment systems, rewriting a service that is holding load right now, infrastructure you answer for with night shifts. This is no longer vibecoding in the sense in which the word is usually used — it is engineering work in which generation stands in the same row as the debugger and the profiler, and just as surely decides nothing on its own.&lt;/p&gt;

&lt;p&gt;The difference, in short, is not in the tool but in what happens before it and after it.&lt;/p&gt;

&lt;h2&gt;
  
  
  About the musician at his friends' gig
&lt;/h2&gt;

&lt;p&gt;When an engineer, in reply to "but you're vibecoding," starts explaining that he is not vibecoding at all but doing real, difficult work, he is arguing about a word — and from the outside that looks not like an argument but like a defence of status, whatever he may actually have meant.&lt;/p&gt;

&lt;p&gt;I like the comparison with music here. Picture a musician from a good orchestra who has come to a concert by his friends' band: they got together six months ago, they play in a garage, they miss half the notes, the drummer rushes. And there he is all evening, wincing and telling everyone around him how bad this is and how much he dislikes it. Formally he is absolutely right — they really do play worse than he does. In human terms he has simply ruined the evening for people who took nothing away from him.&lt;/p&gt;

&lt;p&gt;The people putting together their first application from a description today are the friends in the garage. They are not laying claim to a seat in the orchestra and they are not taking it away from anyone. Some of them will come a year from now to learn to play seriously, and it will be precisely because they had a good evening, not because someone explained to them how bad they were.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to answer it
&lt;/h2&gt;

&lt;p&gt;So when you are told that you are vibecoding, the best thing you can do is not to take offence. Smile and agree: yes, and I do it brilliantly and professionally. After that the conversation either moves into a normal professional channel or ends — and both options are better than an argument about terms.&lt;/p&gt;

&lt;p&gt;Along the way I have forbidden myself three things. Not to prove that my work is difficult: if its difficulty is not visible in the result, no words will explain it, and if it is visible, there is nothing left to explain. Not to correct other people's terminology: the phrase "this isn't vibecoding, it's engineering with code generation" has never yet made anyone smarter, but it reliably ends the conversation. And most importantly — not to devalue someone else's interest: a person showing you a thing they put together over the weekend is not waiting for a review, what they need is for someone to notice that they made it at all, and that is worth saying out loud even if inside there is terrifying code and three hardcoded passwords.&lt;/p&gt;

&lt;p&gt;The last of these matters more to me than the first two. Interest is the rarest and most fragile resource in our profession: it is put out with a single condescending comment, and lighting it again afterwards almost never works.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for an engineer
&lt;/h2&gt;

&lt;p&gt;The skill does not disappear, it shifts. Value moves out of the speed of typing code — where, frankly, there was never much of it to begin with — and over to where there was always the most of it: understanding the domain, framing a task so that it can be solved at all, seeing the boundaries, and answering for what runs in prod at three in the morning.&lt;/p&gt;

&lt;p&gt;That, incidentally, is exactly the same set the team lead is responsible for — &lt;a href="https://mind.mansur.expert/en/kpis-that-dont-turn-into-a-stick/" rel="noopener noreferrer"&gt;there was a separate post&lt;/a&gt; about it. And it is exactly what cannot be generated: no model knows what downtime counts as acceptable at your company, why the last migration was rolled back at the final moment, and whom you need to talk to before touching this service.&lt;/p&gt;

&lt;p&gt;And how a working everyday tool gets assembled out of generation — projects, goals, routine — I will tell separately, closer to December.&lt;/p&gt;

&lt;h2&gt;
  
  
  In short
&lt;/h2&gt;

&lt;p&gt;⬜ Vibecoding has gone mainstream; arguing with that is too late and pointless&lt;br&gt;&lt;br&gt;
⬜ The tool does not level, it amplifies: the better you understand systems, the greater the return&lt;br&gt;&lt;br&gt;
⬜ Generation covers the middle of the work, but not the framing, the context, the review or the responsibility&lt;br&gt;&lt;br&gt;
⬜ Entertainment and engineering work share one word — they are different occupations, and both are fine&lt;br&gt;&lt;br&gt;
⬜ If you are told "you're vibecoding," agree and add that you do it professionally&lt;br&gt;&lt;br&gt;
⬜ Do not review someone else's interest: notice it and praise them for having made the thing&lt;/p&gt;

</description>
      <category>ai</category>
      <category>coding</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>KPIs That Don't Turn Into a Stick</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Sun, 27 Sep 2026 14:42:03 +0000</pubDate>
      <link>https://dev.to/fattakhov/kpis-that-dont-turn-into-a-stick-44oj</link>
      <guid>https://dev.to/fattakhov/kpis-that-dont-turn-into-a-stick-44oj</guid>
      <description>&lt;p&gt;You now have several teams, each with a lead, and you want the system to run predictably — without diving into every team every day. The first reflex is to "hand out KPIs." And that is exactly the place where it's easy to ruin everything.&lt;/p&gt;

&lt;p&gt;It all breaks against one regularity the economist Charles Goodhart formulated half a century ago: &lt;strong&gt;once a metric becomes a target, it stops being a good metric.&lt;/strong&gt; Not because people are dishonest, but because you can almost always work on the number directly — bypassing the thing it was introduced for. Give every lead a share of uptime and you'll get cautious releases and arguments about whose incident it was, not a more reliable product.&lt;/p&gt;

&lt;p&gt;What follows is how I built KPIs for team leads on a gamedev project with six teams while working around that trap. The main question there turned out to be not "which metrics do we take" but "what is a lead responsible for in the first place."&lt;/p&gt;

&lt;p&gt;This is a spin-off of the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;"How to bring order to a project"&lt;/a&gt; series — this time about management.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it started
&lt;/h2&gt;

&lt;p&gt;Six teams: the mobile client, the launcher, the game server and the rest. At the C-level we wanted transparency — to see how things were going without climbing into every team daily.&lt;/p&gt;

&lt;p&gt;But there was also a concrete trigger, not just manageability in general. In several teams, churn was eating the focus on the things that never look urgent: ANRs and crashes. Nobody worked on them deliberately — not because they couldn't, but because it was nobody's zone. A feature has a deadline and an owner; a stability number has neither. So beyond transparency, the KPIs had to place responsibility: not "who's to blame for the crash," but "whose zone of attention is this at all."&lt;/p&gt;

&lt;p&gt;The starting point was the CTO's own KPI. It already existed — put together quickly, but it existed, and that turned out to matter more than the precision of the wording: there was a single outcome everyone worked toward. The mistake would have been to slice it into shares and hand them down — more on that trap below. What was valuable was something else: it set a direction, and from there you had to answer your own question — what is this particular lead responsible for within that direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the C-level wants this — honestly
&lt;/h2&gt;

&lt;p&gt;The motivation from the top usually sounds like "transparency and manageability," but it's worth unpacking all the way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Seeing the state of the teams without micromanagement.&lt;/strong&gt; Not "what is everyone busy with" but "is the system healthy" — and noticing degradation before an incident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Predictability.&lt;/strong&gt; Promises to the business rest on something measurable, not on a feeling that "it seems fine."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scaling yourself.&lt;/strong&gt; A CTO with six teams physically cannot be the entry point to every decision — lead KPIs are a way to delegate responsibility, not tasks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And right here is the honest trap you have to see in advance: &lt;strong&gt;the easiest way to "decompose the CTO's KPI" is to cut it into pieces and hand them down.&lt;/strong&gt; The CTO's goal is "uptime and product quality" — so give each lead a share of uptime. It looks logical and works badly: the lead gets a metric they only partly influence, starts optimizing the number instead of the system, and defends themselves with reports instead of doing the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the leads think about it
&lt;/h2&gt;

&lt;p&gt;Leads have their own set of motivations — and fears, and ignoring them is the most expensive option of all:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What a lead gets from honest KPIs:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;clarity: "what is expected of me" stops being telepathy — these are literally the &lt;a href="https://mind.mansur.expert/en/people-bus-factor-expectations-and-day-one/" rel="noopener noreferrer"&gt;expectations spoken out loud&lt;/a&gt; from the people track, only at the level of roles;&lt;/li&gt;
&lt;li&gt;protection: written criteria are insurance against arbitrary evaluation "by mood" and against rules changed after the fact;&lt;/li&gt;
&lt;li&gt;arguments: "we need one more person" sounds weak; "we won't hold the stated recovery time without a second on-call" sounds strong.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What a lead is afraid of:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;that the metrics will become a stick: any red number is grounds for an inquest;&lt;/li&gt;
&lt;li&gt;that they'll have to work "for the number": producing reporting instead of value;&lt;/li&gt;
&lt;li&gt;that the goals will be set above their head — from a world where their teams don't exist.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything below is, in essence, an answer to those fears.&lt;/p&gt;

&lt;h2&gt;
  
  
  The key turn: a lead owns the system, not the outcome
&lt;/h2&gt;

&lt;p&gt;The central decision everything rests on: &lt;strong&gt;a lead is responsible not for the final business metric, but for the system and the processes that make the outcome inevitable.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CTO KPI → processes → team → outcome
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The outcome is a consequence of a working system, not the subject of a monthly evaluation. The QA lead doesn't own "zero bugs in production" — they own the regression suite being current, a verdict existing before a release, and defects that escaped to production being reviewed so they close the hole in the system. The infra lead doesn't own "zero incidents" — they own alerts being meaningful, runbooks existing, and MTTR going down.&lt;/p&gt;

&lt;p&gt;And this is the answer to the question "how do I make everyone work on product quality rather than on the CTO's KPI": &lt;strong&gt;don't translate the outcome downward — ask about the health of the system that produces that outcome.&lt;/strong&gt; You can "work on" an outcome number in a report; you can't do that to a working system — it either exists or it doesn't.&lt;/p&gt;

&lt;p&gt;The evaluation is then phrased through the system as well, in three levels:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Good&lt;/strong&gt; — the system works and is developing; the outcome is a stable consequence of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Satisfactory&lt;/strong&gt; — the system exists but has gaps: coverage is incomplete, it runs irregularly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unsatisfactory&lt;/strong&gt; — there is no system; the work is reactive, held together by manual labor and heroics.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Notice: "unsatisfactory" is not "bad numbers," it's "no system." The difference is fundamental — and it takes the fear of the stick away.&lt;/p&gt;

&lt;h2&gt;
  
  
  The construction: standing KPIs plus KPIs of the month
&lt;/h2&gt;

&lt;p&gt;The second finding that made life much easier was to split KPIs into two kinds:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Standing (S)&lt;/strong&gt; — the hygiene of the role: the same bar every month. Monitoring covers the services, the regression suite is current, RCAs get written for incidents, releases go through a verdict. They don't change, they aren't renegotiated — they simply have to be met.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;KPIs of the month (M)&lt;/strong&gt; — three or four goals set for the month: close a specific piece of tech debt, raise feature coverage, lower MTTR, run a Game Day. They live for one cycle and are set anew at the monthly review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The standing ones give stability ("what normal means"), the monthly ones give movement ("what we're improving right now"). Without S, monthly goals turn into firefighting; without M, standing ones turn into stagnation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Metrics: ground them in facts, not in wishes
&lt;/h2&gt;

&lt;p&gt;The selection rules I arrived at:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A metric has to be readable from a live system&lt;/strong&gt; — from monitoring, the tracker, the app store. A metric the lead counts by hand at the end of the month will be dead within a quarter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thresholds are set from facts, not from dreams.&lt;/strong&gt; Before writing down a response-time threshold, I pulled the real percentiles out of monitoring. It turned out that almost all the services were already inside the corridor and one stood out — and the threshold became honest: "pull up the outlier," not "everyone improve everything." A threshold taken from thin air demotivates worse than no threshold at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Few metrics.&lt;/strong&gt; Three or four standing plus three or four monthly per team is the ceiling. Past that, the evaluation system costs more than it's worth.&lt;/li&gt;
&lt;li&gt;Examples of "system → its metrics" pairs by domain: QA — the currency of the regression suite and the trend of defects escaping to production; infrastructure — recovery time and the share of alerts that actually get a response (an alert everyone ignores does more harm than no alert at all); backend — SLOs and error-budget burn; the mobile client — crash-free and ANR, whose thresholds the app store sets for you.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The stabilization pattern: when the goal is far away
&lt;/h2&gt;

&lt;p&gt;The most common reason KPIs "don't land": the goal is objectively unreachable this month. Between the current availability level and the target one there isn't a step but several quarters of work. Making the target level the monthly bar means condemning the lead to an "unsatisfactory" — and to cynicism.&lt;/p&gt;

&lt;p&gt;The pattern that works is &lt;strong&gt;burn-down plus a ladder toward the horizon&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;the absolute goal is a reference point&lt;/strong&gt;, not a monthly bar: it's honestly written down as the horizon;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;the monthly bar is a step&lt;/strong&gt;: noticeably reduce the number of errors, close their main source, climb one rung of the ladder you laid out from the current level to the horizon;&lt;/li&gt;
&lt;li&gt;what gets evaluated is the step: taken — "good," even if the horizon is still far away.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is also the answer to "how do we agree": the conversation with the lead stops being a haggle about the unreachable and becomes planning for the next rung.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to present it and how to agree
&lt;/h2&gt;

&lt;p&gt;The procedural part is short, but everything depends on it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Warn in advance.&lt;/strong&gt; The leads knew from ordinary conversations, long before the first card, that KPIs were being prepared and that we'd sit down to discuss them soon. It's a cheap step that removes half the tension: the system doesn't fall on them as a surprise, it arrives at an expected moment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't hand down a finished thing.&lt;/strong&gt; The frame (S/M, the levels, the pattern) is brought by the CTO; filling the card in is the subject of a conversation with the lead. Thresholds are discussed from the actual numbers, out loud, one on one. This is exactly the &lt;a href="https://mind.mansur.expert/en/people-bus-factor-expectations-and-day-one/" rel="noopener noreferrer"&gt;spoken expectations&lt;/a&gt;, in both directions — the lead also says what they need in order for the system to work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The CTO sets the goals of the month&lt;/strong&gt; — the final word is theirs, this isn't a democracy; but goals set without a conversation don't work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A monthly review&lt;/strong&gt; — by the card, half an hour: how the standing ones are doing, what happened to the goals of the month, what we set for the next one. A red metric is grounds for the question "what's in the way," not for a reprimand. One or two cycles in that mode and the fear of the stick goes away.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Be careful with money.&lt;/strong&gt; Tying this to a bonus from the first month breaks the whole construction: a metric that has become money stops being a measurement instantly — that's the same Goodhart from the opening, only in his fastest form. First a quarter or two of simply keeping the rhythm and building trust in the numbers.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What came out of it
&lt;/h2&gt;

&lt;p&gt;Not once did I hear a flat "no." That doesn't mean the cards were accepted in silence: every item was worked through one on one — and not once at launch, but continuously, at every review. That is the central point of the process itself: not to hand it down and forget it, but to talk and to agree until the person has their own understanding of why the metric is there. The absence of loud resistance is the result of those conversations, not a sign that nobody cared.&lt;/p&gt;

&lt;p&gt;There were corrections along the way, and always of one kind. Either a wording that could be read two ways, or a number that couldn't be reached this month. The first ones we rewrote. The second we replaced with tracking a trend: not "hold the threshold" but "move toward it" — with a check at the next review. Not once did we have to throw a metric out entirely: the argument was always about the threshold, never about the meaning.&lt;/p&gt;

&lt;p&gt;What changed noticeably:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ANRs and crashes became the leads' concern, not the CTO's.&lt;/strong&gt; Exactly what the whole thing was started for: the tracking settled on the teams' side — the lead looks at the numbers and arrives with a plan, instead of waiting for a question from above;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;QA stopped needing to be pulled in by hand.&lt;/strong&gt; Before, the lead had to be prodded both about the release process and about taking part in postmortems. Once the cards existed, the lead and the team started joining in on their own;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;responsibility stopped being a matter of guesswork.&lt;/strong&gt; It became visible who owns what — and that understanding became shared, rather than living "in the CTO's head".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I don't have a case where the scheme wouldn't fit. There is a condition without which it doesn't work: agreeing in a way that leaves everyone understanding the goal. Without that, the cards turn into work for the sake of work — and it won't be the scheme's fault.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alternatives — and when they're better
&lt;/h2&gt;

&lt;p&gt;An honest section: lead KPIs are not the only instrument.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OKRs&lt;/strong&gt; — if the task is not "keep the system healthy" but "move toward ambitious goals": inspiring objectives, measurable key results, 70% attainment is normal. Weaker for hygiene, stronger for pushes. Nothing stops you from combining them: the S part as KPIs, the M part in the spirit of OKRs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team health checks&lt;/strong&gt; (a regular self-assessment against a checklist: code, deployment, tests, mood) — softer, without thresholds, good as a quarterly reflection. Not enough if the C-level needs manageability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Just agreeing on expectations without numbers&lt;/strong&gt; — works in a small team with strong trust; stops working roughly when there are more than three leads and "keeping it in your head" no longer works.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The choice is simple: the more teams you have and the more expensive failures are, the closer you move to a formal system. But it's always worth starting with the conversation about expectations; KPIs are its written form, not a replacement for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist
&lt;/h2&gt;

&lt;p&gt;⬜ The lead owns the system, not the outcome; the outcome is a consequence&lt;br&gt;&lt;br&gt;
⬜ The frame: standing KPIs (hygiene) plus three or four goals of the month; a monthly review&lt;br&gt;&lt;br&gt;
⬜ Three levels of evaluation, through the maturity of the system, not through "red numbers"&lt;br&gt;&lt;br&gt;
⬜ Metrics are read from live systems, thresholds are grounded in facts&lt;br&gt;&lt;br&gt;
⬜ A distant goal is the horizon; the monthly bar is a step (burn-down / one rung)&lt;br&gt;&lt;br&gt;
⬜ The frame is brought by the CTO, the content is a conversation with the lead from actual numbers&lt;br&gt;&lt;br&gt;
⬜ A red metric is the question "what's in the way," not a reprimand&lt;br&gt;&lt;br&gt;
⬜ A number that's unreachable this month is replaced by a trend, not by an "unsatisfactory"&lt;br&gt;&lt;br&gt;
⬜ Don't tie money to the metrics until the system has lived through a couple of quarters&lt;/p&gt;

</description>
      <category>leadership</category>
      <category>management</category>
      <category>productivity</category>
      <category>career</category>
    </item>
    <item>
      <title>People: Bus Factor, Expectations and Day One</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Fri, 18 Sep 2026 11:20:01 +0000</pubDate>
      <link>https://dev.to/fattakhov/people-bus-factor-expectations-and-day-one-2hl4</link>
      <guid>https://dev.to/fattakhov/people-bus-factor-expectations-and-day-one-2hl4</guid>
      <description>&lt;p&gt;Part of the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;"How to bring order to a project"&lt;/a&gt; series — the people track. It comes after infrastructure and architecture not because it matters less, but the opposite: it's the most important and the slowest track. A config can be moved into git in a day; trust, habits and the understanding of "who owns what" take months to rebuild.&lt;/p&gt;

&lt;p&gt;Let me repeat the track's thesis from the hub, because it's the main one: &lt;strong&gt;code with a bus factor of 1 is not an asset, it's a hostage.&lt;/strong&gt; And the other way around: a team where knowledge flows freely can fix any code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1. Bus factor: find the hostages
&lt;/h2&gt;

&lt;p&gt;The source material is the "who knows what" map from the &lt;a href="https://mind.mansur.expert/en/state-snapshot-first-week-on-a-new-project/" rel="noopener noreferrer"&gt;state snapshot&lt;/a&gt;. If the "who knows it" column shows the same name for five services, you have five hostages and one irreplaceable person — who, by the way, isn't enjoying it either: they can't take a vacation or move to new work.&lt;/p&gt;

&lt;p&gt;The treatment is three moves, in order of rising cost:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Documentation on touch&lt;/strong&gt; — the cheapest: figured out a piece — leave a passport or a note (this is already built into the &lt;a href="https://mind.mansur.expert/en/architecture-write-it-down-before-rewriting/" rel="noopener noreferrer"&gt;architecture track&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pairing on tasks&lt;/strong&gt; — a second person on tasks in a "solo" zone; slower, but the knowledge transfers live&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ownership rotation&lt;/strong&gt; — the most expensive, for critical zones: a service officially gets a second owner&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Don't try to push the bus factor to two everywhere — life is too short. Prioritize by the risk map: first what is both critical and lives in one head.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2. Ownership: assemble from facts, don't draw it
&lt;/h2&gt;

&lt;p&gt;The classic mistake is to sit down and draw a beautiful responsibility matrix the way it "should be." A matrix like that dies within a week, because it doesn't match reality.&lt;/p&gt;

&lt;p&gt;The working recipe is to assemble it &lt;strong&gt;from facts&lt;/strong&gt;: pull half a year of closed tracker tasks and the git history of the repositories, and see who actually does the work and who actually decides. That's how our RACI matrix across all services was born: for each one — who executes (R), who's accountable for the result (A), who's consulted (C), who's informed (I). Not "as designed" but "as is" — with all the uncomfortable discoveries, like a service whose executor sits in one team while accountability sits nowhere in particular.&lt;/p&gt;

&lt;p&gt;From there the matrix works both ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a mismatch of "one does it, another answers for it" becomes a conversation and a decision instead of a silent conflict;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;every&lt;/strong&gt; service gets a name next to it — and that name comes from reality.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Updates happen on touch, like everything in this series: the owner changed — the row changes in the same merge request as the passport.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3. Delegate — and spell out expectations
&lt;/h2&gt;

&lt;p&gt;The most personal part of the track. When you're bringing order to a project, it's easy to fall into a trap: you know how to do it right, so you do everything important yourself. The outcome is predictable — you become the very bus factor of 1 you're fighting.&lt;/p&gt;

&lt;p&gt;You have to delegate. But delegation has an entry price that often goes unpaid: &lt;strong&gt;spelled-out expectations.&lt;/strong&gt; A familiar scenario: you hand over a task, the person works for a month, brings the result — and you expected something else. Nobody's guilty: you didn't say, they didn't ask, both filled the gaps with assumptions. A month burned, trust dented on both sides.&lt;/p&gt;

&lt;p&gt;What to spell out at handover — four things, literally as a list:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The result&lt;/strong&gt;: what "done" looks like — not the process, the end state. Better in writing, two sentences.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The deadline and checkpoints&lt;/strong&gt;: not "show me in a month," but "in a week we look at the direction." An early checkpoint is cheap insurance against "I expected something else": turning the work around after a week costs days; after a month it costs a month.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The boundaries of freedom&lt;/strong&gt;: what the person decides on their own, and what they bring to you. Without this they either ping you about everything, or silently make decisions at your level.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The why&lt;/strong&gt;: the task's context. A person who knows "why" makes the right micro-decisions where the instructions are silent.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And the symmetric rule: expectations are spelled out &lt;strong&gt;in both directions&lt;/strong&gt;. Ask what the person needs for the task to move — access, context, uninterrupted time. Half of the "didn't get it done" cases turn out, on review, to be "waited three weeks for access and was too shy to nag."&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4. Onboarding and one-on-ones
&lt;/h2&gt;

&lt;p&gt;Onboarding is not "handed over a laptop and added to the chats." Two metrics worth measuring it by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;environment in a day&lt;/strong&gt; — by the end of day one the person runs the project locally;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;a production task in a week&lt;/strong&gt; — within the first week they ship something real, however tiny.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For that to work, the onboarding documents must exist before the hire, and there are exactly three levels — no more needed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;company&lt;/strong&gt; — what the product is, which teams exist, who leads them, where to look;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;engineering&lt;/strong&gt; — how we work: a task's lifecycle from discussion to deploy, the standards, the decision process;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;glossary&lt;/strong&gt; — the domain dictionary: internal terms and abbreviations "everybody just knows."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And separately — &lt;strong&gt;one-on-ones for the early period.&lt;/strong&gt; Regular short private meetings: weekly for the first month or two, less often later. Not a status update (the tracker exists for that), but "how is it going for you": what's confusing, what's in the way, what surprised you. In their first weeks a fresh person sees your oddities better than any veteran — then they get used to it and that vision is gone. One-on-ones are how you collect it in time. They're also the natural place to re-sync the expectations from step 3 — while a mismatch costs days, not months.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5. Hiring: an honest footnote
&lt;/h2&gt;

&lt;p&gt;Hiring gets one step in this plan, and a short one — not because the topic is small, but because it's a &lt;strong&gt;separate big one&lt;/strong&gt;. The market is full of people; picking the one who will strengthen your particular team is a discipline of its own that doesn't fit into a section.&lt;/p&gt;

&lt;p&gt;What does fit is one rule that costs little and gives a lot: &lt;strong&gt;write the job post from the project's real tasks, not from a list of technologies.&lt;/strong&gt; Not "Python, Docker, K8s, 5 years of experience," but "here are three tasks you'll work on in your first quarter." A text like that filters by itself: the people who respond are the ones interested in the tasks, not the ones who matched the keywords. And the interview gets a ready-made script — you talk about those three tasks.&lt;/p&gt;

&lt;p&gt;The rest — the funnel, evaluation, probation — deserves its own article; it's in the backlog.&lt;/p&gt;

&lt;h2&gt;
  
  
  What came out of it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The irreplaceable stopped being irreplaceable — and breathed easier themselves&lt;/li&gt;
&lt;li&gt;Every service has a name from reality, not from the org chart&lt;/li&gt;
&lt;li&gt;Delegated tasks arrive at the intended result — because the intended result was spelled out up front&lt;/li&gt;
&lt;li&gt;A new person is useful in their first week, and their fresh perspective is collected, not lost&lt;/li&gt;
&lt;li&gt;Job posts answer "what will I actually do here," and that's who applies&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The track's checklist
&lt;/h2&gt;

&lt;p&gt;⬜ Bus factor from the snapshot map; treatment by cost: docs → pairing → rotation&lt;br&gt;&lt;br&gt;
⬜ Responsibility matrix assembled from facts (tracker + git), not drawn&lt;br&gt;&lt;br&gt;
⬜ Every service has a name; "does vs. answers" mismatches resolved&lt;br&gt;&lt;br&gt;
⬜ Delegation = result + checkpoints + boundaries of freedom + the "why"&lt;br&gt;&lt;br&gt;
⬜ Expectations spelled out both ways, in writing&lt;br&gt;&lt;br&gt;
⬜ Onboarding: environment in a day, a production task in a week&lt;br&gt;&lt;br&gt;
⬜ Three documents: company / engineering / glossary — existing before the hire&lt;br&gt;&lt;br&gt;
⬜ One-on-ones early on: weekly for the first months, not status but "how is it going"&lt;br&gt;&lt;br&gt;
⬜ Job posts — from real tasks, not from a list of technologies&lt;/p&gt;

&lt;p&gt;The other tracks of the series are on the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;map&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mind.mansur.expert/en/people-bus-factor-expectations-and-day-one/" rel="noopener noreferrer"&gt;mind.mansur.expert&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>leadership</category>
      <category>management</category>
      <category>teamwork</category>
      <category>career</category>
    </item>
    <item>
      <title>Architecture: Write It Down Before Rewriting</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Wed, 16 Sep 2026 07:13:17 +0000</pubDate>
      <link>https://dev.to/fattakhov/architecture-write-it-down-before-rewriting-2lcl</link>
      <guid>https://dev.to/fattakhov/architecture-write-it-down-before-rewriting-2lcl</guid>
      <description>&lt;p&gt;Part of the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;"How to bring order to a project"&lt;/a&gt; series — the architecture track. Same context: a growing gamedev project, about twenty services on different stacks, an architecture that "grew historically." Nobody designed it badly — nobody designed it at all: services appeared as needs arose, boundaries were drawn by circumstance, decisions lived in chats and heads.&lt;/p&gt;

&lt;p&gt;That's a normal phase — up to a point. The signals that the point has arrived:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;incidents increasingly happen &lt;strong&gt;at the seams&lt;/strong&gt; between services, not inside them;&lt;/li&gt;
&lt;li&gt;every change starts with an argument about "whose area is this anyway";&lt;/li&gt;
&lt;li&gt;estimates keep growing, because "we need to be careful there, we don't know what we might break";&lt;/li&gt;
&lt;li&gt;the only answer to "why is it built this way?" is "historically."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The track's main thesis: &lt;strong&gt;order in architecture starts not with rewriting, but with writing down.&lt;/strong&gt; Half of the "architecture problems" dissolve once the system is honestly described as it is. The other half turns from "scary to touch" into tasks with a price and a plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1. Service passports
&lt;/h2&gt;

&lt;p&gt;The first tool is boring to the point of indecency: every service gets a passport. One page:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# &amp;lt;service&amp;gt;
Why it exists: one sentence.
Area of responsibility: what it does.
Boundaries: what it does NOT do, even when "it would be convenient here."
Stack, port, status (production / deprecated / experiment).
Owner: a name, not "the team in general."
Operations: how it deploys, where the configs are, where it logs.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three rules that keep passports alive instead of becoming a wiki graveyard:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The owner writes it, not the architect.&lt;/strong&gt; A passport is "I'm responsible for this and here's how it works," not "management told me to document it."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Half an hour, no more.&lt;/strong&gt; A passport is not design documentation. If it takes a day, you're writing something else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Updated on touch.&lt;/strong&gt; Changed the boundaries or the deploy method — update the passport in the same merge request. Dedicated "documentation days" don't exist; they don't work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Filling in the passports is diagnostics in itself: you keep catching moments of "wait, why is this even here?" — and those are the right questions, there was just nobody to ask them before. A couple of our services honestly moved to deprecated status during passportization — it turned out habit was the only thing keeping them alive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2. The system map
&lt;/h2&gt;

&lt;p&gt;The passports add up to a map — two artifacts:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The registry&lt;/strong&gt; — a table of all services: type, stack, status, owner. Boring, but this is where you see that there aren't "about fifteen" services, but exactly nineteen, three of which are alive only in legends.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The call graph&lt;/strong&gt; — in text, per service: whom it calls, over which protocol, and why. For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;shop → accounts service    HTTP REST   ownership check at purchase
shop → game API            HTTP REST   progress levels (callback)
shop ← payment webhooks    HTTPS       events from payment providers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No pretty arrows at the start — a table is enough. And it has a property a slide-deck diagram never will: it lives in git next to the code, so it can be reviewed and diffed. Changed an integration — the map's row changes in the same merge request; a mismatch is caught in review, not a year later.&lt;/p&gt;

&lt;p&gt;The rule is the same as in the &lt;a href="https://mind.mansur.expert/en/state-snapshot-first-week-on-a-new-project/" rel="noopener noreferrer"&gt;state snapshot&lt;/a&gt;: trust reality, not documentation. A map that shows the desired state is worse than no map — people make wrong decisions with full confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3. Boundaries: where a service ends
&lt;/h2&gt;

&lt;p&gt;The most expensive architecture bugs live not inside services but between them — where two services consider the same thing "theirs." The signs of broken boundaries are recognizable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;two services write to the same database table;&lt;/li&gt;
&lt;li&gt;the same business logic is duplicated in two places "for speed";&lt;/li&gt;
&lt;li&gt;a service reaches into a neighbor's internals past its API, "because it's simpler";&lt;/li&gt;
&lt;li&gt;the first half hour of every incident goes into figuring out whose area it is.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our classic was the payment flow: historically payments started in the accounts service, then the shop grew its own purchases and webhooks — and the logic sprawled across two homes. Every payment incident opened with archaeology: which of the two services was supposed to handle this event? Duplicated handlers masked the problem right until the day they both handled the same event.&lt;/p&gt;

&lt;p&gt;The cure is a boundaries document. For every service, explicitly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;inside&lt;/strong&gt;: what the service is solely responsible for&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;outside&lt;/strong&gt;: what's adjacent but belongs to someone else (and to whom exactly)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;interface&lt;/strong&gt;: how others interact with it — and that's the only door&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The order of the fix matters: &lt;strong&gt;document first, code second.&lt;/strong&gt; First the team agrees on paper where payment events live (in our case — webhooks in the shop, accounts only read status via API), and only then the code is brought to the agreement. The reverse doesn't work: code without an agreement sprawls back within a quarter.&lt;/p&gt;

&lt;p&gt;The goal isn't beauty but two practical properties: &lt;strong&gt;no overlaps&lt;/strong&gt; (every entity has one owner) and &lt;strong&gt;growth control&lt;/strong&gt; (a new feature first gets an address — which service it goes to and why — instead of settling wherever).&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4. Decisions on paper
&lt;/h2&gt;

&lt;p&gt;Architecture "in heads" dies when people leave. Two document genres stuck with us, and the difference between them matters:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RFC — before a change.&lt;/strong&gt; Half a page, by template:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Problem: what hurts, in numbers or incidents.
Options: 2–3, including "do nothing."
Decision: which one and why.
Cost: work, risks, what about rollback.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lifecycle is simple: Draft → Review → Approved/Rejected → Implemented. A new RFC is automatically announced in the team channel — the review is open, and objections at the Draft stage cost one comment instead of a production rework. A live example: the proposal to replace direct calls between game services with a message bus went through three review iterations — and the final design (with a journal and delivery guarantees) had almost nothing in common with the first draft. All those iterations happened &lt;strong&gt;before&lt;/strong&gt; the first line of code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ADRs and research — around a decision.&lt;/strong&gt; A short record: what was decided, why, what was considered and rejected. Plus research for specific risks — "how to back up the analytics database," "the plan for moving data off the game host." The difference from an RFC: nothing is being proposed here; knowledge is being recorded — so the same research doesn't get redone a year later.&lt;/p&gt;

&lt;p&gt;And one rule on top of both genres: &lt;strong&gt;before any architecture work — check the index first.&lt;/strong&gt; Don't reinvent what's already decided; don't start a migration if there's an open document about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5. Tech debt with a price tag
&lt;/h2&gt;

&lt;p&gt;Everyone has a tech debt list; it only works with two extra columns:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Debt&lt;/th&gt;
&lt;th&gt;What it blocks&lt;/th&gt;
&lt;th&gt;Cost to fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Shop and accounts share payment logic&lt;/td&gt;
&lt;td&gt;every payment incident ×2 in time; new providers go into two places&lt;/td&gt;
&lt;td&gt;~3 weeks: boundaries + moving the webhooks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No backups for the analytics DB&lt;/td&gt;
&lt;td&gt;losing all analytics on a disk failure&lt;/td&gt;
&lt;td&gt;research done, ~1 week to implement&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Blocks" is the main column. Debt that blocks nothing is not debt — it's a trait of the system; cross it out and stop carrying it. Debt that makes every feature in a service cost ×2 is the first candidate for work — and now that's visible not only to the tech lead, but to the product manager who sets priorities.&lt;/p&gt;

&lt;p&gt;The list's sources are the &lt;a href="https://mind.mansur.expert/en/state-snapshot-first-week-on-a-new-project/" rel="noopener noreferrer"&gt;state snapshot&lt;/a&gt; and incident reviews. Debt is paid down two ways: small — by the touch rule (touching a service — leave it cleaner), large — through an RFC, like any architecture change. Once a quarter the list is revised: some items closed by touch, some stopped blocking, some grew into an RFC.&lt;/p&gt;

&lt;h2&gt;
  
  
  How not to slide into bureaucracy
&lt;/h2&gt;

&lt;p&gt;The honest risk of this track: getting carried away and demanding an RFC for renaming a variable. The antidote is explicit thresholds. For us, an RFC is mandatory when a change:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;moves a &lt;strong&gt;boundary&lt;/strong&gt; between services (or creates a new service);&lt;/li&gt;
&lt;li&gt;touches &lt;strong&gt;data&lt;/strong&gt;: a shared DB schema, a migration, a storage change;&lt;/li&gt;
&lt;li&gt;touches &lt;strong&gt;money&lt;/strong&gt;: payment flows, billing;&lt;/li&gt;
&lt;li&gt;is irreversible or expensive to roll back.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else is a regular merge request with a regular review. And the second antidote: all documents live next to the code or in one wiki with a single index, get updated on touch, and fit on a page. A document nobody can find or finish reading — that's what bureaucracy is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it buys you: evolution without revolution
&lt;/h2&gt;

&lt;p&gt;With the map, the boundaries and the decision process in place, big rebuilds stop being leaps of faith. From the same project's practice — three rebuilds, all through RFCs and along the boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The most loaded service was rewritten in another language&lt;/strong&gt; — entirely behind its own boundary. The API contract was pinned in the passport and the boundaries document, so "rewrite the internals" became a local task: the other seventeen services noticed nothing. Without a pinned contract this would have been "rewriting the system."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The shop was extracted from the "main API" that did everything&lt;/strong&gt; — first the boundaries document got the line "purchases, cases, storefront — a separate service," then the code moved along that boundary. Extraction by document rather than intuition is also a ready checklist: what moves, what stays, where the interface is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data was separated from compute onto different hosts&lt;/strong&gt; — following a research document with a migration plan, in stages, with rollback points. Not "overnight on Saturday," but a boring sequence of steps, each of which could be stopped.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these rebuilds was a "big rewrite." Each was a local operation with a clear price and a rollback. That's the point of the track: &lt;strong&gt;architectural order is not a pretty diagram on the wall — it's cheap changes.&lt;/strong&gt; The diagram is a by-product.&lt;/p&gt;

&lt;h2&gt;
  
  
  The track's checklist
&lt;/h2&gt;

&lt;p&gt;⬜ A passport for every service: why, area, boundaries, owner, operations&lt;br&gt;&lt;br&gt;
⬜ Owners write the passports; updates on touch, in the same merge request&lt;br&gt;&lt;br&gt;
⬜ Service registry + call graph in text, in git; verified against reality&lt;br&gt;&lt;br&gt;
⬜ Boundaries document: inside / outside / interface; one owner per entity&lt;br&gt;&lt;br&gt;
⬜ Fixing boundaries: agreement on paper first, code second&lt;br&gt;&lt;br&gt;
⬜ RFC before a change (problem → options → decision → cost); open review in the channel&lt;br&gt;&lt;br&gt;
⬜ Explicit RFC thresholds: boundaries, data, money, irreversibility — the rest is a regular MR&lt;br&gt;&lt;br&gt;
⬜ ADRs and research: record knowledge, "index first"&lt;br&gt;&lt;br&gt;
⬜ Tech debt with "what it blocks" and "cost" columns; quarterly revision&lt;br&gt;&lt;br&gt;
⬜ Small debt — by touch; large — through an RFC&lt;/p&gt;

&lt;p&gt;The other tracks of the series are on the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;map&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mind.mansur.expert/en/architecture-write-it-down-before-rewriting/" rel="noopener noreferrer"&gt;mind.mansur.expert&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>documentation</category>
      <category>backend</category>
      <category>devops</category>
    </item>
    <item>
      <title>Infrastructure and Deployment: Order by Iteration</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Tue, 15 Sep 2026 07:23:01 +0000</pubDate>
      <link>https://dev.to/fattakhov/infrastructure-and-deployment-order-by-iteration-2jga</link>
      <guid>https://dev.to/fattakhov/infrastructure-and-deployment-order-by-iteration-2jga</guid>
      <description>&lt;p&gt;Part of the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;"How to bring order to a project"&lt;/a&gt; series — the infrastructure track. Same starting point: a growing gamedev project, nearly two dozen services on different stacks — Python, Go, C++, a mobile client — and infrastructure that grew faster than it matured. Every service deployed differently, configs lived on servers, and a new service was born by copy-pasting the oldest one.&lt;/p&gt;

&lt;p&gt;Attacking this with a "big refactoring" is a road to nowhere: the business won't let you stop shipping features for a month. What works is the iterative approach: small steps, each paying off immediately and breaking nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 0. Snapshot and light
&lt;/h2&gt;

&lt;p&gt;Two prerequisites, without which infrastructure changes turn into roulette:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A state snapshot&lt;/strong&gt; — the service map, who knows what, top risks. How to take one in a week — &lt;a href="https://mind.mansur.expert/en/state-snapshot-first-week-on-a-new-project/" rel="noopener noreferrer"&gt;a separate article&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability&lt;/strong&gt; — logs, metrics, alerts. Also &lt;a href="https://mind.mansur.expert/en/how-developers-can-monitor-production-and-why-it-matters/" rel="noopener noreferrer"&gt;covered&lt;/a&gt;. The rule is simple: turn on the lights first, then move the furniture.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now the actual infrastructure steps, in the order they pay off.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1. CI/CD: templates instead of copy-paste
&lt;/h2&gt;

&lt;p&gt;The classic disease: every service has its own CI config, written in its own era by its own person. You fix the pipeline in one repo — in the other twenty it stays broken.&lt;/p&gt;

&lt;p&gt;The recipe that worked for me: &lt;strong&gt;a dedicated repository of CI templates&lt;/strong&gt;. Inside:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;core-templates.yml&lt;/code&gt; — image builds, test runs, timings&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;notify-templates.yml&lt;/code&gt; — notifications to the team messenger&lt;/li&gt;
&lt;li&gt;ready-made pipelines for the typical cases: a basic service, a service with secrets, a cluster deployment&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A service plugs it in with one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;include&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;project&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;infra/ci-templates"&lt;/span&gt;
    &lt;span class="na"&gt;file&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/pipeline-basic.yml"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And gets the standard conveyor: on a merge request — test image build → tests → candidate build; on the main branch — the production image with a tag and &lt;code&gt;latest&lt;/code&gt;. Everything goes through shared templates, so a fix or an improvement in one place rolls out to every service with its next pipeline.&lt;/p&gt;

&lt;p&gt;Start with one service — the most frequently deployed one: the template gets polished there, the rest catch up as you touch them.&lt;/p&gt;

&lt;p&gt;👉 The side effect matters more than the main one: deployment stops being one person's knowledge. Anyone can press the button in CI — and that's no longer scary, because before the button there are tests and a process that's identical for everyone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2. Configs into git, servers into code
&lt;/h2&gt;

&lt;p&gt;The second source of chaos: configuration living directly on servers. Edited by hand, versioned nowhere, lost when a server moves.&lt;/p&gt;

&lt;p&gt;There are two separate recipes here, and it pays not to mix them:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Static configs&lt;/strong&gt; — what changes at deploy time: compose files, env templates, environment parameters. They belong in a configuration repository: change → merge request → it's clear who, what and why → rollout through CI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Runtime configs&lt;/strong&gt; — what changes the system's behavior on the fly: feature flags, prices, game parameters. These get their own GitOps repository with a schema and validation: values live in YAML, services read them through a small SDK. A product manager changes a parameter with a merge request, not with a ticket to the dev team — and the change history comes for free.&lt;/p&gt;

&lt;p&gt;The servers themselves — only through Ansible roles (or any other IaC). The rule is hard: &lt;strong&gt;no hand edits on hosts&lt;/strong&gt;. Not "just for a minute," not "I'll move it into the role later." Every manual edit is a future surprise during a migration or a recovery.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3. Secrets: out of CI, into a vault
&lt;/h2&gt;

&lt;p&gt;A separate pain: secrets smeared across CI variables, env files on servers and personal notes. The recipe I arrived at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Deploy a &lt;strong&gt;secrets manager&lt;/strong&gt; (Vault, Infisical — the class of tool matters more than the brand) with test/production environments.&lt;/li&gt;
&lt;li&gt;CI variables keep &lt;strong&gt;exactly five values&lt;/strong&gt;: the vault address, the machine-identity ID and secret, the project ID and the path. That's it.&lt;/li&gt;
&lt;li&gt;Everything else — registry credentials, deploy SSH keys, notification webhooks — the pipeline fetches from the vault itself, per environment.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A technical detail that saved hours of debugging: every secret is written to its own file via &lt;code&gt;base64 -d&lt;/code&gt;, byte for byte, with no shell interpretation. Multiline SSH keys and certificates stop breaking on escaping once and for all.&lt;/p&gt;

&lt;p&gt;The win isn't only security: rotating a secret is an edit in one place, not archaeology across twenty repositories. And when someone leaves the team, you know exactly what to reissue and where.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4. Deployment: one entrance, different backends
&lt;/h2&gt;

&lt;p&gt;Deployment in the CI templates is standard too, with two backends for different scales:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Simple services&lt;/strong&gt; live on hosts under docker compose: the deploy job takes the host list and the SSH key from the vault, connects and updates the service. Boring and predictable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cluster services&lt;/strong&gt; ship to Kubernetes via Helm: the repo holds &lt;code&gt;values.yaml&lt;/code&gt; + &lt;code&gt;values.production.yaml&lt;/code&gt;, the image tag comes from the commit SHA, the job waits for the rollout and logs its progress. Merge requests build a candidate with its own tag — deployable to a test environment with the very same template.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both roads end the same way: a notification in the team messenger — who deployed what and where. It's cheaper than any "release book": the deploy history assembles itself in the channel.&lt;/p&gt;

&lt;p&gt;👉 The general principle: &lt;strong&gt;a service has no deployment method of its own&lt;/strong&gt;. It has parameters (hosts or helm values) — the method is shared. A new deployment method appears in the templates, or it doesn't appear at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5. Backups: order isn't there until you can restore
&lt;/h2&gt;

&lt;p&gt;A test that puts a lot into perspective: "the server is dead for good — how many hours until everything runs again?" If the answer starts with "well…" — the track has further to go.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Database backups — on schedule, to a separate location, with an alert on "backup didn't happen"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recovery is tested by hand&lt;/strong&gt; at least once a quarter: a backup that has never been restored is a lottery ticket, not a backup&lt;/li&gt;
&lt;li&gt;Everything else must be restorable from git: roles bring up the server, CI builds and ships the services, configs arrive from repositories&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When this genuinely works, "moving a server" turns from a special operation with a midnight call into a half-day procedure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6. The reference service template
&lt;/h2&gt;

&lt;p&gt;Once CI is standardized and configs are in git, you can finally pin down "what a proper service looks like." My reference consists of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dockerfile&lt;/strong&gt; — multi-stage, with separate targets for tests and production (the CI templates are built for exactly that)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repo structure&lt;/strong&gt; — code, tests, migrations, a README with a two-command local start&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JSON logs&lt;/strong&gt; with a request_id passed through — so the service plugs into centralized logging right away&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A healthcheck endpoint&lt;/strong&gt; — so the orchestrator and monitoring see liveness without magic&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The CI templates plugged in&lt;/strong&gt; — a pipeline out of the box&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A service passport&lt;/strong&gt; — one page: why it exists, its area of responsibility, boundaries, who owns it, how it deploys&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The passport is the underrated part. Passports add up to a living map of the system: a service registry with ports, stacks and owners, plus a "who calls whom" graph. With twenty services, that map answers half of the architecture questions before they're asked.&lt;/p&gt;

&lt;p&gt;A new service is created from the template in minutes. Old ones are not rewritten en masse: the rule "every touch brings the service closer to the reference" pulls everything alive up to the standard within months — and what's dead gets to die honestly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What came out of it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Deploying any service is one button in CI, the same process for everyone&lt;/li&gt;
&lt;li&gt;Configs and servers are restorable from git — "moving a server" went from special operation to procedure&lt;/li&gt;
&lt;li&gt;A new service comes from the template in minutes, with logs, metrics and a pipeline from day one&lt;/li&gt;
&lt;li&gt;Secrets rotate in one place; someone leaving the team is a procedure, not a panic&lt;/li&gt;
&lt;li&gt;The deploy history assembles itself in the notification channel — who shipped what and where&lt;/li&gt;
&lt;li&gt;Infrastructure knowledge stopped living in one head: passports + templates + roles are readable by any team member&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The track's checklist
&lt;/h2&gt;

&lt;p&gt;⬜ CI templates in a dedicated repo; the first service plugged in&lt;br&gt;&lt;br&gt;
⬜ Tests are mandatory on merge requests; no deploys bypass CI&lt;br&gt;&lt;br&gt;
⬜ Static configs — in a configuration repository&lt;br&gt;&lt;br&gt;
⬜ Runtime parameters — a GitOps repo with a schema and an SDK&lt;br&gt;&lt;br&gt;
⬜ Servers — only through IaC roles; no hand edits on hosts&lt;br&gt;&lt;br&gt;
⬜ Secrets — in a vault; CI keeps only the connection parameters&lt;br&gt;&lt;br&gt;
⬜ Deployment — shared, from templates (SSH-compose / Helm), with a chat notification&lt;br&gt;&lt;br&gt;
⬜ Database backups on schedule + a quarterly restore test&lt;br&gt;&lt;br&gt;
⬜ Service reference: Dockerfile, structure, JSON logs, healthcheck, passport&lt;br&gt;&lt;br&gt;
⬜ Rule: every touch brings an old service closer to the reference&lt;/p&gt;

&lt;p&gt;The next tracks of the series — architecture, people, processes — are on the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;map&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mind.mansur.expert/en/infrastructure-and-deployment-order-by-iteration/" rel="noopener noreferrer"&gt;mind.mansur.expert&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>infrastructure</category>
      <category>cicd</category>
      <category>docker</category>
    </item>
    <item>
      <title>Processes: A Living Rhythm with Formal Points</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Mon, 14 Sep 2026 09:48:48 +0000</pubDate>
      <link>https://dev.to/fattakhov/processes-a-living-rhythm-with-formal-points-4ecm</link>
      <guid>https://dev.to/fattakhov/processes-a-living-rhythm-with-formal-points-4ecm</guid>
      <description>&lt;p&gt;The final track of the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;"How to bring order to a project"&lt;/a&gt; series — processes. It came last not because it matters least, but because it leans on everything before it: without &lt;a href="https://mind.mansur.expert/en/how-developers-can-monitor-production-and-why-it-matters/" rel="noopener noreferrer"&gt;observability&lt;/a&gt;, status updates degrade into retelling feelings — you can only say "we're fine" while looking at metrics; without &lt;a href="https://mind.mansur.expert/en/infrastructure-and-deployment-order-by-iteration/" rel="noopener noreferrer"&gt;reliable deployment&lt;/a&gt;, a regular release cadence simply won't happen — every rollout stays manual work everyone wants to postpone.&lt;/p&gt;

&lt;p&gt;And straight to the main prejudice. Process is not a boring, faceless instrument of corporate management. &lt;strong&gt;A living process changes&lt;/strong&gt; — it adapts to the team, drops what's useless, grows what's needed: that's how it should be. But for all its liveliness, it defines &lt;strong&gt;formal points and artifacts&lt;/strong&gt;: the task is filed, the decision is written down, the status is spoken out loud. These points pin the work down and keep it from getting lost.&lt;/p&gt;

&lt;p&gt;Yes, in the moment it's faster to DM someone than to file a ticket. But the DM is only faster today: a month later that agreement exists nowhere — not in the plan, not in the history, not in the other person's head. A formal point feels like a detour, but over the distance it &lt;strong&gt;shortens the path to a delivered result&lt;/strong&gt;. That's the whole bet: a good process is not more paperwork — it's fewer decisions on the fly and less lost work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The task lifecycle: one explicit path
&lt;/h2&gt;

&lt;p&gt;Every task in a team should travel one route everyone knows. Ours looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Discussion
    ↓
Architecture discussion (when the task calls for it)
    ↓
Decision on paper — an RFC or ADR
    ↓
Decision review
    ↓
Planning in the tracker
    ↓
Development: code → merge request → review → deploy
    ↓
Watching production
    ↓
Documentation update
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Looks like bureaucracy — in reality, for a small team most steps take minutes: an "architecture discussion" is five minutes at a whiteboard, a "decision on paper" is half a page. And "development" is a single step even though it contains code, MR, review and deploy: once the &lt;a href="https://mind.mansur.expert/en/infrastructure-and-deployment-order-by-iteration/" rel="noopener noreferrer"&gt;infrastructure track&lt;/a&gt; is done, that's one well-worn conveyor, not four separate events.&lt;/p&gt;

&lt;p&gt;The point of the ladder isn't meetings — it's this: &lt;strong&gt;no step gets skipped silently.&lt;/strong&gt; You may consciously decide "no RFC needed here" — you may not fail to notice that a decision was never written down. Two rules from practice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;major architecture changes never skip the RFC step — ever;&lt;/li&gt;
&lt;li&gt;work is not done until the documentation is updated. Not "I'll write it later" — not done.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Planning: sprints or kanban is the wrong question
&lt;/h2&gt;

&lt;p&gt;The right question is: &lt;strong&gt;what's the nature of your work?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Building a &lt;strong&gt;feature with a deadline&lt;/strong&gt; you can plan out — take sprints: a fixed window, a clear scope, and at the end you can see whether you hit it.&lt;/li&gt;
&lt;li&gt;Handling a &lt;strong&gt;flow of tasks&lt;/strong&gt; — support, testing, bugfixes, small improvements — take kanban: tasks stream through columns, throughput matters, not "did we make it by Friday."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both live peacefully in one team: feature work in sprints, the flow on a kanban board. It's all individual; the tool serves the work, not the other way around. Only two things are universal: &lt;strong&gt;a short cycle&lt;/strong&gt; (a week or two — beyond that plans lie) and &lt;strong&gt;a visible backlog&lt;/strong&gt; — anyone on the team can open it and see what's next.&lt;/p&gt;

&lt;h2&gt;
  
  
  Status updates and dailies: not optional
&lt;/h2&gt;

&lt;p&gt;The most underrated part of process — and the one where I hold my firmest position: &lt;strong&gt;status updates are mandatory, always.&lt;/strong&gt; Even if the team is two people. Even as plain text in chat. Always.&lt;/p&gt;

&lt;p&gt;Without regular sync the same thing happens every time — I've watched it more than once: people don't know about each other's problems, don't know whom to go to, quietly stall for weeks — and everything falls apart. Not because anyone works badly, but because nobody sees the whole picture.&lt;/p&gt;

&lt;p&gt;The working mechanics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A status — always.&lt;/strong&gt; Short and regular: what I did, what's next, what's blocking me. For a two-person team three lines of text are enough — but every day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The daily is for discussion&lt;/strong&gt;, not for reporting to the boss: you hear about a neighbor's problem — you say "I've had something similar, come over."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Problems get solved after, in the right circle.&lt;/strong&gt; Something half-hour-sized surfaces on the daily — we don't solve it with everyone: we note it, and after the call exactly the two or three people who need it gather. The daily stays short; the problem gets a real deep-dive.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's the cheapest process of all — ten minutes a day — with the highest return: it's the only one that keeps a finger on the pulse.&lt;/p&gt;

&lt;h2&gt;
  
  
  Release cadence: predictable and boring
&lt;/h2&gt;

&lt;p&gt;This section is suspiciously short — because releases are a science of their own: windows, frequency, staging, rollbacks, the fear of Friday deploys. It deserves — and will get — a separate article.&lt;/p&gt;

&lt;p&gt;For this track one principle is enough: &lt;strong&gt;a release should be boring.&lt;/strong&gt; A clear scheme (mine: candidate tag → staging, release tag → production), identical for every service, with a notification in the channel. If the team's pulse rises before a release, that's not a cadence — it's an arrhythmia, and you fix it in the &lt;a href="https://mind.mansur.expert/en/infrastructure-and-deployment-order-by-iteration/" rel="noopener noreferrer"&gt;infrastructure track&lt;/a&gt;, not with courage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Incidents: blameless reviews
&lt;/h2&gt;

&lt;p&gt;When production goes down — and it goes down for everyone — process beats heroism. The minimal skeleton:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;pin the timeline: when it started, what coincided (most often — a deploy);&lt;/li&gt;
&lt;li&gt;fix first, analyze after — not the other way around;&lt;/li&gt;
&lt;li&gt;a review &lt;strong&gt;with no blame&lt;/strong&gt; and &lt;strong&gt;one action item&lt;/strong&gt;: not "we'll be more careful," but a concrete change — an alert, a CI check, a checklist line;&lt;/li&gt;
&lt;li&gt;the postmortem goes into a registry; before the next review, first search the history for a similar case.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;How to set this up in a small team with one or two people — the minimal artifacts and where to grow from there — will also become a separate article, with a real incident walkthrough.&lt;/p&gt;

&lt;h2&gt;
  
  
  The meeting audit
&lt;/h2&gt;

&lt;p&gt;The final step of the track is the most enjoyable: open the calendar and honestly answer, for every recurring meeting, what it produces. The rule is simple: &lt;strong&gt;a meeting either produces decisions, or syncs people, or dies.&lt;/strong&gt; "It's historical" is not a third option — it's a diagnosis from the &lt;a href="https://mind.mansur.expert/en/architecture-write-it-down-before-rewriting/" rel="noopener noreferrer"&gt;architecture track&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;What survives the audit almost always: the daily (it syncs — see above) and incident reviews (they produce decisions). What dies first: the weekly hour-long "all-hands sync" where everyone waits for their minute and stays silent for the other fifty-nine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The series is complete
&lt;/h2&gt;

&lt;p&gt;With this article, the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;"How to bring order to a project"&lt;/a&gt; map is closed — all six tracks now have their articles:&lt;/p&gt;

&lt;p&gt;⬜ → ✅ &lt;a href="https://mind.mansur.expert/en/state-snapshot-first-week-on-a-new-project/" rel="noopener noreferrer"&gt;The state snapshot&lt;/a&gt; — a week of fixing nothing&lt;br&gt;&lt;br&gt;
⬜ → ✅ &lt;a href="https://mind.mansur.expert/en/how-developers-can-monitor-production-and-why-it-matters/" rel="noopener noreferrer"&gt;Observability&lt;/a&gt; — turn on the lights&lt;br&gt;&lt;br&gt;
⬜ → ✅ &lt;a href="https://mind.mansur.expert/en/infrastructure-and-deployment-order-by-iteration/" rel="noopener noreferrer"&gt;Infrastructure and deployment&lt;/a&gt; — order by iteration&lt;br&gt;&lt;br&gt;
⬜ → ✅ &lt;a href="https://mind.mansur.expert/en/architecture-write-it-down-before-rewriting/" rel="noopener noreferrer"&gt;Architecture&lt;/a&gt; — write it down first&lt;br&gt;&lt;br&gt;
⬜ → ✅ &lt;a href="https://mind.mansur.expert/en/people-bus-factor-expectations-and-day-one/" rel="noopener noreferrer"&gt;People&lt;/a&gt; — bus factor and expectations&lt;br&gt;&lt;br&gt;
⬜ → ✅ Processes — you are here&lt;/p&gt;

&lt;p&gt;A reminder of what "order has arrived" feels like: deploying is not scary, a new person is useful in their first week, an incident is a procedure rather than a panic, and "why is it like this here?" has a written answer. You'll never reach perfection — it's enough that every week feels a little calmer than the last.&lt;/p&gt;

&lt;p&gt;The series is finished; the topics aren't: separate articles on release cadence, small-team incident handling and hiring are coming. Stay around.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mind.mansur.expert/en/processes-a-living-rhythm-with-formal-points/" rel="noopener noreferrer"&gt;mind.mansur.expert&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>leadership</category>
      <category>management</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
    <item>
      <title>The State Snapshot: Your First Week on a New Project</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Mon, 14 Sep 2026 07:33:47 +0000</pubDate>
      <link>https://dev.to/fattakhov/the-state-snapshot-your-first-week-on-a-new-project-20oh</link>
      <guid>https://dev.to/fattakhov/the-state-snapshot-your-first-week-on-a-new-project-20oh</guid>
      <description>&lt;p&gt;When you join a running project — as a lead, an architect, or part-time — your hands itch to start fixing things right away. There it is, the crooked deploy; the database without indexes; the service nobody has touched since last year. Stop. For the first week you fix nothing at all.&lt;/p&gt;

&lt;p&gt;This article covers "Track 0" from the &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;project order map&lt;/a&gt;: how to take an honest snapshot of a project's state in one week. I'll show it on a real example — a growing gamedev project I joined to bring order to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a snapshot first, not a fix
&lt;/h2&gt;

&lt;p&gt;Three reasons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You don't yet know what matters.&lt;/strong&gt; The first "obvious" problem almost never turns out to be the main one. Fixing it means spending your credit of trust on something secondary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Any fix without the full picture is a gamble.&lt;/strong&gt; The system lived for years before you; what looks like a bug may be the prop holding everything up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The snapshot is a result in itself.&lt;/strong&gt; A week later you hold something the project has never had: a complete picture. Often it's the first document of its kind in the company's history.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 1. Talk to everyone
&lt;/h2&gt;

&lt;p&gt;The first thing I did on the gamedev project was a one-on-one call with every single person: developers, DevOps, support, product. Not a meeting — one-on-one, 30–45 minutes each.&lt;/p&gt;

&lt;p&gt;The questions are simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What annoys you most about the project?&lt;/li&gt;
&lt;li&gt;What's scary to touch, and why?&lt;/li&gt;
&lt;li&gt;What would you fix first if it were up to you?&lt;/li&gt;
&lt;li&gt;What do you alone know? What happens if you leave for a month?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The discovery that repeats on every project: &lt;strong&gt;everybody already knows everything.&lt;/strong&gt; People carry precise lists of problems in their heads for years — where it hurts, what will fall apart next, why Friday deploys are forbidden. Nobody has ever collected that knowledge in one place. You don't discover problems — you collect them.&lt;/p&gt;

&lt;p&gt;The side effect is as valuable as the main one: the team sees that the new person listens first and breaks nothing. It's the cheapest way to earn trust — and it works exactly once, in your first week.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2. The service map
&lt;/h2&gt;

&lt;p&gt;In parallel with the calls — a table. Boring, but honest:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Where it lives&lt;/th&gt;
&lt;th&gt;Who knows it&lt;/th&gt;
&lt;th&gt;How it deploys&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;…&lt;/td&gt;
&lt;td&gt;…&lt;/td&gt;
&lt;td&gt;…&lt;/td&gt;
&lt;td&gt;…&lt;/td&gt;
&lt;td&gt;…&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The rule: &lt;strong&gt;don't trust the documentation — trust reality.&lt;/strong&gt; Look at the servers, the cron jobs, what is actually running. On my project a couple of services from the "architecture diagram" turned out to be long dead, while two live ones the diagram knew nothing about surfaced: a script on a host and a bot deployed "just for a minute" a year ago.&lt;/p&gt;

&lt;p&gt;In the "who knows it" column you'll keep seeing the same name. That's not a column — that's an alarm.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3. The scary places
&lt;/h2&gt;

&lt;p&gt;The calls and the map produce a special list — the things people describe as "better not touch it." For every item, ask "why" until you hit the bottom:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;sometimes the fear hides a real landmine (no tests, no rollback, the only person who knew it has left);&lt;/li&gt;
&lt;li&gt;sometimes it's just a legend: it crashed once, scared everyone, and has been avoided ever since.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Separating landmines from legends is half the risk work done. Legends dissolve with a single experiment on staging; landmines go into the top risks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4. Top 5 risks on one page
&lt;/h2&gt;

&lt;p&gt;The week's finale is one page, written in plain language, no jargon:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If server X dies, the game won't start for anyone, and recovery will take an unknown amount of time, because the backup has never been tested.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Five items like that, sorted by pain. Not fifty — five. Keep the rest in your working notes.&lt;/p&gt;

&lt;p&gt;I showed that page to the team &lt;strong&gt;before&lt;/strong&gt; showing it to management. First, they fixed my factual mistakes. Second, the document reached management as "ours," not as "the new guy reporting on everyone." That matters: the snapshot is a tool for working together, not an audit denunciation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you have a week later
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A service map that matches reality&lt;/li&gt;
&lt;li&gt;The scary places list, split into landmines and legends&lt;/li&gt;
&lt;li&gt;Top 5 risks the team has signed off on&lt;/li&gt;
&lt;li&gt;People's trust — you listened, wrote things down, and broke nothing&lt;/li&gt;
&lt;li&gt;Most importantly: &lt;strong&gt;a basis for ordering the steps.&lt;/strong&gt; From here on you fix what actually fires most often, not what caught your eye first&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every other track starts from this snapshot: &lt;a href="https://mind.mansur.expert/en/how-developers-can-monitor-production-and-why-it-matters/" rel="noopener noreferrer"&gt;observability&lt;/a&gt; turns the lights on where the map showed darkness, and the &lt;a href="https://mind.mansur.expert/en/infrastructure-and-deployment-order-by-iteration/" rel="noopener noreferrer"&gt;infrastructure steps&lt;/a&gt; follow the risk list instead of gut feeling.&lt;/p&gt;

&lt;h2&gt;
  
  
  The week's checklist
&lt;/h2&gt;

&lt;p&gt;⬜ One-on-one calls with everyone — 4 questions, your own notes&lt;br&gt;&lt;br&gt;
⬜ Service map: what / why / where / who knows it / how it deploys&lt;br&gt;&lt;br&gt;
⬜ Verify the map against reality (servers, crons, processes)&lt;br&gt;&lt;br&gt;
⬜ The "scary places" list → split into landmines and legends&lt;br&gt;&lt;br&gt;
⬜ Top 5 risks on one page → show the team → then upward&lt;br&gt;&lt;br&gt;
⬜ Fix nothing. At all. For a week.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mind.mansur.expert/en/state-snapshot-first-week-on-a-new-project/" rel="noopener noreferrer"&gt;mind.mansur.expert&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>leadership</category>
      <category>management</category>
      <category>career</category>
      <category>devops</category>
    </item>
    <item>
      <title>How Developers Can Monitor Production — and Why It Matters</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Sun, 13 Sep 2026 08:21:47 +0000</pubDate>
      <link>https://dev.to/fattakhov/how-developers-can-monitor-production-and-why-it-matters-9pm</link>
      <guid>https://dev.to/fattakhov/how-developers-can-monitor-production-and-why-it-matters-9pm</guid>
      <description>&lt;p&gt;When we write code, it often feels like the main thing is to make it work locally. But reality is different: the &lt;em&gt;real&lt;/em&gt; life of a service begins not on your laptop, but in production. That’s where it faces load, unpredictable users, and dozens of unexpected situations.  &lt;/p&gt;

&lt;p&gt;Many developers think of monitoring and logs as “something for DevOps.” Let them look at the graphs and figure out why the service crashed. But the truth is: &lt;strong&gt;observability is a developer’s tool first&lt;/strong&gt;. It helps you quickly understand what went wrong and fix a bug before it turns into a midnight call from support.  &lt;/p&gt;

&lt;p&gt;In this post I’ll explain:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🧐 why metrics and dashboards are important,
&lt;/li&gt;
&lt;li&gt;🔧 how developers can actually use them,
&lt;/li&gt;
&lt;li&gt;🚨 and why without observability you’re basically “flying blind” in production.
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What Is Observability (in simple words)
&lt;/h2&gt;

&lt;p&gt;Any service lives in two worlds:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🌍 the code world, where you’re in control with &lt;code&gt;print&lt;/code&gt; or breakpoints,
&lt;/li&gt;
&lt;li&gt;🌍 and the production world, where the system runs under load, with real users and unpredictable inputs.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Observability&lt;/strong&gt; is how you understand &lt;em&gt;what’s happening inside&lt;/em&gt; without looking directly into the code.  &lt;/p&gt;

&lt;p&gt;The three “pillars” are:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;📜 &lt;strong&gt;Logs&lt;/strong&gt; — text outputs (errors, requests, events). Answer: &lt;em&gt;“what happened?”&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;📊 &lt;strong&gt;Metrics&lt;/strong&gt; — numbers (request counts, response times, memory usage). Answer: &lt;em&gt;“how is it working?”&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;🧵 &lt;strong&gt;Traces&lt;/strong&gt; — following one request across multiple services. Answer: &lt;em&gt;“why is it working this way?”&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Together, they give you a complete picture: find bottlenecks, debug errors, and see how changes affect production.  &lt;/p&gt;




&lt;h2&gt;
  
  
  The Main Tools (and Why These)
&lt;/h2&gt;

&lt;p&gt;There are many monitoring products out there — from Datadog and New Relic to Dynatrace. In practice, most projects I see use the &lt;strong&gt;Grafana Labs + Prometheus ecosystem&lt;/strong&gt;.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;⏱ &lt;strong&gt;Prometheus&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The de-facto standard for metrics collection.
&lt;/li&gt;
&lt;li&gt;Simple model: the service exposes metrics via HTTP, Prometheus “scrapes” them.
&lt;/li&gt;
&lt;li&gt;Powerful query language (PromQL).
&lt;/li&gt;
&lt;li&gt;Tons of exporters: MySQL, PostgreSQL, Nginx, Docker, etc.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;📈 &lt;strong&gt;Grafana&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Universal visualization tool.
&lt;/li&gt;
&lt;li&gt;Works not only with Prometheus but also with Loki, Elastic, Postgres, and more.
&lt;/li&gt;
&lt;li&gt;Great dashboards, alerts, annotations.
&lt;/li&gt;
&lt;li&gt;Huge marketplace of ready-to-use dashboards.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;📜 &lt;strong&gt;Loki&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Centralized logging system with Prometheus-like syntax.
&lt;/li&gt;
&lt;li&gt;Much lighter than Elasticsearch.
&lt;/li&gt;
&lt;li&gt;Perfectly integrated with Grafana — jump from a metric to the exact log line.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🧵 &lt;strong&gt;Tempo / Jaeger&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Distributed tracing.
&lt;/li&gt;
&lt;li&gt;Shows the full path of a request across multiple services — essential for microservices.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why developers love this stack:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
✔️ open source — no license cost, easy to start&lt;br&gt;&lt;br&gt;
✔️ low entry barrier — a basic dashboard in a couple of hours&lt;br&gt;&lt;br&gt;
✔️ one ecosystem — everything integrates&lt;br&gt;&lt;br&gt;
✔️ scalable — from pet projects to highload  &lt;/p&gt;




&lt;h2&gt;
  
  
  How Developers Can Actually Use It
&lt;/h2&gt;

&lt;p&gt;Here are some everyday scenarios:  &lt;/p&gt;

&lt;p&gt;🔎 &lt;strong&gt;1. Found a bug you can’t reproduce&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;User says: &lt;em&gt;“The button doesn’t work.”&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Search in Loki by &lt;code&gt;request_id&lt;/code&gt; or exception text.
&lt;/li&gt;
&lt;li&gt;You see the stacktrace → problem identified.
👉 Result: no more “works on my machine.”
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;⚡ &lt;strong&gt;2. Service slows down after a release&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Grafana shows latency doubled after yesterday’s deploy.
&lt;/li&gt;
&lt;li&gt;Drill down into endpoint metrics → find the culprit.
&lt;/li&gt;
&lt;li&gt;Check logs for parameters.
👉 Result: you know exactly what broke.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🧮 &lt;strong&gt;3. Optimizing SQL queries&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Add metrics for query count and duration.
&lt;/li&gt;
&lt;li&gt;Grafana shows the heaviest queries.
&lt;/li&gt;
&lt;li&gt;Compare before/after optimization.
👉 Result: real data, not gut feeling.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🚨 &lt;strong&gt;4. Reacting to alerts yourself&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rule: “if 5xx &amp;gt; 5% over 5 minutes → send alert.”
&lt;/li&gt;
&lt;li&gt;Notifications in Slack/Telegram.
👉 Result: developers react instantly, not a day later via support.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🛠️ &lt;strong&gt;5. Your own “developer dashboard”&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A few panels: error rates, latency, top endpoints.
👉 Result: all essentials in one place.
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  A Real Project Example
&lt;/h2&gt;

&lt;p&gt;I joined a growing gaming project (FastAPI + MySQL) as a part-time engineer. It was unstable — clients often complained that the game wouldn’t load.  &lt;/p&gt;

&lt;p&gt;What I found:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;❌ no log collection,
&lt;/li&gt;
&lt;li&gt;❌ Grafana only had node-exporter metrics nobody watched.
The project was flying blind.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 1. Logs&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Set up Loki with JSON format.
&lt;/li&gt;
&lt;li&gt;Added request/response logs, duration, endpoint filtering.
&lt;/li&gt;
&lt;li&gt;Super fast search via Docker plugin.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 2. Metrics&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Added HTTP latency, DB transaction times, response statuses, concurrent requests.
&lt;/li&gt;
&lt;li&gt;Finally saw the full load picture.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 3. Problem discovery&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MySQL queries crashing → fixed with connection pooling.
&lt;/li&gt;
&lt;li&gt;httpx connection pool overloading → fixed.
&lt;/li&gt;
&lt;li&gt;Long queries → Percona Monitoring and Management → cleaned up DB load.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 4. Alerts&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;500 errors after release.
&lt;/li&gt;
&lt;li&gt;RabbitMQ queue overflow.
&lt;/li&gt;
&lt;li&gt;Network issues.
&lt;/li&gt;
&lt;li&gt;Payment provider outages → switch to backup quickly.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The result:&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;✅ from black box → transparent system
&lt;/li&gt;
&lt;li&gt;✅ developers could debug themselves
&lt;/li&gt;
&lt;li&gt;✅ support and ops stopped firefighting
&lt;/li&gt;
&lt;li&gt;✅ downtime minimized → saved money
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What About Tracing?
&lt;/h2&gt;

&lt;p&gt;Tracing is important, especially in microservices. But for basics, I use a simpler approach: passing a &lt;strong&gt;&lt;code&gt;request_id&lt;/code&gt; across all services&lt;/strong&gt;.  &lt;/p&gt;

&lt;p&gt;That lets me track a request in Loki end-to-end across the ecosystem. It’s not as powerful as Jaeger or Tempo, but it covers 80% of a developer’s needs.  &lt;/p&gt;




&lt;h2&gt;
  
  
  A Minimal Observability Checklist
&lt;/h2&gt;

&lt;p&gt;If your project has nothing yet, start small — 20% effort for 80% results:  &lt;/p&gt;

&lt;p&gt;✅ &lt;strong&gt;Logs&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Centralized (Loki/EFK)
&lt;/li&gt;
&lt;li&gt;JSON format
&lt;/li&gt;
&lt;li&gt;Add &lt;code&gt;request_id&lt;/code&gt; and duration
&lt;/li&gt;
&lt;li&gt;Filter by endpoint/error level
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;✅ &lt;strong&gt;Metrics&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;HTTP latency
&lt;/li&gt;
&lt;li&gt;Response statuses (2xx/4xx/5xx)
&lt;/li&gt;
&lt;li&gt;Concurrent requests
&lt;/li&gt;
&lt;li&gt;DB transaction times
&lt;/li&gt;
&lt;li&gt;Queue load
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;✅ &lt;strong&gt;Database&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MySQL/Postgres exporter
&lt;/li&gt;
&lt;li&gt;Track slow queries, deadlocks
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;✅ &lt;strong&gt;Alerts&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;5xx error threshold
&lt;/li&gt;
&lt;li&gt;Queue overflow
&lt;/li&gt;
&lt;li&gt;Service outages
&lt;/li&gt;
&lt;li&gt;DB overload
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;✅ &lt;strong&gt;Team Interface&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Base Grafana dashboard: errors, latency, load
&lt;/li&gt;
&lt;li&gt;A “developer dashboard” with API essentials
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;👉 With just this, your project stops being a black box. Everything else — tracing, business dashboards, analytics — can be added gradually.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mind.mansur.expert/en/how-developers-can-monitor-production-and-why-it-matters/" rel="noopener noreferrer"&gt;mind.mansur.expert&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>observability</category>
      <category>devops</category>
      <category>monitoring</category>
      <category>backend</category>
    </item>
    <item>
      <title>How to Bring Order to a Project: A Track Map</title>
      <dc:creator>Mansur Fattakhov</dc:creator>
      <pubDate>Sun, 13 Sep 2026 08:21:43 +0000</pubDate>
      <link>https://dev.to/fattakhov/how-to-bring-order-to-a-project-a-track-map-4db2</link>
      <guid>https://dev.to/fattakhov/how-to-bring-order-to-a-project-a-track-map-4db2</guid>
      <description>&lt;p&gt;Every now and then I join a project where "everything works, but nobody dares to touch it." Deployment is done by hand from memory, the architecture grew historically, knowledge lives in people's heads, hiring happens by accident. That's normal: this is what almost any project looks like when it grows faster than it matures.&lt;/p&gt;

&lt;p&gt;You don't bring order to a project like that in one heroic push. What works is different: split the chaos into tracks, take small steps within each track, and make every step pay off immediately. Below is my complete plan — what to do, in what order, and why. Each track will get its own deep-dive article; this is the map.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to use this plan
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Snapshot first, then light, then changes.&lt;/strong&gt; Don't fix anything until you can see what's going on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tracks run in parallel&lt;/strong&gt;, at different speeds — push the one that hurts right now.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A step must fit into a week.&lt;/strong&gt; If it doesn't, it's not a step, it's a project; keep cutting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every touch leaves a trace&lt;/strong&gt;: touched a service — bring it closer to the standard; figured out a piece of the system — write it down.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Track 0. State snapshot
&lt;/h2&gt;

&lt;p&gt;One week, fix nothing — only look and write down.&lt;/p&gt;

&lt;p&gt;⬜ Service map: what exists, where it lives, who deploys it and how&lt;br&gt;&lt;br&gt;
⬜ The "scary places" list — things people describe as "better not touch it"&lt;br&gt;&lt;br&gt;
⬜ Who knows what: which knowledge has a single carrier&lt;br&gt;&lt;br&gt;
⬜ Top 5 risks — one page, in plain human language&lt;/p&gt;

&lt;p&gt;💬 The main artifact of this stage is not a document but a picture in your head. The document is a way to verify it: show it to the team, let them correct you.&lt;/p&gt;

&lt;p&gt;📖 Full article: &lt;a href="https://mind.mansur.expert/en/state-snapshot-first-week-on-a-new-project/" rel="noopener noreferrer"&gt;the state snapshot — your first week on a new project&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Track 1. Observability — turn on the lights
&lt;/h2&gt;

&lt;p&gt;⬜ Centralized, structured logs&lt;br&gt;&lt;br&gt;
⬜ HTTP and database metrics&lt;br&gt;&lt;br&gt;
⬜ 3–5 alerts for the things that actually wake you up at night&lt;br&gt;&lt;br&gt;
⬜ A request_id passed through every service&lt;/p&gt;

&lt;p&gt;💬 This track goes first because it makes every other track cheaper: any change is visible, any incident takes minutes to investigate instead of "going by gut feeling."&lt;/p&gt;

&lt;p&gt;📖 Full article: &lt;a href="https://mind.mansur.expert/en/how-developers-can-monitor-production-and-why-it-matters/" rel="noopener noreferrer"&gt;how developers can monitor production — and why it matters&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Track 2. Infrastructure and deployment
&lt;/h2&gt;

&lt;p&gt;⬜ CI/CD on the most frequently deployed service&lt;br&gt;&lt;br&gt;
⬜ Configuration in git, not on servers&lt;br&gt;&lt;br&gt;
⬜ A reference service template: structure, Dockerfile, healthcheck, pipeline&lt;br&gt;&lt;br&gt;
⬜ Rule: new services only from the template; old ones catch up as you touch them&lt;/p&gt;

&lt;p&gt;💬 The goal of this track is for deployment to stop being an event and stop being one person's knowledge. The success marker: a Friday deploy scares no one.&lt;/p&gt;

&lt;p&gt;📖 Full article: &lt;a href="https://mind.mansur.expert/en/infrastructure-and-deployment-order-by-iteration/" rel="noopener noreferrer"&gt;infrastructure and deployment — order by iteration&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Track 3. Architecture
&lt;/h2&gt;

&lt;p&gt;⬜ Draw the system "as is" — honestly, without prettifying&lt;br&gt;&lt;br&gt;
⬜ Find the boundaries: what is genuinely a separate service and what got glued together by accident&lt;br&gt;&lt;br&gt;
⬜ Start a decision log — short records, half a page: what was decided and why&lt;br&gt;&lt;br&gt;
⬜ Tech debt — an explicit list with a price tag: what it blocks, what it slows down, what a fix costs&lt;/p&gt;

&lt;p&gt;💬 The most common mistake is to start by rewriting. Start by writing down: half of the "architecture problems" dissolve once the system is described as it actually is.&lt;/p&gt;

&lt;p&gt;📖 Full article: &lt;a href="https://mind.mansur.expert/en/architecture-write-it-down-before-rewriting/" rel="noopener noreferrer"&gt;architecture — write it down before rewriting&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Track 4. People
&lt;/h2&gt;

&lt;p&gt;⬜ Bus factor from the Track 0 map: wherever knowledge lives in one head — pair people up or document first&lt;br&gt;&lt;br&gt;
⬜ Onboarding: working environment within a day, first production task within a week&lt;br&gt;&lt;br&gt;
⬜ Hiring: write the job post from the project's real tasks, not from a list of technologies&lt;br&gt;&lt;br&gt;
⬜ Explicit ownership: every service has a name next to it&lt;/p&gt;

&lt;p&gt;💬 Order among people matters more than order in code: code with a bus factor of 1 is not an asset, it's a hostage. And the other way around — a team where knowledge flows freely can fix any code.&lt;/p&gt;

&lt;p&gt;📖 Full article: &lt;a href="https://mind.mansur.expert/en/people-bus-factor-expectations-and-day-one/" rel="noopener noreferrer"&gt;people — bus factor, expectations and day one&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Track 5. Processes
&lt;/h2&gt;

&lt;p&gt;⬜ Planning that survives a week: a short cycle, a visible backlog&lt;br&gt;&lt;br&gt;
⬜ A release rhythm — predictable, boring, documented&lt;br&gt;&lt;br&gt;
⬜ Incidents: blameless reviews with one action item each&lt;br&gt;&lt;br&gt;
⬜ Meetings — audit them: each one either produces decisions or dies&lt;/p&gt;

&lt;p&gt;💬 Processes come last not because they don't matter, but because without light (Track 1) and hands on the wheel (Track 2) any process is theater.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to tell the order has arrived
&lt;/h2&gt;

&lt;p&gt;Not by pretty dashboards. By how the team feels:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;deploying is not scary,&lt;/li&gt;
&lt;li&gt;a new person is useful in their first week,&lt;/li&gt;
&lt;li&gt;an incident is a procedure, not a panic,&lt;/li&gt;
&lt;li&gt;the question "why is it like this here?" has a written answer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You'll never reach perfection — and you don't need to. It's enough that every week feels a little calmer than the last one.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This is the map of the series: links in the tracks will come alive as articles are published. Want a specific track covered sooner — tell me, the queue is flexible.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mind.mansur.expert/en/how-to-bring-order-to-a-project-track-map/" rel="noopener noreferrer"&gt;mind.mansur.expert&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>devops</category>
      <category>leadership</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
