<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Samson Tanimawo</title>
    <description>The latest articles on DEV Community by Samson Tanimawo (@samson_tanimawo).</description>
    <link>https://dev.to/samson_tanimawo</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3830227%2F02ea1ab7-513f-4426-b63d-9120142bc431.png</url>
      <title>DEV Community: Samson Tanimawo</title>
      <link>https://dev.to/samson_tanimawo</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/samson_tanimawo"/>
    <language>en</language>
    <item>
      <title>Building a Career in SRE: From Junior to Staff</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Wed, 12 Aug 2026 01:27:37 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/building-a-career-in-sre-from-junior-to-staff-4naa</link>
      <guid>https://dev.to/samson_tanimawo/building-a-career-in-sre-from-junior-to-staff-4naa</guid>
      <description>&lt;p&gt;I've been in SRE for about 10 years. Started as a junior, made my way through mid, senior, and now staff levels. Here's what I learned about growing in this career path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Junior → Mid
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The main shift:&lt;/strong&gt; learn to operate without constant guidance.&lt;/p&gt;

&lt;p&gt;As a junior, you'll be handed tickets and shown how to handle them. Your job is to do them well and ask questions. At mid-level, you should be able to take an ambiguous problem and produce a reasonable first attempt without help.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signs you're ready to move up:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You can handle a routine on-call shift without senior help&lt;/li&gt;
&lt;li&gt;You can write a post-mortem that doesn't need major revisions&lt;/li&gt;
&lt;li&gt;You're improving your team's processes without being asked&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What I wish I'd focused on:&lt;/strong&gt; fundamentals. Linux, networking, the tools your team actually uses. Not the fancy stuff. The boring deep knowledge pays off forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mid → Senior
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The main shift:&lt;/strong&gt; from executing well to choosing well.&lt;/p&gt;

&lt;p&gt;Mids know how to solve problems. Seniors know &lt;em&gt;which&lt;/em&gt; problems to solve. You stop asking 'how do I implement this' and start asking 'is this the right thing to implement.'&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signs you're ready:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're turning down work that doesn't make sense, with good reasons&lt;/li&gt;
&lt;li&gt;Other engineers come to you for design input&lt;/li&gt;
&lt;li&gt;You're anticipating failures, not just responding to them&lt;/li&gt;
&lt;li&gt;You own an area of the system without anyone telling you to&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What I wish I'd focused on:&lt;/strong&gt; writing and communication. The best senior SREs are also the best writers. Not fancy writing. Clear, boring, structured writing. It's the highest-leverage skill at this level.&lt;/p&gt;

&lt;h2&gt;
  
  
  Senior → Staff
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The main shift:&lt;/strong&gt; from individual contribution to organizational leverage.&lt;/p&gt;

&lt;p&gt;Seniors ship features. Staff engineers ship &lt;em&gt;teams&lt;/em&gt;. At staff level, your biggest impact is making other engineers more effective — through design leadership, process improvements, mentorship, and systems thinking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signs you're ready:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You spend more time in strategic work than tactical work&lt;/li&gt;
&lt;li&gt;Your ideas show up in other teams' roadmaps&lt;/li&gt;
&lt;li&gt;You're the person leadership asks to sanity-check big decisions&lt;/li&gt;
&lt;li&gt;You own organizational problems, not just technical ones&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What I wish I'd focused on:&lt;/strong&gt; saying no gracefully. Staff engineers who can't say no get pulled into everything. Staff engineers who can say no with good reasons become force multipliers.&lt;/p&gt;

&lt;h2&gt;
  
  
  The whole journey
&lt;/h2&gt;

&lt;p&gt;Here's the pattern: every level up is about expanding scope and reducing supervision. Junior: small scope, lots of help. Mid: medium scope, less help. Senior: large scope, peer-level. Staff: org-wide scope, shaping others.&lt;/p&gt;

&lt;p&gt;The biggest mistake is trying to be promoted by working harder at the current level. You get promoted by demonstrating the &lt;em&gt;next&lt;/em&gt; level's behavior before you're asked to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The quiet truth
&lt;/h2&gt;

&lt;p&gt;Most SRE career growth doesn't come from learning new technology. It comes from learning people, communication, and judgment. The technical floor gets you to mid-level. Everything above that is about humans.&lt;/p&gt;

&lt;p&gt;If you're stuck at senior and wondering why you're not hitting staff, it's almost certainly not a technical gap. It's a leverage or communication gap. Work on that. It changes everything.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>career</category>
      <category>growth</category>
    </item>
    <item>
      <title>Incident Automation: What to Automate, What to Leave to Humans</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Tue, 11 Aug 2026 13:19:18 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/incident-automation-what-to-automate-what-to-leave-to-humans-1ece</link>
      <guid>https://dev.to/samson_tanimawo/incident-automation-what-to-automate-what-to-leave-to-humans-1ece</guid>
      <description>&lt;p&gt;Incident response automation is a trap. Some things should be automated. Some things absolutely should not be. Getting the line wrong is worse than automating nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to automate
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Alert enrichment.&lt;/strong&gt; Before a human sees an alert, automate pulling in related data: recent deploys, dependent service health, historical correlation. Save the human 10 minutes of context-gathering.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Known-good remediations.&lt;/strong&gt; If an alert always has the same fix (restart service X, clear cache Y), automate the fix. But: require a human confirmation for the first 30 days before full auto.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Communication scaffolding.&lt;/strong&gt; When an incident starts, auto-create the Slack channel, invite the on-call, post an initial status template, update the status page with a placeholder. Humans fill in the details.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Post-incident paperwork.&lt;/strong&gt; Auto-generate a post-mortem template with timeline pulled from chat and monitoring. Humans edit and refine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Routine handoff.&lt;/strong&gt; When on-call rotates, auto-summarize what's been happening and who's been paged. Saves the incoming engineer 15 minutes of catching up.&lt;/p&gt;

&lt;h2&gt;
  
  
  What not to automate
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Root cause analysis.&lt;/strong&gt; AI or rules-based systems can suggest causes, but the final call has to be human. The 'wrong' root cause ends up in the post-mortem and misleads future responders.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Impact assessment.&lt;/strong&gt; 'How many users are affected' needs context only humans have. Automation will miss business-critical customer segments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Executive communication.&lt;/strong&gt; Your VP of customer success doesn't want a templated bot message. They want 'here's what's happening, here's what we're doing, here's when I'll update you next' from a human.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Deciding severity.&lt;/strong&gt; Yes, automate the initial guess. But a human has to confirm. Severity drives organizational response, and 'a bot marked this as sev-3' will be questioned the moment it matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. The critical decision during an outage.&lt;/strong&gt; Should we roll back? Should we fail over? Should we scale up? These are judgment calls with consequences. Don't hand them to a bot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The principle
&lt;/h2&gt;

&lt;p&gt;Automate the mechanical, human the judgmental. If the task has a clear right answer and no downside if it's wrong, automate. If it requires context, accountability, or judgment, leave it to humans.&lt;/p&gt;

&lt;h2&gt;
  
  
  The measurement
&lt;/h2&gt;

&lt;p&gt;After adding automation, ask two questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Are humans faster to resolution?&lt;/li&gt;
&lt;li&gt;Are humans feeling more in control, or less?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If resolution is faster but humans feel like they've lost visibility, you over-automated. Pull back.&lt;/p&gt;

&lt;h2&gt;
  
  
  The emotional piece
&lt;/h2&gt;

&lt;p&gt;Humans run incidents because humans are accountable. Automation that undermines that accountability creates downstream problems. Build tools that make humans faster and more confident, not tools that replace them. There is a big difference.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>automation</category>
      <category>incident</category>
    </item>
    <item>
      <title>Infrastructure Drift: Detecting and Preventing It</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Tue, 11 Aug 2026 01:45:12 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/infrastructure-drift-detecting-and-preventing-it-5gg4</link>
      <guid>https://dev.to/samson_tanimawo/infrastructure-drift-detecting-and-preventing-it-5gg4</guid>
      <description>&lt;p&gt;Your Terraform says the firewall rule is A. The cloud console says it's B. Someone changed it manually and didn't tell anyone. This is infrastructure drift, and it's a reliability killer.&lt;/p&gt;

&lt;h2&gt;
  
  
  How drift happens
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The emergency edit.&lt;/strong&gt; At 2 AM, someone clicks a button in the console to fix an outage. They promise to 'put it in Terraform later.' They don't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The helpful colleague.&lt;/strong&gt; Someone from a different team edits a shared resource because it was blocking them. They didn't know it was Terraform-managed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The console-first culture.&lt;/strong&gt; Some teams never adopt IaC. They edit everything manually, and then one day someone runs Terraform plan and sees 200 unexpected changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The partial migration.&lt;/strong&gt; You migrated some resources to Terraform but not others. Nobody documented which is which.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why drift is dangerous
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The next Terraform apply will either revert the manual change (surprise!) or fail because of conflicts.&lt;/li&gt;
&lt;li&gt;The manual change often has no audit trail. 'Why is this firewall rule open?' No answer.&lt;/li&gt;
&lt;li&gt;Disaster recovery becomes a guessing game. You can redeploy from Terraform, but you'll lose the manual edits.&lt;/li&gt;
&lt;li&gt;Security audits become impossible because your declared state doesn't match real state.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Detection
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Run &lt;code&gt;terraform plan&lt;/code&gt; on a schedule.&lt;/strong&gt; Daily or more. Alert on any unexpected diffs. The first time I set this up, we had 40 drift items in our main AWS account. Now we have zero because the alert forces cleanup within a day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cloud-native drift detection.&lt;/strong&gt; AWS Config, GCP Asset Inventory. These track changes at the cloud level and can flag manual edits independently of IaC tools.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Git as source of truth.&lt;/strong&gt; If your IaC is in Git and the cloud state differs, Git wins — either manually reconcile or explicitly accept the drift by updating the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prevention
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Remove manual edit permissions.&lt;/strong&gt; This is the nuclear option and the most effective. If engineers can't edit resources manually, drift stops happening. Instead, they have to write Terraform (or equivalent) to make changes.&lt;/p&gt;

&lt;p&gt;Yes, this slows down emergency responses. That's why you build fast-path deployment pipelines for IaC — commit, auto-approve, auto-apply within 5 minutes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write-only service accounts.&lt;/strong&gt; Grant humans read-only in prod. Only CI/CD can write. Breakglass accounts exist for emergencies but require audit trail and review.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cultural shift
&lt;/h2&gt;

&lt;p&gt;The hardest part isn't the tools. It's the mindset shift from 'the console is the source of truth' to 'Git is the source of truth.' Teams used to clicking in the console will resist. Give them better tooling (fast apply pipelines, good Terraform examples) and the resistance fades.&lt;/p&gt;

&lt;p&gt;You'll know you've made it when 'put it in Terraform first' becomes the default response to any infrastructure change request. Until then, drift will bite you.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>terraform</category>
      <category>iac</category>
    </item>
    <item>
      <title>The Engineer Who Owns Nothing: A Cautionary Tale</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Mon, 10 Aug 2026 13:21:02 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/the-engineer-who-owns-nothing-a-cautionary-tale-1n7d</link>
      <guid>https://dev.to/samson_tanimawo/the-engineer-who-owns-nothing-a-cautionary-tale-1n7d</guid>
      <description>&lt;p&gt;I'm going to tell you about an engineer I worked with. Call him Mark. Mark was talented, well-liked, and utterly ineffective. Here's what I learned from watching him.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Mark did
&lt;/h2&gt;

&lt;p&gt;Mark's technical skills were real. He wrote good code. He gave thoughtful design review comments. He spoke well in meetings.&lt;/p&gt;

&lt;p&gt;Mark's problem: he didn't own anything.&lt;/p&gt;

&lt;p&gt;He worked on whatever was in front of him. He fixed bugs in code he didn't write. He helped other engineers with their services. He never said 'this is mine.' Everything was 'we should figure out who handles that.'&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this was a problem
&lt;/h2&gt;

&lt;p&gt;When something broke in production, Mark would help debug — for a few minutes. Then he'd disengage because it 'wasn't really his area.' Nobody ever held him accountable because the code wasn't explicitly assigned to him.&lt;/p&gt;

&lt;p&gt;When planning happened, Mark didn't propose projects. He waited for projects to be assigned to him. Assigned projects rarely came because leaders couldn't predict what he'd actually commit to.&lt;/p&gt;

&lt;p&gt;When reliability work needed doing — alerts to tune, dashboards to fix, runbooks to write — Mark agreed it was important and waited for someone else to do it.&lt;/p&gt;

&lt;p&gt;Mark got good performance reviews for the first two years. He was technically capable, pleasant, and unobjectionable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;At year three, the company hit hard times. Leadership started asking: 'what has this person done? what do they own?' Mark had no answer.&lt;/p&gt;

&lt;p&gt;He got laid off. It wasn't a performance issue in the normal sense — he hadn't done anything wrong. He just hadn't owned anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lesson for SREs
&lt;/h2&gt;

&lt;p&gt;SRE is especially vulnerable to this trap. Everything is shared infrastructure. It's tempting to be the person who 'helps out everywhere' without ever claiming a specific service.&lt;/p&gt;

&lt;p&gt;Don't. Claim something.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Own an SLO&lt;/li&gt;
&lt;li&gt;Own a runbook library&lt;/li&gt;
&lt;li&gt;Own the post-mortem process&lt;/li&gt;
&lt;li&gt;Own the alerting hygiene for a service&lt;/li&gt;
&lt;li&gt;Own the capacity planning model&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It doesn't have to be big. But it has to be explicitly yours, with consequences if it fails and credit if it succeeds.&lt;/p&gt;

&lt;h2&gt;
  
  
  The meta-lesson
&lt;/h2&gt;

&lt;p&gt;Ownership isn't the same as being busy. You can be in a million meetings and own nothing. You can write 10,000 lines of code and own nothing. Ownership means: when something in your area breaks, you are the first person called, and you are the person who gets credit when it works.&lt;/p&gt;

&lt;p&gt;Pick something this week. Ask your manager: 'can I explicitly own X?' Write it down. Defend it.&lt;/p&gt;

&lt;p&gt;Don't be Mark.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>culture</category>
      <category>ownership</category>
    </item>
    <item>
      <title>Error Budget Policies That Hold Leadership Accountable</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Sun, 09 Aug 2026 23:39:50 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/error-budget-policies-that-hold-leadership-accountable-2hkc</link>
      <guid>https://dev.to/samson_tanimawo/error-budget-policies-that-hold-leadership-accountable-2hkc</guid>
      <description>&lt;p&gt;Error budgets are useless without a policy. 'We're out of error budget' should trigger consequences. If it doesn't, you don't have an error budget — you have a vanity metric.&lt;/p&gt;

&lt;p&gt;Here's a policy that actually works.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four states
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Healthy (&amp;lt; 70% of budget used).&lt;/strong&gt; Business as usual. Feature development proceeds at full speed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Watch (70-90% used).&lt;/strong&gt; Feature velocity continues but new risky changes require explicit sign-off from an SRE. No gate, just attention.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Constrained (90-100% used).&lt;/strong&gt; Feature freezes. Only reliability work and critical bug fixes until we're back below 90%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Breached (&amp;gt; 100% used).&lt;/strong&gt; Incident-level response. Leadership informed. Post-mortem for why we blew through. Feature work stays frozen until we recover &lt;em&gt;and&lt;/em&gt; identify systemic causes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part most policies miss
&lt;/h2&gt;

&lt;p&gt;The feature freeze in 'constrained' state is the part that actually changes behavior. Everything else is documentation. Without consequences, teams ignore the budget.&lt;/p&gt;

&lt;p&gt;The freeze has to be &lt;em&gt;real&lt;/em&gt;. Leadership can't override it for a 'really important feature' — that's exactly the time the freeze matters. The only exception is a legitimate emergency fix, and those should be rare.&lt;/p&gt;

&lt;h2&gt;
  
  
  Selling this to leadership
&lt;/h2&gt;

&lt;p&gt;Executives hate feature freezes. They see it as slowing the business. Counter-argument: feature freezes during budget exhaustion &lt;em&gt;protect&lt;/em&gt; the business. Shipping features onto broken infrastructure creates more breakage, which burns more budget, which is a doom loop.&lt;/p&gt;

&lt;p&gt;Frame it as: 'the feature freeze is a safety valve. When it triggers, it's because something's wrong and we need to fix it before making it worse.'&lt;/p&gt;

&lt;p&gt;Also: a good policy lets you spend the budget aggressively when you have it. Feature teams should be encouraged to experiment, deploy fast, and take risks when you're at 30% budget used. The freeze is only for when the safety margin is gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The review cadence
&lt;/h2&gt;

&lt;p&gt;Weekly error budget review, 15 minutes max. Who attended: SRE lead, engineering manager, maybe a PM. Decisions: are we in healthy/watch/constrained? Any actions for the coming week?&lt;/p&gt;

&lt;p&gt;Monthly broader review with leadership. Trends over time. Investment decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The escalation
&lt;/h2&gt;

&lt;p&gt;If a team enters 'constrained' state three times in a quarter, that's a systemic issue. Escalate to engineering leadership with a proposal: either invest in reliability or accept a lower SLO formally.&lt;/p&gt;

&lt;h2&gt;
  
  
  The endgame
&lt;/h2&gt;

&lt;p&gt;A mature organization uses error budget policy to balance feature velocity against reliability automatically. Nobody is negotiating individual decisions. The framework does the work.&lt;/p&gt;

&lt;p&gt;Getting there takes 6-12 months of discipline. The first few freezes will feel painful. After that, they become routine, and something surprising happens: you stop having them as often. The policy is working.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>slo</category>
      <category>leadership</category>
    </item>
    <item>
      <title>Dependency Injection for Observability</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Sun, 09 Aug 2026 13:42:41 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/dependency-injection-for-observability-1iol</link>
      <guid>https://dev.to/samson_tanimawo/dependency-injection-for-observability-1iol</guid>
      <description>&lt;p&gt;Want your code to be easy to observe? Use dependency injection for observability concerns. Sounds dry. Hear me out.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;Your code calls &lt;code&gt;log.info(...)&lt;/code&gt; directly. In tests, you can't verify what was logged. In prod, if you want to change the logger, you're grepping the codebase. If you want to add tracing, you're editing every call site.&lt;/p&gt;

&lt;p&gt;Same for metrics. Same for tracing. Same for error reporting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;Pass the observer in. Functions take a &lt;code&gt;logger&lt;/code&gt;, &lt;code&gt;metrics&lt;/code&gt;, or &lt;code&gt;tracer&lt;/code&gt; as an argument (or constructor dependency). The function doesn't know what's behind it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;HandleOrder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="n"&gt;Order&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;deps&lt;/span&gt; &lt;span class="n"&gt;Deps&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;deps&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Metrics&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Inc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"order.received"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;deps&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Logger&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"processing order"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ID&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In tests, pass in a mock that captures all calls&lt;/li&gt;
&lt;li&gt;In prod, pass in the real logger&lt;/li&gt;
&lt;li&gt;Changing backends (Datadog → Prometheus) is one wire-up change&lt;/li&gt;
&lt;li&gt;Adding tracing is one new field in &lt;code&gt;Deps&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why most codebases don't do this
&lt;/h2&gt;

&lt;p&gt;It feels verbose. 'Why do I have to thread a logger through every function?' Engineers hate boilerplate.&lt;/p&gt;

&lt;p&gt;The alternative is a global singleton. Easy to use, impossible to test cleanly, nightmare to refactor.&lt;/p&gt;

&lt;p&gt;The boilerplate is worth it. Especially for observability, where you &lt;em&gt;will&lt;/em&gt; want to swap implementations later.&lt;/p&gt;

&lt;h2&gt;
  
  
  The subtle win
&lt;/h2&gt;

&lt;p&gt;Dependency-injected observability forces you to think about &lt;em&gt;what&lt;/em&gt; you're observing. When you have to explicitly pass the logger, you notice that a function is calling it 8 times. Is that too much? Is the logging doing real work? Would one structured log at the end of the function be better?&lt;/p&gt;

&lt;p&gt;Functions with injected dependencies tend to have better observability because the developer had to look at it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The starting point
&lt;/h2&gt;

&lt;p&gt;You don't need to refactor everything at once. Pick one critical path — the checkout flow, the auth path. Refactor just that path to use dependency injection for observability. See if tests and debugging get easier.&lt;/p&gt;

&lt;p&gt;If yes, expand. If no, you found something else is the bottleneck.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger idea
&lt;/h2&gt;

&lt;p&gt;Observability is code. Treat it with the same architectural discipline you treat the rest of your codebase. Dependency injection is one tool. There are others (context objects, middleware, decorators). Pick one, apply it consistently, and your future-you will be able to observe your code in ways that are impossible with ad-hoc &lt;code&gt;log.info&lt;/code&gt; calls everywhere.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>observability</category>
      <category>patterns</category>
    </item>
    <item>
      <title>Load Balancer Tuning: Lessons from Production</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Sat, 08 Aug 2026 23:34:11 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/load-balancer-tuning-lessons-from-production-1k21</link>
      <guid>https://dev.to/samson_tanimawo/load-balancer-tuning-lessons-from-production-1k21</guid>
      <description>&lt;p&gt;Load balancers are the silent infrastructure. You don't think about them until they start dropping connections at 2 AM. Here are the settings that have bitten me in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Connection timeouts
&lt;/h2&gt;

&lt;p&gt;Default connection idle timeout on AWS ALB is 60 seconds. Your app might have requests that legitimately take 90 seconds. Result: the ALB drops the connection mid-request, your user sees a 502, and your logs show nothing because the app was still processing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule:&lt;/strong&gt; set the LB idle timeout higher than your longest legitimate request. For most APIs, 120-300 seconds. Verify it matches your app's own timeout.&lt;/p&gt;

&lt;h2&gt;
  
  
  Health check intervals
&lt;/h2&gt;

&lt;p&gt;Default health check interval is usually 30 seconds, with 2 failures before marking unhealthy. That's up to 60 seconds of traffic sent to a dying instance before it's removed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule:&lt;/strong&gt; 10-second intervals, 2 failures. 20 seconds to remove a bad instance is much better. Yes, slightly more load on your service. Worth it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Unhealthy threshold
&lt;/h2&gt;

&lt;p&gt;Marking an instance unhealthy after 2 consecutive failures is usually right. Marking it healthy again after 2 consecutive successes is usually wrong — it lets half-broken instances come back prematurely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule:&lt;/strong&gt; require 3-5 consecutive successes to mark healthy. Slower recovery, more stable routing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Slow start / connection draining
&lt;/h2&gt;

&lt;p&gt;When a new instance comes online, it's often cold — empty caches, no warmed connections. If the LB sends it full traffic immediately, it performs badly and might get marked unhealthy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule:&lt;/strong&gt; enable slow start (AWS calls it 'slow start mode' on ALB). Ramp traffic over 30-60 seconds. Worth it for any service that needs warming.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sticky sessions
&lt;/h2&gt;

&lt;p&gt;Sticky sessions feel like a solution. They're usually a problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule:&lt;/strong&gt; avoid them. If you need them, your app has shared state that should be in a database or cache, not in-process. The one exception: WebSocket connections, where stickiness is unavoidable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cross-zone load balancing
&lt;/h2&gt;

&lt;p&gt;Disabled by default on some LBs. With cross-zone disabled, an instance in AZ-A only serves traffic from AZ-A. If AZ-A has fewer instances, that AZ's traffic gets uneven distribution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule:&lt;/strong&gt; enable it unless you have a specific reason not to. Costs a small amount of cross-AZ traffic but gives you even load distribution.&lt;/p&gt;

&lt;h2&gt;
  
  
  The meta-lesson
&lt;/h2&gt;

&lt;p&gt;Load balancer defaults are designed for 'will work for everybody, badly.' Any given workload needs tuning.&lt;/p&gt;

&lt;p&gt;Spend one afternoon reviewing your LB configs. Ask: 'is this default right for my app?' At least half the defaults will be wrong. Tune them. Future you will thank you during the next incident.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>loadbalancer</category>
      <category>performance</category>
    </item>
    <item>
      <title>Capacity Planning for Startups</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Sat, 08 Aug 2026 13:32:21 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/capacity-planning-for-startups-1doe</link>
      <guid>https://dev.to/samson_tanimawo/capacity-planning-for-startups-1doe</guid>
      <description>&lt;p&gt;Capacity planning sounds like enterprise spreadsheet work. For a startup, it's 'don't get embarrassed when traffic spikes, don't go broke overprovisioning.' Here's the pragmatic version.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 3 questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. What does normal look like right now?&lt;/strong&gt; Peak RPS, p99 latency, CPU/memory utilization at peak. If you don't know these, stop and measure. You cannot plan without baseline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. What's the next expected spike?&lt;/strong&gt; A launch. A press mention. A marketing campaign. The Black Friday of your industry. Put these on a calendar.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. How long does it take to add capacity?&lt;/strong&gt; Minutes (autoscaling)? Hours (VM provisioning)? Weeks (vendor contracts)?&lt;/p&gt;

&lt;p&gt;Your capacity plan is the gap between expected spike and response time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The startup hack: overprovision early
&lt;/h2&gt;

&lt;p&gt;At startup scale, overprovisioning is cheap. An extra $5k/month of slack is trivial compared to the embarrassment of 'we went down during the launch.'&lt;/p&gt;

&lt;p&gt;Run at 30-40% peak utilization. Yes, that's wasteful. It's also a 3x buffer for unexpected spikes. Worth it until you're big enough to care about the efficiency.&lt;/p&gt;

&lt;h2&gt;
  
  
  The autoscaling reality
&lt;/h2&gt;

&lt;p&gt;Autoscaling is great but has limits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cold start times mean it can't handle traffic that doubles in 30 seconds&lt;/li&gt;
&lt;li&gt;Provisioning limits mean you can only add X instances per minute&lt;/li&gt;
&lt;li&gt;Downstream dependencies (databases, queues) usually don't autoscale&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Test your autoscaling &lt;em&gt;before&lt;/em&gt; you need it. The first time I depended on autoscaling in production, it worked. The second time, it didn't, because our database connection pool was capped. That was the real bottleneck.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 3 bottlenecks to check
&lt;/h2&gt;

&lt;p&gt;For every scaling test, verify:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Stateless service capacity.&lt;/strong&gt; Usually easy — just add more instances.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database capacity.&lt;/strong&gt; Connection counts, query latency, replication lag. Usually the real bottleneck.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Third-party dependencies.&lt;/strong&gt; Rate limits on external APIs, email providers, payment processors. A sudden 10x spike usually hits someone's rate limit.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The launch checklist
&lt;/h2&gt;

&lt;p&gt;Before any planned spike:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Overprovision by 3x what you think you need&lt;/li&gt;
&lt;li&gt;Pre-warm caches and connection pools&lt;/li&gt;
&lt;li&gt;Confirm your paging rotation is ready&lt;/li&gt;
&lt;li&gt;Prepare a rollback plan for the feature being launched&lt;/li&gt;
&lt;li&gt;Schedule the launch during your team's awake hours, not off-hours&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The real lesson
&lt;/h2&gt;

&lt;p&gt;For startups, the goal of capacity planning is not efficiency. It's confidence. If you have to spend a little more to avoid panic during growth, spend it. Optimize for efficiency later, when you have a year of traffic history to work from.&lt;/p&gt;

&lt;p&gt;Right now, your job is to stay standing. Do that first.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>capacity</category>
      <category>startup</category>
    </item>
    <item>
      <title>How We Handled Our First Major Outage (And Survived)</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Sat, 08 Aug 2026 01:27:13 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/how-we-handled-our-first-major-outage-and-survived-3nap</link>
      <guid>https://dev.to/samson_tanimawo/how-we-handled-our-first-major-outage-and-survived-3nap</guid>
      <description>&lt;p&gt;Three years ago we had our first real outage. Six hours of downtime. Thousands of angry users. Multiple executives on the call. Here's what we did right, what we did wrong, and what we'd do differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we did right
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Communicated immediately.&lt;/strong&gt; The moment we knew we had a problem, we updated the status page and emailed our biggest customers personally. Not when we had answers. When we had a question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Had a single incident commander.&lt;/strong&gt; One person making calls. Not a committee. When the CEO tried to direct technical work, the IC politely rerouted and told her where her help was actually needed (talking to customers).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Took care of our people.&lt;/strong&gt; During hour 4, I ordered food. During hour 5, I forced the primary engineer off the call for 20 minutes to walk outside. Long incidents destroy people. You have to feed them and force them to rest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Wrote it down as we went.&lt;/strong&gt; We had a shared doc with a live timeline. When the post-mortem came, we had every decision captured.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we did wrong
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Tried to fix the root cause during the incident.&lt;/strong&gt; For the first 2 hours, we were digging into &lt;em&gt;why&lt;/em&gt; the database was struggling. We should have been mitigating (rolling back) first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Let too many people 'help.'&lt;/strong&gt; By hour 3, we had 12 engineers in the call. Half of them were useless. The IC should have kicked people out sooner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Gave optimistic estimates.&lt;/strong&gt; 'We'll be back in 30 minutes.' We were not back in 30 minutes. That miscommunication was worse than saying 'unknown.'&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Didn't prepare the executive communication.&lt;/strong&gt; The CEO had to answer customer questions in real time with no script. We should have drafted talking points for her after hour 1.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we'd do differently
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Mitigate first, investigate second. Always.&lt;/li&gt;
&lt;li&gt;Cap the number of active engineers at 4 during an incident. Others go on standby.&lt;/li&gt;
&lt;li&gt;Default to 'unknown' for estimates. Only give a number when we're sure.&lt;/li&gt;
&lt;li&gt;Assign someone explicitly to 'executive liaison.' Their job is to keep the C-suite informed without interrupting the technical team.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The aftermath
&lt;/h2&gt;

&lt;p&gt;The post-mortem was brutal and cathartic. We identified 14 action items. We actually did 11 of them over the next quarter.&lt;/p&gt;

&lt;p&gt;The outage was the best thing that happened to our reliability culture. It turned reliability from 'a thing SRE owns' into 'a thing everyone takes seriously.' I wouldn't wish a 6-hour outage on anyone, but I also wouldn't trade the lessons.&lt;/p&gt;

&lt;h2&gt;
  
  
  The final lesson
&lt;/h2&gt;

&lt;p&gt;Your first major outage will happen. Prepare for it by running game days. The game days will feel silly until the real thing happens, at which point every muscle you trained will kick in.&lt;/p&gt;

&lt;p&gt;Incident response is a skill. Skills need practice. Practice now.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>incident</category>
      <category>culture</category>
    </item>
    <item>
      <title>The Economics of Reliability: When to Invest, When to Accept Risk</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Fri, 07 Aug 2026 13:19:02 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/the-economics-of-reliability-when-to-invest-when-to-accept-risk-19hg</link>
      <guid>https://dev.to/samson_tanimawo/the-economics-of-reliability-when-to-invest-when-to-accept-risk-19hg</guid>
      <description>&lt;p&gt;Reliability is not a virtue. It's an investment. Too little and you lose customers. Too much and you can't afford to ship. The question is: where's the right balance?&lt;/p&gt;

&lt;h2&gt;
  
  
  The error budget framing
&lt;/h2&gt;

&lt;p&gt;The SRE book covers this well. Pick an SLO (say 99.9% uptime). That's 43 minutes of budget per month. If you're at 99.95%, you have budget to spend on risky things. If you're at 99.85%, you need to stop shipping risk.&lt;/p&gt;

&lt;p&gt;This works. But it doesn't answer 'what SLO should I pick?' Let me give you a framework.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. What do users expect?&lt;/strong&gt; A consumer banking app needs 4 nines or more. A developer tool can get away with 3. A beta product can live with 99%. Ask users (or watch churn numbers).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. What does an outage cost?&lt;/strong&gt; Dollars of lost revenue + dollars of customer churn + hours of engineering time. For a checkout-heavy product, an hour of downtime might cost $500k. For a B2B internal tool, $5k.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. What does the next 9 cost?&lt;/strong&gt; Going from 99% to 99.9% might cost $50k of engineering work. Going from 99.9% to 99.99% often costs $500k or more. Each 9 is 10x harder.&lt;/p&gt;

&lt;h2&gt;
  
  
  The math
&lt;/h2&gt;

&lt;p&gt;Invest in reliability up to the point where the next $1 invested saves less than $1 of outage cost over the amortization period.&lt;/p&gt;

&lt;p&gt;If moving from 99% to 99.9% costs $50k and would save $200k over a year in reduced outage damage, invest. Easy call.&lt;/p&gt;

&lt;p&gt;If moving from 99.9% to 99.99% costs $500k and saves $100k, don't invest. Accept the risk, and spend the engineering time on something with better ROI.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hidden cost
&lt;/h2&gt;

&lt;p&gt;Over-investing in reliability has a hidden cost: team velocity. Teams that chase 99.999% uptime spend so much on tests, canaries, staging environments, and approval gates that they ship slowly. Competitors with 99.5% reliability but 5x your velocity will win the market.&lt;/p&gt;

&lt;p&gt;Reliability that kills velocity is bad reliability. Measure both.&lt;/p&gt;

&lt;h2&gt;
  
  
  The political reality
&lt;/h2&gt;

&lt;p&gt;The hardest part isn't the math. It's defending 'we're not going to fix this' when a VP demands reliability improvements. You need explicit agreement, in writing, on the target SLOs — so 'we're not fixing this' is 'we agreed on 99.9% and we're at 99.92%, which is within budget.'&lt;/p&gt;

&lt;h2&gt;
  
  
  The pragmatic answer
&lt;/h2&gt;

&lt;p&gt;Most teams I've worked with should pick:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;99.9% for production-critical services&lt;/li&gt;
&lt;li&gt;99% for internal tools&lt;/li&gt;
&lt;li&gt;Best-effort for dev environments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then measure, spend the error budget, and stop arguing about it. The math is usually clearer than the politics.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>reliability</category>
      <category>strategy</category>
    </item>
    <item>
      <title>Why Your Status Page Should Be Boring</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Fri, 07 Aug 2026 01:28:56 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/why-your-status-page-should-be-boring-m9i</link>
      <guid>https://dev.to/samson_tanimawo/why-your-status-page-should-be-boring-m9i</guid>
      <description>&lt;p&gt;A good status page is boring. Calm design, minimal copy, clear current state. If your status page feels exciting, something is wrong.&lt;/p&gt;

&lt;p&gt;Here's what I've learned from running status pages for three different products.&lt;/p&gt;

&lt;h2&gt;
  
  
  What users actually want from a status page
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Is it me or is it you?&lt;/strong&gt; The #1 question. Answer it in the first 3 seconds of landing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. If it's you, what exactly is broken?&lt;/strong&gt; Not 'we're experiencing issues.' Specifically: 'API endpoint /v2/checkout is returning 500 errors for ~15% of requests.'&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. How long until it's fixed?&lt;/strong&gt; Even 'unknown' is better than no estimate. 'Investigating' with a last-updated timestamp beats silence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Should I keep retrying?&lt;/strong&gt; If you're broken and expect to stay broken for a while, tell users to back off. Your support queue will thank you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What users don't want
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Corporate-speak ('we're aware of a potential service degradation')&lt;/li&gt;
&lt;li&gt;Vague promises ('working to resolve as quickly as possible')&lt;/li&gt;
&lt;li&gt;Technical jargon they can't parse&lt;/li&gt;
&lt;li&gt;Delayed acknowledgments (updating the page 20 minutes after an outage)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The update cadence rules
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Acknowledge within 5 minutes of detection&lt;/li&gt;
&lt;li&gt;Update every 15-30 minutes during active investigation&lt;/li&gt;
&lt;li&gt;Mark as monitoring as soon as mitigation is in place, even if cause is unknown&lt;/li&gt;
&lt;li&gt;Mark as resolved when you're confident it's fixed — not before&lt;/li&gt;
&lt;li&gt;Always do a final post-incident summary&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The hard part: being honest
&lt;/h2&gt;

&lt;p&gt;The temptation is to minimize language. 'A small number of users' when it's actually 20%. 'Minor issue' when it's a real outage.&lt;/p&gt;

&lt;p&gt;Don't. Users trust a status page that's honest with them. The first time you get caught minimizing, you've lost credibility that takes years to earn back.&lt;/p&gt;

&lt;h2&gt;
  
  
  The automation trap
&lt;/h2&gt;

&lt;p&gt;Automated status pages that say 'all systems operational' while your product is clearly broken are worse than no page. Users lose trust in the entire signal.&lt;/p&gt;

&lt;p&gt;If you automate, automate the detection to trigger human review. Don't automate the reassurance.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a boring status page looks like
&lt;/h2&gt;

&lt;p&gt;Uptime over 30 days. Current state of each service (green/yellow/red). A list of recent incidents with their post-mortems linked.&lt;/p&gt;

&lt;p&gt;That's it. No marketing copy. No animated elements. Boring.&lt;/p&gt;

&lt;p&gt;Boring is trustworthy. Make your status page boring.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>statuspage</category>
      <category>communication</category>
    </item>
    <item>
      <title>Building Trust with Product Teams as an SRE</title>
      <dc:creator>Samson Tanimawo</dc:creator>
      <pubDate>Thu, 06 Aug 2026 13:09:42 +0000</pubDate>
      <link>https://dev.to/samson_tanimawo/building-trust-with-product-teams-as-an-sre-9h1</link>
      <guid>https://dev.to/samson_tanimawo/building-trust-with-product-teams-as-an-sre-9h1</guid>
      <description>&lt;p&gt;SRE teams that fight with product teams don't get things done. SRE teams that get along with product teams get surprising amounts of reliability work done by product engineers themselves. Here's how to build that trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start by making them faster, not slower
&lt;/h2&gt;

&lt;p&gt;The common SRE pattern: introduce yourself by adding gates. 'You need to do X before you can deploy.' 'Your service needs Y before production.' Product teams immediately see you as friction.&lt;/p&gt;

&lt;p&gt;Better pattern: introduce yourself by removing friction. 'I noticed your CI takes 20 minutes. I can get it to 6. Interested?' Now you're useful. The gates come later, and you have trust to spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speak their language
&lt;/h2&gt;

&lt;p&gt;Product engineers care about: shipping, feature quality, user complaints. They don't care about: error budgets, SLIs, observability stacks.&lt;/p&gt;

&lt;p&gt;Translate. Instead of 'we're over our SLO,' say 'users are seeing errors on checkout — here's what's hitting them and what it's costing.' The facts are the same. The reception is totally different.&lt;/p&gt;

&lt;h2&gt;
  
  
  Share ownership of incidents
&lt;/h2&gt;

&lt;p&gt;When a product team's code breaks prod, resist the urge to fix it yourself. Instead: be in the room, coach them through it, let them own the fix.&lt;/p&gt;

&lt;p&gt;Yes, it's slower. Yes, sometimes they'll ask awkward questions. That's exactly the point. They're learning. After 3 incidents, they'll be better engineers and more grateful to your team.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give credit publicly
&lt;/h2&gt;

&lt;p&gt;When a product team does reliability work well, say so. Publicly. In eng all-hands. On the CEO's Slack thread.&lt;/p&gt;

&lt;p&gt;This sounds performative. It's not. It's you saying 'reliability is valued here, and we notice when people invest in it.' Other teams see it and start wanting credit too. You're using recognition as a reliability multiplier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Absorb the hit sometimes
&lt;/h2&gt;

&lt;p&gt;Sometimes a product team will say 'we don't have time for the runbook right now, can you just do it?' Say yes. Write the runbook. Don't make a thing of it.&lt;/p&gt;

&lt;p&gt;Do this too often and you're a service team. Do this never and you're hostile. Do this sometimes, strategically, and you're building long-term trust. Read the room.&lt;/p&gt;

&lt;h2&gt;
  
  
  The end goal
&lt;/h2&gt;

&lt;p&gt;After a year of this, product teams start coming to you &lt;em&gt;before&lt;/em&gt; launches. 'We're about to ship X — any reliability concerns?' That's when you know you've won. They see you as a partner, not a checkpoint.&lt;/p&gt;

&lt;p&gt;SRE culture work is slow. It compounds. Invest in it from day one.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Written by Dr. Samson Tanimawo&lt;/strong&gt;&lt;br&gt;
BSc · MSc · MBA · PhD&lt;br&gt;
Founder &amp;amp; CEO, Nova AI Ops. &lt;a href="https://novaaiops.com" rel="noopener noreferrer"&gt;https://novaaiops.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>culture</category>
      <category>collaboration</category>
    </item>
  </channel>
</rss>
