<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: David Yanacek</title>
    <description>The latest articles on DEV Community by David Yanacek (@dyanacek).</description>
    <link>https://dev.to/dyanacek</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4099263%2F74921cc6-bd89-4c85-8bbf-0ee3cce44c4f.jpeg</url>
      <title>DEV Community: David Yanacek</title>
      <link>https://dev.to/dyanacek</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dyanacek"/>
    <language>en</language>
    <item>
      <title>It was never about coding</title>
      <dc:creator>David Yanacek</dc:creator>
      <pubDate>Tue, 22 Sep 2026 16:00:00 +0000</pubDate>
      <link>https://dev.to/dyanacek/it-was-never-about-coding-3gof</link>
      <guid>https://dev.to/dyanacek/it-was-never-about-coding-3gof</guid>
      <description>&lt;p&gt;This is part one of a two-part post, split so it isn't unwieldy. This part covers the &lt;em&gt;what&lt;/em&gt; and &lt;em&gt;why&lt;/em&gt;; part two will cover the &lt;em&gt;how&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Furt5vjiqf64wfqn43yxx.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Furt5vjiqf64wfqn43yxx.jpeg" alt="Astronaut meme: one astronaut looks at Earth and asks 'so it's not about the code?', the other, pointing a gun, replies 'never has been'" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;"With the AI doing the coding, do we just spend all day reviewing its output?" The rapid onset of impressively capable coding agents continues to raise this question in conversations with peers, to &lt;a href="https://dev.to/dyanacek/coding-after-coders-the-end-of-computer-programming-as-we-know-it-3k44-temp-slug-4613962"&gt;press interviews&lt;/a&gt;, to &lt;a href="https://dev.to/dyanacek/new-podcast-episode-of-code-with-jason-with-david-3dcf-temp-slug-3433673"&gt;podcasts&lt;/a&gt;. I remain an optimist, convinced that this is the magical technology I've always wanted - a way to get more done, in less time, with &lt;em&gt;higher&lt;/em&gt; polish than before. I can reach further into my backlog than ever before, reaching for the nice-to-have tasks that were triaged out of every sprint plan.&lt;/p&gt;

&lt;p&gt;While I used to spend a lot of time interacting directly with code - writing, reviewing, debugging, fixing, refactoring, and operating (dealing with the consequences of mistakes), I have always viewed that as only part of my job: a means to an end, not the end. So then what is my job?&lt;/p&gt;

&lt;h2&gt;
  
  
  What's in a name?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwgp6bymghcfah6veofh3.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwgp6bymghcfah6veofh3.jpeg" alt="A single long-stemmed red rose with green leaves on a white background" width="800" height="1165"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When I tell people what I do, sometimes I say " &lt;strong&gt;programmer&lt;/strong&gt;" to have a short, understandable, and non-pretentious sounding description of my job to people who don't code. But that's not my job title. The thought of being called a " &lt;strong&gt;coder&lt;/strong&gt;" somehow feels almost derogatory, because of how shallow or superficial it is as a description of what I do. It's one of the things I was hired to do, but it's not the full picture at all, often not the hard part, and not my job title. So what is my job?&lt;/p&gt;

&lt;p&gt;What about &lt;strong&gt;Software Architect&lt;/strong&gt;? Well I'm not a Software Architect. I love hearing Real Architects describe the concepts around buildings - their design, aesthetic, married with function and usability for the people who occupy them, with adaptability toward future unknown use. My friend of 23 years and college roommate for 3 of those, Steve, is an architect, and whenever we are near or in a building, I love it when he goes on an explanation of the things about buildings that are not obvious to a non-architect like me. He describes the customer - the people in or around the building - and how the design is meant to be intuitive, helpful, but grounded in the reality of the tradeoffs of physics and cost to build. A lot of the Architect title sounds like what I do day to day when it comes to designing systems and how they meet my customer's needs, but still, Software Architect isn't my job.&lt;/p&gt;

&lt;p&gt;Some people describe the job as &lt;strong&gt;Software Engineering&lt;/strong&gt;. While I'm not technically an engineer from a &lt;a href="https://engineerscanada.ca/become-an-engineer/use-of-professional-title-and-designations" rel="noopener noreferrer"&gt;professional license standpoint&lt;/a&gt;, there are plenty of parallels to what real engineers do. For example, Civil Engineers don't &lt;em&gt;build&lt;/em&gt; bridges. They design, they plan, they model, they test, they measure, they feel the weight of responsibility for safety and resilience, scale, and they reflect on the process that they go through along the way to make sure it's right. That same friend Steve also has a Civil Engineering degree in addition to being an Architect, which is a &lt;a href="https://memes.yarn.co/yarn-clip/34450659-0de3-430e-90e4-dc837c3459e6" rel="noopener noreferrer"&gt;hell of a combination&lt;/a&gt;. With both architecture and engineering degrees, he works across both worlds: the artistic, focusing on the beauty of design and the lives of the people who occupy the buildings, and also the unglamorous but life-safety behind-the-scenes reality of what it takes to prove that a design can be brought into reality. While I can draw connections to the testing, simulation, and planning work I do in the digital world, Software Engineer is not my job either.&lt;/p&gt;

&lt;p&gt;So what &lt;em&gt;is&lt;/em&gt; my job? My job title is &lt;strong&gt;Software Development Engineer&lt;/strong&gt;. Isn't that the same thing as Software Engineer? Well it has an extra word in there, so I guess not. That extra word, "Development", is a process. The "Engineer" term is a modifier on "Development". So that means I'm engineering the &lt;em&gt;process of&lt;/em&gt; software development. But that sounds like a pedantic unimportant difference, like Michael Scott correcting Dwight Schrute's title as being &lt;a href="https://www.youtube.com/watch?v=ozMs8Np84Ew" rel="noopener noreferrer"&gt;Assistant &lt;em&gt;TO&lt;/em&gt; the Regional Manager&lt;/a&gt;. But if I think about what I've been doing for the last 20 years, the work I've done around process has been more lasting and impactful than the code. For a given project, I'm working on design, implementation, testing, and operations. But the most impactful and durable parts of that work are the in-betweens where I'm making those things easier the next time around. If something didn't go quite right the first time around, I make sure that I, and my team, or organization, doesn't make that same mistake or hit that roadblock the next time around. Essentially, when instead of engineering the software, I'm engineering the &lt;em&gt;process&lt;/em&gt; of engineering the software.&lt;/p&gt;

&lt;p&gt;Of course this distinction of "SDE" vs "SWE" is a convenient play on words. Even though role names and responsibility vary across companies, I think these roles are actually equivalent and all focus on all of these aspects, from designing to building, to maintaining. And all throughout that, people in these roles spend time and effort trying to improve their process to take things they've learned and somehow "codify" for themselves and for their teams so that things are better the next time around. (Wait, the word "codify" also looks similar to "coding"! Foreshadowing!)&lt;/p&gt;

&lt;h2&gt;
  
  
  What does software development look like today?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7woarlkjlicew4x1cdq7.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7woarlkjlicew4x1cdq7.jpeg" alt="A sprawling automated factory of pipes, tanks, and machinery" width="512" height="512"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Building software today is a lot like an automated factory. An idea gets turned into a requirements and design, which gets turned into code, which gets tested, merged, and deployed. But it's a little different than a linear assembly line, where a conveyor belt moves a product (or in this case a "feature") from one stage to the next. Like a conveyor belt style, this one has tons of quality gates. When a feature, design choice, or important decision needs to be addressed by a human, the feature is &lt;strong&gt;set aside&lt;/strong&gt; until someone answers, and then the feature resumes its path on the assembly line. In the meantime, the rest of the features move forward on the assembly line. But maybe unlike an assembly line, this one has &lt;strong&gt;loops&lt;/strong&gt;. The in-progress code change gets sent back for rework a ton of times until the tests pass and all of the quality gates approve it. It also has &lt;strong&gt;merges&lt;/strong&gt; where features (commits) come together into the mainline branch, and they're further tested and released together. If these tests fail, there's a different loop in the assembly line that figures out what to do about the failure. It &lt;strong&gt;triages&lt;/strong&gt; , and figures out which commit broke the tests or the deployment, and finds the best way to unblock the assembly line. It might either rush a patch while more changes pile up, or find and kick out the offending change and let the rest of the feature development advance.&lt;/p&gt;

&lt;p&gt;In a sense, this is a lot like a normal agentic loop. In a nutshell, agents are super simple - a while() loop that tries something, sees if it's done, and if not, tries something a little different. But it's more complicated than that. We can't just take a coding agent and say, "Hey design and build this feature. Make no mistakes." The agent won't follow my rules, might make decisions or assumptions I don't want it to make, could break my application, or any number of things. So this factory is all about having the right deterministic checks and gates in place so that it can't proceed until its work is correct (verified by tests and testing agents, and sometimes humans), and done the way I want it (verified by agents and sometimes humans).&lt;/p&gt;

&lt;p&gt;The machinery to drive the factory is rapidly changing, with new tools sprouting up every day. It seems like every 3rd person I meet has made their own. There are a ton of open source ones out there, from &lt;a href="https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04" rel="noopener noreferrer"&gt;Gas Town&lt;/a&gt; to more unopinionated building-block tools like &lt;a href="https://kiro.dev/crew/" rel="noopener noreferrer"&gt;Kiro Crew&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;So where do I fit into this as an SDE? When coding agents first came out, it started with me pressing "yes, approve this tool use", and copy-pasting error messages from one tool to the next. Sure it was an improvement over doing everything myself sometimes, compared to copy-pasting the error into stack overflow and fixing the error myself. But it wasn't a huge step function multiplier advancement in productivity.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fczdqossj38t4uourttvc.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fczdqossj38t4uourttvc.gif" alt="A drinking bird toy bobbing down to peck a keyboard key over and over" width="480" height="360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;To get that big improvement, I - and my team - had to change from babysitting the agents to nurturing them. We spend our time focusing on building and tuning the machinery, so that we can be in the way of the assembly line only when we truly want and need to be to make sure the right things happen.&lt;/p&gt;

&lt;p&gt;Unlike physical factories that make physical things, software factories are easy to change. There's no need to move heavy machinery around the factory floor. Improving the factory can be as simple as adding a bullet point to an AGENTS.md file. Or better yet, having an agent run periodically to find rough spots in the process from looking at past runs, and propose changes to the factory code and configuration to improve it.&lt;/p&gt;

&lt;h2&gt;
  
  
  So then what is my job?
&lt;/h2&gt;

&lt;p&gt;Every day I'm defining, designing, evaluating, and shipping things for my customers. I have more time than before to listen to their needs, respond to their questions. I have more time to be curious and dig in about the operational performance of my systems, and look around corners. But the biggest flywheel comes from continuously improving the software factory.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.clareliguori.com/writing" rel="noopener noreferrer"&gt;Clare Liguori&lt;/a&gt; wrote the &lt;a href="https://kiro.dev/topics/frontier-engineering/" rel="noopener noreferrer"&gt;Frontier engineering manifesto&lt;/a&gt; that breaks down how we at Amazon approach software development in this new world. It describes a shift in mindset and effort, toward focusing on the process of making sure the coding agents are held to the standards that you set for yourself. It's a good read. Principal Engineers across Amazon weighed in on it, but Clare is the mastermind who brought it all together in this compact and dense form. Here's a short version.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You are the architect, not the typist.&lt;/strong&gt; Write intent, direction, and targeted review of what you want.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maximize agent time, minimize your involvement.&lt;/strong&gt; Rethink when you need to be involved in a coding agent's session. Challenge yourself to have multiple sessions running in parallel, and structure their tasks so they won't constantly nag.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build your codebase for agents.&lt;/strong&gt; A new human team member onboards to your team once and maybe every few months or years. With agents, you have team members onboarding 1,000 times a day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give agents a fast feedback loop.&lt;/strong&gt; Give your agents the tools they need to run fast, local tests that are as end-to-end as possible. Give the agent a web browser tool to do its own evaluation of the interfaces it touches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execution is cheap. Direction is everything.&lt;/strong&gt; Collaborate back and forth with the agent at the important points, like on design and definition.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat code as disposable.&lt;/strong&gt; It's way easier to discard and replace than ever before.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hold AI output to human standards.&lt;/strong&gt; With a well designed factory, you can get higher quality code than before.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trust the boundaries, not the agent.&lt;/strong&gt; Put the agent in a sandbox with a thoughtful boundary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use agents for everything, not just code.&lt;/strong&gt; I even used an agent to &lt;a href="https://dev.to/dyanacek/amazon-quick-found-my-lost-coffee-mug-35k7-temp-slug-1660558"&gt;find my lost coffee mug&lt;/a&gt;, which was missing for weeks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Continuously tune your agent setup.&lt;/strong&gt; I view every steering I have to do as a possible defect in my factory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As a Software &lt;em&gt;Development&lt;/em&gt; Engineer, I build and tune the factory. I add review agents and improve their steering, direct the agent to write better tests, review their output (but strategically at just the right time and just the right way), look at the process to see where I'm wasting time or where the agent is getting stuck. But what does this mean? The Frontier engineering manifesto describes the types of things I spend time on at a high-level, but for blog form, stay tuned for part two!&lt;/p&gt;

</description>
      <category>software</category>
      <category>ai</category>
      <category>coding</category>
      <category>programming</category>
    </item>
    <item>
      <title>Hello, world!</title>
      <dc:creator>David Yanacek</dc:creator>
      <pubDate>Fri, 28 Aug 2026 16:37:05 +0000</pubDate>
      <link>https://dev.to/dyanacek/hello-world-ik9</link>
      <guid>https://dev.to/dyanacek/hello-world-ik9</guid>
      <description>&lt;p&gt;I'm late to the party! I just found out about dev.to thanks to &lt;a class="mentioned-user" href="https://dev.to/hacksore"&gt;@hacksore&lt;/a&gt; . Seems neat! I've been writing content on my &lt;a href="https://blog.dyanacek.com/" rel="noopener noreferrer"&gt;personal blog&lt;/a&gt;, and the &lt;a href="https://aws.amazon.com/builders-library/" rel="noopener noreferrer"&gt;Amazon Builders' Library&lt;/a&gt;, but maybe this is the place to be now too. I don't have a comment feed on my blog or anything so maybe I'll cross-post blog posts here so we can chat about things.&lt;/p&gt;

</description>
      <category>discuss</category>
      <category>webdev</category>
    </item>
    <item>
      <title>DuneOps Part 2: Crossing bloodlines</title>
      <dc:creator>David Yanacek</dc:creator>
      <pubDate>Sat, 16 May 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/dyanacek/duneops-part-2-crossing-bloodlines-3fi3</link>
      <guid>https://dev.to/dyanacek/duneops-part-2-crossing-bloodlines-3fi3</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Finod5ui4tw6zhur1jwbd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Finod5ui4tw6zhur1jwbd.png" alt="Paul Atreides with blue Fremen eyes in Dune Part Two" width="800" height="442"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In last week's &lt;a href="https://dev.to/dyanacek/dune-as-an-operational-model-om1"&gt;Dune as an operational model&lt;/a&gt; post, I combined two of my obsessions — Dune and Ops — into one allegory. This week, I talk about how we can evolve beyond the existing operational models, thanks to AI. Which of course is ironic because in Dune, society rejected AI through the Butlerian Jihad. But anyway...&lt;/p&gt;

&lt;p&gt;As a short recap to what became somewhat of a manifesto, I talk about how harsh environments like the planet Arrakis lead to hardened fighters. In ops, a hands-on environment leads to better operators. And in operational models where those same people are empowered to automate and solve their problems, and given the space to be altruistic and expand their tools for others to use, that yields the best operational outcome for customers and the most efficiency for a company.&lt;/p&gt;

&lt;p&gt;And yes, as I write this I'm listening to the soundtrack from Dune Part 2. Thank you Hans Zimmer for this beautiful gift to the world.&lt;/p&gt;

&lt;h2&gt;
  
  
  Platform Engineering: The Next Generation
&lt;/h2&gt;

&lt;p&gt;This post looks to the future of how AI transforms Platform Engineering teams in exciting ways.&lt;/p&gt;

&lt;p&gt;Of course AI helps all of the models. The latest service I've been working on — &lt;a href="https://aws.amazon.com/blogs/mt/announcing-general-availability-of-aws-devops-agent/" rel="noopener noreferrer"&gt;AWS DevOps Agent&lt;/a&gt; — is being used by teams of all configurations, from customer support teams, to frontline ops, to SRE, to DevOps, and even Platform Engineering. It troubleshoots typical alarm-driven "find the root cause", but also helps with adhoc ops like researching a specific customer's reported issue, to scanning across teams' stuff for best practices, like finding missing or misconfigured alarm configuration.&lt;/p&gt;

&lt;p&gt;But in Platform Engineering, the thing that AI can help with is particularly interesting to me. A Platform Engineering team is one who builds tools that help other teams build, test, deploy, secure, and operate their systems with lower effort and with better outcomes. These teams try to push the ownership boundary as much as possible to simplify things for other teams. But it's hard to do and is a balancing act. Each team's environment varies just enough that building software to automate the act of applying, configuring, and using a tool in every bespoke environment is simply impractical.&lt;/p&gt;

&lt;p&gt;I've worked on a team like this before. We built the web service framework that does all of the request/reply serialization, authentication, admission control, and that kind of stuff for teams across Amazon and AWS. By defining the API model, we could expand that out to take care of all that stuff, and even generate SDKs in a bunch of languages, all automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trading off solving the problem for customization
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh8rzyv08yis5swnfall0.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh8rzyv08yis5swnfall0.webp" alt="" width="800" height="531"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But there was only so far we could go with the abstraction. Sure, we had features like caching that people could configure, but we couldn't decide to add caching automatically without knowing more about the guarantees behind their APIs and the expectations of their clients. If we wanted to pursue that, we could provide a sort of "configuration wizard" but it would ultimately be a tedious 20 questions exercise. Everyone would have to decide to bother using the wizard, and then the hard part of testing and verification would be an exercise left to the reader.&lt;/p&gt;

&lt;p&gt;In other cases, Platform Engineering teams have been able to push the decisions and ownership quite far. I've seen a couple of solutions that automatically set up certain types of system alarms (CPU, file descriptors, etc) and try to tune their thresholds, but that only works for customers who have specific tool setups, and specific architectures. Customization on top of these would have limited knobs that had to be coded up, but more knobs leads to more complexity, and pushes customers to just do the whole thing themselves. And it leaves everyone to set up their own application health alarms, so adding infra alarms isn't that hard to do once you're in there setting up alarms anyway.&lt;/p&gt;

&lt;p&gt;It's a tough problem. On one hand, a Platform Engineering team can either stop short of solving the complete problem and give checklists and tons of configuration knobs that everyone has to sort through, understand, and apply. On the other extreme, a Platform Engineering team can solve the whole problem, but only if the customer is using very specific technologies, tools, and architectures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Flexibility without configuration knob hell
&lt;/h2&gt;

&lt;p&gt;The missing ingredient in both extremes is adaptability. Adaptability means that the tool can handle and reason about differences that it wasn't coded to handle. An agentic solution can say, "Oh, your service has 2 load balancers behind one DNS record? Great, let's set up alarms for all of them." But adaptability also needs customizability. But that's also solved with agents through steering, skills, and deterministic tools that gate the agent. If the agent didn't realize that one of those load balancers behind the DNS record is for a blue/green deployments, you can write a couple bullet points in a markdown file to guide it to understand that it's normal for one to be weighted to 0% traffic, and to adjust the alarms accordingly.&lt;/p&gt;

&lt;p&gt;This really works. Last week I was looking at a retrospective where a system had a misconfigured alarm that would have helped the team get an earlier start on the problem. Some types of alarms are unfortunately easy to misconfigure, either pointing to the wrong underlying resource, or using the wrong statistic like "avg" where it was supposed to be "sum", or that kind of thing. Well to test my "agents are great" hypothesis, I tweaked my own test service to have the same alarm misconfiguration, and prompted AWS DevOps Agent to audit my alarms. And sure enough, it found the misconfiguration, and even recommended some others to add. Since I configure alarms as IaC, I fed its report to Kiro and had it fix and add the missing ones. To organize the alarms the way I like, I described my alarm strategy of rollups with composite alarms — something that I hadn't written down before because it's complicated and varies a little bit by every application — and it shifted the alarms to set them up better than I had it in the first place.&lt;/p&gt;

&lt;p&gt;AI pushes the boundary on "how much of the full problem can you own" so much further than before. Instead of dumping pull requests of code upgrades to every team, you can trigger their load test suite. Sure, this was possible before, but it was impractical because there was no "run my bespoke load test suite and interpret the results in an isolated test environment" API. At least that's not a W3C standard that I'm familiar with. But with agents that can adapt, and try repeatedly until they get what they need, doing it is practical at scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Crossing bloodlines
&lt;/h2&gt;

&lt;p&gt;Ops remains a cultural problem, and for the longest time people try to treat the cultural problem as a technology problem and fail to solve it. So while this new agentic AI stuff is a very powerful new technology, Platform Engineering teams still need to consider the DuneOps philosophy: strength comes from experience. When I was on a Platform Engineering team, the longer I was on it, the less of an intuition I had for where the real problems were for my customers. Sure, I'd talk to as many customers as I could, attend ops meetings and read retrospectives of where things went wrong, and all that. But my personal hands-on experience of the pain of ops was a snapshot in time that kept moving further into the past.&lt;/p&gt;

&lt;p&gt;So maybe there's a new DuneOps model for the next generation of Platform Engineering. AI makes it so those teams can reach even further into the end solution — at scale — than ever before. Maybe this gives those teams just enough time to take on the ops for a production system that looks more like what their customers run. Sure, Platform Engineering teams operate the systems that they build, but those tend to have different operational requirements and be built differently. This way they still have the charter to build for everyone, but they gain the firsthand understanding that they are (or aren't) building the right thing for everyone. Wait, but this sounds like it has properties of an SRE team or a DevOps team or a Frontline Ops team! A sort of combination of parts of each!&lt;/p&gt;

&lt;p&gt;In Dune, the Bene Gesserit worked in the shadows over 90 generations on a eugenics program, crossing bloodlines of the major houses with the end goal of producing a human with a mind so powerful that it could see past space and time — a being they called the Kwisatz Haderach.&lt;/p&gt;

&lt;p&gt;DuneOps is exactly this. With AI, you can now cherry-pick genes from each operational model and combine them, to borrow strengths of other operational models to offset each of any inherent weaknesses in the core model you choose. (And if you think your operational model lacks weaknesses, read last week's &lt;a href="https://dev.to/dyanacek/dune-as-an-operational-model-57m-temp-slug-2894378"&gt;Dune as an operational model&lt;/a&gt; and then "@" me on twitter or whatever.)&lt;/p&gt;

&lt;p&gt;And so just like how Paul Atreides defeated Emperor Shaddam and his Sardaukar by harnessing his Harkonnen heritage, AI is the leverage we needed to be able to borrow from and combine the operational models to take them further than we could ever before.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Dune as an operational model</title>
      <dc:creator>David Yanacek</dc:creator>
      <pubDate>Sun, 10 May 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/dyanacek/dune-as-an-operational-model-om1</link>
      <guid>https://dev.to/dyanacek/dune-as-an-operational-model-om1</guid>
      <description>&lt;h2&gt;
  
  
  Dune as an allegory for ops
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb1uk7ulaoqm3vovh4y8j.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb1uk7ulaoqm3vovh4y8j.jpeg" alt="Fremen fighters in the desert of Arrakis" width="780" height="438"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I am &lt;em&gt;obsessed&lt;/em&gt; with the new Dune movies by Denis Villeneuve. The filmmaking, the music by Hans Zimmer, the restraint and pacing, sweeping visuals, the use of space - it's a masterpiece. The subsequent books get a little...interesting for me, but I made it through them. Internet person &lt;a href="https://www.instagram.com/naturebased0/" rel="noopener noreferrer"&gt;NatureBased&lt;/a&gt; publishes a steady stream of Instagram reels to fill in the gaps in my understanding of the lore.&lt;/p&gt;

&lt;p&gt;One recurring part of the lore of the Dune universe is how some groups of people become so much more powerful fighters than others. There are two groups in particular: the Sardaukar, who are the Emperor Shaddam's personal army, and the Fedaykin, who are Fremen fighters. Frank Herbert, the author of Dune, writes that both factions of fighters are stronger because of the environment from which they are forged. The Sardaukar are raised on the planet Salusa Secundus, which has a harsh environment that trains them to reject weakness. The Fremen live on the planet Dune, which is an even harsher environment, leading the fighters to be even more powerful and fanatical.&lt;/p&gt;

&lt;p&gt;This idea about being "forged by harsh environments" has amusing parallels to my experience being a software developer and operator. I've learned to be a better operator by being in an intense operational environment. This in turn has made me a better software developer. It's taught me to avoid making the same mistake twice, but it's also taught me a certain operational paranoia that makes me question everything in terms of its operational safety. So I'll call this idea - becoming a better operator and better developer through a harsh ops experience - to be known as DuneOps.&lt;/p&gt;

&lt;h2&gt;
  
  
  What are the different operational models?
&lt;/h2&gt;

&lt;p&gt;I've seen, participated in, and talked with people from engineers to CTOs about different operational models - the pros and cons, applicability to a given company or environment, and there's no consensus around which is universally "the best". Different situations call for different approaches.&lt;/p&gt;

&lt;p&gt;To start with some bookkeeping, I'll give my personal definition of SRE, DevOps, Frontline, and Platform Engineering. These definitions aren't going to be nuanced and they'll be imprecise, but at least they'll serve as ragebait for engagement on social media. Consider these to be caricatures of each, so if I draw the forehead of your operational model to be comically large, just know it's all in good fun.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On one extreme, you have &lt;strong&gt;Frontline Ops&lt;/strong&gt;. This is just having a separate ops team of people whose job it is to do pure ops. Think carry the pager and take all of the alarms. They unblock CICD pipelines, patch hosts, you name it. The dev team throws everything over the fence to the front line team.&lt;/li&gt;
&lt;li&gt;On the opposite end, you have &lt;strong&gt;Platform Engineering&lt;/strong&gt;. This is a central team who writes tools and systems that everyone can use, but they don't own any aspect of the operations of others' systems. They might build CI/CD tools or best practice scanners or golden path abstractions that hopefully solve real problems that people have day to day.&lt;/li&gt;
&lt;li&gt;Somewhere in between is &lt;strong&gt;Site Reliability Engineering (SRE)&lt;/strong&gt;. I've heard some interpretations where these teams do take the first-pass at every alarm. But they build their way out of pain, are experts in best practices and how to set up everything from alarms to pipelines, and they have strict contracts with the dev teams so that if the systems become unwieldy, the dev team agrees to pause everything and burn down the ops backlog.&lt;/li&gt;
&lt;li&gt;And then there's &lt;strong&gt;DevOps&lt;/strong&gt;. Here the dev team is on the hook for everything, and they all carry the first-responder pager. If their pipeline is flakey, they own cleaning it up. If their alarms are either too sensitive or not sensitive enough, they're the ones who need to figure out the right way to deal with it. If their ops are too manual, they need to automate it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Of course you don't have to pick one model. Platform Engineering supplements all of these approaches. Some systems might shared require specific ownership that lends itself to SRE, like a Kubernetes cluster, database service, web server, or API gateway.&lt;/p&gt;

&lt;h3&gt;
  
  
  DevOps
&lt;/h3&gt;

&lt;p&gt;I've spent the vast majority of my career on teams that do DevOps - where developers do all of their own ops. As developers, we can build our way out of any problem. We're all at least a little bit prima donnas so when ops gets too painful, we bring pitchforks to our backlog prioritization meetings and fix it or we walk. We can see the pager pain of our team compared to other teams, so that transparency is helpful incentive even for managers. Fewer people will transfer to a team that has excessive ops, and more people will leave such teams.&lt;/p&gt;

&lt;p&gt;I swear by the DevOps model, and I think in the age of agentic development tools, developers will become more generalist and wear more hats, not fewer. I'm part developer, working on code. I'm part people manager - especially the more senior I get - because plenty of getting things done involves involving and coaching the right people and getting everyone marching in the same (or compatible) direction. I'm part customer support, because by hearing straight from customers I can tell where the product I'm working on is bad or what it's missing. I'm part sales and marketing, so I can make sure the product sounds good and has easy and intuitive getting started and attractive features to buyers. I'm part product manager, since I tend to work on products where I'm also the customer. And of course I'm part operator since I can architect and build the most amazing sounding system in the world, but if I miss some important detail that matters a ton in practice, then the system will fail and nobody will care about how it works in theory. With agentic coding, this gives us more time and need to wear these multiple hats. So I think where we're headed is into more of this DevOps model vs one of specialization.&lt;/p&gt;

&lt;p&gt;At least in the environments I've been in, DevOps has led to the best long-term quality of the services we build, and the best operational outcomes for customers. When operational pain creeps up, the team naturally trades off other work to respond to resolving the pain. Since the team is in touch directly with customers, they hear about things that are bad about the operational or customer experience, even if the metrics don't measure or represent the pain well. And most importantly, they teach developers about what actually matters. You can have the most amazing theoretical architecture in the world, but if your database replication doesn't catch up even during a load spike and node down situation, then your customers are going to see latency or outages.&lt;/p&gt;

&lt;p&gt;But of course there are multiple sides to a coin (typically two). DevOps' downside is you never have enough time. The ops backlog is infinite. There's always some alarm that you can tune to "cry wolf" less often, or another alarm to be made more sensitive to detect problems earlier. In DevOps, when you build things to automate away the pain, you build it for yourself. It's the Ayn Rand model of operations. Actually I never read any of her stuff so I could be totally off base. What I mean is that teams are responsible for solving all of their problems themselves. They're on the hook for everything, from security to availability to customer service. Hopefully someone helps provide tools (like a platform engineering team), but each team is left cobbling the tools together and deals with the consequences of the gaps in the tooling.&lt;/p&gt;

&lt;p&gt;Within DevOps teams, sometimes an altruistic engineer comes along and makes their bespoke automation tool reusable for other teams, but that requires vision from the engineer, and an understanding manager who gives them room to force multiple instead of overrotating on that one team's goals. I've been that altruistic engineer and have spent a ton of time cheerleading the other ones around, providing covering fire when needed. Sometimes these people find themselves rotating into a central team to spend all of their time on their side project with the goal of helping everyone in the company. And since AWS' business is all about making ops easier for customers, sometimes these people rotate onto a product team who aims to solve that problem for everyone.&lt;/p&gt;

&lt;p&gt;DevOps' other downside is that you can miss things. Yes, even an experienced DevOps engineer like myself has somewhat recently missed having an auto-rollback alarm in a pipeline to kick off a roll back of a deployment before anyone even gets paged, leading to the oncall having to get paged and trigger the rollback themselves. It wasn't the end of the world, but it left me facepalming wondering how I missed something so obvious.&lt;/p&gt;

&lt;h3&gt;
  
  
  Platform Engineering
&lt;/h3&gt;

&lt;p&gt;To complement the DevOps model, you need a great platform engineering team. These teams get ahead of the problems faced by all DevOps teams, and build abstractions that makes everyone's lives easier. They also try to help teams see what things they might be missing. To figure out what to build, they can perform "tool harvesting" where they go around and look at all of the bespoke tools that people have made, look for patterns, and make the general version. To make it so DevOps teams don't miss things, platform engineering teams also make "scanners" or "auditors" or "recommenders". For example if a pipeline is missing an auto-rollback alarm, the pipeline system will reach out to you and let you know you're missing it. For certain types of misconfigurations, the pipeline can even refuse to deploy until you fix it.&lt;/p&gt;

&lt;p&gt;I've also worked on a Platform Engineering team, building a web service framework called Coral. This framework handles request and response serialization, validation, protocols, authentication, rate limiting, SDK generation, and whatever else we can make more convenient for service owners and consumers. I joined it because I was fed up with service teams providing REST APIs and leaving actually calling the service from a typed language like Java as an exercise for the reader. When I heard we were making a framework that would let service owners have their cake and let consumers eat it too by auto-generating those clients, I said, "sign me up".&lt;/p&gt;

&lt;p&gt;This web service framework team was interesting. In some ways, we had zero oncall burden. After all, we were making a framework, not running a service. On the other hand, our ops burden was the highest of any team I'd been on. Why? Because we fielded every question from every Amazon developer who got stuck and suspected the framework. Sure, we had rough edges in the framework, and gaps in our documentation, but any time someone saw a stack trace with our framework name somewhere in it (every stack trace), there was a chance that they would do the low-effort thing of just asking us to help. We'd be as patient and helpful as humanly possible in helping people figure out their own problems in their own code, but we also had plenty of sharp edges that we had to smooth out and documentation to improve.&lt;/p&gt;

&lt;p&gt;In a sense were a DevOps team operating this framework. When ops spiked, we'd find the common patterns in the questions we got, and would deprioritize new feature development until we got it under control. Every customer question we answered without having that answer be available and discoverable in our internal documentation was a defect, and we worked to drive that down to zero. We also spent time on general education on how to help oneself and "how to help us help you", frequently citing this brilliant manifesto &lt;a href="http://www.catb.org/~esr/faqs/smart-questions.html" rel="noopener noreferrer"&gt;"How to Ask Questions The Smart Way"&lt;/a&gt; by &lt;a href="https://twitter.com/esrtweet" rel="noopener noreferrer"&gt;Eric Steven Raymond&lt;/a&gt; and &lt;a href="http://linuxmafia.com/~rick/" rel="noopener noreferrer"&gt;Rick Moen&lt;/a&gt;. (Eric and Rick, THANK YOU.)&lt;/p&gt;

&lt;p&gt;The biggest downside is that you lack a first-hand signal about whether you're solving the real problems for customers. Every year I spent on the platform engineering team, the less confident I was that I was solving the biggest source of pain from customers feature-wise. One solution is to rotate in and out of a team like this. When peers have been super passionate about an idea and abstraction to solve a problem broadly, I encouraged them to rotate into our platform engineering team for a while (long enough to see it through and iterate on it). And when they're done, some stay but others rotate back out into the fray. I know plenty of folks who have stayed on a platform engineering team working on problems that take a really, really long time to solve and adapt to changing times, and they're doing great at it. So it isn't necessary to rotate out, but it's very necessary to rotate in.&lt;/p&gt;

&lt;p&gt;The second downside to platform engineering is scaling ops. At the end of the day if it costs someone nothing to ask a question, and costs a great deal to answer that question, that's an asymmetric scaling problem that turns into a DDoS. Scaling support for that leads to the loss of signal about what's bad in your stuff, reducing overall quality.&lt;/p&gt;

&lt;h3&gt;
  
  
  Site Reliability Engineering (SRE)
&lt;/h3&gt;

&lt;p&gt;While I haven't been on a &lt;a href="https://en.wikipedia.org/wiki/No_true_Scotsman" rel="noopener noreferrer"&gt;"true" SRE team&lt;/a&gt; before, I've been on a team that seemed a lot like one. This was my first job out of college at Amazon on a team called Website Hosting. We ran the web server environment for the fleets that rendered HTML for amazon.com and its regional variants. Unlike the rest of Amazon, we weren't exactly DevOps. We didn't run the code we wrote. I didn't write the website rendering framework. I didn't write the perl/mason code that decided what to actually show on the page. I ran the servers. So I needed to make sure we forecasted how many we'd need, estimating traffic and efficiency of each page render. I ran the CI/CD pipeline that everyone fence-threw their releases onto. And I dealt with single-server failures and crashes, and OS upgrades and performance tuning. After all, someone had to do this stuff. It couldn't be done by a shared oncall rotation of hundreds of developers who knew nothing about the server environment.&lt;/p&gt;

&lt;p&gt;The upsides were great in this model. Our only job was to automate the operations of this fleet, so we weren't distracted by any features to build. We were connected directly to the pain of what mattered, so we were the perfect product managers to decide what to build. And while we ran server ops for a countable number of fleets, we weren't on the hook for everything in the company. But we owned just enough fleets that everything we built needed to be configurable and extensible - making it an easy extension to make our tools usable by other teams too. Essentially we could be a platform engineering team that was slowed down by operating an unrelated system to the tools we built, but it sharpened our focus into the pragmatic. In fact one of the alarm aggregation systems we built on this team got adopted across all of Amazon, and eventually became (essentially) the CloudWatch composite alarms feature to roll up many alarms into one page. We realized we needed this feature when every web server paged us individually resulting in a loop of some 500 back-to-back pages, to which the oncall responded by ripping the battery out of the pager and hurling it against the wall. (In retrospect either of those actions would have likely succeeded, but why not both?)&lt;/p&gt;

&lt;p&gt;There were certainly downsides to this model. We were hired as software developers, but we were doing the shit work writing and running scripts while our peers did the glamorous work building features and frameworks. We spent our day writing Greasemonkey browser scripts for provisioning servers, perl scripts for killing runaway processes, and VBScript to automate Excel-driven forecasting. And of course someone had to run those scripts and processes, so that was us too. Don't get me wrong, I love solving real practical problems of any shape or form, but we were doing this while we watched our peers build distributed systems, protocols, and frameworks in C++. Of course we thought big and worked on the long term solutions, like a replacement for the Excel sheet involving some cool managed R environment with a time series database. But it wasn't glamorous.&lt;/p&gt;

&lt;p&gt;We also got paged all the time. Whenever there was an "order drop" (the best metric for "is there a problem with the website right now" is "have people stopped buying things"), our team was on the default engagement list because sometimes the problem was with our web servers or our load balancers in front of them. Even if it wasn't our fault, we could help because we were really good at reading the tea leaves (the logs and metrics and traces, even though they weren't called that yet) to figure out which microservice in the ball of wax architecture was the culprit. Engineers tended not to stay super-long on this team before they rotated out to other teams. But it was a hell of a Dune-like environment for us Fedaykin Fremen fighters to grow stronger in ops.&lt;/p&gt;

&lt;h3&gt;
  
  
  Frontline Ops
&lt;/h3&gt;

&lt;p&gt;The only model that I can't say I've personally participated in is frontline ops. I suppose being a bank teller in high school was essentially this, since we were doing manual transactions for people. My attempt at reducing ops was to teach customers how to get ATM cards. But we were actually attached to sales to try to bring in more money to the bank, so that part of the job was not my jam.&lt;/p&gt;

&lt;p&gt;But I've experienced frontline ops plenty and have studied it somewhat. Often frontline ops comes in the unenviable position of customer support. One day I tried to order a replacement physical pager so I could go on call, but the reordering system was broken. Long story short, I found that there was a re-ordering edge case that some people hit, and the resolution to the problem was to email Brenda at a 3p pager vendor directly. The support team for this kind of corporate system support had gotten over a hundred of these tickets in a couple years, but none of the awareness of the problem had bubbled up to the team who owned pager reordering. Once I pointed out the problem the team who owned paging fixed its integration with the 3p right quite quickly. Everyone involved was very helpful and great owners; it was the structure that removed the signal from the team who could stop the pain. Sure frontline ops teams produce reports and data and ticket trends so central teams can know what to fix. This is even described in the SRE books. But I find it's too easy for those best intentions mechanisms to break down, or for nuance to be lost in translation.&lt;/p&gt;

&lt;p&gt;The place I've seen frontline ops used the most effectively is when scaling and automation is not yet possible, from certain physical or network infrastructure ops, to customer support. Every time I shadow customer support I get hit with a wave of learning about everything that sucks. But it's energizing. I feed off of that with optimism that we can fix things. I just need to go back to the well often so that I get fill up with the next problem to carry the banner for. And every time I talk with physical or network infrastructure folks I run into people who both fearlessly save the day and who put in place the most careful operational processes I have seen.&lt;/p&gt;

&lt;p&gt;Frontline ops engineers are forged from the harshest environment of anyone, creating the strongest operators. I've worked with many support engineers who have converted to software developer roles or solution architects, and they bring a level of customer obsession and operational excellence that nobody else on the team has.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping the best while avoiding the worst
&lt;/h2&gt;

&lt;p&gt;DuneOps is all about understanding the environment that you're forged from, so you can take the parts that naturally improve you, and compensate for the parts of the environment that make some skill atrophe. Here are some things to consider to compensate for each operational model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DevOps works best when you incentivize organizational altruism, making tools that you actively share with other teams to make their lives better too. Cross-team altruism needs to be part of the job description so that managers know that they can still put together a high performance rating or promotion case for someone who spends a lot of their time on it. No two software developers are the same; some spend time on ops automation, others focus on testing, and others focus on micro-optimization. But it's easy to have a role guideline talk about "shipping features" and forget to encourage altruism.&lt;/li&gt;
&lt;li&gt;Platform Engineering works best when it is connected directly to what's real. That means actively "tool harvesting". If a team spent their precious cycles to build some automation, that means they've solved a real problem that could apply broadly, and should be amplified or brought into the fold. It also means encouraging mobility so that people who lean toward doing that force multiplying can do some kind of rotation into the platform engineering team to bring in direct understanding of what it's like to build using the platform engineering team's stuff.&lt;/li&gt;
&lt;li&gt;SRE works best when you aren't carrying the pager for other teams' crappy code and writing contracts around how much pain is acceptable before service teams fix their shit code. And it works best when the team has license to see things through and not spend all their time writing Greasemonkey bandaids and one-off scripts. They need to be able to be a Platform Engineering team and tool provider for other teams to use too or else they're stuck doing only the crap work.&lt;/li&gt;
&lt;li&gt;Frontline ops works best when their extreme expertise is sought out regularly and issues addressed urgently. These engineers have the best understanding of how things actually work and fail, and how customers are disappointed. It's easy for teams to settle in and ignore this, but watching it is essential.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  DuneOps
&lt;/h2&gt;

&lt;p&gt;DuneOps can emerge from any of these operational models - DevOps or SRE or Frontline ops. In Dune, the most powerful fighters are forged by harsh environments. And so in ops, the best operators are born from harsh, real-world operations. They need to be able to make mistakes and learn, see what works and what doesn't, and find out what actually matters by being accountable for the outcomes.&lt;/p&gt;

&lt;p&gt;No matter what operational model you go with, I find it's important that dev and ops teams:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Stay grounded in reality.&lt;/strong&gt; The closer you can put the responsibility for building central solutions to the place where the real ops is happening, the more complete the outcome.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put operational ownership with service ownership.&lt;/strong&gt; The people whose software is broken should have their pager's go off first, not second.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incentivize organizational altruism.&lt;/strong&gt; Build a culture where people share their solutions to improve others' ops as well as theirs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep blame focused on structure and not people.&lt;/strong&gt; We didn't talk about it in this post but it's super important and is covered in &lt;a href="https://youtu.be/7MrD4VSLC_w?si=CGpmJ9dO36l7YtGO&amp;amp;t=638" rel="noopener noreferrer"&gt;this talk&lt;/a&gt; if you're interested.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Epilogue
&lt;/h2&gt;

&lt;p&gt;I'm sure folks will disagree with parts of this. I'm happy to chat about it (see social media links) and keep updating this as I collect more perspectives. And please know and believe that I'm not trying to offend anyone with any of these takes or colorful language. All of this stuff is extremely hard, and I respect everyone who been involved with any of this kind of stuff. After all, we are all brothers and sisters of the Fremen, forged out of the harsh desert of Dune.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>ai</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
