<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: VIREN PATEL</title>
    <description>The latest articles on DEV Community by VIREN PATEL (@viren_patel_4351db2f23035).</description>
    <link>https://dev.to/viren_patel_4351db2f23035</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3887969%2F8df83c5d-11ba-4f5f-b142-b998fcffbc3a.jpeg</url>
      <title>DEV Community: VIREN PATEL</title>
      <link>https://dev.to/viren_patel_4351db2f23035</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/viren_patel_4351db2f23035"/>
    <language>en</language>
    <item>
      <title>Building a Production Observability Pipeline with AWS, Prometheus and Grafana</title>
      <dc:creator>VIREN PATEL</dc:creator>
      <pubDate>Tue, 06 Oct 2026 00:06:12 +0000</pubDate>
      <link>https://dev.to/viren_patel_4351db2f23035/building-a-production-observability-pipeline-with-aws-prometheus-and-grafana-43c4</link>
      <guid>https://dev.to/viren_patel_4351db2f23035/building-a-production-observability-pipeline-with-aws-prometheus-and-grafana-43c4</guid>
      <description>&lt;h2&gt;
  
  
  1. The Problem
&lt;/h2&gt;

&lt;p&gt;In a production environment, knowing that an application is deployed isn't enough — you also need to know whether the devices and services running at the edge are actually healthy. In my case, that meant thousands of &lt;strong&gt;edge devices&lt;/strong&gt; deployed across many customer locations, each running local software that depends on network connectivity, hardware health, and correct configuration.&lt;/p&gt;

&lt;p&gt;A device can silently go offline, lose connectivity to its local network, run low on storage, or fall out of compliance with the fleet's expected OS version — and none of that is visible unless something is actively watching for it. We needed a way to answer, at any moment, for any location: &lt;em&gt;"Is this device alive, and is it healthy?"&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Requirements
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Every device should report its liveness on a predictable cadence (a &lt;strong&gt;heartbeat&lt;/strong&gt;), scoped to a specific location.&lt;/li&gt;
&lt;li&gt;The ingestion path must be resilient to bursts, retries, and duplicate/late messages without losing data.&lt;/li&gt;
&lt;li&gt;Failed or malformed messages must not be silently dropped — they need a dead-letter path for investigation.&lt;/li&gt;
&lt;li&gt;In addition to "is it alive," we wanted richer telemetry (CPU, memory, storage, battery, compliance) without requiring every device to push it directly.&lt;/li&gt;
&lt;li&gt;All ingested data needed to land in a time-series store so it could be queried, graphed, and alerted on.&lt;/li&gt;
&lt;li&gt;Operators needed dashboards for fleet-wide and per-location visibility, plus alerts when something goes wrong — without needing to read logs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Architecture
&lt;/h2&gt;

&lt;p&gt;End-to-End System Diagram&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo4dutran3bvpclnjih62.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo4dutran3bvpclnjih62.jpeg" alt=" " width="799" height="436"&gt;&lt;/a&gt;&lt;br&gt;
Two ingestion paths feed the same metrics backend:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Direct heartbeats&lt;/strong&gt; — each edge device pushes a small heartbeat payload on its own cadence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scheduled telemetry pull&lt;/strong&gt; — a Lambda polls an external device management API on a fixed schedule to enrich the picture with hardware and compliance data the devices don't push themselves.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Both paths converge on the same time-series backend, so a single set of dashboards and alerts covers "is it alive" and "is it healthy."&lt;/p&gt;
&lt;h2&gt;
  
  
  4. AWS Components
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;API Gateway&lt;/strong&gt; + Lambda authorizer&lt;/td&gt;
&lt;td&gt;Authenticates and routes incoming heartbeat requests, scoped per location&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lambda (enqueue)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Validates the heartbeat payload and pushes it onto SQS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SQS + DLQ&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Buffers and batches heartbeats; isolates malformed/failed messages after repeated delivery failures instead of dropping them&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lambda (dequeue)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Consumes SQS batches, converts each heartbeat into Prometheus-format metrics, and writes them to the metrics backend&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lambda (scheduled telemetry sync)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Runs on a fixed EventBridge schedule, pulls device and management telemetry from an external device-management API, and writes it as metrics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;S3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stores a raw, timestamped backup of every metrics payload as an audit trail, with lifecycle expiry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SSM Parameter Store&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Holds service credentials and sync-state (e.g. last successful poll timestamp) outside of source code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CloudWatch Alarms + SNS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Watches queue age, DLQ depth, and Lambda error rates — the "is the pipeline itself healthy" layer, separate from device health&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Keeping the ingestion pipeline serverless meant we didn't need to run or patch any always-on infrastructure just to accept heartbeats — it scales automatically with the size of the fleet and the queue absorbs bursts.&lt;/p&gt;
&lt;h2&gt;
  
  
  5. Heartbeat Design
&lt;/h2&gt;

&lt;p&gt;The heartbeat payload is intentionally small and cheap to send. At minimum, a device only needs to identify itself and increment a counter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"serialNumber"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"XXXXXXXXXXXXXXXXXXXXXXXXXX"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"count"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Optionally, a device can include local network connectivity state, since a healthy device that's lost its local network connection is a different failure mode from an offline device:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"serialNumber"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"XXXXXXXXXXXXXXXXXXXXXXXXXX"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"count"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"isConnected"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"lastPublished"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2025-07-29T20:12:26Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"connectionType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"local"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"messageCount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few deliberate design choices here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The device serial number is validated against the authenticated caller&lt;/strong&gt;, not trusted blindly — a device can't report a heartbeat under an identity it isn't authorized for.&lt;/li&gt;
&lt;li&gt;The schema is strict (&lt;code&gt;additionalProperties: false&lt;/code&gt;) — unexpected fields fail validation loudly instead of being silently ignored.&lt;/li&gt;
&lt;li&gt;The payload has no dependency on any specific transport; it's just JSON over HTTPS, which keeps the device-side implementation trivial.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  6. Metrics and Monitoring
&lt;/h2&gt;

&lt;p&gt;Rather than storing heartbeats as raw events in a database, every heartbeat is immediately converted into &lt;strong&gt;Prometheus-format gauges&lt;/strong&gt; at ingestion time. That decision shaped everything downstream: querying "which devices haven't reported in the last N minutes" becomes a standard PromQL query instead of a custom aggregation job.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8naub2u9hvc7rkjvnwze.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8naub2u9hvc7rkjvnwze.gif" alt=" " width="600" height="359"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two families of metrics exist:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Liveness metrics&lt;/strong&gt; (from the heartbeat path): heartbeat counter, network connection status, last publish time, message count — labeled by device serial number.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deep telemetry metrics&lt;/strong&gt; (from the scheduled sync path): CPU utilization and temperature, memory and storage capacity, battery health, network diagnostics, OS version and compliance status, enrollment state.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because both are just labeled Prometheus gauges in the same store, a dashboard or alert doesn't need to care which pipeline produced a given metric.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Prometheus / VictoriaMetrics
&lt;/h2&gt;

&lt;p&gt;We chose &lt;strong&gt;VictoriaMetrics&lt;/strong&gt; as the metrics backend because it speaks the Prometheus remote-write/import protocol (so nothing upstream needs to know it isn't "real" Prometheus) while being cheaper to run at scale for a high cardinality of devices and locations.&lt;/p&gt;

&lt;p&gt;The dequeue Lambda batches metrics and writes them via VictoriaMetrics' &lt;code&gt;/api/v1/import/prometheus&lt;/code&gt; endpoint, gzip-compressed, over HTTPS with basic auth. The scheduled telemetry sync does the same, but chunks its output into sub-1.5MB batches (to stay under payload limits) with retry and exponential backoff, since a single sync run can cover the entire device fleet.&lt;/p&gt;

&lt;p&gt;Locally, the same VictoriaMetrics image runs in Docker so the ingestion path can be tested end-to-end without touching production metrics.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Grafana Dashboards
&lt;/h2&gt;

&lt;p&gt;VictoriaMetrics is queried by Grafana for both fleet-wide and per-location visibility:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;fleet overview&lt;/strong&gt; dashboard shows aggregate liveness across all locations/devices — how many are currently reporting heartbeats, how many have gone silent, and overall compliance/health trends.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-location dashboards&lt;/strong&gt; drill into a single location's devices — heartbeat recency, network connectivity, hardware telemetry (CPU/memory/storage/battery) — so an operator investigating a specific customer complaint can go straight to the relevant view.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because everything is labeled consistently by device serial number, the same panels/queries generalize across dashboards just by changing the label filter.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Alerting
&lt;/h2&gt;

&lt;p&gt;Alerting is defined as Grafana alert rules evaluated directly against the VictoriaMetrics data — for example, a device with no heartbeat metric update within an expected window, or a rising rate of network disconnections at a location. This keeps alerting logic close to the same PromQL used for the dashboards, rather than as a separate system with its own definition of "healthy."&lt;/p&gt;

&lt;p&gt;This is layered with &lt;strong&gt;CloudWatch alarms&lt;/strong&gt; that watch the pipeline itself rather than the devices — SQS queue age, DLQ message visibility, and Lambda error rates — because a healthy-looking dashboard doesn't help if the ingestion pipeline is actually the thing that's broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Challenges and Trade-offs
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Two ingestion paths, one data model.&lt;/strong&gt; Direct heartbeats and polled telemetry arrive on completely different cadences and shapes, but both had to converge into the same labeled metric format so dashboards and alerts didn't need to special-case the source.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batching vs. payload limits.&lt;/strong&gt; The scheduled telemetry sync can produce a large volume of metrics in one run across the whole fleet; naively posting them in one request risks hitting the metrics backend's payload limits, so batching with retry/backoff was necessary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trusting device-reported data.&lt;/strong&gt; A device claiming a given serial number can't be taken at face value — it has to be checked against what the caller is actually authorized for, otherwise one compromised or misconfigured device could pollute another device's metrics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Losing data safely.&lt;/strong&gt; Anything that fails to enqueue or process shouldn't just vanish — the DLQ and the raw S3 backup of every metrics payload exist specifically so a bad batch can be replayed or inspected after the fact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two definitions of "healthy."&lt;/strong&gt; Device health (via Grafana/VictoriaMetrics) and pipeline health (via CloudWatch/SNS) are deliberately separate concerns, alerting through different channels, because conflating them makes it harder to tell "the fleet is unhealthy" from "our monitoring is broken."&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  11. What We Learned
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Converting data into Prometheus format as early as possible in the pipeline (at ingestion, not at query time) simplified everything downstream — dashboards, alerts, and ad-hoc debugging all speak the same query language.&lt;/li&gt;
&lt;li&gt;Treating the metrics pipeline itself as a first-class thing to monitor (via CloudWatch) is just as important as monitoring the devices it reports on.&lt;/li&gt;
&lt;li&gt;Keeping the heartbeat payload minimal paid off — it made the device-side integration trivial and kept the ingestion Lambda's job simple and fast.&lt;/li&gt;
&lt;li&gt;A raw payload backup (S3) turned out to be cheap insurance — being able to replay or inspect exactly what was received, even non-fatally, saved investigation time more than once.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  12. Possible Improvements
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Push more of the "deep telemetry" fields into liveness-style alerting (e.g. combining "hasn't heartbeated" with "was last seen non-compliant") for higher-signal alerts.&lt;/li&gt;
&lt;li&gt;Explore anomaly-detection style alerting rather than fixed thresholds for fleet-wide trends.&lt;/li&gt;
&lt;li&gt;Reduce the polling cadence gap between direct heartbeats and scheduled telemetry sync so both liveness and deep health data are closer to real-time.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  13. Conclusion
&lt;/h2&gt;

&lt;p&gt;Observability at the edge isn't just about collecting data — it's about designing the smallest possible signal (a heartbeat) that reliably answers "is this thing alive," and complementing it with richer telemetry pulled on our own terms rather than depending on every device to push everything. By normalizing both into the same time-series format early, we got a single, consistent place — Grafana on top of VictoriaMetrics — to answer both "is the fleet healthy" and "is our monitoring pipeline itself healthy."&lt;/p&gt;

</description>
      <category>aws</category>
      <category>devops</category>
      <category>monitoring</category>
      <category>sre</category>
    </item>
    <item>
      <title>Building a Reusable AWS SAM CI/CD Pipeline for Multi-Account Deployments</title>
      <dc:creator>VIREN PATEL</dc:creator>
      <pubDate>Tue, 14 Jul 2026 18:56:04 +0000</pubDate>
      <link>https://dev.to/viren_patel_4351db2f23035/building-a-reusable-aws-sam-cicd-pipeline-for-multi-account-deployments-167g</link>
      <guid>https://dev.to/viren_patel_4351db2f23035/building-a-reusable-aws-sam-cicd-pipeline-for-multi-account-deployments-167g</guid>
      <description>&lt;p&gt;&lt;strong&gt;Introduction &amp;amp; The Problem&lt;/strong&gt;&lt;br&gt;
I am an Intermediate Software Engineer at one of New Zealand’s global QSR enterprises.&lt;/p&gt;

&lt;p&gt;One challenge kept showing up in my day-to-day work: every new AWS SAM project in our team came with yet another deployment pipeline.&lt;/p&gt;

&lt;p&gt;Different repositories used different AWS CLI and AWS SAM CLI versions. Some pipelines deployed to a single account, others to multiple accounts. Even small improvements to the deployment process had to be repeated repo by repo.&lt;/p&gt;

&lt;p&gt;After maintaining this for a while, I realized we weren’t solving deployment problems anymore. We were solving the same tooling problem over and over again.&lt;/p&gt;

&lt;p&gt;I wanted every project to use one reusable deployment engine instead of maintaining dozens of nearly identical pipelines. So, this article is aimed at teams that deploy multiple AWS SAM applications and want to standardize deployments across AWS accounts without duplicating CI/CD logic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Architecture Overview&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because this article is about solving a repeated deployment problem, I am showing two architecture views:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The common fragmented model many teams start with.&lt;/li&gt;
&lt;li&gt;The standardized deployment-engine model I implemented with sam-pipeline.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Architecture 1: Fragmented Team-by-Team Deployments (Current State)&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fql5rjumvadzisqs8dll5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fql5rjumvadzisqs8dll5.png" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In this model, each repository carries its own deployment logic, tool versions, and account assumptions. It works initially, but maintenance scales badly and drift becomes unavoidable.&lt;/p&gt;

&lt;p&gt;Architecture 2: Shared Deployment Engine with sam-pipeline (Solution)&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff0tsci5d7r56pt3zj3do.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff0tsci5d7r56pt3zj3do.png" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here, every project uses the same deployment engine. Runtime setup, SAM build/deploy flow, and multi-account behavior are standardized, which reduces duplication and makes deployments more predictable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Building the Docker Deployment Image&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;To solve this consistently, I built sam-multi-account-pipeline, a reusable deployment image that teams can pull from Amazon ECR Public.&lt;/p&gt;

&lt;p&gt;The idea is simple: developers use a shared public ECR image to deploy AWS SAM stacks across multiple AWS accounts and regions, without rebuilding deployment logic in every repository.&lt;/p&gt;

&lt;p&gt;Image: public.ecr.aws/z7q5l7x5/sam-pipeline:1.0.0&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Publishing to Amazon ECR Public&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I published the image to Amazon ECR Public so anyone can consume it directly in their CI/CD pipeline.&lt;/p&gt;

&lt;p&gt;Complete setup and usage instructions are available in the repository:&lt;br&gt;
&lt;a href="https://github.com/Viren1993-web/sam-multi-account-pipeline" rel="noopener noreferrer"&gt;https://github.com/Viren1993-web/sam-multi-account-pipeline&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Multi-Account Deployment with AssumeRole&lt;br&gt;
Here’s how multi-account deployment works in this model:&lt;/p&gt;

&lt;p&gt;For every AWS account you want to deploy to, create a deployment role in that account and grant it the permissions required by your SAM template resources.&lt;/p&gt;

&lt;p&gt;The pipeline then assumes that role and performs the deployment securely in the target account. In my repository, I’ve included an example using a DeployerGithub role to demonstrate this pattern.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I Wanted to Achieve&lt;/strong&gt;&lt;br&gt;
I wasn’t trying to build “another CI script.” I wanted a deployment foundation the team could trust and reuse everywhere.&lt;/p&gt;

&lt;p&gt;My goals were:&lt;/p&gt;

&lt;p&gt;One consistent deployment engine for all SAM projects.&lt;br&gt;
Support for both single-account and multi-account deployments.&lt;br&gt;
No repository-specific setup drift for AWS/SAM tooling.&lt;br&gt;
Easy onboarding for new services.&lt;br&gt;
A safer security model using role assumption instead of long-lived credentials.&lt;br&gt;
In short, I wanted deployment to be boring, predictable, and easy to maintain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where This Fits Alongside AWS Services&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;AWS provides excellent services such as AWS CodePipeline and CloudFormation Change Sets for building and managing deployment workflows. My goal wasn't to replace these services, but to solve a different challenge: standardizing how AWS SAM applications are deployed across multiple repositories and AWS accounts.&lt;/p&gt;

&lt;p&gt;By packaging the deployment tooling into a reusable, versioned Docker image, every project uses the same AWS CLI, AWS SAM CLI, and deployment process regardless of the CI/CD platform. This approach reduces duplication, simplifies maintenance, and provides a consistent deployment experience whether the pipeline runs in Bitbucket Pipelines, GitHub Actions, Jenkins, or even AWS CodePipeline itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Supporting Node.js and Python&lt;/strong&gt;&lt;br&gt;
One practical requirement was runtime flexibility. Some of our services are Node.js, others are Python, and the deployment pipeline needed to support both without branching into separate systems.&lt;/p&gt;

&lt;p&gt;So I designed the image and scripts to be runtime-aware:&lt;/p&gt;

&lt;p&gt;Node.js projects can use the right Node version.&lt;br&gt;
Python projects can use the right Python version.&lt;br&gt;
The SAM deployment process stays consistent.&lt;br&gt;
That means teams can keep their service language choices while still using a common deployment pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Repository Walkthrough&lt;/strong&gt;&lt;br&gt;
The repository is structured to keep responsibilities clear:&lt;/p&gt;

&lt;p&gt;Core pipeline orchestration logic.&lt;br&gt;
Scripts for runtime setup and SAM build/deploy execution.&lt;br&gt;
Example projects (Node.js and Python) showing expected usage.&lt;br&gt;
Tests for parsing, validation, and deployment flow behavior.&lt;br&gt;
Sample CI workflows to demonstrate real integration.&lt;br&gt;
This structure helped keep the project approachable, especially for teams adopting it for the first time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to Use the Pipeline&lt;/strong&gt;&lt;br&gt;
Using the pipeline follows a simple pattern:&lt;/p&gt;

&lt;p&gt;Configure your CI workflow to use the public ECR image.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;public.ecr.aws/z7q5l7x5/sam-pipeline:1.0.0&lt;/span&gt;

&lt;span class="na"&gt;pipelines&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;branches&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;main&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;step&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;script&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./deploy.sh&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Provide deployment variables (accounts, regions, runtime language, stack name, working directory, and optional SAM arguments).&lt;br&gt;
Ensure target AWS accounts have the deploy role and trust policy configured.&lt;br&gt;
Trigger the pipeline.&lt;br&gt;
After that, the container handles runtime setup, account role assumption, and SAM build/deploy flow in a standardized way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benefits &amp;amp; Results&lt;/strong&gt;&lt;br&gt;
Moving to this model gave immediate benefits:&lt;/p&gt;

&lt;p&gt;Standardization: same deployment behavior across repositories.&lt;br&gt;
Lower maintenance: no repeated fixes in every pipeline.&lt;br&gt;
Faster onboarding: new projects plug into an existing pattern.&lt;br&gt;
Better security posture: role assumption with temporary credentials.&lt;br&gt;
More confidence in releases: fewer environment/toolchain surprises.&lt;br&gt;
The biggest gain was reducing operational noise. We now spend more time shipping features and less time fixing deployment plumbing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lessons Learned&lt;/strong&gt;&lt;br&gt;
A few lessons stood out while building this:&lt;/p&gt;

&lt;p&gt;Reusability matters more than cleverness in CI/CD tooling.&lt;br&gt;
Most deployment pain comes from inconsistency, not complexity.&lt;br&gt;
Security and developer productivity can improve together.&lt;br&gt;
Good examples are just as important as good code.&lt;br&gt;
Treat deployment tooling like a product: version it, document it, test it.&lt;br&gt;
Also, making one good shared pipeline is usually easier than maintaining ten almost-identical ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Future Improvements&lt;/strong&gt;&lt;br&gt;
There are still a few things I want to improve:&lt;/p&gt;

&lt;p&gt;Better deployment reporting per account/region.&lt;br&gt;
Optional pre-deployment checks and guardrails.&lt;br&gt;
Stronger observability around failures and retries.&lt;br&gt;
More advanced runtime customization for edge cases.&lt;br&gt;
Cleaner developer experience for local dry-run validation.&lt;br&gt;
The long-term goal is to keep scaling this without increasing operational complexity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Conclusion&lt;/strong&gt;&lt;br&gt;
The biggest lesson from this work is that the real value is not just a Docker image, but a reusable deployment model.&lt;/p&gt;

&lt;p&gt;Instead of maintaining similar deployment logic across many repositories, teams can standardize on one runtime and one SAM execution flow, then securely deploy to multiple accounts using AssumeRole.&lt;/p&gt;

&lt;p&gt;If you want quick adoption, you can use my public image directly:&lt;br&gt;
public.ecr.aws/z7q5l7x5/sam-pipeline:1.0.0&lt;/p&gt;

&lt;p&gt;If your organization requires tighter governance, use the same architecture and publish your own private deployment image.&lt;/p&gt;

&lt;p&gt;Either way, the outcome is the same: less deployment drift, lower maintenance costs, and faster delivery across environments/ multi-account/ regions.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>serverless</category>
      <category>awssam</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
