<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rafagross</title>
    <description>The latest articles on DEV Community by Rafagross (@rafagross).</description>
    <link>https://dev.to/rafagross</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1041701%2F478e8253-afa1-41a1-be42-69758a77c15f.png</url>
      <title>DEV Community: Rafagross</title>
      <link>https://dev.to/rafagross</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rafagross"/>
    <language>en</language>
    <item>
      <title>AWS Organizations in Production: Structure, SCP Evaluation, and Landing Zone Gotchas</title>
      <dc:creator>Rafagross</dc:creator>
      <pubDate>Sun, 27 Sep 2026 05:20:57 +0000</pubDate>
      <link>https://dev.to/rafagross/aws-organizations-in-production-structure-scp-evaluation-and-landing-zone-gotchas-5c0j</link>
      <guid>https://dev.to/rafagross/aws-organizations-in-production-structure-scp-evaluation-and-landing-zone-gotchas-5c0j</guid>
      <description>&lt;p&gt;Most teams adopt AWS Organizations for consolidated billing and stop there. The real value is the control plane: a tree of accounts where guardrails attached at the top apply to everything below. That same tree is also where most "why is this denied?" tickets come from, because the evaluation rules are less intuitive than IAM alone.&lt;/p&gt;

&lt;p&gt;This post covers the structure, how policies are evaluated, and the gotchas worth knowing before you run it in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Account Is the Boundary
&lt;/h2&gt;

&lt;p&gt;An AWS account is the hard boundary for resources, security, and billing. IAM manages identities &lt;em&gt;inside&lt;/em&gt; an account. Organizations manages the accounts themselves.&lt;/p&gt;

&lt;p&gt;That distinction drives the whole multi-account model: separate accounts give you blast-radius isolation, clean cost attribution, and a place to attach controls that no one inside the account can override.&lt;/p&gt;

&lt;h2&gt;
  
  
  Structure
&lt;/h2&gt;

&lt;p&gt;An organization has:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One management account.&lt;/strong&gt; It owns the organization and is the payer account. It should run org and billing tasks only. No workloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One root.&lt;/strong&gt; The top of the tree. Policies attached here apply to every account.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Organizational units (OUs).&lt;/strong&gt; Containers for accounts. OUs can nest up to five levels deep under the root.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Member accounts.&lt;/strong&gt; Where workloads, logging, and security tooling live.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Policies.&lt;/strong&gt; Attached at the root, an OU, or an individual account.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fywtmf5pi85idpcbwwfkq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fywtmf5pi85idpcbwwfkq.png" alt="AWS Organizations account hierarchy: management account, root, Security, Infrastructure and Workloads OUs" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A common baseline layout:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;OU&lt;/th&gt;
&lt;th&gt;Accounts&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;td&gt;Log Archive, Security Tooling (Audit)&lt;/td&gt;
&lt;td&gt;Centralized logs, delegated admin for security services&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infrastructure&lt;/td&gt;
&lt;td&gt;Network, Shared Services&lt;/td&gt;
&lt;td&gt;Transit, DNS, shared tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workloads&lt;/td&gt;
&lt;td&gt;Prod, NonProd&lt;/td&gt;
&lt;td&gt;Application accounts, grouped by lifecycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sandbox&lt;/td&gt;
&lt;td&gt;Per-developer or per-team&lt;/td&gt;
&lt;td&gt;Experimentation with looser controls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Suspended&lt;/td&gt;
&lt;td&gt;Accounts pending closure&lt;/td&gt;
&lt;td&gt;Deny-all quarantine&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Tip:&lt;/strong&gt; Design OUs around the controls and lifecycle you want to apply, not around the org chart. Teams get reorganized; "prod needs stricter guardrails than sandbox" does not change.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Enable All Features
&lt;/h2&gt;

&lt;p&gt;Organizations has two feature sets. Consolidated-billing mode only merges invoices. &lt;strong&gt;All features&lt;/strong&gt; adds the authorization policies (SCPs, RCPs), management policies (tag, backup, EC2), trusted access for service integrations, and delegated administrators. If you are building a landing zone, you need all features.&lt;/p&gt;

&lt;h2&gt;
  
  
  Policy Types That Matter Day to Day
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Policy&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;SCP&lt;/strong&gt; (service control policy)&lt;/td&gt;
&lt;td&gt;Sets the maximum permissions for IAM users and roles in member accounts. Grants nothing on its own.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;RCP&lt;/strong&gt; (resource control policy)&lt;/td&gt;
&lt;td&gt;Sets the maximum permissions on resources in member accounts (S3, KMS, Secrets Manager, SQS, and many more), regardless of who is calling. Grants nothing on its own.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tag policy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Standardizes tag keys and values so cost allocation and ABAC stay consistent.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Backup policy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Deploys AWS Backup plans centrally across accounts and Regions.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;EC2 policy&lt;/strong&gt; (declarative)&lt;/td&gt;
&lt;td&gt;Enforces EC2, VPC, and EBS baselines org-wide, such as IMDS defaults, AMI and EBS snapshot block public access, serial console access, and VPC Block Public Access. The configuration is maintained even as the service adds new APIs.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9crbszm8q5gzl06lf1eg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9crbszm8q5gzl06lf1eg.png" alt="SCP inheritance and request evaluation flow across identity policies, SCPs and RCPs" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How Inheritance Actually Works
&lt;/h2&gt;

&lt;p&gt;This is the part that trips people up.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;Deny&lt;/strong&gt; is simple: attach it anywhere and it applies to everything below that point.&lt;/p&gt;

&lt;p&gt;An &lt;strong&gt;Allow&lt;/strong&gt; is not inherited the way people expect. For an action to be permitted in an account, an SCP must allow it at &lt;strong&gt;every&lt;/strong&gt; level from the root down to the account: the root, each OU in the path, and the account itself. If any level is missing the Allow, the action is implicitly denied, even if an IAM policy in the account grants &lt;code&gt;*:*&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Root            FullAWSAccess        -&amp;gt; allows *
  OU Workloads  AllowOnlyApprovedSvc -&amp;gt; allows ec2:*, s3:*, ... (no dynamodb:*)
    Prod acct   FullAWSAccess        -&amp;gt; allows *

Result in Prod: dynamodb:* is denied. The OU level never allowed it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why the default &lt;code&gt;FullAWSAccess&lt;/code&gt; SCP exists at every level. If you replace it with an allow-list at the OU, that allow-list becomes the ceiling for every account underneath.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Most common SCP incident:&lt;/strong&gt; Someone detaches &lt;code&gt;FullAWSAccess&lt;/code&gt; from an OU while testing an allow-list SCP, and every account under that OU loses access to services that weren't on the list. Prefer deny-list SCPs (explicit &lt;code&gt;Deny&lt;/code&gt; statements on top of &lt;code&gt;FullAWSAccess&lt;/code&gt;) unless you have a specific reason to allow-list.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  How a Request Is Evaluated
&lt;/h2&gt;

&lt;p&gt;For a principal in a member account, access requires an Allow from every applicable layer:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Identity-based policy&lt;/strong&gt; on the user or role must allow the action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SCPs&lt;/strong&gt; in the account's path must allow it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RCPs&lt;/strong&gt; in the path must allow it for the target resource (if the service supports RCPs).&lt;/li&gt;
&lt;li&gt;Permission boundaries, session policies, and resource-based policies still apply as usual.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;An &lt;strong&gt;explicit Deny at any layer wins.&lt;/strong&gt; SCPs and RCPs only ever reduce what is possible; they never grant access.&lt;/p&gt;

&lt;p&gt;A practical consequence: when you debug an &lt;code&gt;AccessDenied&lt;/code&gt; in a member account and the IAM policy looks correct, check the SCPs on every node in the account's path, not just the ones attached directly to the account.&lt;/p&gt;

&lt;h2&gt;
  
  
  Landing Zone Pattern
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fulbbzeflorebqrt6k01p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fulbbzeflorebqrt6k01p.png" alt="Multi-account landing zone pattern and operating lifecycle" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The goal is to separate the governance plane from workloads:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Management account:&lt;/strong&gt; org and billing only. Tightly restricted human access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security OU:&lt;/strong&gt; Log Archive receives org-wide CloudTrail and Config data. Security Tooling (Control Tower calls it Audit) is the delegated administrator for GuardDuty, Security Hub, and similar services.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infrastructure OU:&lt;/strong&gt; networking and shared services.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workloads OU:&lt;/strong&gt; prod and non-prod split so guardrails can differ.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;AWS Control Tower&lt;/strong&gt; orchestrates this on top of Organizations: Account Factory for provisioning, managed controls (preventive via SCPs/RCPs, detective via AWS Config, proactive via CloudFormation hooks), and drift detection. If you manage infrastructure as code, Account Factory for Terraform (AFT) provisions accounts through a Terraform pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Operating Lifecycle
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Identity:&lt;/strong&gt; IAM Identity Center for human access, cross-account roles for automation, least privilege throughout.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provisioning:&lt;/strong&gt; Account Factory, AFT, or the Organizations API. Never click-ops accounts in production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Baseline:&lt;/strong&gt; logging, Config, security services, and networking applied from day one, typically via StackSets or Control Tower.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails:&lt;/strong&gt; preventive, detective, and proactive controls, plus drift detection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lifecycle:&lt;/strong&gt; move accounts between OUs as their purpose changes. To retire one, move it to a Suspended OU with a deny-all SCP, then close it. A closed account can be reopened during the post-closure period (90 days).&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Production Gotchas
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SCPs and RCPs never restrict the management account.&lt;/strong&gt; Anything you run there is outside your guardrails. That is the main reason to keep workloads out of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SCPs don't affect service-linked roles.&lt;/strong&gt; AWS services acting through service-linked roles are not limited by your SCPs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SCPs do apply to delegated administrator accounts.&lt;/strong&gt; They are member accounts. Make sure your guardrails don't block the security services you delegated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OUs nest at most five levels deep.&lt;/strong&gt; Deep hierarchies are also harder to reason about during an incident. Keep it shallow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delegate security services to the Audit/Security Tooling account&lt;/strong&gt; instead of administering them from the management account.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Limits:&lt;/strong&gt; SCP documents max out at 10,240 characters, and each root, OU, or account can have up to 10 SCPs attached. Plan statement consolidation early.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Quick Inventory Commands
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Root ID and enabled policy types&lt;/span&gt;
aws organizations list-roots

&lt;span class="c"&gt;# All accounts in the org&lt;/span&gt;
aws organizations list-accounts

&lt;span class="c"&gt;# OUs under a parent (root or OU)&lt;/span&gt;
aws organizations list-organizational-units-for-parent &lt;span class="nt"&gt;--parent-id&lt;/span&gt; r-examplerootid

&lt;span class="c"&gt;# SCPs attached to a specific target&lt;/span&gt;
aws organizations list-policies-for-target &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--target-id&lt;/span&gt; 123456789012 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter&lt;/span&gt; SERVICE_CONTROL_POLICY
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run these from the management account or a delegated administrator.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_introduction.html" rel="noopener noreferrer"&gt;What is AWS Organizations?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_getting-started_concepts.html" rel="noopener noreferrer"&gt;AWS Organizations terminology and concepts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies_scps.html" rel="noopener noreferrer"&gt;Service control policies (SCPs)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies_rcps.html" rel="noopener noreferrer"&gt;Resource control policies (RCPs)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies_ec2.html" rel="noopener noreferrer"&gt;EC2 policies&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_best-practices_mgmt-acct.html" rel="noopener noreferrer"&gt;Best practices for the management account&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_ous_best_practices.html" rel="noopener noreferrer"&gt;Best practices for organizational units&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_reference_limits.html" rel="noopener noreferrer"&gt;Quotas for AWS Organizations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/controltower/latest/userguide/what-is-control-tower.html" rel="noopener noreferrer"&gt;What is AWS Control Tower?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/controltower/latest/userguide/aws-multi-account-landing-zone.html" rel="noopener noreferrer"&gt;AWS multi-account landing zone&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aws</category>
      <category>devops</category>
      <category>security</category>
      <category>cloud</category>
    </item>
    <item>
      <title>EC2 Instance Unreachable via SSM Session Manager</title>
      <dc:creator>Rafagross</dc:creator>
      <pubDate>Sun, 13 Sep 2026 03:12:04 +0000</pubDate>
      <link>https://dev.to/rafagross/ec2-instance-unreachable-via-ssm-session-manager-1fe9</link>
      <guid>https://dev.to/rafagross/ec2-instance-unreachable-via-ssm-session-manager-1fe9</guid>
      <description>&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;This runbook resolves the situation where an EC2 instance is running and visible in the console but either does not appear in the Systems Manager Fleet Manager inventory, or returns a connection error when you attempt to start a Session Manager session.&lt;/p&gt;




&lt;h2&gt;
  
  
  When to Use This Runbook
&lt;/h2&gt;

&lt;p&gt;Use this runbook when you observe any of the following:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Session Manager shows "Instance not connected" or "Unable to start session"&lt;/li&gt;
&lt;li&gt;The instance does not appear in &lt;strong&gt;Systems Manager → Fleet Manager → Managed Nodes&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ssm:StartSession&lt;/code&gt; returns &lt;code&gt;TargetNotConnected&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;A previously working instance stopped responding to SSM after a restart, IAM change, or network modification&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Step 1: Verify Instance State
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; EC2 → Instances → [Instance ID]&lt;/p&gt;

&lt;p&gt;Confirm the instance is in &lt;code&gt;running&lt;/code&gt; state and that both &lt;strong&gt;Status checks&lt;/strong&gt; show 2/2 passed.&lt;/p&gt;

&lt;p&gt;If status checks are failing, stop here and use the EC2 Status Check runbook. SSM is irrelevant if the instance itself is impaired.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expected result:&lt;/strong&gt; Instance state = &lt;code&gt;running&lt;/code&gt;, 2/2 status checks passed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 2: Check SSM Managed Node Status
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; Systems Manager → Fleet Manager → Managed nodes&lt;/p&gt;

&lt;p&gt;Search for the instance ID. If it does not appear, or shows &lt;code&gt;Connection Lost&lt;/code&gt;, the issue is one of: missing IAM role, stopped SSM agent, or blocked network path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expected result:&lt;/strong&gt; Instance appears with status &lt;code&gt;Online&lt;/code&gt;. If missing or &lt;code&gt;Connection Lost&lt;/code&gt;, continue to Step 3.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3: Verify IAM Instance Profile
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; EC2 → Instances → [Instance ID] → Security tab → IAM Role&lt;/p&gt;

&lt;p&gt;The instance must have an IAM role attached with the &lt;code&gt;AmazonSSMManagedInstanceCore&lt;/code&gt; managed policy (or equivalent custom policy granting the minimum SSM actions).&lt;/p&gt;

&lt;p&gt;Minimum required IAM actions if using a custom policy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ssm:UpdateInstanceInformation
ssm:ListInstanceAssociations
ssm:DescribeInstanceProperties
ssm:DescribeDocumentParameters
ssmmessages:CreateControlChannel
ssmmessages:CreateDataChannel
ssmmessages:OpenControlChannel
ssmmessages:OpenDataChannel
ec2messages:AcknowledgeMessage
ec2messages:DeleteMessage
ec2messages:FailMessage
ec2messages:GetEndpoint
ec2messages:GetMessages
ec2messages:SendReply
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; IAM propagation delay:&lt;br&gt;
IAM role changes applied to a running instance take effect on next SSM agent heartbeat, typically within 2-3 minutes. If you just updated the role, wait 5 minutes and re-check Fleet Manager before continuing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Expected result:&lt;/strong&gt; IAM role attached with &lt;code&gt;AmazonSSMManagedInstanceCore&lt;/code&gt; or equivalent. If missing, attach the role and wait 5 minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Check SSM Agent on the Instance
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; EC2 → Instances → [Instance ID] → Actions → Monitor and troubleshoot → Get system log&lt;/p&gt;

&lt;p&gt;Look for lines referencing &lt;code&gt;amazon-ssm-agent&lt;/code&gt;. A healthy agent shows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight systemd"&gt;&lt;code&gt;&lt;span class="err"&gt;amazon-ssm-agent.service:&lt;/span&gt; &lt;span class="err"&gt;active&lt;/span&gt; &lt;span class="err"&gt;(running)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Errors to look for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Failed to start Amazon SSM Agent&lt;/code&gt; — agent start failure, likely OS-level issue&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Error connecting to endpoint&lt;/code&gt; — network path problem (proceed to Step 5)&lt;/li&gt;
&lt;li&gt;No SSM lines at all — agent not installed or not running&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the agent is not running, you can use EC2 Run Command with &lt;code&gt;AWS-RunShellScript&lt;/code&gt; to restart it — but only if the instance is already registered (partial connectivity). If it's fully unreachable, use EC2 Instance Connect or serial console (if available) to restart manually.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expected result:&lt;/strong&gt; SSM agent running. No connection errors.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Diagnose the Network Path
&lt;/h2&gt;

&lt;p&gt;SSM Session Manager requires outbound HTTPS (port 443) from the EC2 instance to three regional endpoints:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Endpoint&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ssm.{region}.amazonaws.com&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Agent registration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ssmmessages.{region}.amazonaws.com&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Session data channel&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ec2messages.{region}.amazonaws.com&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;EC2 message delivery&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;For private instances (no NAT Gateway, VPC endpoint required):&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; VPC → Endpoints&lt;/p&gt;

&lt;p&gt;Verify that all three interface endpoints exist, are associated with the correct VPC, and have status &lt;code&gt;available&lt;/code&gt;. Check that the endpoint security group allows inbound HTTPS from the instance's security group.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; VPC → Security Groups → [Endpoint SG]&lt;/p&gt;

&lt;p&gt;Inbound rule required:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;Type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;HTTPS (443)&lt;/span&gt;
&lt;span class="py"&gt;Source&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;[Instance security group ID]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;For instances with NAT Gateway:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; VPC → Route Tables → [Instance subnet's route table]&lt;/p&gt;

&lt;p&gt;Confirm a &lt;code&gt;0.0.0.0/0&lt;/code&gt; route pointing to the NAT Gateway exists and the NAT Gateway is in &lt;code&gt;available&lt;/code&gt; state.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt;&lt;br&gt;
If the instance has a public IP and is in a public subnet with an IGW route, outbound port 443 directly to AWS endpoints is sufficient. No VPC endpoints needed in that case.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Expected result:&lt;/strong&gt; Either VPC endpoints present and &lt;code&gt;available&lt;/code&gt;, or valid NAT/IGW route exists. Instance SG allows outbound 443.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 6: Check Instance Security Group Outbound Rules
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; EC2 → Instances → [Instance ID] → Security tab → Security groups → [SG ID] → Outbound rules&lt;/p&gt;

&lt;p&gt;Confirm outbound HTTPS is allowed. The default AWS SG allows all outbound traffic. If rules have been tightened:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;Type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;HTTPS (443)&lt;/span&gt;
&lt;span class="py"&gt;Destination&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;0.0.0.0/0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OR (preferred, more restrictive):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;Type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;HTTPS (443)&lt;/span&gt;
&lt;span class="py"&gt;Destination&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;[VPC endpoint prefix list or specific endpoint IPs]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Expected result:&lt;/strong&gt; Outbound 443 allowed from instance security group.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 7: Force SSM Agent Re-Registration (Last Resort)
&lt;/h2&gt;

&lt;p&gt;If all above checks pass but the instance still doesn't appear in Fleet Manager, the agent's registration may be stale (common after AMI snapshots or instance cloning).&lt;/p&gt;

&lt;p&gt;Use EC2 Run Command with &lt;code&gt;AWS-RunShellScript&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Amazon Linux 2 / AL2023&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl stop amazon-ssm-agent
&lt;span class="nb"&gt;sudo rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; /var/lib/amazon/ssm/registration
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl start amazon-ssm-agent
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl status amazon-ssm-agent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wait 2 minutes, then re-check Fleet Manager.&lt;/p&gt;




&lt;h2&gt;
  
  
  Validation Checks
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;How to Verify&lt;/th&gt;
&lt;th&gt;Expected Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Instance in Fleet Manager&lt;/td&gt;
&lt;td&gt;SSM → Fleet Manager → search instance ID&lt;/td&gt;
&lt;td&gt;Status: Online&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session starts&lt;/td&gt;
&lt;td&gt;SSM → Session Manager → Start session → select instance&lt;/td&gt;
&lt;td&gt;Terminal opens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent heartbeat&lt;/td&gt;
&lt;td&gt;CloudWatch Logs → /aws/ssm/amazon-ssm-agent&lt;/td&gt;
&lt;td&gt;Recent heartbeat log entries&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Rollback
&lt;/h2&gt;

&lt;p&gt;No destructive changes are made by this runbook. If you attached a new IAM role that you want to remove:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Console path:&lt;/strong&gt; EC2 → Instances → [Instance ID] → Actions → Security → Modify IAM role → No role&lt;/p&gt;




&lt;h2&gt;
  
  
  Escalation
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Next Action&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;All steps pass but instance still unreachable&lt;/td&gt;
&lt;td&gt;Open AWS Support case — provide instance ID, region, VPC endpoint IDs, and CloudWatch agent logs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent crashes immediately on start&lt;/td&gt;
&lt;td&gt;Check OS disk space (&lt;code&gt;df -h&lt;/code&gt;), memory, and &lt;code&gt;/var/log/amazon/ssm/amazon-ssm-agent.log&lt;/code&gt; for errors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instance in private subnet, no VPC endpoints, no NAT&lt;/td&gt;
&lt;td&gt;Network path is broken by design — work with network team to add VPC endpoints or NAT&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




</description>
      <category>aws</category>
      <category>devops</category>
      <category>linux</category>
      <category>cloud</category>
    </item>
  </channel>
</rss>
