DEV Community

AdminPackStudio
AdminPackStudio

Posted on

Windows Autopilot ESP Stuck? The Complete Runway Guide: Hash, Profile Assignment, Blocking Apps, Network/TLS, and CA Deadlocks

Windows Autopilot ESP Stuck? The Complete Runway Guide

Short version: when Windows Autopilot stalls on the Enrollment Status Page (ESP), the cause almost always sits on one of five layers: registration (hash), profile assignment, ESP blocking design, network/TLS, or Conditional Access. Work through them in that order and collect evidence with read-only tools before anyone resets the device.

This is the long version of an earlier, shorter triage checklist. That one is a single page for the moment a laptop is stuck. This one is the runway: what each layer does, how it fails, what you'll see, how to prove it, and how to stop it coming back on the next dock day.

No Microsoft affiliation. Operational guidance for admins who are authorized to manage their tenant and devices. Check current Microsoft Learn docs for your OS build and cloud, because endpoints and ESP behavior change over time.


Table of contents

  1. How Autopilot and ESP actually fit together
  2. Before you touch anything: capture the state
  3. Layer 1: Hardware hash and Autopilot registration
  4. Layer 2: Deployment profile assignment
  5. Layer 3: ESP blocking apps, policies, and timeouts
  6. Layer 4: Network, proxy, and TLS inspection
  7. Layer 5: The Conditional Access enrollment deadlock
  8. Hybrid join sidebar: when ESP takes the blame for AD
  9. Read-only diagnostics toolkit
  10. Symptom to layer lookup table
  11. Escalation note template
  12. Making dock day boring: prevention habits
  13. FAQ

How Autopilot and ESP actually fit together

People say "Autopilot is stuck" for problems in very different places, so it helps to split the flow into stages:

Stage What happens What has to be true
OOBE network Device gets online Reachable internet, no captive portal, sane clock
Profile download Device asks the Autopilot service "who am I?" Hash registered and a deployment profile assigned to this device
Identity / join User signs in (user-driven) or device authenticates (self-deploying / pre-provisioning) Join type matches profile; TPM attestation works where required
MDM enrollment Device enrolls into Intune License, MDM user scope, CA doesn't block enrollment
ESP device phase Device-targeted policies, certificates, and blocking apps install Assignments reach the device; content downloads; apps detect correctly
ESP account phase User-targeted policies and apps Assignments reach the user; CA allows token issuance for the session
Desktop User can work ESP released

Two things to keep in mind for the rest of this guide:

  • ESP only waits for what it's told to wait for. If a blocking app or policy never succeeds, or never targets the device, ESP keeps waiting until it times out.
  • A failure in an early stage often shows up later. A profile assignment gap can look like an ESP hang, and a CA policy can look like a broken app. The layered order exists so you don't fix the wrong thing.

Before you touch anything: capture the state

The most expensive mistake in Autopilot troubleshooting is resetting a device before you know why it failed. A reset removes local evidence, and if the cause is assignment, network, or CA, the next attempt fails the same way.

Capture these first:

  • [ ] Serial number and model
  • [ ] Exact ESP text and phase: "Device preparation", "Device setup", or "Account setup", plus the item it's stuck on
  • [ ] Elapsed time and the timeout configured in the ESP profile
  • [ ] Network type: corporate wired, corporate Wi-Fi, home, hotel/guest, branch
  • [ ] Error code if the ESP shows one (photo is fine)
  • [ ] Did a sister device succeed on the same network with the same profile today?

During OOBE, Shift+F10 usually opens a command prompt (unless you've disabled it by policy). That's enough to run the read-only checks in the diagnostics section.


Layer 1: Hardware hash and Autopilot registration

What it does

Autopilot identifies the device from its hardware hash. No registered hash means no profile, which means a generic, consumer-style OOBE however good your Intune configuration is.

How it fails

Failure What you see
OEM / partner never uploaded this order Generic OOBE on brand-new hardware
Hash captured but CSV import failed or was never done Device absent from Autopilot devices list
Motherboard replaced after registration Device behaves as unregistered or mismatched
Duplicate / stale objects from a prior life Wrong profile, wrong Group Tag, confusing assignment
Group Tag typo (CORP-STD vs CORP_STD) Registered but never lands in the dynamic group

How to prove it

  • Look the serial up in the Autopilot devices list. Is it there, and is there exactly one entry?
  • Check Group Tag / Order ID against the dynamic group rule character for character.
  • Check the device's profile status. "Not assigned" after a sync points to Layer 2, not Layer 1.

How to fix it

Register the hash (OEM/partner, or capture and import in your normal process), wait for the service to sync, then confirm a profile is assigned before restarting OOBE. Resetting an unregistered device won't make a profile appear.

Prevention: one source of truth that maps purchase order, asset tag, serial, and hash, plus a pre-ship check that every device appears with the correct tag.


Layer 2: Deployment profile assignment

What it does

The deployment profile tells the device which mode to use (user-driven, self-deploying, pre-provisioning), which join type (Entra or hybrid), and how OOBE should behave. The profile has to reach the device before the user signs in.

The device-group vs user-group trap

This causes more "Autopilot is haunted" tickets than anything else:

  • Deployment profiles should target device groups that contain Autopilot device objects, often through a dynamic rule on Group Tag or ZTDId.
  • Assigning a deployment profile to a user group doesn't work the way people expect, because during OOBE there's no user yet. The device has to be in scope on its own.
  • Overlapping assignments (the device is in two groups with two profiles) give unpredictable results. Aim for one device, one winning profile.

Dynamic group lag

Dynamic membership is not instant. Right after registration a device may not be in the group yet, and then the profile isn't assigned yet. Testing too early causes plenty of false failures. Check membership and profile status first, then start OOBE.

Checklist

  • [ ] Device is a member of the targeted device group (check the group, not just the rule)
  • [ ] Profile status for the device shows assigned, and it's the profile you expected
  • [ ] Mode and join type match the scenario (self-deploying needs TPM 2.0 and no user affinity; hybrid needs directory prerequisites)
  • [ ] ESP profile targets the same population as the deployment profile
  • [ ] Profile names carry intent and version, e.g. AP-UserDriven-Entra-Std-v3

Tip: if the device showed the right branding and sign-in page, Layer 2 probably worked. If it showed a generic OOBE or the wrong join behavior, stay here.


Layer 3: ESP blocking apps, policies, and timeouts

What it does

ESP holds the desktop until selected apps and policies finish. Done well, users get a ready-to-work PC. Done badly, it turns every flaky installer into a dock-day outage.

Common ESP blockers

Blocker Why it stalls ESP
Too many blocking apps Every extra app adds download time and another chance to fail
Win32 detection rule that never matches App installs fine, Intune thinks it didn't, ESP waits
Requirement rules (OS build, architecture, disk) excluding the device App never applies, so "blocking" means waiting forever
Dependencies / supersedence chains One failure deep in the chain stops everything
Installer needs user context or a reboot nobody planned Exit codes misread; hard reboot mid-ESP
Mixed app types colliding Different install channels fighting during provisioning
Timeout shorter than real-world p95 Fine in the lab on gigabit, fails on branch Wi-Fi

Design rules that keep ESP sane

  1. Block only what's needed before the desktop. Usually security agent, Company Portal, and the bare minimum of line-of-business apps. A starter cap of around five proven apps is a reasonable first production ring.
  2. New apps are never ESP-blocking on day one. Deploy as required/available outside ESP until installs are reliable, then promote.
  3. Detection must reflect what the installer actually leaves behind. Validate on a clean device, not your packaging VM.
  4. Run a timing study. Ten pilot devices across office, home, and branch networks. Record minutes to ESP complete. If your p95 goes past the timeout, shrink the blocking set or fix the app before you raise the timeout.
  5. Device ESP vs account ESP. Know which phase each blocking item lands in. User-targeted items stall the account phase, device-targeted items stall the device phase.

When ESP names a specific app

Don't raise the timeout first. Check:

  • Install status for that app on that device
  • The install command, return codes, and the detection rule against what's actually on disk or in the registry
  • Requirements and dependencies
  • The Intune Management Extension logs (see diagnostics)

Fix the package, test it outside ESP, and only then put it back in the blocking list.


Layer 4: Network, proxy, and TLS inspection

Why network problems look like ESP problems

ESP downloads policy, certificates, and Win32 content from cloud endpoints. If the network quietly interferes, ESP sees "not installed yet", not "network broken". The classic sign is: works on the lab bench, fails at the branch or on home Wi-Fi.

Network failure modes

Problem Signal
Captive portal / guest Wi-Fi Online for the first sign-in, then content stops after the session expires
Proxy requiring user authentication OOBE and system-context downloads can't do interactive proxy auth
TLS / SSL inspection Certificate mismatch on endpoints that expect the real Microsoft chain; enrollment or content download fails without a clear message
Firewall category blocks Some endpoints allowed, CDN / content endpoints blocked
DNS filtering Name resolution fails for content or attestation hosts
Clock skew Token and certificate validation fails in ways that look random
TPM attestation blocked (self-deploying / pre-provisioning) Device can't reach the TPM manufacturer certificate endpoints and attestation fails early
Low bandwidth + large blocking payload Not "broken", just slower than the ESP timeout

What to do

  • [ ] Compare with a known-good network (wired corporate or a phone hotspot). If the same device succeeds, the problem is the path, not the profile.
  • [ ] Use the current Microsoft network endpoint guidance for Autopilot, Entra, Intune, and Win32 content delivery in your cloud, and allowlist through your normal change process.
  • [ ] Exclude enrollment and content endpoints from TLS inspection where your security process allows. Inspection appliances break more provisioning than almost anything else.
  • [ ] Keep provisioning on a network segment without captive portals or per-user proxy auth.
  • [ ] Check the device clock / time sync before you chase token errors.
  • [ ] Include at least one branch device in every ESP timing study.

If several devices enroll fine but then fail the same Win32 installs, fix delivery before you rewrite detection rules again.


Layer 5: The Conditional Access enrollment deadlock

The chicken-and-egg

A device has to enroll before it can be compliant. Compliance evaluation also needs time after enrollment. A broad CA policy such as "All cloud apps → require compliant device" can block the sign-ins and token requests the device needs during enrollment and the ESP account phase. You get a loop: can't finish enrollment → can't become compliant → CA blocks → can't finish enrollment.

What it looks like

  • ESP reaches the account setup phase, then fails or prompts for sign-in again
  • Errors mention the device not being compliant, managed, or registered
  • Sign-in logs show CA failures for the user during the provisioning window
  • It started right after someone enforced a CA policy that worked fine in report-only

How to prove it

  • Pull the user's sign-in logs for the provisioning window. Which policy applied, and which grant control failed?
  • Use What If with the user, the app, and the device platform to see which policies would apply.
  • Check whether the policy is report-only or on. If report-only shows it would have failed, you've found the likely future deadlock before it happens.

How to break the loop safely

  • [ ] Follow current Microsoft guidance on how enrollment-related apps should be handled in "require compliant device" policies. Don't invent broad exclusions under pressure.
  • [ ] Keep new device-compliance CA policies in report-only until Autopilot success rates are stable.
  • [ ] Watch MFA and registration requirements (for example, "register or join devices" user actions) so they don't fight the OOBE sign-in.
  • [ ] Keep break-glass accounts excluded and monitored, as with every CA rollout.
  • [ ] Document every exclusion with an owner and a review date.

Rule: keep CA enforcement one step behind your Autopilot success rate. Tighten after provisioning is boring, not before.


Hybrid join sidebar: when ESP takes the blame for AD

Hybrid Autopilot adds failure domains: the on-prem directory, sync, and a line of sight to domain controllers. Plenty of "ESP hangs" are really hybrid prerequisites failing quietly.

Check before blaming ESP:

  • [ ] Directory sync healthy
  • [ ] Service connection point (SCP) configured correctly
  • [ ] Provisioning network can reach domain controllers when the domain-join step runs
  • [ ] Offline domain join connector / components healthy, with rights to create computer objects in the target OU
  • [ ] Naming template respects AD / NetBIOS limits

If your apps don't strictly need hybrid join on day one, it's worth reviewing whether Entra join could be the default with hybrid as a documented exception. That's an architecture decision for your org, not a quick fix in the middle of an incident.


Read-only diagnostics toolkit

Everything in this section reads state or collects logs. None of it enrolls, resets, wipes, or changes policy. Only run it on devices you're authorized to support.

1. Join and MDM state: dsregcmd /status

# Read-only: summarize join + MDM signals
$raw = dsregcmd /status | Out-String
$pick = 'AzureAdJoined','DomainJoined','TenantName','DeviceId','MDMUrl','AzureAdPrt'
foreach ($k in $pick) {
    if ($raw -match "(?m)^\s*$k\s*:\s*(.*)$") {
        '{0,-14} {1}' -f $k, $Matches[1].Trim()
    }
}
Enter fullscreen mode Exit fullscreen mode

How to read it:

  • AzureAdJoined : NO after the user signed in → join stage failed (Layer 2 / identity)
  • MDMUrl empty → not enrolled into MDM (license, MDM user scope, or CA)
  • AzureAdPrt : NO during the account phase → token problems; look at CA and network

2. Autopilot and MDM event logs

# Read-only: recent Autopilot + MDM enrollment events
$logs = @(
  'Microsoft-Windows-ModernDeployment-Diagnostics-Provider/Autopilot',
  'Microsoft-Windows-DeviceManagement-Enterprise-Diagnostics-Provider/Admin'
)
foreach ($l in $logs) {
  Get-WinEvent -LogName $l -MaxEvents 30 -ErrorAction SilentlyContinue |
    Select-Object TimeCreated, Id, LevelDisplayName,
      @{n='Msg';e={ ($_.Message -split "`n")[0] }}
}
Enter fullscreen mode Exit fullscreen mode

Look for profile download results, enrollment errors, and repeated failures around the time the ESP stalled.

3. Win32 app logs (Intune Management Extension)

Win32 installs and detection results are logged under:

C:\ProgramData\Microsoft\IntuneManagementExtension\Logs\
Enter fullscreen mode Exit fullscreen mode

IntuneManagementExtension.log (and related logs in that folder) shows download progress, exit codes, and detection results. A "detection rule not satisfied" after a successful install exit code is the classic Layer 3 detection bug.

4. Built-in diagnostic bundle

Windows includes an MDM diagnostics tool that collects logs into an archive for offline review:

mdmdiagnosticstool.exe -area "DeviceEnrollment;DeviceProvisioning;Autopilot" -cab C:\Temp\ap-diag.cab
Enter fullscreen mode Exit fullscreen mode

It writes a log bundle and doesn't change device configuration. Attach the bundle to the escalation instead of resetting first.

5. Local snapshot to CSV

For dock days with many devices, a small read-only script that writes join state, tenant, MDM URL, and OS version to CSV makes it easy to compare a failing device with a working one. The Autopilot pack below includes one (Get-LocalAutopilotEspSnapshot.ps1). It only reads local signals and writes an optional CSV.


Symptom to layer lookup table

Symptom Most likely layer First check
Generic OOBE, no org branding 1 Registration / 2 Assignment Serial in Autopilot devices; profile status
Wrong join type or mode 2 Assignment Overlapping profiles; device group membership
Stuck in "Device preparation" 4 Network / TPM Known-good network test; attestation reachability
Stuck in "Device setup" on an app 3 ESP blocker IME log; detection rule; requirements
Stuck in "Device setup" on policies/certs 3 ESP / 4 Network Assignment of the blocking items; content reachability
Fails in "Account setup" with sign-in loop 5 CA deadlock Sign-in logs; What If; report-only results
Works on HQ wired, fails on branch/home 4 Network Proxy, TLS inspection, captive portal, bandwidth
Hybrid device fails at domain join Hybrid sidebar Directory sync, SCP, DC line of sight
Pre-provisioning (technician phase) red screen 3 / 4 Apps in technician phase; network; TPM

Escalation note template

Paste this into the ticket so the next person doesn't start from zero:

AUTOPILOT / ESP ESCALATION
Serial / model:
Autopilot registered (Y/N), Group Tag:
Device group member (Y/N), deployment profile assigned:
Mode / join type:
ESP phase + item stuck on:
Elapsed time vs ESP timeout:
Network type (and known-good network test result):
dsregcmd: AzureAdJoined / MDMUrl / AzureAdPrt:
CA sign-in log result (policy + grant failed):
Logs attached (IME / diag cab):
Sister device on same profile+network succeeded? (Y/N)
Actions taken so far (no reset yet? Y/N):
Enter fullscreen mode Exit fullscreen mode

Making dock day boring: prevention habits

  • Pre-ship registration check: every device present, one object, correct tag.
  • One primary profile covering most of the fleet; exceptions get a second profile, not seven.
  • ESP blocking cap with a promotion rule: only apps that have installed reliably outside ESP.
  • Timing study on every major change (new blocking app, new site, new OS build).
  • Provisioning network standard: no captive portal, no user-auth proxy, TLS inspection exclusions approved.
  • CA changes go through report-only and are checked against Autopilot sign-ins before enforcement.
  • Written reset criteria: a reset/retry happens only after the layers above are checked and evidence is collected, and always under your org's change and device policy.

FAQ

Why is my Autopilot device stuck on "Setting up your device for work"?
Usually an ESP blocking app or policy that never succeeds or never targets the device. Check which phase it's in, which item is pending, and the Intune Management Extension log for that app.

Why does Autopilot show a normal OOBE instead of my company's?
The device either isn't registered (no hash) or has no deployment profile assigned yet. Check the serial in Autopilot devices, device group membership, and profile status.

Should the deployment profile be assigned to users or devices?
Devices. The profile has to reach the device before anyone signs in, so target a device group containing the Autopilot device objects.

Can Conditional Access break Autopilot?
Yes. A "require compliant device" policy that also applies to the sign-ins needed during enrollment can create a loop. Validate with sign-in logs and What If, and keep new compliance policies in report-only until Autopilot is stable.

Can TLS inspection cause ESP failures?
Yes. Inspection can break certificate validation for enrollment and content endpoints. Exclude them according to current Microsoft guidance and your security process.

Should I just increase the ESP timeout?
Only after you've fixed failing apps and confirmed content delivery. A longer timeout on a broken app just makes users wait longer for the same failure.

When is resetting the device the right call?
When registration, assignment, ESP design, network, and CA all check out, the evidence is collected, and the device is in a bad local state. Do it under your org's policy and confirm hash and profile are still correct before the next OOBE.


Want the full Autopilot runbooks?

This guide gives you the triage order. If you'd like the rest written down, the Windows Autopilot & ESP Ops Pack ($29) from Admin Pack Studio has five markdown modules:

  • Registration hygiene (OEM/hash paths, Group Tag strategy, duplicate cleanup, hand-off template)
  • Deployment profiles that stick (decision matrix, assignment rules, hybrid gates, change-control snippet)
  • ESP design and timing (blocking-set rules, configuration checklist, timing study template)
  • Cloud vs hybrid join ops (decision prompts, migration path, exec one-pager)
  • Seven break/fix cards (generic OOBE, ESP hangs, failing Win32, hybrid join, branch-only failures, CA deadlock, pre-provisioning)

It also includes the read-only local snapshot script mentioned above. The script reads local signals and writes an optional CSV. It does not enroll, reset, wipe, or change policy.

👉 https://cashflow4375.gumroad.com/l/hlanei

If your problem is more about enrollment and compliance than Autopilot itself, the Intune & M365 Admin Starter Pack covers enrollment hygiene, baseline compliance with report-only CA phasing, and break/fix cards. It's $19 during launch week (normally $29): https://cashflow4375.gumroad.com/l/joonf

The checklist above works fine without either pack. Questions and war stories welcome in the comments.

— Admin Pack Studio


Not affiliated with or endorsed by Microsoft. Windows, Intune, Autopilot, and Entra are trademarks of their respective owners. Operational guidance for admins authorized to manage their tenant and devices. Test in a pilot first; provided AS-IS, no warranty.

Top comments (0)