MFA fatigue is the attack that works because people are tired. The attacker already has the password, usually from a phishing kit or a credential dump, and MFA is the only thing left in the way. So they hammer it. Push notification after push notification, sometimes at two in the morning, until the person on the other end approves one just to make their phone shut up. That single tap is the whole breach. It’s how some of the most publicised compromises of the last few years started, and there’s nothing exotic about it. It’s a denial-of-patience attack.
The good news is that it’s loud. An attacker bombing someone with MFA prompts leaves a very recognisable shape in your sign-in logs, and you can find it with the same handful of KQL operators from the primer. This post builds that detection from nothing, one line at a time. If you haven’t read the primer, it’s the place to start; this one assumes you know what a pipe does.
What the attack looks like in the data
Every MFA prompt that gets denied or times out lands in SigninLogs with a result code. The one that matters here is 500121, which is Entra’s code for an MFA challenge that failed: the user declined it, ignored it, or it timed out. One of those on its own is nothing. Someone fat-fingered a prompt or left their phone in the kitchen.
The attack shape is different: a burst of them, against one account, in a short window. Ten denied prompts in fifteen minutes isn’t someone making tea. That clustering is the entire detection.
Start by looking, not detecting
Same discipline as always. Before writing anything clever, look at what the failures actually look like in your own tenant:
SigninLogs
| where TimeGenerated > ago(7d)
| where ResultType == "500121"
| take 20
Run that and read the rows. You’ll see who’s failing MFA in an ordinary week, which apps trigger it, and roughly how often it happens when nothing is wrong. That baseline matters, because the difference between noise and an attack is volume against time, and you can’t judge volume until you’ve seen normal.
Now add the shape
The attack is a cluster, so we count failures per user inside small time windows. bin() chops time into buckets, and summarize counts inside them:
SigninLogs
| where TimeGenerated > ago(1d)
| where ResultType == "500121"
| summarize FailedPrompts = count() by UserPrincipalName, bin(TimeGenerated, 15m)
| where FailedPrompts >= 8
| sort by FailedPrompts desc
Read it as a sentence. Take the last day of sign-ins, keep only failed MFA challenges, count them per user per fifteen-minute window, and show me anyone with eight or more in a single window. That’s the whole detection, and it’s five lines.
The threshold is yours to tune. Eight in fifteen minutes is a reasonable starting point, but your baseline from the first query is the real guide. Set it low enough to catch a determined attacker, high enough that the person with a flaky Authenticator setup doesn’t page you every morning.
The refinement that makes it serious
Here’s the version worth promoting to an analytics rule, and the thinking behind it. A burst of denied prompts is suspicious. A burst of denied prompts followed by a success is the actual disaster, because it means the bombing worked. Someone got tired and tapped approve.
let Failures = SigninLogs
| where TimeGenerated > ago(1d)
| where ResultType == "500121"
| summarize FailedPrompts = count(), LastFail = max(TimeGenerated)
by UserPrincipalName
| where FailedPrompts >= 8;
SigninLogs
| where TimeGenerated > ago(1d)
| where ResultType == "0"
| join kind=inner Failures on UserPrincipalName
| where TimeGenerated > LastFail
| project UserPrincipalName, FailedPrompts, LastFail,
SuccessAt = TimeGenerated, IPAddress, Location, AppDisplayName
Two new ideas, both small. let names a query so you can reuse it, here the set of accounts that got bombed. Then join matches those accounts against successful sign-ins (ResultType == "0") that happened after the last failure. What comes out is the list you actually care about: accounts that were hammered and then let someone in, with the IP, location and app of the sign-in that got through.
If that query ever returns a row, it’s not a hunt anymore. It’s an incident, and that account needs its sessions revoked and its credentials reset before you finish your coffee.
What it misses, because every detection misses something
Honesty about the gaps is what separates a detection you trust from one you hope about. This one won’t see an attacker who paces themselves below your threshold, and it won’t catch number-matching bypasses or token theft, which don’t generate denied prompts at all. It also assumes push MFA; if you’ve already moved to phishing-resistant methods for privileged accounts, whole classes of this attack stop applying to those users, which is rather the point of moving.
None of that makes the query less worth running. It means this is one detection, not a strategy. The strategy is number matching turned on, phishing-resistant MFA for the accounts that matter most, and this query watching for the accounts still on push.
Promote it slowly, same as always
Run it in the logs for a week or two first. Watch what it catches, tune the threshold and the window to your tenant, and only then wire it into a scheduled analytics rule with an automation that flags the account. The rule you understand at 2am is the rule that was built slowly at 2pm. That was true in the primer and it doesn’t stop being true when the attack gets a scarier name.
The wider lesson is the one worth keeping. Good detections are rarely clever. This one is a count, a threshold and a join, and it catches an attack that has embarrassed some very large organisations. The gap was never the difficulty of the query. It was that nobody had opened the logs and asked.
Originally published at alanconlon.com, where I post practical security, Azure and data write-ups every other Tuesday.
Top comments (0)