DEV Community

Cover image for Getting Started with Sentinel Hunting: A KQL Primer for People Who Keep Putting It Off
Alan Conlon
Alan Conlon

Posted on Originally published at alanconlon.com on

Getting Started with Sentinel Hunting: A KQL Primer for People Who Keep Putting It Off

Most people who run Sentinel don’t really hunt in it. They wire up the connectors, switch on a pile of built-in analytics rules, and then treat the whole thing as a box that pings them when something’s wrong. Which is fine, until the day you actually need to go looking for something the built-in rules never thought to ask about. That’s when you open the logs, stare at the query bar, and realise you’ve been avoiding KQL for two years.

I put it off myself for a long time. It looked like another query language to learn on top of everything else, and the examples online were either trivially basic or written by someone showing off. The truth is somewhere in between, and the useful middle is much smaller than it looks. You can get genuinely productive with about five operators. Everything else you pick up when you need it.

Start by just looking at one table

The mistake most people make is reaching for a detection before they’ve ever looked at the raw data. Don’t. Open the logs and run the most boring query there is:

SigninLogs
| take 10

Enter fullscreen mode Exit fullscreen mode

That’s it. take 10 grabs ten rows so you can see what’s actually in the table, the column names, what the values look like and where the useful stuff lives. Sign-in logs are huge and nested, and you’ll save yourself a lot of grief by seeing the shape of them before you try to be clever. Do the same with whatever table you care about. Look first, query second.

The pipe is the whole idea

KQL reads top to bottom, and every line starts with a pipe. Each step takes the table that came out of the line above and does one thing to it. That’s the entire mental model. You’re passing a table down a chain and narrowing or reshaping it as you go.

SigninLogs
| where TimeGenerated > ago(24h)
| where ResultType != "0"
| take 20

Enter fullscreen mode Exit fullscreen mode

Read it as a sentence. Take the sign-in logs, keep the last 24 hours, keep only the failures, show me twenty. ResultType != "0" is the failures, because zero means success. The quotes matter too, the column holds text even though the values look like numbers. One of those things nobody tells you and everybody just has to learn.

Filter early, always

Here’s the one habit worth building from day one: cut the data down before you do anything expensive with it. Put your where clauses high up, especially the time filter. Sentinel charges you for what you scan and makes you wait for it, so a query that filters to the last hour before it starts counting things will be faster and cheaper than one that counts everything and trims afterwards.

It’s the difference between this:

SigninLogs
| where TimeGenerated > ago(1h)
| summarize count() by UserPrincipalName

Enter fullscreen mode Exit fullscreen mode

and the version that summarises the whole table first and then filters. Same answer. Wildly different cost. Get the time bound in early and it becomes second nature.

Summarize is where it gets useful

where finds rows. summarize answers questions. The moment you want to know how many or how often rather than just show me, you’re reaching for summarize.

SigninLogs
| where TimeGenerated > ago(24h)
| where ResultType != "0"
| summarize FailedAttempts = count() by UserPrincipalName
| sort by FailedAttempts desc

Enter fullscreen mode Exit fullscreen mode

That counts failed sign-ins per user over the last day and puts the worst offenders at the top. It’s about four lines and it’s already more useful than half the dashboards I’ve seen. You can see where this goes: repeated failures clustered on one account in a short window is exactly the shape of someone being attacked, and you found it with five operators and no special tooling.

Don’t trust a detection you can’t read

This is the part that matters more than the syntax. There’s a temptation, once you find a clever query online, to paste it into an analytics rule and move on. Resist it. A detection you don’t understand is worse than no detection, because it gives you confidence you haven’t earned. When it fires at 2am you need to know what it actually checked, and when it stays silent you need to know whether that means you’re safe or whether the query was quietly broken the whole time.

So build them slowly. Get the query returning sensible rows in the logs first. Look at what it catches and what it misses. Only then promote it to a scheduled rule. The boring, incremental version is the one that works.

Where to go next

Five operators get you a long way: take, where, summarize, sort, and project (which trims the columns down to the ones you care about). Build a few queries by hand against the tables you already have: failed sign-ins, sign-ins from somewhere you don’t operate, accounts that suddenly got handed an admin role. The detections that catch real attacks are rarely exotic. They’re usually just someone who took the time to look.

That’s the whole secret, honestly. The people who get value out of Sentinel aren’t the ones who memorised the language. They’re the ones who actually opened the logs and started asking questions.


Originally published at alanconlon.com, where I post practical security, Azure and data write-ups every other Tuesday.

Top comments (0)