DEV Community

OpsKit
OpsKit

Posted on

10 PowerShell scripts that catch Windows server problems before your users do

Every Windows admin eventually writes the same five scripts from scratch: clean the temp folders, rotate the backups, watch the services, check the certificates, and offboard someone who left. Then the server gets rebuilt, and you write them again.

I got tired of that loop and packaged the ten scripts I kept re-writing into one small kit. Here are five of them with the part that actually matters — the logic, not the boilerplate — plus what each one has caught in practice.

1. Disk cleanup that tells you what it freed

The usual trap: Remove-Item -Recurse on C:\Windows\Temp while Windows Update is mid-download. Age-gate the files instead:

$cutoff = (Get-Date).AddDays(-$MinFileAgeDays)   # default 3 days
$files = Get-ChildItem -Path $dir -Recurse -File -Force -ErrorAction SilentlyContinue |
    Where-Object { $_.LastWriteTime -lt $cutoff }
Enter fullscreen mode Exit fullscreen mode

Run it with -WhatIf first, then check the CSV: every run appends Timestamp, Host, FreedMB. On a terminal-server farm this is usually 4–10 GB per box on the first pass, then a few hundred MB a week.

2. Certificate expiry before the 3 a.m. page

Nobody notices the TLS cert until the webhook stops firing. A raw check needs no modules:

$ssl = New-Object System.Net.Security.SslStream($tcp.GetStream(), $false, { $true })
$ssl.AuthenticateAsClient($host)
$days = ([datetime]$ssl.RemoteCertificate.NotAfter - (Get-Date)).Days
Enter fullscreen mode Exit fullscreen mode

Wire it to a host list, make it exit 1 when anything is under 30 days, and drop it in Task Scheduler. It has never been the interesting part of anyone's job, and it is always the thing that pages you.

3. A watchdog with a restart cap

Restart-Service in a loop is how you turn a 5-minute outage into a restart storm. The cap is the point:

$restartsToday = @(Import-Csv $LogPath |
    Where-Object { $_.Service -eq $name -and $_.Action -eq 'Restarted' `
                   -and [datetime]$_.Timestamp -ge (Get-Date).Date }).Count
if ($restartsToday -ge $MaxRestartsPerDay) { Write-Log $name 'Suppressed'; continue }
Enter fullscreen mode Exit fullscreen mode

Three restarts and then it stops touching the service and writes Suppressed to the log. A dead service that stays dead is a signal; a service in a crash loop that you keep poking is noise.

4. Share permissions: find the Everyone group

if ($leaf -eq 'Everyone' -and $rule.FileSystemRights -match 'FullControl|Modify') {
    $flag += "RISK: $identity -> $perms"
}
Enter fullscreen mode Exit fullscreen mode

Run against Get-SmbShare on every file server and you will, statistically, find at least one share where someone granted Everyone FullControl "temporarily" in 2019. The audit CSV gives you the share path and the exact ACE, so the remediation conversation is short.

5. Offboarding that leaves a trail

Disabling the account is the easy 10%. The parts people forget:

Set-ADUser -Identity $user -Replace @{ pwdLastSet = '0' }   # password dies with the session
# kill the user's processes, robocopy the home dir to an archive, append one audit line:
# Timestamp | Ticket | Operator | User | LastLogon | Actions | HomeBackup
Enter fullscreen mode Exit fullscreen mode

Every action writes one CSV row with the ticket number, so when someone asks "when exactly did we cut off jsmith?", the answer is a Select-String, not archaeology.

What ties them together

Three rules, applied everywhere:

  1. -WhatIf before destructive. Every script that deletes, disables, or restarts supports SupportsShouldProcess.
  2. CSV audit for every mutation. If it changed state, it logged a timestamped row to %ProgramData%\OpsKit\.
  3. No obfuscation, no agents, no telemetry. Plain commented PowerShell you can read before you run it — which is the only kind of script worth putting on a production box.

The honest pitch

These ten — cleanup, backup rotation, service watchdog, cert expiry, event-log alerting, inventory export, share-permission audit, password-age audit, patch reporting, and offboarding — are the pack I'd have bought years ago instead of re-writing. Two ways in: the whole kit is $7 (OpsKit — 10 Windows Server Admin PowerShell Scripts), or if offboarding is your only pain point, the offboarding script alone is $2 (OpsKit Offboard). Prefer to try first? disk-cleanup and ssl-expiry-check are MIT-licensed and free in the powershell-admin-scripts repo. Everything is tested by dry-run and live execution, 30-day refund if it's not for you.

(Full disclosure: yes, that's my product. The five scripts above work standalone if you'd rather copy them — no hard feelings.)

What would you add? The one I'm still missing is a DNS stale-record sweeper.

Top comments (0)