
In the last post three HP boxes became a Windows Server 2025 domain: AD, DNS, DHCP failover, Group Policy, BitLocker, Storage Spaces. That was day one: everything installed, everything green, everything “works”.
This post is about day two, the part where you stop admiring the green checkmarks and start asking uncomfortable questions. What time is it, according to whom? Which services are running that nobody asked for? What happens if a server simply doesn’t boot tomorrow? Spoiler: one of those questions had a very funny answer.
Same setup as before: I decide, and my AI co-admin (Claude, driving PowerShell over WinRM) does the inventory, runs the changes and checks the result after every step. Read first, change second, verify third. Boring on purpose.
🔍 Step zero: inventory before opinions
Before touching anything, a read-only sweep over all three DCs: OS build, RAM, power plan, installed roles, running services, SMB config, LSA protection, LDAP signing, time source, Defender, event log sizes, NIC DNS order, page file, disks, firewall profiles. Plus the domain itself: password policy, functional level, LAPS schema, DNS scavenging, GPO list.
Fun fact: the first version of that inventory script returned nothing at all. No error, no output, just silence. WinRM (via pywinrm) ships PowerShell as a Base64-encoded UTF-16 command line, and Windows caps command lines at 8191 characters. A 3.5 KB script becomes ~9.5 KB on the wire and just… vanishes. Split it in three, and suddenly the servers have a lot to say.
🕰️ The clock that answered to nobody
My favourite finding of the day:
PS> w32tm /query /source
Free-running System ClockThat was on HP1, the PDC emulator, the one domain controller whose whole job in the time hierarchy is to be the domain’s reference clock. Every domain member syncs from a DC, every DC syncs from the PDC… and the PDC synced from its own quartz crystal and good intentions. The entire domain was keeping perfectly consistent, perfectly unanchored time.
Why that matters: Kerberos tolerates 5 minutes of clock skew by default. Drift past that and logons start failing with errors that look like anything except “your clock is wrong”. The fix is three lines, only on the PDC:
w32tm /config /manualpeerlist:"ptbtime1.ptb.de,0x8 ptbtime2.ptb.de,0x8 ptbtime3.ptb.de,0x8" `
/syncfromflags:manual /reliable:yes /update
Restart-Service w32time
w32tm /resync /rediscoverPTB is the German national metrology institute, so the domain now takes its time from the people who literally define the second in this country. HP1 reports stratum 2, the other two DCs follow HP1. Hierarchy restored. The iLOs got the same treatment later: they had no NTP at all and two of them believed they lived in “Eastern Europe (DST disabled)”.
🌐 DNS client order: why HP2 panicked when HP1 rebooted
Remember the note from last time, that HP2 rejected logons for about seven minutes while HP1 was rebooting? The inventory explained it: HP2’s NIC had exactly one DNS server configured, HP1. No HP1, no DNS, no SRV records, no finding a DC, no Kerberos. HP1 itself pointed only at 127.0.0.1.
The best-practice order for a DC’s own NIC is: a partner DC first, another partner second, loopback last. Partner first so a freshly booted DC doesn’t wait on its own not-yet-started DNS service. Loopback last so it still resolves when everyone else is down. HP3 already had that, and now all three do.
🗺️ AD Sites: renaming “Default-First-Site-Name”
Every forest starts with a site called Default-First-Site-Name and no subnets. With three DCs in one rack, one site is exactly right. Intra-site replication uses change notification (~15 seconds), and adding more sites would only make things slower.
Two small changes anyway: the site was renamed to Homelab (rename, never delete and recreate; AD tracks it by GUID, and Netlogon re-registers the _sites SRV records on its own), and 192.168.10.0/24 was assigned to it. Every client can now map its IP to a site, which avoids Netlogon 5807 events for unknown subnets. As you’ll see below, it also ties into how Windows Update works after WSUS.
🖨️ Services nobody ordered
Windows Server with Desktop Experience ships with a surprising number of services that make sense on a laptop and none on a domain controller in a rack. Disabled on all three:
- Print Spooler: the big one. PrintNightmare, and the “printer bug” that lets an attacker coerce a DC into authenticating somewhere. Microsoft itself recommends disabling it on DCs. Nobody prints from a domain controller. Nobody.
- Audio (
Audiosrv,AudioEndpointBuilder): my DCs don’t sing. - Geolocation (
lfsvc): they also know exactly where they are. In the rack. - Radio management (
RmSvc), push notifications (WpnService), telemetry (DiagTrack), network discovery publishing (FDResPub,fdPHost).
Honest disclaimer: this does not make the servers measurably faster. HP1 has 64 GB RAM with 60 GB free, so a handful of idle services is not a bottleneck. It’s attack surface and noise reduction, not a turbo button. Anyone selling you “disable 40 services for 30% more performance” on a modern server is selling you a placebo.
🔐 Hardening, the unglamorous list
- Password & lockout policy: minimum length was 7 (a default straight out of 2003). Now 12, and account lockout after 10 failed attempts for 15 minutes, where before it was never. Gotcha: the Default Domain Policy’s security template and the domain object have to agree. Change only the domain object and the PDC quietly reverts it on the next policy refresh.
- Windows LAPS: schema extended, clients’ OU allowed to write their own password attributes, GPO linked. Each client now gets its own random 16-character local admin password, rotated every 30 days and stored encrypted in AD, readable only by domain admins. It’s built into Windows now, with no separate MSI like the old LAPS.
- LDAP: channel binding enforced. The full “require signing” switch is the next step, after checking the logs for unsigned binds (there were none).
- LSA Protection (
RunAsPPL): LSASS runs as a protected process, which makes the classic Mimikatz-style credential dumping a lot harder. Set on all three, then activated with a rolling reboot, one DC at a time and never in parallel (lesson learned last time). Verified the boring way: after each boot, Wininit logs event 12, “LSASS.exe was started as a protected process with level: 4”. The file server’s reboot waited until I had closed the notes database I had open on its share. Rebooting a box under your own open files is a special kind of self-own. - Event logs: the Directory Service log was 1 MB. One megabyte. That covers roughly “what happened since lunch”. Now Security 1 GB, Directory Service 256 MB, System/Application 128 MB each. Disk is cheap, and forensics without logs isn’t possible.
- DNS scavenging: enabled with 7+7 days aging. Without it, every DHCP client that ever registered keeps its A record forever, and after a year your DNS is a museum.
- WinRM for clients: a GPO enables WinRM on the clients’ OU, with the firewall open to the home LAN only (
192.168.10.0/24), no Basic auth and no unencrypted traffic. It’s the prerequisite for managing the workstations remotely, including Windows Admin Center, which goes onto the admin workstation (it’s unsupported on a DC).
🪦 WSUS is dead. So who downloads the updates now?
In 2024 Microsoft announced that WSUS is deprecated. It’s still in Server 2025 and still supported, but it gets no new features, and the direction is clear. My setup already switched to plain Windows Update + Group Policy: deferral days for clients, fixed maintenance windows for servers, one DC patched on Saturday as the canary.
But WSUS did two jobs, and approval rules only cover one. The other was bandwidth: download each update once from the internet, then serve it to a thousand clients over the LAN. What replaces that?
Delivery Optimization. It’s built into Windows, and it’s basically BitTorrent for patches (with less piracy and more Kerberos). An update is split into pieces. The first client pulls pieces from Microsoft’s CDN. Every other client asks its peers first: “anyone on my network already have piece 17?” If yes, it comes over the LAN; if not, from the internet. The more clients a network has, the more of the download stays local. 1000 clients no longer means 1000× the download on your internet link.
The knobs that matter:
DODownloadMode = 1(LAN): peers on the same NAT/network. That’s what my client GPO uses.DODownloadMode = 2(Group) withDOGroupIdSource = 1(AD site): clients only share with peers in the same AD site. So a branch office doesn’t pull pieces over the slow WAN from headquarters. This is why clean sites and subnets suddenly matter again, and why the site rename above wasn’t purely cosmetic.- Microsoft Connected Cache: the grown-up version for large environments. A local cache node per location that fetches content once and serves everyone, the spiritual successor of WSUS’s “download once” without the WSUS database and its care and feeding.
What you do lose: hand-approving individual updates. That now lives in policy (deferrals, rings, deadlines) or, in the enterprise, in Intune/Windows Autopatch. For a homelab with two clients, all of this is gloriously overkill. The two laptops share a few hundred MB per month. But it’s nice to know the mechanism.
💾 Windows Server Backup: how it actually works
Three DCs replicate each other, so AD itself is redundant. But replication is not a backup: a deleted OU replicates just as fast as a created one. (The AD Recycle Bin from last time covers the “oops, deleted” case. It doesn’t cover “the server won’t boot”.) So HP1 and HP2 now run a nightly Windows Server Backup to their second, so far empty, 2 TB SSD.
Under the hood:
- VSS first. The Volume Shadow Copy Service freezes a consistent snapshot of the volumes. AD’s database (
ntds.dit) has a VSS writer, so the copy is transactionally clean even while the DC keeps working. No downtime, no maintenance mode. - Block-level copy into VHDX. The snapshot is copied block by block into VHDX files on the target volume. After the first full run, later runs only copy changed blocks, which makes the nightly job fast.
- Versions via shadow copies on the target. Older versions are kept as shadow copies on the backup volume itself. When space runs out, the oldest ones are dropped automatically. No retention script needed.
- What’s included: Bare Metal Recovery (C:, the EFI partition, the recovery partition) plus System State (AD database, SYSVOL, registry, boot files, the COM+ and cert stores). The DHCP database lives under
C:\Windows\System32\dhcp, so it’s in there too.
The three restore scenarios:
- Single files:
wbadmin get versions, then restore individual files or folders from any version while the server runs. - AD objects gone wrong: boot into Directory Services Restore Mode, restore System State (non-authoritative: the DC then catches up via replication), or mark specific objects authoritative with
ntdsutilso they win against the other DCs. - Server dead, disk dead, OS dead: boot the Windows Server ISO → Repair your computer → System Image Recovery, point it at the backup disk, and it rebuilds the partitions and restores C: 1:1. For a DC that is a non-authoritative restore: it comes back with the state of last night, then replicates everything newer from its partners. It has to happen within the tombstone lifetime (180 days), which is never a problem with a nightly job.
The honest caveat: the backup lives inside the same box. That covers a dead OS disk, a botched update, a corrupted AD. If the mainboard dies, the SSD can move into a replacement. But fire, theft or ransomware that encrypts every volume would take the backup along. That’s what copies on a different machine are for (see the next section). And for a DC there’s always plan B: with two healthy partners, reinstalling and re-promoting is often faster than restoring anything.
📦 Backups, part two: getting the family archive off the box
The DCs can always be rebuilt from their partners. The 1.5 TB of family archive on HP3’s parity pool can’t. Parity protects against one dead SSD, not against a deleted folder, ransomware or a pool that decides to have a bad day. So the Linux backup box that already pulls my Nextcloud instances now pulls the NAS shares too.
- A dedicated read-only account.
svc-nas-backupgets Read on the three SMB shares and read-and-execute NTFS rights, and nothing else: no admin, no Backup Operators. (Backup Operators on a DC can readntds.dit, which is to say: the whole domain.) Its random 32-character password lives in a root-only credentials file on the Linux box. - Read-only CIFS automounts. The shares are mounted
rovia fstab withx-systemd.automount. A write test from the backup box fails with “read-only file system”, which is exactly the point: a compromised backup host can’t encrypt the source. - rdiff-backup, one repository per user. The current state sits as plain files, older versions as reverse diffs, 30 days of history. The script refuses to back up a share that is unreachable or empty. Otherwise a network hiccup would be recorded as “user deleted everything”.
The first run took about six hours, and the numbers were a lesson in what really costs time. My 1.2 TB (a few thousand big archives) streamed at a steady ~108 MB/s and was done in three hours. The 236 GB of another family member, spread over 189,000 files, took almost as long. For rdiff-backup, the file count matters far more than the data volume. The second run took 73 seconds.
And yes, I checked the file counts afterwards. They didn’t match. After a mild moment of panic: PowerShell’s Get-ChildItem skips hidden files unless you add -Force. With -Force and minus the excluded junk (Thumbs.db, .DS_Store, Office lock files) the counts matched to the file. Trust, but -Force.
🔐 BitLocker on the servers: for the day the disks leave the house
My threat model here is modest: a disk that dies and goes back under warranty, or a server I sell in five years. Nobody should be able to read an AD database or a family photo from it. So BitLocker with TPM only: no PIN, because servers have to boot unattended.
- Recovery keys in AD, and outside of it. A GPO on the Domain Controllers OU makes backing up the recovery key to AD mandatory. Windows refuses to encrypt if that fails. The keys are stored under the computer object and replicate to every DC: HP3’s keys showed up on HP1 within a minute. But if all three DCs sit at the recovery screen at the same time (say, after a BIOS update on all of them), the key is in AD and AD is not running. So there’s also a copy in the password manager.
- The recovery screen is pre-boot, so the iLO remote console works for it even without an Advanced license. Before firmware updates,
Suspend-BitLocker -RebootCount 1avoids the prompt entirely. - Data volumes auto-unlock, and on the big pool only the used space is encrypted. Otherwise it would take days on parity.
Two plot twists. First, HP1’s TPM reported “owned” but not “ready for storage”, with no SRK provisioned. It’s probably still owned by some earlier install. Initialize-Tpm answered with ClearRequired=True, so it needs a clear. I deliberately did not queue that as a pending BIOS setting via Redfish: it would fire at the next unattended Windows Update reboot at 3 a.m., and the physical-presence prompt would keep the DC waiting at a console nobody is watching. It will be done by hand, with the iLO console open.
Second, timing. HP1 and HP2 were rebooted for LSA protection minutes before the BitLocker feature was installed. Get-WindowsFeature happily reports “Installed”, but manage-bde.exe and the PowerShell module only appear after the next reboot. HP3 happened to reboot a day later for patches and was encrypted right away; the other two will follow after their next reboot.
🔌 iLO: asking the hardware directly
All three boxes (two ProLiant DL20 Gen10 Plus, one MicroServer Gen10 Plus v2) have iLO 5, and iLO 5 speaks Redfish, a REST API over HTTPS. So instead of clicking through three web UIs, one Python script read BIOS settings, firmware inventory, temperatures, fans and the Integrated Management Log from all three.
- Power: already optimal everywhere. Workload profile General Power Efficient Compute, power regulator Dynamic Power Savings, C6 package states. Fans idle at 6–8%. Nothing to tune, and setting “Static High Performance” would only turn electricity into heat for a load that never needs it.
- Firmware: current (System ROM from 2026, iLO 5 v3.21).
- Power meter: reads 0 W. These entry models don’t have metered power supplies. Oh well.
- IML: mostly “link down” events that line up with cabling and switch reboots. One DL20 logged a single uncorrectable PCIe error two months ago that never came back. It stays on the watch list.
- Housekeeping: time zone fixed, NTP pointed at the DCs, consistent
-ilohost names, and A records for the iLOs in the AD zone, so documentation and reality use the same names.
🗜️ Deduplication: measured, then declined
HP3’s ReFS parity pool holds about 1.5 TB of family archive, and Server 2025 has shiny new ReFS dedup + compression built in. Tempting! So before enabling anything: look at the data.
- 77% is
.rar, followed by JPEG, ZIP and MP4. All already compressed. Compressing compressed data yields a rounding error. - Real duplicates (same name, same size, excluding multi-part archives): ~72 GB, about 5%. Mostly copies of old laptops inside copies of old laptops. We’ve all been there.
72 GB saved versus nightly optimization jobs on a pool whose parity writes are already the slow path, plus a chunk store that becomes a single point of “one bad block now hurts many files” on single-parity storage. With 2.2 TB free, the answer is no. Dedup shines on VM disks, VDI and backup targets, not on a pile of RAR archives. Cleaning up the duplicate folders by hand is cheaper than any feature.
🎯 Takeaway
Day one makes things work. Day two makes them trustworthy: a real time source, DNS that survives a reboot, fewer services, longer memory, passwords that rotate by themselves, and a backup you’ve actually thought through before you need it. None of it is glamorous, and almost none of it shows up as a green checkmark. And the scariest finding wasn’t a missing patch. It was a domain whose sense of time came from a quartz crystal with no supervision.
Next up: the TPM clear on HP1 and BitLocker on the first two DCs, and finally enforcing LDAP signing. That one is a single security option in the Default Domain Controllers Policy, and it quietly overrides the registry value I had set first. And then maybe, maybe, some YAML again. I’m starting to miss it. A little.





