
HP3, my HPE MicroServer Gen10 Plus v2, is back on Debian 13 with OpenZFS. Why it left Windows again is a story for another post. This one is about the restore afterwards, and a setting that had been hiding in plain sight for two operating systems and two blog posts.
The plan was boring, the way restores should be: a fresh RAIDZ1 pool over the same three WD Red SA500 SSDs, per-user datasets, then about 1.7 TB copied back from a USB disk with rsync. Same workflow as always: I make the calls, my AI co-admin (Claude) runs the commands and keeps an eye on the numbers. It was the numbers that started looking wrong.
🐌 Seven megabytes per second. Per SSD.
The copy crawled at 25 to 35 MB/s. The estimate said roughly sixteen hours. Even small things like zfs create or zpool status hung for minutes, stuck waiting for the next transaction group to finish syncing. Then zpool iostat -l showed the actual problem:
$ zpool iostat -vl storage 5 1
bandwidth total_wait disk_wait
read write read write read write
storage 3.42K 22.7M 1ms 233ms 1ms 65ms
raidz1-0 3.42K 22.7M 1ms 233ms 1ms 65ms
ata-WD_Red_SA500_..._00458 1.14K 7.57M 1ms 262ms 1ms 68ms
ata-WD_Red_SA500_..._01387 1.14K 7.58M 1ms 258ms 1ms 72ms
ata-WD_Red_SA500_..._02050 1.14K 7.58M 1ms 181ms 1ms 57ms
$ cat /proc/pressure/io
some avg10=94.05 avg60=90.06 avg300=80.24
full avg10=88.82 avg60=84.62 avg300=75.87About 7.5 MB/s per drive at 65 ms disk latency, with the box spending around 90 % of its time waiting for I/O. These are SATA SSDs rated for well over 500 MB/s sequential writes. 65 ms is spinning-rust territory. Worse, actually: a decent HDD does better than that.
🕵️ Suspect number one: TRIM, and I’d been here before
TRIM was the first suspect, and not just in theory. I’d already been bitten by it once. When I first moved these SSDs from Windows to Linux, I installed Debian over disks that had been BitLocker-encrypted. The numbers were just as awful as now, and that time the fix was a TRIM. So when the same symptoms came back, déjà vu kicked in.
Why BitLocker and SSDs are such a bad combination when you reinstall comes down to how an SSD sees the world. It has no idea what a file is. It only knows logical blocks, and a block is “in use” from the moment anything is written to it until someone explicitly says “you can forget this one”. That someone is TRIM (or discard in Linux terms). A filesystem sends it when it deletes files. A fresh mkfs or zpool create on top of old data doesn’t necessarily do it for the whole disk.
BitLocker makes this as bad as it gets. Encrypted data is indistinguishable from random noise, so there is nothing to compress or deduplicate inside the drive. And once a volume has been encrypted and written to for a while, a large part of the flash holds what the SSD has to treat as precious data. Then I wipe the disk with a new OS. The partition table is new, the filesystem is new, but nobody tells the SSD that all those old blocks are garbage now. From its point of view the drive is still full.
That hurts because flash can’t be overwritten in place. It’s written in pages and erased in much larger blocks. A drive that believes it is full has no pre-erased blocks left. For every new write, its garbage collector first has to find a block, copy the still-“valid” pages (the old encrypted noise) somewhere else, erase the block, and only then write the new data. One host write turns into several internal ones (write amplification). Throughput collapses and latency explodes, and the drive wears out faster on top.
The right way, which I’m writing down so future me actually does it: before handing used SSDs to a new filesystem, discard them completely. This destroys everything on the disk, which is the point:
# Before zpool create / mkfs — irreversibly wipes the drive!
blkdiscard /dev/disk/by-id/ata-WD_Red_SA500_2.5_2TB_...
# Check that the drive supports TRIM at all (non-zero DISC-GRAN/DISC-MAX):
lsblk --discard /dev/sdXOn an idle SATA SSD, blkdiscard usually takes seconds to a few minutes for the whole drive. Doing it afterwards with zpool trim on a live pool works too. ZFS only discards the space it doesn’t use. But it’s slower, it competes with real I/O, and it only helps from the moment it runs.
There’s a third option I only noticed later, while digging through the BIOS for something else (see below): the HPE RBSU has SATA Secure Erase and SATA Sanitize built in. They are the firmware version of the same idea. The drive itself throws away everything, including its own mapping tables, and comes back as fresh as on day one, with every block free. For SSDs that are about to be wiped anyway, that’s arguably the cleanest reset there is, and it needs no running OS. Same warning, in bigger letters: it erases the drive completely.
This time I had skipped that step, so we started a full zpool trim storage in the middle of the restore. Forty minutes later it was at 2 %, and the copy was barely faster. TRIM was part of the problem, but it wasn’t the part that turned SSDs into floppy drives. Something more basic was off as well. And once you know what it was, a TRIM crawling along at 2 % starts to make sense too.
💡 Suspect number two: one word in sysfs
$ for d in sda sdb sdc sdd; do echo "$d $(cat /sys/block/$d/queue/write_cache)"; done
sda write through
sdb write through
sdc write through
sdd write through
$ hdparm -W /dev/disk/by-id/ata-WD_Red_SA500_2.5_2TB_...
write-caching = 0 (off)Write through. The drives’ own write cache was switched off. Every SSD has a small DRAM buffer: it takes incoming writes, says “done” and then writes them to flash in large, efficient chunks. Without it, the drive has to finish every single write on the flash itself before it reports back. For an SSD that is about the worst way to work, because flash is written in large pages and erased in even larger blocks. Lots of small, separate writes mean slow writes and extra wear.
Who turned it off? Debian didn’t, it doesn’t touch that setting. The kernel log answers it, two seconds after power-on, long before anything in userspace has run:
$ dmesg | grep -i "write cache"
[ 2.101312] sd 2:0:0:0: [sda] Write cache: disabled, read cache: enabled
[ 2.101316] sd 3:0:0:0: [sdc] Write cache: disabled, read cache: enabled
[ 2.101332] sd 5:0:0:0: [sdd] Write cache: disabled, read cache: enabled
[ 2.101407] sd 4:0:0:0: [sdb] Write cache: disabled, read cache: enabled
[ 3.644946] sd 6:0:0:0: [sde] Write cache: enabled, read cache: enabled ← USB stick
[ 530.822448] sd 7:0:0:0: [sdf] Write cache: enabled, read cache: enabled ← USB diskAll four internal SATA SSDs come up disabled, while both USB devices come up enabled. SATA drives ship with the write cache on, so something switches it off during POST, and that something is the server’s firmware. HPE apparently defaults to “drive write cache off” for SATA drives, and as it turned out, it doesn’t let you change that (more on that below). The evidence points squarely at the platform, not at Linux.
⚡ One command later
$ for d in /dev/disk/by-id/ata-WD_Red_SA500_2.5_2TB_*[0-9]; do hdparm -W1 "$d"; done
setting drive write-caching to 1 (on)
write-caching = 1 (on)
...| Write cache off | Write cache on | |
|---|---|---|
| rsync restore throughput | 25–35 MB/s | 85–95 MB/s |
| I/O pressure (PSI, some avg60) | ~90 % | ~30–45 % |
| Estimated restore time | ~16 hours | ~6 hours |
| Bottleneck | the SSDs | the 2.5″ USB disk I was copying from |
Live, mid-copy, no reboot. The SSDs went from being the problem to waiting on a portable USB hard drive. Once the restore was done, I ran a clean benchmark as well.
📊 The clean benchmark: same disks, one switch
After the restore had finished and the pool was idle again (about 1.3 TB in use, TRIM still paused), I ran the same dd tests as in the original NAS post, twice: once with the drive write cache switched off again, once with it on. This time I took the lessons from that post seriously: a dedicated test dataset with compression=off, a 1 GiB file of random data served from RAM (/dev/shm) so /dev/urandom can’t become the bottleneck, fsync at the end, and O_DIRECT reads so ARC can’t flatter the read numbers.
zfs create -o compression=off storage/bench
head -c 1G /dev/urandom > /dev/shm/rand.bin
dd if=/dev/shm/rand.bin of=/storage/bench/t bs=1M conv=fsync # buffered + fsync
dd if=/dev/shm/rand.bin of=/storage/bench/t bs=1M oflag=direct conv=fsync # O_DIRECT
dd if=/dev/shm/rand.bin of=/storage/bench/t4k bs=4k count=5000 oflag=direct,dsync
dd if=/storage/bench/t of=/dev/null bs=1M iflag=direct # read| Test | Write cache off | Write cache on | Factor |
|---|---|---|---|
| Sequential write 1M, buffered + fsync | 8 MB/s | 547 MB/s | ~68× |
| Sequential write 1M, O_DIRECT | 2 MB/s | 239 MB/s | ~120× |
| 4K write, O_DIRECT + dsync | 292 kB/s | 334 kB/s | ~1× |
| Sequential read 1M, O_DIRECT | 972 MB/s | 865 MB/s | – |
Two megabytes per second. On SSDs. The first attempt used a 4 GiB file like the original post. With the cache off, it managed 3 MB/s and would have taken well over half an hour, so I cut it down to 1 GiB. That’s a result in itself.
What the table actually says:
- Writes go up by a factor of 70 to 120. That’s no tuning gain anymore. That’s the difference between “broken” and “working”.
- Reads don’t change. Of course they don’t, it’s a write cache. The small difference is run-to-run noise.
- 4K synchronous writes don’t improve at all, and that’s correct. With
dsync, every single 4K write ends in a cache flush, so the cache is emptied as fast as it fills. For sync-heavy workloads (databases, NFS with sync, VMs) the answer isn’t the drive cache. It’s a separate log device for ZFS (SLOG) or SSDs with power-loss protection. The write cache isn’t magic. It just stops the drive from doing every write the hardest possible way. - Compared to the original post (138 MB/s buffered, 15.4 MB/s O_DIRECT), the cache-off numbers are even worse today. I suspect that’s the second problem from this post at work: back then the drives were fresh, this time they were full of old BitLocker data and the TRIM hadn’t run yet. Two things missing, and each one makes the other worse.
🔙 The part where I re-read my own blog
Here’s the slightly embarrassing bit. These are the same three SSDs in the same box I’ve written about twice before. Both times the numbers were telling me exactly this, and both times I found a plausible explanation and moved on.
First miss: Building a RAIDZ1 NAS on Debian 13 with OpenZFS. I measured 138 MB/s buffered writes and 15.4 MB/s with O_DIRECT, roughly nine times slower. I even wrote that it “surprised me enough to double- and triple-check it”. Then I explained it with ZFS transaction-group batching: buffered writes get coalesced into full-stripe writes, O_DIRECT bypasses that. That part is true. But the explanation was too comfortable. With the drive cache off, every unbatched write has to wait for the flash itself, and that is what turns “somewhat slower” into “nine times slower”. Even the “good” number deserved a second look: 138 MB/s across three SATA SSDs is not great.
Second miss: Enough YAML for Now, on the Windows side. Storage Spaces parity on the same disks: 88 MiB/s sequential write, 461 random-write IOPS and a 99th-percentile latency of over a second. I blamed parity: “Reads are great. Writes are… parity.” Parity’s read-modify-write cost is real. But I also ran DiskSpd with -Sh, which deliberately disables the drive’s write cache for the test. The benchmark measured “no write cache” by design, and that happened to be exactly how the server ran every day anyway. The worst case and everyday reality were the same, and I never asked why.
Two operating systems, two benchmarks, two reasonable-sounding explanations, one setting nobody checked. The lesson isn’t “benchmarks lie”. It’s that I compared the numbers against my expectations of the software and never against the datasheet of the hardware. A 15 MB/s write on an SSD should have been a red flag no matter which filesystem was on top.
🛡️ Is it safe to turn it on?
That’s the obvious question: if HPE switches it off, maybe there’s a reason.
There is a reason, and it mostly doesn’t apply here. The drive cache is volatile RAM. If the power drops, whatever is in it and not yet on flash is lost. Enterprise SSDs have capacitors for this (“power-loss protection”), consumer and NAS SSDs usually don’t. On top of that, a classic HPE server uses a Smart Array RAID controller with its own battery-backed cache. In that design the drive cache is redundant at best and a liability at worst, so turning it off is the conservative choice. For a vendor that doesn’t know what OS, filesystem or controller you’ll use, “off” is the safe default.
On a modern filesystem, a write cache is safe because the filesystem tells the drive when data has to be on stable storage. At every point that matters, ZFS, ext4 and NTFS send a cache flush and don’t consider anything committed until the drive confirms it. ZFS is copy-on-write on top of that, so a crash leaves you with either the last consistent state or the new one, never a mix. The only real danger is a drive that confirms flushes it never performed (cheap, dodgy firmware), or, on Windows, ticking “Turn off Windows write-cache buffer flushing on the device”. That box really is the “I like living dangerously” checkbox.
And in my case all three HP servers sit behind a UPS anyway. What’s still on the to-do list is making HP3 shut down cleanly when the battery runs low (NUT). A UPS that runs flat while the server keeps writing just postpones the power cut.
💾 SSDs, HDDs, and the PC under your desk
- SSDs suffer the most. Flash is written in pages and erased in blocks. Without a cache the drive can’t collect small writes, so each one costs more time and more wear. That’s how you get 7.5 MB/s.
- HDDs lose too, mostly on random writes, because the drive can no longer reorder requests to minimise head movement. Sequential writes suffer less.
- Normal desktop PCs are almost always fine. Consumer BIOSes don’t touch the setting, and Windows leaves “Enable write caching on the device” on for internal drives. My jump host, an ordinary desktop board, says
write backfor both its 8 TB HDD and its NVMe SSD. NVMe has a volatile write cache as well, and the OS controls it, on by default. - USB drives on Windows are the exception you’ve probably felt. They default to “Quick removal”, where Windows skips write-back caching on its side so you can yank the stick without ejecting it. That’s one reason big copies to USB often feel slower on Windows than on Linux.
Check it yourself:
# Linux
cat /sys/block/*/queue/write_cache
dmesg | grep -i "write cache"
hdparm -W /dev/sdX # SATA
# Windows (PowerShell)
Get-PhysicalDisk | Get-StorageAdvancedProperty # IsDeviceCacheEnabled, IsPowerProtected🔧 Making it stick
hdparm -W1 is volatile: the drive resets it at the next power cycle, and the firmware switches it off again on every boot. So the fix needs to happen at every boot, too. For now that’s a udev rule:
# /etc/udev/rules.d/69-ssd-write-cache.rules
# HPE firmware disables the SATA SSD write cache at boot (write through).
# ZFS and ext4 issue cache flushes, so enabling it is safe.
ACTION=="add|change", KERNEL=="sd[a-z]", SUBSYSTEM=="block", ENV{ID_BUS}=="ata", \
ATTR{queue/rotational}=="0", RUN+="/usr/sbin/hdparm -W1 /dev/%k"It only touches internal, non-rotational ATA drives, so the OS SSD gets it too and USB devices are left alone.
Wouldn’t the BIOS be the better place? In theory, yes: a firmware setting applies before any OS loads, survives a reinstall and also covers installers and rescue systems. So I rebooted into the RBSU (F9) and went looking. Under System Configuration → BIOS/Platform Configuration → Storage Options → SATA Controller Options on the MicroServer Gen10 Plus v2 there are exactly three entries:
- Embedded SATA Configuration: SATA AHCI Support
- SATA Secure Erase
- SATA Sanitize
That’s all. No “Drive Write Cache”, nothing in the performance options either. In AHCI mode the firmware switches the cache off and gives you no way to change that. (Smart Array controllers do have a “Physical Drive Write Cache Policy”, but switching this box into a RAID mode to get that knob would hide the disks from Linux and take the ZFS pool with them. Not an option.) So the udev rule isn’t a seatbelt next to the real fix. It is the fix, applied by the OS at every boot, because nothing else will do it.
Which also means it has to be part of the build, not something remembered afterwards. If the next reinstall forgets this one file, it’s back to 7.5 MB/s, and probably another blog post.
The other two HP servers almost certainly have the same default. They’re next once I’ve decided what happens to them.
🎯 Takeaway
I wrote two posts full of careful benchmarking: compressible versus incompressible data, ARC versus page cache, DiskSpd’s zero-filled test files. They were good lessons, and they were all about the software lying to me. Meanwhile the hardware was being completely honest: one word in sysfs, one line in dmesg, and a TRIM lesson I’d already learned once after the BitLocker install and promptly forgotten. I never looked.
Next time a number looks odd, the first question isn’t “how does my filesystem explain this?” It’s “what is this hardware supposed to do?” A SATA SSD at 15 MB/s isn’t an interesting filesystem quirk. It’s a setting, or a missing TRIM, or in my case both. 🔍💾
My checklist for reusing SSDs on a server from now on:
- Discard first.
blkdiscardevery used SSD before creating the pool (or use the BIOS’s SATA Secure Erase / Sanitize), especially if BitLocker or any other full-disk encryption was on it. - Ship the udev rule with the install on HPE boxes in AHCI mode. The BIOS won’t fix it for you.
- Check the write cache.
dmesg | grep -i "write cache"and/sys/block/*/queue/write_cache. Expect “write back”. - Compare against the datasheet, not against your expectations of the filesystem.
- Only then benchmark, with incompressible data, on an idle pool.
📚 More in this series: the RAIDZ1 NAS
- Building a RAIDZ1 NAS on Debian 13 with OpenZFS: Datasets, Compression, and the Lies dd Told Me
- RAIDZ1 NAS Addendum: How Scrub and Trim Actually Get Scheduled — and the Alert That Wasn’t
- Everybody Smile — We’re Taking a Snapshot! Moving From rdiff-backup to ZFS on the RAIDZ1 NAS
- Second Addendum: ZED Watches My Pool’s Health, Not Its Waistline






