Writing
Incidents
-
Detection Is Easy, Recovery Is Hard
My firewall's watchdog detected every outage and kept failing to reboot the box. Every layer I added broke something new, including one kernel panic. Photos included.
-
When Both Mirror Drives Fail the Same Way
A ZFS mirror lost both NVMe members to the same firmware bug, one night apart, and I fired the second shot myself. What redundancy actually assumes, and what a UPS quietly breaks.
-
Three Bugs Wearing One Trench Coat
Four months of DNS outages that made no sense, because they weren't one problem. A failing port, a 534k entry multiplier, and a recovery command that planted an impostor every time I ran it.
-
My Backups Ran Green Every Night and Backed Up Nothing
A replication task can succeed every night for months while protecting zero bytes. The audit that found it, the ZFS semantics behind it, and what verification means.