Skip to content
← All thoughts

original · 22 Sept 2026

My homelab broke a dozen times. There were two causes..

Three months of Unraid outages that each looked unique, and the two structural problems underneath all of them.

HomelabUnraidPostmortem

I run an Unraid server at home. It holds the family photos, the media library, a few databases, and some things I'd rather not lose. Over three months this year it broke roughly a dozen times, and each time it broke differently enough that I treated it as a new problem.

It wasn't. There were two causes. Everything else was a symptom wearing a costume.

The empty room

The first real outage looked like total data loss. Every container gone. My password manager showed a fresh setup screen. Nextcloud offered to walk me through installation. The photo library was empty.

Nothing had been deleted. Unraid's emhttpd had skipped mounting the user-share filesystem at boot, so /mnt/user was just an empty directory sitting in RAM. Containers started, found nothing where their data should be, and did the reasonable thing: they initialised themselves. Meanwhile the real data sat untouched on the cache pool, a few inches away.

That shape repeated all year. When a mount is missing, this system doesn't refuse to start. It writes to RAM and shows you an empty room, and every layer above it behaves correctly given what it can see. The diagnostic that settles it in one line:

stat -c %i /mnt/user/appdata/somefile
stat -c %i /mnt/cache/appdata/somefile

Different inode numbers mean /mnt/user isn't really mounted. I've since used that check more times than I'd like.

The wedge that never fired the OOM killer

One afternoon I typed "check the homelab" out of habit. The server had been frozen for fifty minutes and I had no idea.

15 GB of RAM, no swap. A VM I'd started for a side project pushed the baseline to 90%, then a bulk phone photo sync fired machine-learning jobs that load two to three gigabyte models on demand. The host layer thrashed itself to a standstill while the containers kept serving traffic. SSH accepted my TCP connection and then never sent a banner. Load average peaked near 200.

The part I didn't expect: the kernel's OOM killer never fired. Not once, across three separate wedges. With no swap, a machine can thrash indefinitely without ever crossing the threshold that triggers the thing designed to rescue it. It just stops, politely, for ten hours.

The worst one I ended by walking over and holding the power button.

Failures that hide

Almost nothing announced itself. This turned out to be the expensive part.

A scheduler died silently and froze a catalogue six days in the past — I found out because something downstream looked stale, not because anything alerted. A reboot reset the server's timezone to US Pacific; the absolute time stayed correct, so nothing looked wrong, while every cron job fired twelve and a half hours off. My notification system kept working perfectly for weeks while a corrupted config file quietly routed everything to a browser bell nobody was watching. It had faithfully recorded 27 alerts I never saw.

The lesson I keep relearning is that a silent failure and a working system are indistinguishable from the outside. If you can't tell them apart without logging in, you don't have monitoring — you have a dashboard.

The backup that ate the database

My favourite failure, in hindsight.

One morning a self-hosted service stopped accepting logins with a SQLite disk I/O error. The cause was the nightly backup script I'd written to protect it. The script ran sqlite3 .backup from the host through one mount path, while the container held the same database open through a different one — a FUSE overlay over the same bytes. SQLite's write-ahead-log locking can't coordinate across two mount views of one file. The backup corrupted the live WAL state of the thing it was backing up.

The script was careful in every way I'd thought to be careful. It verified the dump, refused to ship an empty one, and shipped offsite. It was just wrong about the filesystem underneath it. The fix was to stop the container, copy, and start it again — three seconds of downtime at 5am — and the general rule is now simply: no host process touches a running container's SQLite file.

Only flash is real

A whole category of my fixes didn't survive a reboot, which is a humiliating thing to discover twice.

A container restart policy set with docker update survives restarts but not template recreation. A memory limit applied to a running container does the same. Cron entries written to /etc/cron.d/root get regenerated from flash at every boot, so a watchdog I'd installed to catch silent failures was itself erased by a power cut — and then, of course, the thing it was watching died silently, with nothing left alive to catch it.

On this OS the root filesystem is a RAM disk rebuilt at every boot. If a change isn't written to the flash drive, it didn't happen. Obvious in retrospect; not obvious at 1am.

The two causes

After the third wedge I stopped diagnosing incidents and went looking for what they had in common. There were exactly two things:

  1. 15 GB of RAM, oversubscribed, with no swap. Every freeze traces back to this. No single process was at fault — it was a flat 93% baseline that any small allocation could tip over.
  2. Power-loss reboots, with no UPS. Every unclean shutdown found a different victim: a DNS race at boot, a web server spinning at 500% CPU on stale state, a wiped cron table, a zeroed config file, a cache pool that lost its device assignments.

Everything else — a dozen distinct incidents across three months — was one of those two things wearing a different costume.

What the UPS proved in 72 hours

I'd been told to buy a UPS since June. I bought one in September.

Three days after I plugged it in, the logs showed five power failures. Four of them in a single evening, between 8pm and 10:30. Eighty-two seconds total on battery.

The server didn't notice any of them. No unclean reboot, no filesystem repair, no parity check, no dead scheduler, no corrupted config. Every event logged as low line voltage — the mains at my address sags constantly, and I had simply never known. I'd been reading those sags as "the homelab is unreliable."

Ten days on it's absorbed eighteen power events and about eight minutes on battery, with continuous uptime throughout. The last piece was enabling the UPS to cut its own output after a clean shutdown, so that when mains returns the board's power-on setting boots the machine by itself. Nobody has to be home.

A coda about decimals

I also write an iOS app for managing these servers, and it has a card that displays UPS status. That card had never worked. Not once, for anyone.

The API reports the current load in watts. My server said 77.85. The code expected an integer, and Swift's JSON decoder does not quietly round — it throws. That error landed in the one query deliberately wrapped in a "this server might not have a UPS" catch, so the failure was swallowed and the card just never appeared.

It had shipped broken and stayed broken because no real UPS had ever tested it. Mine was the first. There's something fitting about the server finally teaching the app that watches it.

What I'd tell myself in June

Buy the UPS first. Not because power cuts are dramatic, but because unclean shutdowns are how a stable system becomes a system with a dozen unrelated-looking bugs. Add swap even if you think you have enough RAM, because the failure mode without it isn't a crash you can read afterwards — it's a machine that stops answering and never says why.

And check that your monitoring can tell the difference between "fine" and "not reporting". Mine couldn't, for about a month, and it was the most expensive thing on this list.