// writing · 2026-09-28

A network mount was dead for 11 days while every service wrote to the empty folder underneath

systemd recorded the mount failed once and never retried. Nothing alerted, because nothing was broken from the writers' point of view.

Eleven days. That is how long a network share sat unmounted on a machine that was supposedly backing up to it, without a single error that anyone noticed.

This is the failure mode that makes mounts genuinely dangerous: it does not look like failure. Every process writing to that path got a perfectly ordinary success.

What happened

The node mounted a share that was served by a VM running on the same node. On boot, the mount raced the VM's startup. The VM was not up yet, so the mount timed out.

With plain nofail,_netdev, systemd records the mount as failed, gives up, and never retries. That is the entire bug. nofail means "do not block boot if this is unavailable" — it does not mean "keep trying".

Now the nasty part. The mountpoint directory still exists. It is just an ordinary empty directory on the root filesystem. So every service that writes to that path writes to local disk instead, silently, forever. The backup script ran, reported success, and filled a directory that was doing nothing useful.

Why nothing alerted

Every monitoring check asked "did the backup job succeed?" It did. Nobody asked "is there actually new data on the backup target?" — a different question with a different answer.

This is the general lesson, and it is worth stating plainly: check the outcome, not the exit code. A job that returns zero has told you what it attempted, not what it achieved. The check that catches this is "is the newest file on the target newer than 26 hours", not "did rsync exit 0".

The fix

Make the mount lazy and self-healing with systemd's automount:

# /etc/fstab
//192.168.1.10/share  /mnt/backup  cifs  credentials=/root/.smbcred,nofail,_netdev,x-systemd.automount,x-systemd.mount-timeout=15  0 0
systemctl daemon-reload
systemctl start mnt-backup.automount
ls /mnt/backup          # first access triggers the real mount

With x-systemd.automount the filesystem mounts on first access rather than at boot, so the VM has time to come up — and if an attempt fails, the next access retries. That last word is the whole fix.

Proving it is really mounted

Do not test this by checking whether a path exists. The directory exists either way. Check the mount itself, and make sure the directory underneath is not quietly collecting files:

findmnt /mnt/backup
ls -A /mnt/backup/.. | grep backup     # underlying dir: should stay empty

Then add a freshness check to whatever sends you alerts — assert the newest file on the target is recent, and alert when it is not. Silence is the failure here, so a check that only fires on errors will never tell you about it.

The uncomfortable version

The reason this went unnoticed for eleven days is that backups are the one system where success and failure look identical until you need them. A backup you believe ran is worse than no backup, because you have made plans around it.

Self-Hosted Ops Pack — $29

Includes the backup-freshness verification workflow that catches exactly this: it asserts the newest backup is recent and alerts when it is not.

Get it

← All writing