A Proxmox node that pings, accepts SSH, and still refuses every key
Node out of the cluster, port 8006 accepting TCP, SSH rejecting every key, /etc/pve empty. One missing fstab line was the whole story.
This one wasted an evening, because every symptom pointed somewhere other than the cause. A node came back from a reboot pingable, with sshd running and ports 22, 8006 and 3128 accepting TCP — and yet it was out of the cluster, refused every cluster SSH key, and its web UI never completed TLS.
It was not hung. It was not a network fault. It was not a failing disk. The node had booted perfectly fine onto a read-only root filesystem.
Why it looks like a networking problem
The tell is the combination:
pingworks, andncshows 22, 8006 and 3128 open — the OS booted.- SSH returns
Permission denied (publickey)for every node's key. - Port 8006 accepts TCP but the TLS handshake times out.
pvecm nodesdoes not list the node, and/etc/pveis an empty directory — just.and...
Every one of those reads as "the network or the cluster stack is broken". The actual explanation is upstream of all of it.
The smoking gun
pmxcfs[3226]: [main] crit: unable to create lock
'/var/lib/pve-cluster/.pmxcfs.lockfile': Read-only file system
That is the whole incident in one line. pmxcfs cannot write its lock file,
pve-cluster dies, and everything downstream follows: corosync,
pvestatd, the pveproxy web UI, and the cluster SSH keys.
That last one surprises people. /root/.ssh/authorized_keys is a
symlink into /etc/pve. If /etc/pve is not mounted,
there are no keys, so no key will ever authenticate — no matter how correct it is.
The actual cause
The kernel always mounts root read-only first. systemd-remount-fs.service
then flips it to read-write using the fstab entry for /. If that entry
is missing, root simply stays read-only for the life of the boot.
In our case the line had been dropped during an earlier fstab rewrite that was adding a storage pool. Two load-bearing lines disappeared silently: the root line and the EFI partition line. The other two cluster nodes still had theirs and booted fine — which is exactly what made it look like a fault specific to this node.
Diagnosing it (from the console, since SSH is refused)
findmnt -no SOURCE,OPTIONS / # look for 'ro'
systemctl is-active pve-cluster corosync pvestatd # all failed
ls -la /etc/pve/ # empty = pmxcfs never mounted
journalctl --unit=pve-cluster -b --no-pager | tail -20
One small trap in that last command: journalctl -b -u pve-cluster fails with
"failed to parse boot descriptor '-u'", because -b swallows the next
argument. Put --unit= before -b.
The fix
Non-destructive, and no data movement:
mount -o remount,rw /
systemctl restart corosync pve-cluster pveproxy pvestatd
pvecm status # expect all nodes, Quorate: Yes
Then repair /etc/fstab so it does not come back — and the honest way to do that
is to diff your file against a node that boots correctly, not to type it from memory:
/dev/pve/root / ext4 errors=remount-ro 0 1
UUID=<efi-partition> /boot/efi vfat defaults 0 1
/dev/pve/swap none swap sw 0 0
proc /proc proc defaults 0 0
Follow with systemctl daemon-reload, mount anything you added, and validate
with findmnt --verify.
The rule worth keeping
A node that pings, answers SSH, and refuses every key should make you check whether root is read-write before you investigate the network, the firewall, or the cluster.
Two more things that cost time on the way:
- Guests do not auto-start in this state.
pve-guestsneeds/etc/pve, so every VM and container has to be started by hand. - A later SSH attempt may return
Not allowed at this timefrom OpenSSH 9.8'sPerSourcePenalties. That is rate-limiting caused by the earlier failed attempts, not a second problem. It clears in seconds.
And the one that stings in hindsight: a cluster flap three weeks earlier had been written off as noise during a larger outage. The fstab mtime showed the rewrite predated it. It was almost certainly this same bug, reporting itself quietly the first time.
The Agent-Run Homelab — $39
The full pitfalls chapter, the monitoring stack, and the four n8n workflows that watch for this class of failure.