// writing · 2026-10-01

A Proxmox node that pings, accepts SSH, and still refuses every key

Node out of the cluster, port 8006 accepting TCP, SSH rejecting every key, /etc/pve empty. One missing fstab line was the whole story.

This one wasted an evening, because every symptom pointed somewhere other than the cause. A node came back from a reboot pingable, with sshd running and ports 22, 8006 and 3128 accepting TCP — and yet it was out of the cluster, refused every cluster SSH key, and its web UI never completed TLS.

It was not hung. It was not a network fault. It was not a failing disk. The node had booted perfectly fine onto a read-only root filesystem.

Why it looks like a networking problem

The tell is the combination:

  • ping works, and nc shows 22, 8006 and 3128 open — the OS booted.
  • SSH returns Permission denied (publickey) for every node's key.
  • Port 8006 accepts TCP but the TLS handshake times out.
  • pvecm nodes does not list the node, and /etc/pve is an empty directory — just . and ...

Every one of those reads as "the network or the cluster stack is broken". The actual explanation is upstream of all of it.

The smoking gun

pmxcfs[3226]: [main] crit: unable to create lock
  '/var/lib/pve-cluster/.pmxcfs.lockfile': Read-only file system

That is the whole incident in one line. pmxcfs cannot write its lock file, pve-cluster dies, and everything downstream follows: corosync, pvestatd, the pveproxy web UI, and the cluster SSH keys.

That last one surprises people. /root/.ssh/authorized_keys is a symlink into /etc/pve. If /etc/pve is not mounted, there are no keys, so no key will ever authenticate — no matter how correct it is.

The actual cause

The kernel always mounts root read-only first. systemd-remount-fs.service then flips it to read-write using the fstab entry for /. If that entry is missing, root simply stays read-only for the life of the boot.

In our case the line had been dropped during an earlier fstab rewrite that was adding a storage pool. Two load-bearing lines disappeared silently: the root line and the EFI partition line. The other two cluster nodes still had theirs and booted fine — which is exactly what made it look like a fault specific to this node.

Diagnosing it (from the console, since SSH is refused)

findmnt -no SOURCE,OPTIONS /                       # look for 'ro'
systemctl is-active pve-cluster corosync pvestatd  # all failed
ls -la /etc/pve/                                   # empty = pmxcfs never mounted
journalctl --unit=pve-cluster -b --no-pager | tail -20

One small trap in that last command: journalctl -b -u pve-cluster fails with "failed to parse boot descriptor '-u'", because -b swallows the next argument. Put --unit= before -b.

The fix

Non-destructive, and no data movement:

mount -o remount,rw /
systemctl restart corosync pve-cluster pveproxy pvestatd
pvecm status          # expect all nodes, Quorate: Yes

Then repair /etc/fstab so it does not come back — and the honest way to do that is to diff your file against a node that boots correctly, not to type it from memory:

/dev/pve/root /  ext4  errors=remount-ro  0 1
UUID=<efi-partition> /boot/efi vfat defaults 0 1
/dev/pve/swap none swap sw 0 0
proc /proc proc defaults 0 0

Follow with systemctl daemon-reload, mount anything you added, and validate with findmnt --verify.

The rule worth keeping

A node that pings, answers SSH, and refuses every key should make you check whether root is read-write before you investigate the network, the firewall, or the cluster.

Two more things that cost time on the way:

  • Guests do not auto-start in this state. pve-guests needs /etc/pve, so every VM and container has to be started by hand.
  • A later SSH attempt may return Not allowed at this time from OpenSSH 9.8's PerSourcePenalties. That is rate-limiting caused by the earlier failed attempts, not a second problem. It clears in seconds.

And the one that stings in hindsight: a cluster flap three weeks earlier had been written off as noise during a larger outage. The fstab mtime showed the rewrite predated it. It was almost certainly this same bug, reporting itself quietly the first time.

The Agent-Run Homelab — $39

The full pitfalls chapter, the monitoring stack, and the four n8n workflows that watch for this class of failure.

Get it

← All writing