KamoCRM

Bring the worldchunks drive back into use after it drops

FeatureKlusterServices
Shipped
September 28, 2026 at 3:19 AM UTC
Author
Kamo
Commit
db7591a

On 2026-09-23 the USB earth-tile drive on k3m1 dropped for a second. systemd unmounted it and stopped nfs-server, which nfs-utils makes require every exported path, and nothing mounted it or started the server again when the disk came back. k1m1's hard NFS mount then hung statfs() for four days, until etcd stalled and the platform went down on 2026-09-28. worldchunks-reconcile runs every minute on both nodes and does nothing when all is well. On k3m1 it remounts the drive when it is present but unmounted or dead, and starts nfs-server when the drive is healthy. On k1m1 it remounts the share once the server answers. On both it restarts any container still bound to an old or empty /mnt/worldchunks, at most once per pod per 15 minutes, because until then DaemonService's scenery sampler reads a dead or empty directory. Installed on k1m1 and k3m1 by hand, like k3m1-ext. Verified live: a stopped mount and nfs-server came back in 36s, an unmounted k1m1 share in 55s, and a test pod left on a swapped-out filesystem was restarted onto the new one.

All changes

Like what you see shipping?

All of it arrives in your workspace on its own. Start on the free plan and read this page again in a month.

Start Free ForeverView Pricing