xcp-ng server crashed/rebooted due to issues with drbd/linstor?
-
Looks to be similar to last time
Time,Node,Component / Service,Logged Event & Impact 23:17:13,xen01,LINSTOR/Controller,Resource Created: linstor-csi provisions a short-lived volume: pvc-395da199-d455-4004-ac0d-f6c6016ebf4e. 23:17:16,xen01,kernel (drbd),Sync Begins: xen01 connects to xen05 as a SyncSource and begins synchronizing the volume to it. 23:17:49,xen01,LINSTOR/Controller,Teardown Requested: The CSI plugin changes its mind or finishes its task, requesting sequential deletion (DelRsc) of the volume replicas across the cluster. 23:17:51,xen01,LINSTOR/Controller,Source Disk Removed: Linstor removes the volume from xen01 itself (Toggle Disk ... removing disk). 23:17:52,xen01,kernel (drbd),Source Goes Diskless: xen01 detaches the underlying storage block (UpToDate -> Detaching -> Diskless). It alerts peers it can no longer satisfy read requests, calling drbd_bm_resize with capacity == 0. 23:17:53,xen05,LINSTOR/Satellite, While the cluster connection is actively tearing down, xen05's local background DeviceManager wakes up and detects the volume. It attempts an online re-alignment to round up the block layer size to match the 4MiB extent rule (26220040 KiB to 26222592 KiB). 23:17:54,xen05,kernel (drbd),Fatal Crash: xen05's DRBD driver tries to resize and count its synchronization bitmap block pointers. Because its peer source (xen01) dropped to zero capacity mid-flight, DRBD hits memory corruption. It triggers a General Protection Fault at drbd_bm_count_bits+0x223/0x300 and panics the kernel. -
My first idea was to have velero use a different storage class. Something like
apiVersion: storage.k8s.io/v1 kind: StorageClass metadata: name: linstor-ephemeral-retain parameters: linstor.csi.linbit.com/allowRemoteVolumeAccess: "true" linstor.csi.linbit.com/disklessOnRemaining: "true" linstor.csi.linbit.com/disklessStoragePool: DfltDisklessStorPool linstor.csi.linbit.com/placementCount: "3" provisioner: linstor.csi.linbit.com **reclaimPolicy: Retain** volumeBindingMode: WaitForFirstConsumerThen have a cronjob collected the
Releasedpvcs every hour. The problem is that can not work, as velero uses the storage class of the backed up pvc. So I would have to change every linstor storage class toRetain, and would effectively disable the automatic garbage collection for the cluster. Which might be fine, but would love to avoid it.My second idea is to eliminate the need for the resizing (another work around to hopefully get away from this race condition/bug)
I am also not seeing any way to modify the storageclass to round up or pad the pv. This would be to attempt to eliminate the need fordrbd_bm_resize.
Currently looking into Kyverno Mutating Webhook to dynamically pad these volumes on the fly to see if that would help me out at all.
All pvs are already whole numbers of Gi.
Velero calculates the size of the temporary clone pvc based on the exact byte-count of the used/allocated snapshot metadata, not the #Gi specification of the original pvc.
Just confirmed with a test

-
I was just moving around the velero backup schedules when it decided to run some again even though it was not time. Crashed xen04 this time. Deleting the possible offending backup schedule until this bug can be identified and fixed.
Uploaded new bug bundle to nextcloud too.
-
Logs from xen04 look like they are the same too. Which makes sense.
20:01:53.805: The Linstor DeviceManager on xen04 triggered an asynchronous alignment for a newly created volume (pvc-fac0a18b-2c23-4fe8-b005-b4f492cec92f), padding it from 26220040 KiB to 26222592 KiB. 20:01:54.034: Simultaneously, an orchestrator (like Velero) began tearing the volume down, shifting the DRBD state to conn( Connected -> TearDown ) and pdsk( Diskless -> DUnknown ). 20:01:58.484: The DRBD worker thread evaluated the shrinking bitmap while the alignment was trying to alter it. This caused the general protection fault: 0000 [#1] SMP NOPTI at instruction RIP: e030:drbd_bm_count_bits+0x223/0x300 [drbd], crashing the kernel. -
That timeline is the clearest version of this yet. xen04 going down the same way as xen05 makes the race look real to me rather than coincidental, unless I am reading the ordering wrong.
The fault lands insidedrbd_bm_count_bits, so in the DRBD module itself, which makes me think this is an upstream LINBIT question as much as an XCP-ng one. I do not know enough about how that module gets packaged here to say who owns the fix though. Might be worth another mention to @Team-Storage so someone can open your bundles with your post #10 next to them.
On the padding idea I am out of my depth: no idea whether LINSTOR can be told to pre-round a volume at creation, but asking LINBIT directly seems worth a try since it would sidestepdrbd_bm_resizeentirely.
-
@jonathon are you using your LINSTOR/DRBD stack in the form of licensed XOSTOR? Your issue looks a case for pro support. Did you consider opening a case with pro support?
-
@MajorP93 No we do not have pro support at the moment
hmmm. Really seems like a build with
9.2.19could fix all of this.
How can I request for a build @poddingue ?https://github.com/LINBIT/drbd/blob/drbd-9.2.19/ChangeLog#L61
9.2.17- Fix a kernel crash triggered by a crafted/invalid netlink message
https://github.com/LINBIT/drbd/blob/drbd-9.2.19/ChangeLog#L24
9.2.19- Fix an AB-BA deadlock between online resize and activity-log transactions
- Fix several races during connection teardown that could crash or
hang (ack_sender requeue, pending ping work, in-progress lb-tcp
connect)
-
On the version question, 8.3 ships DRBD 9.2.18 right now (that's what
kmod-drbd.specsays on the 8.3 branch), so you'd be asking for a single point bump rather than a big jump.
That probably helps you. Of your two links though, I think the netlink crash fix is a 9.2.17 entry, which would mean it's already in what you're running.
The 9.2.19 line that looks closest to your crash to me is the connection-teardown races one rather than the AB-BA deadlock, since a deadlock would hang rather than throw a GPF, but you've read these traces far more closely than I have so correct me if that's backwards.
I can't request a build myself, that's a @Team-Storage call rather than mine, though your post #13 timeline next to those changelog lines is a stronger case than most version-bump requests turn up with. -
@poddingue I'm fully up to date, but only at version
9.2.16[14:57 ovbh-pprod-xen02 ~]# drbdadm --version DRBDADM_BUILDTAG=GIT-hash:\ 71c8bcff6ea77a022b272a7eba649a774251bac4\ build\ by\ @buildsystem\,\ 2025-11-03\ 10:21:36 DRBDADM_API_VERSION=2 DRBD_KERNEL_VERSION_CODE=0x090210 DRBD_KERNEL_VERSION=9.2.16 DRBDADM_VERSION_CODE=0x092100 DRBDADM_VERSION=9.33.0 [14:57 ovbh-pprod-xen02 ~]# drbd-reactor --version drbd-reactor 1.9.0 [14:57 ovbh-pprod-xen02 ~]#
There are no updates that I can install
[14:08 ovbh-pprod-xen02 ~]# yum check-update Loaded plugins: fastestmirror Loading mirror speeds from cached hostfile Excluding mirror: updates.xcp-ng.org * xcp-ng-base: mirrors.xcp-ng.org Excluding mirror: updates.xcp-ng.org * xcp-ng-updates: mirrors.xcp-ng.org zabbix | 3.0 kB 00:00:00 zabbix-non-supported | 2.9 kB 00:00:00 zabbix/x86_64/primary_db | 114 kB 00:00:00 zabbix-agent.x86_64 7.0.28-release1.el7 zabbix -
For now I have upgraded all hosts to use
--enablerepo=xcp-ng-linstor-testingto getkmod-drbd.x86_64 0:9.2.18-2.0.xcpng8.3
That way it is there to be picked up if the host crashes again.
https://github.com/xcp-ng-rpms/drbd/blob/master/SPECS/kmod-drbd.spec
https://koji.xcp-ng.org/buildinfo?buildID=5747[12:15 ovbh-pprod-xen04 ~]# yum upgrade --enablerepo=xcp-ng-linstor-testing Loaded plugins: fastestmirror Loading mirror speeds from cached hostfile Excluding mirror: updates.xcp-ng.org * xcp-ng-base: mirrors.xcp-ng.org Excluding mirror: updates.xcp-ng.org * xcp-ng-updates: mirrors.xcp-ng.org grafana/signature | 810 B 00:00:00 grafana/signature | 3.0 kB 00:00:00 !!! xcp-ng-linstor-testing/signature | 473 B 00:00:00 xcp-ng-linstor-testing/signature | 3.0 kB 00:00:00 !!! zabbix | 3.0 kB 00:00:00 zabbix-non-supported | 2.9 kB 00:00:00 (1/3): grafana/primary_db | 879 kB 00:00:00 (2/3): xcp-ng-linstor-testing/primary_db | 2.4 kB 00:00:00 (3/3): zabbix/x86_64/primary_db | 114 kB 00:00:00 Resolving Dependencies --> Running transaction check ---> Package kmod-drbd.x86_64 0:9.2.16-1.0.xcpng8.3 will be updated ---> Package kmod-drbd.x86_64 0:9.2.18-2.0.xcpng8.3 will be an update ---> Package zabbix-agent.x86_64 0:7.0.27-release1.el7 will be updated ---> Package zabbix-agent.x86_64 0:7.0.28-release1.el7 will be an update --> Finished Dependency Resolution Dependencies Resolved ============================================================================================================================================================================================================================================================== Package Arch Version Repository Size ============================================================================================================================================================================================================================================================== Updating: kmod-drbd x86_64 9.2.18-2.0.xcpng8.3 xcp-ng-linstor-testing 2.8 M zabbix-agent x86_64 7.0.28-release1.el7 zabbix 660 k Transaction Summary ============================================================================================================================================================================================================================================================== Upgrade 2 Packages Total download size: 3.4 M Is this ok [y/d/N]: y Downloading packages: Delta RPMs disabled because /usr/bin/applydeltarpm not installed. (1/2): zabbix-agent-7.0.28-release1.el7.x86_64.rpm | 660 kB 00:00:00 (2/2): kmod-drbd-9.2.18-2.0.xcpng8.3.x86_64.rpm | 2.8 MB 00:00:01 -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- Total 2.8 MB/s | 3.4 MB 00:00:01 Running transaction check Running transaction test Transaction test succeeded Running transaction Updating : kmod-drbd-9.2.18-2.0.xcpng8.3.x86_64 1/4 Updating : zabbix-agent-7.0.28-release1.el7.x86_64 2/4 Cleanup : kmod-drbd-9.2.16-1.0.xcpng8.3.x86_64 3/4 Cleanup : zabbix-agent-7.0.27-release1.el7.x86_64 4/4 Verifying : zabbix-agent-7.0.28-release1.el7.x86_64 1/4 Verifying : kmod-drbd-9.2.18-2.0.xcpng8.3.x86_64 2/4 Verifying : kmod-drbd-9.2.16-1.0.xcpng8.3.x86_64 3/4 Verifying : zabbix-agent-7.0.27-release1.el7.x86_64 4/4 Updated: kmod-drbd.x86_64 0:9.2.18-2.0.xcpng8.3 zabbix-agent.x86_64 0:7.0.28-release1.el7 Complete!
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login