XenOrchestra not showing VM Disks on Pool (on single Server working) - XCP-ng Center is showing them
-
@poddingue Bug report filed as requested — https://github.com/xcp-ng/xcp/issues/825 — and tagging @Team-Storage per your suggestion.
Full evidence bundle is attached to the issue (versions, sweep output,
vhd-utilvsxecomparison, SMlog). Summary of what I found:One correction to the mechanism, and I think it matters. The recap describes
is-a-snapshotbeing flipped totrue. On my system that isn't what's happening —is-a-snapshotisfalseon every affected VDI. The field being wrongly written issnapshot-of, which is getting populated on base disks that aren't snapshots at all. XO's disappearing-disks symptom is consistent with either (it filters on a non-emptysnapshot-of), but if the storage team is hunting for a badis-a-snapshotwrite, that may be the wrong field. Every affected VDI here looks like:is-a-snapshot: false <-- correct snapshot-of: <populated with an unrelated VDI's UUID> <-- wrongA VDI that is a snapshot of itself. The clearest single artifact:
uuid: 806f7f42-083f-4a40-b3f1-0700d00bab5a name-label: WinSrv2022SHB_Disk1_Data is-a-snapshot: false snapshot-of: 806f7f42-083f-4a40-b3f1-0700d00bab5a <-- itself snapshot-time: 20260709T11:19:15Z sm-config: vhd-parent: c86e3247-... <-- bears no relation to the snapshot-of valueNo valid code path produces
snapshot-of = self. Whatever writes this field isn't validating the target.It's still actively corrupting new VDIs — this is not just legacy damage. That self-referential VDI was created 2026-07-09, a week after my patch + reboot. Sweeps 9 days apart went from ~180 → 191 affected VDIs on one SR, and a fourth anchor UUID appeared that didn't exist in the first sweep. Newly created VHDs keep landing in the affected set. So "stop it happening again" is the urgent half of the two-part fix, at least in my case.
The bogus targets cluster onto a tiny anchor set, and the anchors point at each other:
Count Anchor 97 937c3945(→a893fdb4)50 a893fdb4(→ea150883)37 ea1508837 806f7f42(→ itself, new since Jul 9)That looks less like corrupted lineage and more like the field being filled from an incorrect/uninitialised source.
On-disk VHDs are completely healthy.
vhd-util checksays valid, parent locators are consistent, GC reports no work. The two VDIs the DB calls parent/child are, on disk, siblings under a common parent. The corruption is purely in the XAPI database — which is good news for recoverability.The
VDI_IN_USEis not a real lock.current-operationsis empty,xe task-listis empty, no tapdisk holds it.VM.startfails because it's walking a snapshot relationship that doesn't exist on disk. Reproduces fromxeon the pool master with XO entirely out of the path — which is why I filed againstxcp-ng/xcprather than the XO tracker.Versions: XCP-ng 8.3.0, xapi 26.1.11 (
xapi-core-26.1.11-1.2),sm-3.2.12-17.9,sm-fairlock-3.2.12-17.9, blktap 3.55.5-9.1, build20260618.I have not attempted to bulk-clear the fields — on-disk data is intact and I'd rather not do a mass write against the XAPI DB on a live SR without guidance. Backing store snapshotted as a safety net.
Happy to run whatever diagnostics would help. And +1 to the hand-grenade feeling — the affected set growing on its own is the part that worries me.
-
@kagbasi-wgsdac , thank you so much for the issue creation and the details, that will help for sure!

-
@poddingue You're most welcome, sir.

-
@poddingue Bug report filed as requested — https://github.com/xcp-ng/xcp/issues/825 — and tagging @Team-Storage per your suggestion.
Full evidence bundle is attached to the issue (versions, sweep output,
vhd-utilvsxecomparison, SMlog). Summary of what I found:One correction to the mechanism, and I think it matters. The recap describes
is-a-snapshotbeing flipped totrue. On my system that isn't what's happening —is-a-snapshotisfalseon every affected VDI. The field being wrongly written issnapshot-of, which is getting populated on base disks that aren't snapshots at all. XO's disappearing-disks symptom is consistent with either (it filters on a non-emptysnapshot-of), but if the storage team is hunting for a badis-a-snapshotwrite, that may be the wrong field. Every affected VDI here looks like:is-a-snapshot: false <-- correct snapshot-of: <populated with an unrelated VDI's UUID> <-- wrongA VDI that is a snapshot of itself. The clearest single artifact:
uuid: 806f7f42-083f-4a40-b3f1-0700d00bab5a name-label: WinSrv2022SHB_Disk1_Data is-a-snapshot: false snapshot-of: 806f7f42-083f-4a40-b3f1-0700d00bab5a <-- itself snapshot-time: 20260709T11:19:15Z sm-config: vhd-parent: c86e3247-... <-- bears no relation to the snapshot-of valueNo valid code path produces
snapshot-of = self. Whatever writes this field isn't validating the target.It's still actively corrupting new VDIs — this is not just legacy damage. That self-referential VDI was created 2026-07-09, a week after my patch + reboot. Sweeps 9 days apart went from ~180 → 191 affected VDIs on one SR, and a fourth anchor UUID appeared that didn't exist in the first sweep. Newly created VHDs keep landing in the affected set. So "stop it happening again" is the urgent half of the two-part fix, at least in my case.
The bogus targets cluster onto a tiny anchor set, and the anchors point at each other:
Count Anchor 97 937c3945(→a893fdb4)50 a893fdb4(→ea150883)37 ea1508837 806f7f42(→ itself, new since Jul 9)That looks less like corrupted lineage and more like the field being filled from an incorrect/uninitialised source.
On-disk VHDs are completely healthy.
vhd-util checksays valid, parent locators are consistent, GC reports no work. The two VDIs the DB calls parent/child are, on disk, siblings under a common parent. The corruption is purely in the XAPI database — which is good news for recoverability.The
VDI_IN_USEis not a real lock.current-operationsis empty,xe task-listis empty, no tapdisk holds it.VM.startfails because it's walking a snapshot relationship that doesn't exist on disk. Reproduces fromxeon the pool master with XO entirely out of the path — which is why I filed againstxcp-ng/xcprather than the XO tracker.Versions: XCP-ng 8.3.0, xapi 26.1.11 (
xapi-core-26.1.11-1.2),sm-3.2.12-17.9,sm-fairlock-3.2.12-17.9, blktap 3.55.5-9.1, build20260618.I have not attempted to bulk-clear the fields — on-disk data is intact and I'd rather not do a mass write against the XAPI DB on a live SR without guidance. Backing store snapshotted as a safety net.
Happy to run whatever diagnostics would help. And +1 to the hand-grenade feeling — the affected set growing on its own is the part that worries me.
Has this issue been validated on a storage server built around Debian 13, LVM and ext4 or just TrueNAS when connected to XCP-ng version 8.3.0. As part of the XAPI DB corruption. Can anyone answer this please or give a clue?
-
@john.c Unfortunately, I cannot answer your question in the affirmative. I only use XCP-ng with TrueNAS.
@Team-Storage have y'all had a chance to take a look at the bug report I filed? This issue is quite serious.
-
@john.c See this post - https://xcp-ng.org/forum/post/105564
-
The July updates batch that went out on 28 July carries
xapi-26.1.11-1.3.xcpng8.3, and @kagbasi-wgsdac has since confirmed on https://github.com/xcp-ng/xcp/issues/825 that three days after patching his disks are all visible again and reverting a snapshot no longer duplicates VDIs.
The blog entry for that batch names the fix as non-snapshotted VBDs staying attached afterVM.revert, a regression fromxapi-26.1.4-3.3: https://xcp-ng.org/blog/2026/07/28/july-2026-updates-1-for-xcp-ng-8-3-lts/. Host reboots are needed. That stops new damage, but it does not unstamp VDIs that were already hit, so if disks are still hidden after you patch you probably still want the repair script at https://xcp-ng.org/forum/post/105564.
My post above also had the mechanism wrong, and @kagbasi-wgsdac corrected it: the field being wrongly written issnapshot-ofon base disks, notis-a-snapshot.
If anyone is still watching disks disappear on a fully updated pool, please say so here, because that would be something new rather than the tail of this one.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login