We are investigating recurring snapshot error 1200 on XCP-ng 8.3.
Snapshot creation and deletion succeeded on two different VMs using QCOW2 disks. However, the subsequent coalesce did not complete normally, and recovery repeatedly failed because it attempted to deactivate a logical volume still held open by the VM’s tapdisk.
For the first VM, we captured the original tapdisk commit failure:
qcow2_commit: error: The device is not writable: Permission denied
For the second VM, we have confirmed the interrupted coalesce and repeated rollback failure, but have not yet captured its original coalesce failure.
A comparison test using VHD disks on the same SR did not show the same failure.
Environment
- XCP-ng 8.3, three-host pool.
- Shared LVM over iSCSI SR on a QNAP TS-h2483XU-RP.
- SR name: 70TB_HDD
- SR UUID: 26d9572f-01c1-9cfd-79de-ca96eab601cf
- iSCSI target address: 172.16.150.2
- All three SR PBDs were reported as attached.
- Storage/GC operations were observed on xcp-hp-31x1.
- Both affected VMs were resident on xcp-hp-c8y.
Host label Shell hostname where different Address Host UUID
xcp-hp-31x xcp-hp-31x1 10.99.99.64 265d19d8-efa5-43cf-910b-6d03dd373bdd
xcp-hp-09v — 10.99.99.65 ea265ea3-31fd-414f-8ec6-07dd1498aff2
xcp-hp-c8y — 10.99.99.66 057656c9-0a32-464d-b99b-a297adcbbc72
Package versions verified on xcp-hp-31x1 and xcp-hp-c8y:
sm-3.2.12-23.5.xcpng8.3.x86_64
blktap-3.55.5-9.5.xcpng8.3.x86_64
All log timestamps below are local CEST, UTC+02:00.
Case 1: Mantak_CentOS7
VM UUID: 4d656c87-01f7-1183-48af-ead05f432274
VDI UUID: 96bc17fd-9188-458e-b0cb-5d6442e65237
Format: QCOW2
Size: 2 TiB
Host: xcp-hp-c8y
CBT was enabled for this disk.
Earlier during the incident, the VM became inaccessible over SSH. A shutdown attempt reported:
Internal error: Object with type VM and id
4d656c87-01f7-1183-48af-ead05f432274/config
does not exist in xenopsd
The previous domain subsequently disappeared and the VM was observed running as a new domain with SSH available. The exact action responsible for that transition is not established in the collected evidence.
A subsequent test snapshot was successfully created:
Snapshot UUID: ac95964f-1156-fbce-7318-28081b8075ff
Name: Mantak_test_2026-10-01
It was removed using xe snapshot-uninstall, which reported:
All objects destroyed
However, live coalesce subsequently failed.
October 1, 14:43:18 — xcp-hp-c8y, tapdisk log:
tapdisk[327527]: received 'commit' message (uuid = 7)
tapdisk[327527]: control: commit 7
tapdisk[327527]: vbd: commit /dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-96bc17fd-9188-458e-b0cb-5d6442e65237
tapdisk[327527]: qcow2_commit: error: The device is not writable: Permission denied
tapdisk[327527]: vbd: commit started (-22)
tapdisk[327527]: sending 'error' message (uuid = 7)
The corresponding command failed with status 22:
/usr/sbin/tap-ctl commit -p 327527 -m 7
-a /dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-96bc17fd-9188-458e-b0cb-5d6442e65237
Invalid argument
ERROR [-22 - tap_cli_commit]
Rollback then failed to deactivate the disk because it was in use.
The cleanup eventually completed at 14:50:07:
Tree bdd09a86-4c7a-440a-bc21-e2534a83dde3 gone
Removed leaf-coalesce from 96bc17fdqcow2
A different tapdisk PID was observed afterward. The exact recovery transition was not fully captured.
Comparison test: VHD disks on the same SR
VM UUID: 835d84d1-df18-1863-155f-d42ce9edbac8
VDI: 3a2299fb-4393-41e1-941d-7fc3f2380dc8
Format: VHD
Size: 100 GiB
Test parent: c002e130-7538-4cfa-a51a-aa603e593777
VDI: 74c1d103-acab-459b-9295-331621c511d5
Format: VHD
Size: 2040 GiB
Test parent: 5757888c-1b7e-48dc-b374-a7773fdce6be
Snapshots were created around 15:00–15:01 on October 1 and snapshot VDIs were removed around 15:02.
GC spent considerable time processing other disks on the same SR. We verified actual I/O progress and multiple successful VHD coalesce operations; it was not simply an inactive GC process.
By the following morning, the GC tree showed both test VDIs without their test parents, consistent with completed coalesce. The exact completion timestamps for these two disks have not yet been extracted.
Case 2: AlmaLinux 10 — second QCOW2 reproduction
VM UUID: d6ae4b31-e3dd-ff98-006e-4700aa586857
VDI UUID: 1bc0e201-1712-4409-b421-1d32f1b08d51
Format: QCOW2
Size: 3 TiB
Host: xcp-hp-c8y
Snapshot UUID: 3ef81ef1-81f0-c742-6bef-5083bc7e8f89
Snapshot VDI UUID: 46c74400-723c-44fd-9dd3-fd8d6de49ffe
Parent VDI UUID: 8ce4378a-421d-4323-b81d-ab93ebc013c3
The snapshot was created on October 1 at approximately 15:14. Its VDI was deleted at 15:15:39.
On October 2, GC was repeatedly attempting recovery approximately once per minute. The QCOW2 parent/child chain remained unchanged.
Master-side log, October 2:
07:03:56 SMGC: [19340] *** UNDO LEAF-COALESCE
07:03:56 SMGC: [19340] Updating qcow2OLD_1bc0e201-1712-4409-b421-1d32f1b08d51, QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51, QCOW2-8ce4378a-421d-4323-b81d-ab93ebc013c3 on slave xcp-hp-c8y
07:04:46 SM: [19340] ***** LVM over iSCSI: EXCEPTION <class 'XenAPI.Failure'>, ['XENAPI_PLUGIN_FAILURE', 'multi', 'CommandException', 'Input/output error']
Relevant traceback path:
LVMSR.load
_undoAllJournals
_handleInterruptedCoalesceLeaf
cleanup.gc_force
scanLocked
scan
_handleInterruptedCoalesceLeaf
_undoInterruptedCoalesceLeaf
_updateSlavesOnUndoLeafCoalesce
host.call_plugin(..., "multi", ...)
Slave-side log, xcp-hp-c8y:
LVMCache.deactivateNoRefcount: no LV qcow2OLD_1bc0e201-1712-4409-b421-1d32f1b08d51
on-slave.action 2: deactivateNoRefcount
['/sbin/lvchange', '-an',
'/dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51']
FAILED in util.pread: (rc 5)
Logical volume .../QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51 in use.
*** lvchange -an failed on attempt #1
...
*** lvchange -an failed on attempt #8
on-slave.deactivateNoRefcount failed
At 07:06, the disk was held open by tapdisk:
COMMAND PID USER FD TYPE DEVICE
tapdisk 351361 root 24u BLK 253,67
XAPI and Xen both showed the VM running:
name-label: AlmaLinux 10
power-state: running
dom-id: 61
xl list -v 61:
AlmaLinux 10 61 16385 4 -b---- ... d6ae4b31-e3dd-ff98-006e-4700aa586857
The guest remained accessible over SSH.
Recovery and current status
A maintenance outage was approved for AlmaLinux 10, and normal VM shutdown was requested to release the disk. No forced tapdisk termination or manual journal deletion was used in the documented procedure.
After the disk was released, GC completed the operation:
Oct 2 07:08:45 Running COW coalesce on 1bc0e201qcow2
Oct 2 07:08:46 Update-on-rename: VDI 1bc0e201... not attached on any slave
Oct 2 07:08:46 Removed vhd-parent from 1bc0e201qcow2
Oct 2 07:08:48 Tree 8ce4378a-421d-4323-b81d-ab93ebc013c3 gone
Oct 2 07:08:48 Removed leaf-coalesce from 1bc0e201qcow2
Restart instructions were provided afterward, but post-restart guest/service validation has not yet been reported.
This recovered the interrupted operation; it does not establish that the underlying live-coalesce defect has been fixed.
Questions
-
Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage?
-
Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open?
-
Is the qcow2OLD_<UUID> naming seen in rollback expected, compared with QCOW2-OLD_<UUID> seen during successful cleanup?
-
Is there a supported update, hotfix, or workaround for this configuration?
-
What additional logs are required to identify the original failure of the second QCOW2 VM?
-
What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups?
Full collected SMlog excerpts and the tapdisk error excerpt are available.