<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[XCP-ng 8.3 — QCOW2 snapshot deletion followed by failed live coalesce and recurring rollback failures on shared iSCSI SR]]></title><description><![CDATA[<p dir="auto">We are investigating recurring snapshot error 1200 on XCP-ng 8.3.<br />
Snapshot creation and deletion succeeded on two different VMs using QCOW2 disks. However, the subsequent coalesce did not complete normally, and recovery repeatedly failed because it attempted to deactivate a logical volume still held open by the VM’s tapdisk.<br />
For the first VM, we captured the original tapdisk commit failure:<br />
qcow2_commit: error: The device is not writable: Permission denied<br />
For the second VM, we have confirmed the interrupted coalesce and repeated rollback failure, but have not yet captured its original coalesce failure.<br />
A comparison test using VHD disks on the same SR did not show the same failure.<br />
Environment</p>
<ul>
<li>XCP-ng 8.3, three-host pool.</li>
<li>Shared LVM over iSCSI SR on a QNAP TS-h2483XU-RP.</li>
<li>SR name: 70TB_HDD</li>
<li>SR UUID: 26d9572f-01c1-9cfd-79de-ca96eab601cf</li>
<li>iSCSI target address: 172.16.150.2</li>
<li>All three SR PBDs were reported as attached.</li>
<li>Storage/GC operations were observed on xcp-hp-31x1.</li>
<li>Both affected VMs were resident on xcp-hp-c8y.<br />
Host label	Shell hostname where different	Address	Host UUID<br />
xcp-hp-31x	xcp-hp-31x1	10.99.99.64	265d19d8-efa5-43cf-910b-6d03dd373bdd<br />
xcp-hp-09v	—	10.99.99.65	ea265ea3-31fd-414f-8ec6-07dd1498aff2<br />
xcp-hp-c8y	—	10.99.99.66	057656c9-0a32-464d-b99b-a297adcbbc72</li>
</ul>
<p dir="auto">Package versions verified on xcp-hp-31x1 and xcp-hp-c8y:<br />
sm-3.2.12-23.5.xcpng8.3.x86_64<br />
blktap-3.55.5-9.5.xcpng8.3.x86_64</p>
<p dir="auto">All log timestamps below are local CEST, UTC+02:00.<br />
Case 1: Mantak_CentOS7<br />
VM UUID:  4d656c87-01f7-1183-48af-ead05f432274<br />
VDI UUID: 96bc17fd-9188-458e-b0cb-5d6442e65237<br />
Format:   QCOW2<br />
Size:     2 TiB<br />
Host:     xcp-hp-c8y<br />
CBT was enabled for this disk.<br />
Earlier during the incident, the VM became inaccessible over SSH. A shutdown attempt reported:<br />
Internal error: Object with type VM and id<br />
4d656c87-01f7-1183-48af-ead05f432274/config<br />
does not exist in xenopsd<br />
The previous domain subsequently disappeared and the VM was observed running as a new domain with SSH available. The exact action responsible for that transition is not established in the collected evidence.<br />
A subsequent test snapshot was successfully created:<br />
Snapshot UUID: ac95964f-1156-fbce-7318-28081b8075ff<br />
Name: Mantak_test_2026-10-01<br />
It was removed using xe snapshot-uninstall, which reported:<br />
All objects destroyed<br />
However, live coalesce subsequently failed.<br />
October 1, 14:43:18 — xcp-hp-c8y, tapdisk log:<br />
tapdisk[327527]: received 'commit' message (uuid = 7)<br />
tapdisk[327527]: control: commit 7<br />
tapdisk[327527]: vbd: commit /dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-96bc17fd-9188-458e-b0cb-5d6442e65237<br />
tapdisk[327527]: qcow2_commit: error: The device is not writable: Permission denied<br />
tapdisk[327527]: vbd: commit started (-22)<br />
tapdisk[327527]: sending 'error' message (uuid = 7)<br />
The corresponding command failed with status 22:<br />
/usr/sbin/tap-ctl commit -p 327527 -m 7 <br />
-a /dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-96bc17fd-9188-458e-b0cb-5d6442e65237</p>
<p dir="auto">Invalid argument<br />
ERROR [-22 - tap_cli_commit]<br />
Rollback then failed to deactivate the disk because it was in use.<br />
The cleanup eventually completed at 14:50:07:<br />
Tree bdd09a86-4c7a-440a-bc21-e2534a83dde3 gone<br />
Removed leaf-coalesce from 96bc17fd<a href="2048.000G///2048.316G%7Ca" target="_blank" rel="noopener noreferrer nofollow ugc">qcow2</a><br />
A different tapdisk PID was observed afterward. The exact recovery transition was not fully captured.<br />
Comparison test: VHD disks on the same SR<br />
VM UUID: 835d84d1-df18-1863-155f-d42ce9edbac8</p>
<p dir="auto">VDI: 3a2299fb-4393-41e1-941d-7fc3f2380dc8<br />
Format: VHD<br />
Size: 100 GiB<br />
Test parent: c002e130-7538-4cfa-a51a-aa603e593777</p>
<p dir="auto">VDI: 74c1d103-acab-459b-9295-331621c511d5<br />
Format: VHD<br />
Size: 2040 GiB<br />
Test parent: 5757888c-1b7e-48dc-b374-a7773fdce6be<br />
Snapshots were created around 15:00–15:01 on October 1 and snapshot VDIs were removed around 15:02.<br />
GC spent considerable time processing other disks on the same SR. We verified actual I/O progress and multiple successful VHD coalesce operations; it was not simply an inactive GC process.<br />
By the following morning, the GC tree showed both test VDIs without their test parents, consistent with completed coalesce. The exact completion timestamps for these two disks have not yet been extracted.<br />
Case 2: AlmaLinux 10 — second QCOW2 reproduction<br />
VM UUID:  d6ae4b31-e3dd-ff98-006e-4700aa586857<br />
VDI UUID: 1bc0e201-1712-4409-b421-1d32f1b08d51<br />
Format:   QCOW2<br />
Size:     3 TiB<br />
Host:     xcp-hp-c8y</p>
<p dir="auto">Snapshot UUID:     3ef81ef1-81f0-c742-6bef-5083bc7e8f89<br />
Snapshot VDI UUID: 46c74400-723c-44fd-9dd3-fd8d6de49ffe<br />
Parent VDI UUID:   8ce4378a-421d-4323-b81d-ab93ebc013c3<br />
The snapshot was created on October 1 at approximately 15:14. Its VDI was deleted at 15:15:39.<br />
On October 2, GC was repeatedly attempting recovery approximately once per minute. The QCOW2 parent/child chain remained unchanged.<br />
Master-side log, October 2:<br />
07:03:56 SMGC: [19340] *** UNDO LEAF-COALESCE<br />
07:03:56 SMGC: [19340] Updating qcow2OLD_1bc0e201-1712-4409-b421-1d32f1b08d51, QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51, QCOW2-8ce4378a-421d-4323-b81d-ab93ebc013c3 on slave xcp-hp-c8y<br />
07:04:46 SM: [19340] ***** LVM over iSCSI: EXCEPTION &lt;class 'XenAPI.Failure'&gt;, ['XENAPI_PLUGIN_FAILURE', 'multi', 'CommandException', 'Input/output error']<br />
Relevant traceback path:<br />
LVMSR.load<br />
_undoAllJournals<br />
_handleInterruptedCoalesceLeaf<br />
cleanup.gc_force<br />
scanLocked<br />
scan<br />
_handleInterruptedCoalesceLeaf<br />
_undoInterruptedCoalesceLeaf<br />
_updateSlavesOnUndoLeafCoalesce<br />
host.call_plugin(..., "multi", ...)<br />
Slave-side log, xcp-hp-c8y:<br />
LVMCache.deactivateNoRefcount: no LV qcow2OLD_1bc0e201-1712-4409-b421-1d32f1b08d51</p>
<p dir="auto">on-slave.action 2: deactivateNoRefcount</p>
<p dir="auto">['/sbin/lvchange', '-an',<br />
'/dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51']</p>
<p dir="auto">FAILED in util.pread: (rc 5)<br />
Logical volume .../QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51 in use.</p>
<p dir="auto">*** lvchange -an failed on attempt #1<br />
...<br />
*** lvchange -an failed on attempt #8<br />
on-slave.deactivateNoRefcount failed<br />
At 07:06, the disk was held open by tapdisk:<br />
COMMAND    PID USER FD  TYPE DEVICE<br />
tapdisk 351361 root 24u BLK  253,67<br />
XAPI and Xen both showed the VM running:<br />
name-label: AlmaLinux 10<br />
power-state: running<br />
dom-id: 61</p>
<p dir="auto">xl list -v 61:<br />
AlmaLinux 10  61  16385  4  -b----  ... d6ae4b31-e3dd-ff98-006e-4700aa586857<br />
The guest remained accessible over SSH.<br />
Recovery and current status<br />
A maintenance outage was approved for AlmaLinux 10, and normal VM shutdown was requested to release the disk. No forced tapdisk termination or manual journal deletion was used in the documented procedure.<br />
After the disk was released, GC completed the operation:<br />
Oct 2 07:08:45 Running COW coalesce on 1bc0e201<a href="3072.000G///3072.473G%7Ca" target="_blank" rel="noopener noreferrer nofollow ugc">qcow2</a><br />
Oct 2 07:08:46 Update-on-rename: VDI 1bc0e201... not attached on any slave<br />
Oct 2 07:08:46 Removed vhd-parent from 1bc0e201<a href="3072.000G///3072.473G%7Ca" target="_blank" rel="noopener noreferrer nofollow ugc">qcow2</a><br />
Oct 2 07:08:48 Tree 8ce4378a-421d-4323-b81d-ab93ebc013c3 gone<br />
Oct 2 07:08:48 Removed leaf-coalesce from 1bc0e201<a href="3072.000G///3072.473G%7Ca" target="_blank" rel="noopener noreferrer nofollow ugc">qcow2</a><br />
Restart instructions were provided afterward, but post-restart guest/service validation has not yet been reported.<br />
This recovered the interrupted operation; it does not establish that the underlying live-coalesce defect has been fixed.</p>
<p dir="auto">Questions</p>
<ol>
<li>
<p dir="auto">Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage?</p>
</li>
<li>
<p dir="auto">Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open?</p>
</li>
<li>
<p dir="auto">Is the qcow2OLD_&lt;UUID&gt; naming seen in rollback expected, compared with QCOW2-OLD_&lt;UUID&gt; seen during successful cleanup?</p>
</li>
<li>
<p dir="auto">Is there a supported update, hotfix, or workaround for this configuration?</p>
</li>
<li>
<p dir="auto">What additional logs are required to identify the original failure of the second QCOW2 VM?</p>
</li>
<li>
<p dir="auto">What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups?<br />
Full collected SMlog excerpts and the tapdisk error excerpt are available.</p>
</li>
</ol>
]]></description><link>https://xcp-ng.org/forum/topic/12508/xcp-ng-8.3-qcow2-snapshot-deletion-followed-by-failed-live-coalesce-and-recurring-rollback-failures-on-shared-iscsi-sr</link><generator>RSS for Node</generator><lastBuildDate>Fri, 02 Oct 2026 07:49:15 GMT</lastBuildDate><atom:link href="https://xcp-ng.org/forum/topic/12508.rss" rel="self" type="application/rss+xml"/><pubDate>Fri, 02 Oct 2026 06:27:47 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to XCP-ng 8.3 — QCOW2 snapshot deletion followed by failed live coalesce and recurring rollback failures on shared iSCSI SR on Fri, 02 Oct 2026 07:48:52 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/forum/user/racom.jiristerba" aria-label="Profile: racom.jiristerba">@<bdi>racom.jiristerba</bdi></a> Hello, this appear to be bugs we are already fixing in the build currently in the testing repository. Please see this post for instruction to install the updates : <a href="https://xcp-ng.org/forum/post/108803">https://xcp-ng.org/forum/post/108803</a></p>
<p dir="auto">These two bugs are: the chain being RO sometime which the tapdisk process doing the coalesce can't access and a bug with the LV being in use when it's trying to refresh the VDI it was coalescing and it's parent following the cleanup. Both of those are fixed in <code>sm-3.2.12-25.1.xcpng8.3</code></p>
]]></description><link>https://xcp-ng.org/forum/post/108831</link><guid isPermaLink="true">https://xcp-ng.org/forum/post/108831</guid><dc:creator><![CDATA[dthenot]]></dc:creator><pubDate>Fri, 02 Oct 2026 07:48:52 GMT</pubDate></item></channel></rss>