XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    XCP-ng 8.3 — QCOW2 snapshot deletion followed by failed live coalesce and recurring rollback failures on shared iSCSI SR

    Scheduled Pinned Locked Moved Unsolved XCP-ng
    2 Posts 2 Posters 28 Views 2 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • R
      racom.jiristerba
      last edited by

      We are investigating recurring snapshot error 1200 on XCP-ng 8.3.
      Snapshot creation and deletion succeeded on two different VMs using QCOW2 disks. However, the subsequent coalesce did not complete normally, and recovery repeatedly failed because it attempted to deactivate a logical volume still held open by the VM’s tapdisk.
      For the first VM, we captured the original tapdisk commit failure:
      qcow2_commit: error: The device is not writable: Permission denied
      For the second VM, we have confirmed the interrupted coalesce and repeated rollback failure, but have not yet captured its original coalesce failure.
      A comparison test using VHD disks on the same SR did not show the same failure.
      Environment

      • XCP-ng 8.3, three-host pool.
      • Shared LVM over iSCSI SR on a QNAP TS-h2483XU-RP.
      • SR name: 70TB_HDD
      • SR UUID: 26d9572f-01c1-9cfd-79de-ca96eab601cf
      • iSCSI target address: 172.16.150.2
      • All three SR PBDs were reported as attached.
      • Storage/GC operations were observed on xcp-hp-31x1.
      • Both affected VMs were resident on xcp-hp-c8y.
        Host label Shell hostname where different Address Host UUID
        xcp-hp-31x xcp-hp-31x1 10.99.99.64 265d19d8-efa5-43cf-910b-6d03dd373bdd
        xcp-hp-09v — 10.99.99.65 ea265ea3-31fd-414f-8ec6-07dd1498aff2
        xcp-hp-c8y — 10.99.99.66 057656c9-0a32-464d-b99b-a297adcbbc72

      Package versions verified on xcp-hp-31x1 and xcp-hp-c8y:
      sm-3.2.12-23.5.xcpng8.3.x86_64
      blktap-3.55.5-9.5.xcpng8.3.x86_64

      All log timestamps below are local CEST, UTC+02:00.
      Case 1: Mantak_CentOS7
      VM UUID: 4d656c87-01f7-1183-48af-ead05f432274
      VDI UUID: 96bc17fd-9188-458e-b0cb-5d6442e65237
      Format: QCOW2
      Size: 2 TiB
      Host: xcp-hp-c8y
      CBT was enabled for this disk.
      Earlier during the incident, the VM became inaccessible over SSH. A shutdown attempt reported:
      Internal error: Object with type VM and id
      4d656c87-01f7-1183-48af-ead05f432274/config
      does not exist in xenopsd
      The previous domain subsequently disappeared and the VM was observed running as a new domain with SSH available. The exact action responsible for that transition is not established in the collected evidence.
      A subsequent test snapshot was successfully created:
      Snapshot UUID: ac95964f-1156-fbce-7318-28081b8075ff
      Name: Mantak_test_2026-10-01
      It was removed using xe snapshot-uninstall, which reported:
      All objects destroyed
      However, live coalesce subsequently failed.
      October 1, 14:43:18 — xcp-hp-c8y, tapdisk log:
      tapdisk[327527]: received 'commit' message (uuid = 7)
      tapdisk[327527]: control: commit 7
      tapdisk[327527]: vbd: commit /dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-96bc17fd-9188-458e-b0cb-5d6442e65237
      tapdisk[327527]: qcow2_commit: error: The device is not writable: Permission denied
      tapdisk[327527]: vbd: commit started (-22)
      tapdisk[327527]: sending 'error' message (uuid = 7)
      The corresponding command failed with status 22:
      /usr/sbin/tap-ctl commit -p 327527 -m 7
      -a /dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-96bc17fd-9188-458e-b0cb-5d6442e65237

      Invalid argument
      ERROR [-22 - tap_cli_commit]
      Rollback then failed to deactivate the disk because it was in use.
      The cleanup eventually completed at 14:50:07:
      Tree bdd09a86-4c7a-440a-bc21-e2534a83dde3 gone
      Removed leaf-coalesce from 96bc17fdqcow2
      A different tapdisk PID was observed afterward. The exact recovery transition was not fully captured.
      Comparison test: VHD disks on the same SR
      VM UUID: 835d84d1-df18-1863-155f-d42ce9edbac8

      VDI: 3a2299fb-4393-41e1-941d-7fc3f2380dc8
      Format: VHD
      Size: 100 GiB
      Test parent: c002e130-7538-4cfa-a51a-aa603e593777

      VDI: 74c1d103-acab-459b-9295-331621c511d5
      Format: VHD
      Size: 2040 GiB
      Test parent: 5757888c-1b7e-48dc-b374-a7773fdce6be
      Snapshots were created around 15:00–15:01 on October 1 and snapshot VDIs were removed around 15:02.
      GC spent considerable time processing other disks on the same SR. We verified actual I/O progress and multiple successful VHD coalesce operations; it was not simply an inactive GC process.
      By the following morning, the GC tree showed both test VDIs without their test parents, consistent with completed coalesce. The exact completion timestamps for these two disks have not yet been extracted.
      Case 2: AlmaLinux 10 — second QCOW2 reproduction
      VM UUID: d6ae4b31-e3dd-ff98-006e-4700aa586857
      VDI UUID: 1bc0e201-1712-4409-b421-1d32f1b08d51
      Format: QCOW2
      Size: 3 TiB
      Host: xcp-hp-c8y

      Snapshot UUID: 3ef81ef1-81f0-c742-6bef-5083bc7e8f89
      Snapshot VDI UUID: 46c74400-723c-44fd-9dd3-fd8d6de49ffe
      Parent VDI UUID: 8ce4378a-421d-4323-b81d-ab93ebc013c3
      The snapshot was created on October 1 at approximately 15:14. Its VDI was deleted at 15:15:39.
      On October 2, GC was repeatedly attempting recovery approximately once per minute. The QCOW2 parent/child chain remained unchanged.
      Master-side log, October 2:
      07:03:56 SMGC: [19340] *** UNDO LEAF-COALESCE
      07:03:56 SMGC: [19340] Updating qcow2OLD_1bc0e201-1712-4409-b421-1d32f1b08d51, QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51, QCOW2-8ce4378a-421d-4323-b81d-ab93ebc013c3 on slave xcp-hp-c8y
      07:04:46 SM: [19340] ***** LVM over iSCSI: EXCEPTION <class 'XenAPI.Failure'>, ['XENAPI_PLUGIN_FAILURE', 'multi', 'CommandException', 'Input/output error']
      Relevant traceback path:
      LVMSR.load
      _undoAllJournals
      _handleInterruptedCoalesceLeaf
      cleanup.gc_force
      scanLocked
      scan
      _handleInterruptedCoalesceLeaf
      _undoInterruptedCoalesceLeaf
      _updateSlavesOnUndoLeafCoalesce
      host.call_plugin(..., "multi", ...)
      Slave-side log, xcp-hp-c8y:
      LVMCache.deactivateNoRefcount: no LV qcow2OLD_1bc0e201-1712-4409-b421-1d32f1b08d51

      on-slave.action 2: deactivateNoRefcount

      ['/sbin/lvchange', '-an',
      '/dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51']

      FAILED in util.pread: (rc 5)
      Logical volume .../QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51 in use.

      *** lvchange -an failed on attempt #1
      ...
      *** lvchange -an failed on attempt #8
      on-slave.deactivateNoRefcount failed
      At 07:06, the disk was held open by tapdisk:
      COMMAND PID USER FD TYPE DEVICE
      tapdisk 351361 root 24u BLK 253,67
      XAPI and Xen both showed the VM running:
      name-label: AlmaLinux 10
      power-state: running
      dom-id: 61

      xl list -v 61:
      AlmaLinux 10 61 16385 4 -b---- ... d6ae4b31-e3dd-ff98-006e-4700aa586857
      The guest remained accessible over SSH.
      Recovery and current status
      A maintenance outage was approved for AlmaLinux 10, and normal VM shutdown was requested to release the disk. No forced tapdisk termination or manual journal deletion was used in the documented procedure.
      After the disk was released, GC completed the operation:
      Oct 2 07:08:45 Running COW coalesce on 1bc0e201qcow2
      Oct 2 07:08:46 Update-on-rename: VDI 1bc0e201... not attached on any slave
      Oct 2 07:08:46 Removed vhd-parent from 1bc0e201qcow2
      Oct 2 07:08:48 Tree 8ce4378a-421d-4323-b81d-ab93ebc013c3 gone
      Oct 2 07:08:48 Removed leaf-coalesce from 1bc0e201qcow2
      Restart instructions were provided afterward, but post-restart guest/service validation has not yet been reported.
      This recovered the interrupted operation; it does not establish that the underlying live-coalesce defect has been fixed.

      Questions

      1. Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage?

      2. Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open?

      3. Is the qcow2OLD_<UUID> naming seen in rollback expected, compared with QCOW2-OLD_<UUID> seen during successful cleanup?

      4. Is there a supported update, hotfix, or workaround for this configuration?

      5. What additional logs are required to identify the original failure of the second QCOW2 VM?

      6. What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups?
        Full collected SMlog excerpts and the tapdisk error excerpt are available.

      dthenotD 1 Reply Last reply
      Reply Quote 0
      • dthenotD
        dthenot Vates 🪐 XCP-ng Team @racom.jiristerba
        last edited by

        @racom.jiristerba Hello, this appear to be bugs we are already fixing in the build currently in the testing repository. Please see this post for instruction to install the updates : https://xcp-ng.org/forum/post/108803

        These two bugs are: the chain being RO sometime which the tapdisk process doing the coalesce can't access and a bug with the LV being in use when it's trying to refresh the VDI it was coalescing and it's parent following the cleanup. Both of those are fixed in sm-3.2.12-25.1.xcpng8.3

        1 Reply Last reply
        Reply Quote 0
        • poddingueP poddingue marked this topic as a question

        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

        With your input, this post could be even better 💗

        Register Login
        • First post
          Last post