Categories

  • All news regarding Xen and XCP-ng ecosystem

    145 Topics
    5k Posts
    olivierlambertO
    Tested in my home lab, no issues so far (Intel CPU).
  • Everything related to the virtualization platform

    1k Topics
    15k Posts
    R
    We are investigating recurring snapshot error 1200 on XCP-ng 8.3. Snapshot creation and deletion succeeded on two different VMs using QCOW2 disks. However, the subsequent coalesce did not complete normally, and recovery repeatedly failed because it attempted to deactivate a logical volume still held open by the VM’s tapdisk. For the first VM, we captured the original tapdisk commit failure: qcow2_commit: error: The device is not writable: Permission denied For the second VM, we have confirmed the interrupted coalesce and repeated rollback failure, but have not yet captured its original coalesce failure. A comparison test using VHD disks on the same SR did not show the same failure. Environment XCP-ng 8.3, three-host pool. Shared LVM over iSCSI SR on a QNAP TS-h2483XU-RP. SR name: 70TB_HDD SR UUID: 26d9572f-01c1-9cfd-79de-ca96eab601cf iSCSI target address: 172.16.150.2 All three SR PBDs were reported as attached. Storage/GC operations were observed on xcp-hp-31x1. Both affected VMs were resident on xcp-hp-c8y. Host label Shell hostname where different Address Host UUID xcp-hp-31x xcp-hp-31x1 10.99.99.64 265d19d8-efa5-43cf-910b-6d03dd373bdd xcp-hp-09v — 10.99.99.65 ea265ea3-31fd-414f-8ec6-07dd1498aff2 xcp-hp-c8y — 10.99.99.66 057656c9-0a32-464d-b99b-a297adcbbc72 Package versions verified on xcp-hp-31x1 and xcp-hp-c8y: sm-3.2.12-23.5.xcpng8.3.x86_64 blktap-3.55.5-9.5.xcpng8.3.x86_64 All log timestamps below are local CEST, UTC+02:00. Case 1: Mantak_CentOS7 VM UUID: 4d656c87-01f7-1183-48af-ead05f432274 VDI UUID: 96bc17fd-9188-458e-b0cb-5d6442e65237 Format: QCOW2 Size: 2 TiB Host: xcp-hp-c8y CBT was enabled for this disk. Earlier during the incident, the VM became inaccessible over SSH. A shutdown attempt reported: Internal error: Object with type VM and id 4d656c87-01f7-1183-48af-ead05f432274/config does not exist in xenopsd The previous domain subsequently disappeared and the VM was observed running as a new domain with SSH available. The exact action responsible for that transition is not established in the collected evidence. A subsequent test snapshot was successfully created: Snapshot UUID: ac95964f-1156-fbce-7318-28081b8075ff Name: Mantak_test_2026-10-01 It was removed using xe snapshot-uninstall, which reported: All objects destroyed However, live coalesce subsequently failed. October 1, 14:43:18 — xcp-hp-c8y, tapdisk log: tapdisk[327527]: received 'commit' message (uuid = 7) tapdisk[327527]: control: commit 7 tapdisk[327527]: vbd: commit /dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-96bc17fd-9188-458e-b0cb-5d6442e65237 tapdisk[327527]: qcow2_commit: error: The device is not writable: Permission denied tapdisk[327527]: vbd: commit started (-22) tapdisk[327527]: sending 'error' message (uuid = 7) The corresponding command failed with status 22: /usr/sbin/tap-ctl commit -p 327527 -m 7 -a /dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-96bc17fd-9188-458e-b0cb-5d6442e65237 Invalid argument ERROR [-22 - tap_cli_commit] Rollback then failed to deactivate the disk because it was in use. The cleanup eventually completed at 14:50:07: Tree bdd09a86-4c7a-440a-bc21-e2534a83dde3 gone Removed leaf-coalesce from 96bc17fdqcow2 A different tapdisk PID was observed afterward. The exact recovery transition was not fully captured. Comparison test: VHD disks on the same SR VM UUID: 835d84d1-df18-1863-155f-d42ce9edbac8 VDI: 3a2299fb-4393-41e1-941d-7fc3f2380dc8 Format: VHD Size: 100 GiB Test parent: c002e130-7538-4cfa-a51a-aa603e593777 VDI: 74c1d103-acab-459b-9295-331621c511d5 Format: VHD Size: 2040 GiB Test parent: 5757888c-1b7e-48dc-b374-a7773fdce6be Snapshots were created around 15:00–15:01 on October 1 and snapshot VDIs were removed around 15:02. GC spent considerable time processing other disks on the same SR. We verified actual I/O progress and multiple successful VHD coalesce operations; it was not simply an inactive GC process. By the following morning, the GC tree showed both test VDIs without their test parents, consistent with completed coalesce. The exact completion timestamps for these two disks have not yet been extracted. Case 2: AlmaLinux 10 — second QCOW2 reproduction VM UUID: d6ae4b31-e3dd-ff98-006e-4700aa586857 VDI UUID: 1bc0e201-1712-4409-b421-1d32f1b08d51 Format: QCOW2 Size: 3 TiB Host: xcp-hp-c8y Snapshot UUID: 3ef81ef1-81f0-c742-6bef-5083bc7e8f89 Snapshot VDI UUID: 46c74400-723c-44fd-9dd3-fd8d6de49ffe Parent VDI UUID: 8ce4378a-421d-4323-b81d-ab93ebc013c3 The snapshot was created on October 1 at approximately 15:14. Its VDI was deleted at 15:15:39. On October 2, GC was repeatedly attempting recovery approximately once per minute. The QCOW2 parent/child chain remained unchanged. Master-side log, October 2: 07:03:56 SMGC: [19340] *** UNDO LEAF-COALESCE 07:03:56 SMGC: [19340] Updating qcow2OLD_1bc0e201-1712-4409-b421-1d32f1b08d51, QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51, QCOW2-8ce4378a-421d-4323-b81d-ab93ebc013c3 on slave xcp-hp-c8y 07:04:46 SM: [19340] ***** LVM over iSCSI: EXCEPTION <class 'XenAPI.Failure'>, ['XENAPI_PLUGIN_FAILURE', 'multi', 'CommandException', 'Input/output error'] Relevant traceback path: LVMSR.load _undoAllJournals _handleInterruptedCoalesceLeaf cleanup.gc_force scanLocked scan _handleInterruptedCoalesceLeaf _undoInterruptedCoalesceLeaf _updateSlavesOnUndoLeafCoalesce host.call_plugin(..., "multi", ...) Slave-side log, xcp-hp-c8y: LVMCache.deactivateNoRefcount: no LV qcow2OLD_1bc0e201-1712-4409-b421-1d32f1b08d51 on-slave.action 2: deactivateNoRefcount ['/sbin/lvchange', '-an', '/dev/VG_XenStorage-26d9572f-01c1-9cfd-79de-ca96eab601cf/QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51'] FAILED in util.pread: (rc 5) Logical volume .../QCOW2-1bc0e201-1712-4409-b421-1d32f1b08d51 in use. *** lvchange -an failed on attempt #1 ... *** lvchange -an failed on attempt #8 on-slave.deactivateNoRefcount failed At 07:06, the disk was held open by tapdisk: COMMAND PID USER FD TYPE DEVICE tapdisk 351361 root 24u BLK 253,67 XAPI and Xen both showed the VM running: name-label: AlmaLinux 10 power-state: running dom-id: 61 xl list -v 61: AlmaLinux 10 61 16385 4 -b---- ... d6ae4b31-e3dd-ff98-006e-4700aa586857 The guest remained accessible over SSH. Recovery and current status A maintenance outage was approved for AlmaLinux 10, and normal VM shutdown was requested to release the disk. No forced tapdisk termination or manual journal deletion was used in the documented procedure. After the disk was released, GC completed the operation: Oct 2 07:08:45 Running COW coalesce on 1bc0e201qcow2 Oct 2 07:08:46 Update-on-rename: VDI 1bc0e201... not attached on any slave Oct 2 07:08:46 Removed vhd-parent from 1bc0e201qcow2 Oct 2 07:08:48 Tree 8ce4378a-421d-4323-b81d-ab93ebc013c3 gone Oct 2 07:08:48 Removed leaf-coalesce from 1bc0e201qcow2 Restart instructions were provided afterward, but post-restart guest/service validation has not yet been reported. This recovered the interrupted operation; it does not establish that the underlying live-coalesce defect has been fixed. Questions Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage? Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open? Is the qcow2OLD_<UUID> naming seen in rollback expected, compared with QCOW2-OLD_<UUID> seen during successful cleanup? Is there a supported update, hotfix, or workaround for this configuration? What additional logs are required to identify the original failure of the second QCOW2 VM? What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups? Full collected SMlog excerpts and the tapdisk error excerpt are available.
  • 3k Topics
    29k Posts
    Z
    The internet check was timing out on the IPv6 address for xen-orchestra.com IPv4 was fine the whole time. curl -4 -I --max-time 15 https://xen-orchestra.com/ # HTTP/2 200 curl -6 -I --max-time 15 https://xen-orchestra.com/ # couldn't connect What fixed it was disabling IPv6 at boot, then rebooting so xo-server starts with it already off: printf '%s\n' 'net.ipv6.conf.all.disable_ipv6=1' 'net.ipv6.conf.default.disable_ipv6=1' | sudo tee /etc/sysctl.d/99-disable-ipv6.conf After that, xoa check was green, including internet connectivity and no more check for upgrade issues. I'm hoping this fixes my XO Config Cloud Backup. I will report back if not. 17/17 - Internet connectivity: Error: HTTP connection has timed out at ClientRequest.<anonymous> (/usr/local/lib/node_modules/xoa-cli/node_modules/http-request-plus/index.js:61:25) at ClientRequest.emit (node:events:519:28) at ClientRequest.patchedError [as emit] (file:///usr/local/lib/node_modules/xoa-cli/index.mjs:31:17) at TLSSocket.emitRequestTimeout (node:_http_client:927:9) at Object.onceWrapper (node:events:633:28) at TLSSocket.emit (node:events:531:35) at TLSSocket.patchedError [as emit] (file:///usr/local/lib/node_modules/xoa-cli/index.mjs:31:17) at Socket._onTimeout (node:net:604:8) at listOnTimeout (node:internal/timers:585:17) at process.processTimers (node:internal/timers:521:7) { url: 'https://xen-orchestra.com/', originalUrl: 'http://xen-orchestra.com/' } XOA Update sometimes failed too but hitting Refresh a couple times got it going. 10/1/2026, 10:03:28 AM: All up to date 10/1/2026, 10:26:26 AM: Start updating... 10/1/2026, 10:26:32 AM: HTTP connection has timed out 10/1/2026, 10:27:14 AM: Start updating... 10/1/2026, 10:27:14 AM: stable channel selected 10/1/2026, 10:27:14 AM: All up to date
  • Our hyperconverged storage solution

    54 Topics
    824 Posts
    J
    @poddingue Turns out it is all secondary-to-secondary out of sync. So that makes it less intense, though if primary dies I wonder how it will resolve this, or if it will become split brained. Still unsure of how it happened, but the ones I manually cleaned up have not come back. Going to continue to manually clear them up. If it happens again I will have an alert setup to notify me, and I have all the xcp-ng logs and everything to be able to see what happened. If that happens I will post here with details and logs for the resource so we can see how it occurs. [image: image.jpeg]
  • 37 Topics
    136 Posts
    J
    @AtaxyaNetwork Merci pour tes recherches ! Oui "cd_label" serait cool comme ajout au plugin ce qui permet sur les distro type Fedora/Redhat de ne pas avoir de boot_command à gérer