Categories

  • All news regarding Xen and XCP-ng ecosystem

    145 Topics
    5k Posts
    gduperreyG
    Thank you everyone for your tests and your feedback! The updates are live now: https://xcp-ng.org/blog/2026/10/05/october-2026-updates-1-for-xcp-ng-8-3-lts/ We decided not to release the xen update that was initially included in the testing phase of this update cycle. It introduced changes related to support for Windows VMs with more than 64 vCPUs; a feature that is not yet fully supported and requires further work to be completely implemented. We therefore chose to postpone it to avoid any regressions or user confusion, aiming to release a properly finalized and documented version in the near future. This version should have no impact on your test environments, and they can continue to operate as is. Any future xen update should have a higher version number and should install over the current one. However, if you do not wish to keep this update on your test environments, you can downgrade using the command yum downgrade xen-*.
  • Everything related to the virtualization platform

    1k Topics
    15k Posts
    dthenotD
    @racom.jiristerba Sorry, I missed answering your questions the first time Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage? Yes, it's a known issue, most of them will be fixed with the latest update, you might need to activate the VDI again but a migration during RPU will be enough Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open? It's because in the case of VHD, it's the case but with QCOW2 we don't stop the tapdisk process accessing the VDI, it's a change that was missed Is the qcow2OLD_<UUID> naming seen in rollback expected, compared with QCOW2-OLD_<UUID> seen during successful cleanup? It's a pre-existing bug that I already have on my TODO list Is there a supported update, hotfix, or workaround for this configuration? Installing the latest release sm-3.2.12-25.1 during updates (and maybe launching a xe sr-scan uuid=<SR UUI> after updating it so it auto-resolve the undo) should be enough What additional logs are required to identify the original failure of the second QCOW2 VM? I don't think we need any more logs since the errors I'm seeing should already be fixed. If you have any more issues after installing the newest packages, I will take another look What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups? The sr-scan after updating should do it automatically, it shouldn't need any manipulation. Looking at the storage logs in /var/log/SMlog for any irregularities could help to see problems.
  • 3k Topics
    29k Posts
    A
    Data point from a 2-host XCP-ng 8.3 pool that supports @tosh's diagnosis, plus a no-patch workaround that gave us ~2.7x until the TCP_NODELAY fix lands. Setup XCP-ng 8.3, xapi-core / vhd-tool 26.1.16-1.2.xcpng8.3, dom0 kernel 4.19.0+1 Two hosts, local ext SRs on NVMe (Samsung 970 EVO Plus / WD Red SN700) Dedicated migration network: 2.5 GbE (Intel i226-V, igc), MTU 9000, no errors Offline storage migration (halted VM, VM.migrate_send) of 2 VDIs (disk + snapshot), 16 GiB virtual / 7773 MiB allocated each; sparse_dd ... -prezeroed ... -dest-proto nbd over TLS Baseline (stock) 132 s per VDI, i.e. ~59 MiB/s on the wire, flat line, same in both directions Link at ~20 % of line rate, source SR read latency 0.3 ms, dom0 total <= 0.5 vCPU, no single dom0 vCPU above 0.08 (RRD, 60 s averages) Sending host with turbo or capped at 2.5 GHz gave the same 132 s, so not CPU-bound here Workaround: disable delayed ACKs on the receiver, only for the migration-network route # on the destination host; adapt prefix/dev/src to the output of: ip route show <migration-net> ip route change 10.10.10.0/24 dev xenbr1 proto kernel scope link src 10.10.10.1 quickack 1 With the receiver ACKing immediately, Nagle on the sender only waits one RTT instead of the ~40 ms delayed-ACK timer. Result (same VM, same direction, same data) stock quickack 1 on receiver per VDI 132 s / 132 s 50 s / 46 s throughput ~59 MiB/s ~155-169 MiB/s whole migration 4:50 2:00 dom0 on the receiver went up to ~1.1 vCPU total, still nothing saturated. Caveats One run per arm, one pool. I did not capture ss -ti during the runs, so the notsent:512 signature is inferred, not observed here. It does not address the reply gating @TeddyAstie described, it only removes the delayed-ACK half. Consistent with that, we land at ~2.7x, below the ~4.3x reported above for TCP_NODELAY (different rig, so not directly comparable). Not persistent: lost on reboot and when xapi re-plugs the PIF. We re-apply it from a small systemd unit at boot as a stopgap and will remove it once vhd-tool ships TCP_NODELAY. Scope is only connections routed via the migration network (storage and live migration). More ACK packets on that link, nothing else changes. Is there a PR or issue for the TCP_NODELAY patch in xapi-project/xen-api that we can follow? I could not find one.
  • Our hyperconverged storage solution

    54 Topics
    824 Posts
    J
    @poddingue Turns out it is all secondary-to-secondary out of sync. So that makes it less intense, though if primary dies I wonder how it will resolve this, or if it will become split brained. Still unsure of how it happened, but the ones I manually cleaned up have not come back. Going to continue to manually clear them up. If it happens again I will have an alert setup to notify me, and I have all the xcp-ng logs and everything to be able to see what happened. If that happens I will post here with details and logs for the resource so we can see how it occurs. [image: image.jpeg]
  • 37 Topics
    136 Posts
    J
    @AtaxyaNetwork Merci pour tes recherches ! Oui "cd_label" serait cool comme ajout au plugin ce qui permet sur les distro type Fedora/Redhat de ne pas avoir de boot_command à gérer