Categories

  • All news regarding Xen and XCP-ng ecosystem

    145 Topics
    5k Posts
    TeddyAstieT
    @bvitnik Various reasons : In some niche cases, it can cause compatibility issues with guests. Viridian exposes additionnal attack surface e.g VSA-2025-006 and VSA-2026-031 and given that this is generally unused for non-Windows guests, there is no good reason to expose it to the guest For Windows guests, use of the dedicated Windows template is strongly encouraged, as it brings additional viridian flags and other tunings for better compatibility and performance.
  • Everything related to the virtualization platform

    1k Topics
    15k Posts
    dthenotD
    @racom.jiristerba Sorry, I missed answering your questions the first time Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage? Yes, it's a known issue, most of them will be fixed with the latest update, you might need to activate the VDI again but a migration during RPU will be enough Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open? It's because in the case of VHD, it's the case but with QCOW2 we don't stop the tapdisk process accessing the VDI, it's a change that was missed Is the qcow2OLD_<UUID> naming seen in rollback expected, compared with QCOW2-OLD_<UUID> seen during successful cleanup? It's a pre-existing bug that I already have on my TODO list Is there a supported update, hotfix, or workaround for this configuration? Installing the latest release sm-3.2.12-25.1 during updates (and maybe launching a xe sr-scan uuid=<SR UUI> after updating it so it auto-resolve the undo) should be enough What additional logs are required to identify the original failure of the second QCOW2 VM? I don't think we need any more logs since the errors I'm seeing should already be fixed. If you have any more issues after installing the newest packages, I will take another look What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups? The sr-scan after updating should do it automatically, it shouldn't need any manipulation. Looking at the storage logs in /var/log/SMlog for any irregularities could help to see problems.
  • 3k Topics
    29k Posts
    A
    Data point from a 2-host XCP-ng 8.3 pool that supports @tosh's diagnosis, plus a no-patch workaround that gave us ~2.7x until the TCP_NODELAY fix lands. Setup XCP-ng 8.3, xapi-core / vhd-tool 26.1.16-1.2.xcpng8.3, dom0 kernel 4.19.0+1 Two hosts, local ext SRs on NVMe (Samsung 970 EVO Plus / WD Red SN700) Dedicated migration network: 2.5 GbE (Intel i226-V, igc), MTU 9000, no errors Offline storage migration (halted VM, VM.migrate_send) of 2 VDIs (disk + snapshot), 16 GiB virtual / 7773 MiB allocated each; sparse_dd ... -prezeroed ... -dest-proto nbd over TLS Baseline (stock) 132 s per VDI, i.e. ~59 MiB/s on the wire, flat line, same in both directions Link at ~20 % of line rate, source SR read latency 0.3 ms, dom0 total <= 0.5 vCPU, no single dom0 vCPU above 0.08 (RRD, 60 s averages) Sending host with turbo or capped at 2.5 GHz gave the same 132 s, so not CPU-bound here Workaround: disable delayed ACKs on the receiver, only for the migration-network route # on the destination host; adapt prefix/dev/src to the output of: ip route show <migration-net> ip route change 10.10.10.0/24 dev xenbr1 proto kernel scope link src 10.10.10.1 quickack 1 With the receiver ACKing immediately, Nagle on the sender only waits one RTT instead of the ~40 ms delayed-ACK timer. Result (same VM, same direction, same data) stock quickack 1 on receiver per VDI 132 s / 132 s 50 s / 46 s throughput ~59 MiB/s ~155-169 MiB/s whole migration 4:50 2:00 dom0 on the receiver went up to ~1.1 vCPU total, still nothing saturated. Caveats One run per arm, one pool. I did not capture ss -ti during the runs, so the notsent:512 signature is inferred, not observed here. It does not address the reply gating @TeddyAstie described, it only removes the delayed-ACK half. Consistent with that, we land at ~2.7x, below the ~4.3x reported above for TCP_NODELAY (different rig, so not directly comparable). Not persistent: lost on reboot and when xapi re-plugs the PIF. We re-apply it from a small systemd unit at boot as a stopgap and will remove it once vhd-tool ships TCP_NODELAY. Scope is only connections routed via the migration network (storage and live migration). More ACK packets on that link, nothing else changes. Is there a PR or issue for the TCP_NODELAY patch in xapi-project/xen-api that we can follow? I could not find one.
  • Our hyperconverged storage solution

    54 Topics
    824 Posts
    J
    @poddingue Turns out it is all secondary-to-secondary out of sync. So that makes it less intense, though if primary dies I wonder how it will resolve this, or if it will become split brained. Still unsure of how it happened, but the ones I manually cleaned up have not come back. Going to continue to manually clear them up. If it happens again I will have an alert setup to notify me, and I have all the xcp-ng logs and everything to be able to see what happened. If that happens I will post here with details and logs for the resource so we can see how it occurs. [image: image.jpeg]
  • 37 Topics
    136 Posts
    J
    @AtaxyaNetwork Merci pour tes recherches ! Oui "cd_label" serait cool comme ajout au plugin ce qui permet sur les distro type Fedora/Redhat de ne pas avoir de boot_command à gérer