• Migrating an offline VM disk between two local SRs is slow

    Unsolved Xen Orchestra
    35
    1
    0 Votes
    35 Posts
    10k Views
    A
    Data point from a 2-host XCP-ng 8.3 pool that supports @tosh's diagnosis, plus a no-patch workaround that gave us ~2.7x until the TCP_NODELAY fix lands. Setup XCP-ng 8.3, xapi-core / vhd-tool 26.1.16-1.2.xcpng8.3, dom0 kernel 4.19.0+1 Two hosts, local ext SRs on NVMe (Samsung 970 EVO Plus / WD Red SN700) Dedicated migration network: 2.5 GbE (Intel i226-V, igc), MTU 9000, no errors Offline storage migration (halted VM, VM.migrate_send) of 2 VDIs (disk + snapshot), 16 GiB virtual / 7773 MiB allocated each; sparse_dd ... -prezeroed ... -dest-proto nbd over TLS Baseline (stock) 132 s per VDI, i.e. ~59 MiB/s on the wire, flat line, same in both directions Link at ~20 % of line rate, source SR read latency 0.3 ms, dom0 total <= 0.5 vCPU, no single dom0 vCPU above 0.08 (RRD, 60 s averages) Sending host with turbo or capped at 2.5 GHz gave the same 132 s, so not CPU-bound here Workaround: disable delayed ACKs on the receiver, only for the migration-network route # on the destination host; adapt prefix/dev/src to the output of: ip route show <migration-net> ip route change 10.10.10.0/24 dev xenbr1 proto kernel scope link src 10.10.10.1 quickack 1 With the receiver ACKing immediately, Nagle on the sender only waits one RTT instead of the ~40 ms delayed-ACK timer. Result (same VM, same direction, same data) stock quickack 1 on receiver per VDI 132 s / 132 s 50 s / 46 s throughput ~59 MiB/s ~155-169 MiB/s whole migration 4:50 2:00 dom0 on the receiver went up to ~1.1 vCPU total, still nothing saturated. Caveats One run per arm, one pool. I did not capture ss -ti during the runs, so the notsent:512 signature is inferred, not observed here. It does not address the reply gating @TeddyAstie described, it only removes the delayed-ACK half. Consistent with that, we land at ~2.7x, below the ~4.3x reported above for TCP_NODELAY (different rig, so not directly comparable). Not persistent: lost on reboot and when xapi re-plugs the PIF. We re-apply it from a small systemd unit at boot as a stopgap and will remove it once vhd-tool ships TCP_NODELAY. Scope is only connections routed via the migration network (storage and live migration). More ACK packets on that link, nothing else changes. Is there a PR or issue for the TCP_NODELAY patch in xapi-project/xen-api that we can follow? I could not find one.
  • XCP-ng Shutdown Hangs When an SR Is Unavailable

    Unsolved XCP-ng
    3
    0 Votes
    3 Posts
    104 Views
    poddingueP
    This came up in https://xcp-ng.org/forum/topic/11728 last winter, with the same kind of setup, an ISO SR served by a VM on the host. Olivier's answer there was to eject the CDs from the VMs, unplug the ISO SR's PBD, and only then shut the host down, using xe vm-cd-eject --multiple and xe pbd-unplug uuid=<PBD UUID> for the first two steps (the PBD one is in the CLI reference: https://docs.xcp-ng.org/appendix/cli_reference#pbd-unplug). @mickwilli then added the unplug to his UPS shutdown script, and his host shut down cleanly after that. I haven't tried it with an SMB share myself, so I can't promise it behaves like his NFS one, but your NUT script on the Pi looks like the natural place for those two steps, just before the shutdown command.
  • Cleanup orphans backups

    Unsolved Backup
    10
    0 Votes
    10 Posts
    593 Views
    P
    @johnnezero You can use the docs directly in /rest/v0/docs. You just have to connect and fill in the uuid you need. It will work
  • Remote desktop on Gnome hangs randomly

    Unsolved Hardware
    24
    0 Votes
    24 Posts
    3k Views
    O
    @dinhngtu i've manage to fixed it. Your build adds some extra entries in /var/lib/xcp/state.db that aren't in the version from testing, that broke things when i downgraded the packages. ipv4_dns VIF.ipv6_dns
  • Kioxia CM7 PCIe pass-through crash

    Unsolved Compute
    12
    0 Votes
    12 Posts
    2k Views
    D
    Hello, this issue has been fixed in XAPI 26.1.19-1.1 available in the testing updates. For now you have to configure your VMs manually (https://docs.xcp-ng.org/troubleshooting/common-problems/#pci-passthrough-error-unsupported-msi-delivery-mode-7) but the fix will be activated by default in the future.
  • Internet connectivity - Check XOA failed.

    Unsolved Management
    8
    1 Votes
    8 Posts
    279 Views
    Z
    @poddingue said: Thanks for posting the fix. The same thing came up in an older thread, https://xcp-ng.org/forum/topic/9957, where disabling IPv6 also made xoa check go green. @HamiltonWDS explained a likely reason in https://xcp-ng.org/forum/post/87831: Node tries the addresses with a short timeout and can end up on the IPv6 one. Your curl -6 test helps a lot here, because it shows IPv6 doesn't connect at all on that network, so it's more than a slow path. Turning it off in XOA seems reasonable if you don't use IPv6 there; if you do, I'd guess the router side is where it really needs fixing, though I could be wrong (still haven't migrated to IPv6 myself ). Curious whether it sorts out the Cloud Backup too. XOA is using our management network where ipv6 is not needed and not enabled. Looks like our cloud backup finally ran for the first time in while too.[image: cloudbackup.jpg]
  • [V2V] Without VDDK problems

    Solved Migrate to XCP-ng
    4
    0 Votes
    4 Posts
    145 Views
    J
    @mpiton Thanks for a fantasticly quick answer, and detailed as well. I will test this immediately! Update: I can confirm that this was indeed the problem!
  • XO Tasks - backups just pilling up

    Unsolved Management
    5
    1
    0 Votes
    5 Posts
    236 Views
    poddingueP
    I didn't know, so I went and looked at the code. XO prunes that list in a cleanup step: backup runs are kept for 31 days, and other tasks are kept by count, the newest 1000 (https://github.com/vatesfr/xen-orchestra/blob/b59c8f3d20954aa1d72e92c1e8f2db635d74fcba/@xen-orchestra/mixins/Tasks.mjs#L119-L127). You can change both with tasks.gc.keep and tasks.gc.backupKeepDuration in the xo-server config. As far as I can tell the cleanup runs when xo-server starts, not on a timer, so a long-running XO can collect quite a few before the next restart. My guess for the Cloud Backup entries @acebmxer still sees after 30 days is that they're kept by count rather than by age.
  • Unclear errors

    Unsolved Management
    4
    1
    0 Votes
    4 Posts
    149 Views
    poddingueP
    I tried it on a lab host (XCP-ng 8.3, xapi 26.1.16): calling VM.pool_migrate on a halted VM gives me your error, VM_BAD_POWER_STATE with running, halted, so the two values are the state XAPI wanted and the one it found. On the XO side, a migration inside the same pool with no SR or network mapping goes straight to VM.pool_migrate without checking the power state first: https://github.com/vatesfr/xen-orchestra/blob/b59c8f3d20954aa1d72e92c1e8f2db635d74fcba/packages/xo-server/src/xapi/index.mjs#L942-L981. So a stopped VM in that batch would give this, which fits what you remember. I haven't tried the batch migrate from the UI itself though, so I don't know if the UI is supposed to filter halted VMs out before that. It's a good concrete example for the feedback post anyway, XO could just say "this VM is halted" instead of passing XAPI's raw code through.
  • PCIe Pass-through lanes and lane performance

    Unsolved Compute
    57
    0 Votes
    57 Posts
    10k Views
    pandusenP
    @stormi I am glad to hear there is progress, and I never doubted that there was interest and will from the Dev team, it was just stated, by Oliver himself, that the focus was mostly on the classic hypervisor features and that resources for edge cases were scarce. The numbers might not be spot on, but it was something similar.. Anyways I know you care and that the team always take interest in home-lab scenarios. And that, is why I keep recommending XCP-ng. You are good and motivated people, you care and interact with the community and you have a product that is simple in structure and free of most of the technical bureaucracy that the other ecosystems have. And yes, this particular Intel Arc issue, is not only on XCP-ng, The intel toolbox and drivers for this platform is currently a mess for Linux. And you are right to take home-labbers under your wing and prioritize it. As I'm sure you know, although it might not be of direct importance for revenue, they do have a huge impact on popularity. I dont think Proxmox would have had the traction it has, without them Thank you for replying, and thank you to the team for their hard work!
  • 0 Votes
    20 Posts
    2k Views
    M
    @anthoineb, Thanks for the fix! We plan to install blktap-3.55.5-11.1.xcpng8.3 on all three hypervisors tomorrow. We will fully shut down and start the data VMs one at a time, then verify that all running tapdisk processes use the updated binary. As the stalls are intermittent, we will monitor for recurrence and report back with the results.
  • XOA 6.9 Update

    Xen Orchestra
    9
    1 Votes
    9 Posts
    532 Views
    S
    @MathieuRA We reverted back to 6.8.2 for now. We can wait for patch, otherwise, I have opened a tunnel #38081.
  • 1 Votes
    3 Posts
    144 Views
    D
    @trelane Hello, the bug affects Xen drivers of any 9.0 series (considering the upstream code, it goes back to circa September 2019). 9.2.385 contains the fix, which was merged upstream just this month. The bug also impacts network connections, however to a lesser extent due to its nature. I believe that the bug is also present in XenServer VM Tools 9.6.0 and by extension, older versions of it. We recommend that you switch to the XCP-ng tools whenever possible, also due to its other benefits (fully open source, validated for XCP-ng and optimized for Xen Orchestra).
  • XCP-ng Windows PV tools announcements

    Moved News
    116
    0 Votes
    116 Posts
    45k Views
    olivierlambertO
    Great, thanks for the feedback @yomeyo !
  • HA causes reboot of xcp-ng nodes

    Unsolved Management
    32
    0 Votes
    32 Posts
    2k Views
    tjkreidlT
    @john.c Indeed, John, and I almost forgot that the backup network on each host was actually on a separate, isolated NIC, and not at all on the VLAN.
  • Enable Maintenance Mode = Host Not Enough Memory

    Unsolved XCP-ng
    5
    0 Votes
    5 Posts
    215 Views
    J
    @poddingue Our HA-pool is a 3 host system, yes. I'll add a note about our plans to go paid in that feedback-item. Thanks for the tip. In the meantime, I'm testing and reporting as much as I can. In order to hopefully help the product be better for all. Cheers!
  • [dedicated thread] Dell Open Manage Appliance (OME)

    Solved Compute
    104
    1
    0 Votes
    104 Posts
    65k Views
    S
    @ataxyanetwork hey, thanks for sharing this. I am new to the forum so maybe i am not looking in the right place. Your Blog seems to be offline, so did you link the Appliance here in the forum or is it only available on the blog thx
  • DRBD reactor metrics in k8s

    XOSTOR
    7
    3
    0 Votes
    7 Posts
    455 Views
    J
    @poddingue Turns out it is all secondary-to-secondary out of sync. So that makes it less intense, though if primary dies I wonder how it will resolve this, or if it will become split brained. Still unsure of how it happened, but the ones I manually cleaned up have not come back. Going to continue to manually clear them up. If it happens again I will have an alert setup to notify me, and I have all the xcp-ng logs and everything to be able to see what happened. If that happens I will post here with details and logs for the resource so we can see how it occurs. [image: image.jpeg]
  • Slow VM migration on Linstor SR

    Unsolved XOSTOR
    4
    0 Votes
    4 Posts
    260 Views
    K
    @ronan-a For testing purposes, I did the following: I created a VM on the local SR of a host, installed Linux on it, and configured an ISCSI target. I then added that ISCSI target as an ISCSI SR within xcp-ng pool, and created a small VM on that repository, using XO. While the machine was running, I performed a migration without losing a single ping.I assume this confirms my previous assertion regarding the use of a standard storage system.
  • Piraeus operator 2.12.0 does not work on current xostor version

    Unsolved XOSTOR
    4
    3
    0 Votes
    4 Posts
    188 Views
    poddingueP
    Thanks, @Jonathon !