XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    • Profile
    • Following 0
    • Followers 0
    • Topics 0
    • Posts 1
    • Groups 0
    A Online
    1. Home
    2. AndiD
    3. Posts

    Posts

    Recent Best Controversial
    • RE: Migrating an offline VM disk between two local SRs is slow

      Data point from a 2-host XCP-ng 8.3 pool that supports @tosh's diagnosis, plus a no-patch workaround that gave us ~2.7x until the TCP_NODELAY fix lands.

      Setup

      • XCP-ng 8.3, xapi-core / vhd-tool 26.1.16-1.2.xcpng8.3, dom0 kernel 4.19.0+1
      • Two hosts, local ext SRs on NVMe (Samsung 970 EVO Plus / WD Red SN700)
      • Dedicated migration network: 2.5 GbE (Intel i226-V, igc), MTU 9000, no errors
      • Offline storage migration (halted VM, VM.migrate_send) of 2 VDIs (disk + snapshot), 16 GiB virtual / 7773 MiB allocated each; sparse_dd ... -prezeroed ... -dest-proto nbd over TLS

      Baseline (stock)

      • 132 s per VDI, i.e. ~59 MiB/s on the wire, flat line, same in both directions
      • Link at ~20 % of line rate, source SR read latency 0.3 ms, dom0 total <= 0.5 vCPU, no single dom0 vCPU above 0.08 (RRD, 60 s averages)
      • Sending host with turbo or capped at 2.5 GHz gave the same 132 s, so not CPU-bound here

      Workaround: disable delayed ACKs on the receiver, only for the migration-network route

      # on the destination host; adapt prefix/dev/src to the output of: ip route show <migration-net>
      ip route change 10.10.10.0/24 dev xenbr1 proto kernel scope link src 10.10.10.1 quickack 1
      

      With the receiver ACKing immediately, Nagle on the sender only waits one RTT instead of the ~40 ms delayed-ACK timer.

      Result (same VM, same direction, same data)

      stock quickack 1 on receiver
      per VDI 132 s / 132 s 50 s / 46 s
      throughput ~59 MiB/s ~155-169 MiB/s
      whole migration 4:50 2:00

      dom0 on the receiver went up to ~1.1 vCPU total, still nothing saturated.

      Caveats

      • One run per arm, one pool. I did not capture ss -ti during the runs, so the notsent:512 signature is inferred, not observed here.
      • It does not address the reply gating @TeddyAstie described, it only removes the delayed-ACK half. Consistent with that, we land at ~2.7x, below the ~4.3x reported above for TCP_NODELAY (different rig, so not directly comparable).
      • Not persistent: lost on reboot and when xapi re-plugs the PIF. We re-apply it from a small systemd unit at boot as a stopgap and will remove it once vhd-tool ships TCP_NODELAY.
      • Scope is only connections routed via the migration network (storage and live migration). More ACK packets on that link, nothing else changes.

      Is there a PR or issue for the TCP_NODELAY patch in xapi-project/xen-api that we can follow? I could not find one.

      posted in Xen Orchestra
      A
      AndiD