XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    Migrating an offline VM disk between two local SRs is slow

    Scheduled Pinned Locked Moved Xen Orchestra
    24 Posts 7 Posters 7.4k Views 6 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • olivierlambertO Offline
      olivierlambert Vates 🪐 Co-Founder CEO
      last edited by olivierlambert

      Can you describe exactly the steps that were done so we can double check/compare and understand the why?

      edit: also, are you comparing a live migration vs an offline copy? It's very different, since in live you have to replicate the blocks while the VM on top is running.

      pkgwP 1 Reply Last reply Reply Quote 0
      • pkgwP Offline
        pkgw @olivierlambert
        last edited by

        @olivierlambert This is all offline. Unfortunately I can't describe exactly what was done, since someone else was doing the work and they were trying a bunch of different things all in a row. I suspect that the apparently fast migration is a red herring (maybe a previous attempt left a copy of the disk on the destination SR, and the system noticed that and avoided the actual I/O?) but if there turned out to be a magical fast path, I wouldn't complain!

        1 Reply Last reply Reply Quote 0
        • olivierlambertO Offline
          olivierlambert Vates 🪐 Co-Founder CEO
          last edited by

          You can also try warm migration, which can go a lot faster.

          ForzaF 1 Reply Last reply Reply Quote 1
          • ForzaF Offline
            Forza @olivierlambert
            last edited by Forza

            Using XOA "Disaster Recovery" backup method can be a lot faster than normal offline migration.

            One time I did it, it took approx 10 minutes instead of 2 hours...

            M 1 Reply Last reply Reply Quote 0
            • M Offline
              magicker @Forza
              last edited by

              I think I am seeing a similar issue. Raid1 NVME copy to raid 10 4x2tb HDD on same host

              a 300gb transfer is estimated at 7 hours. (11% done in 50 mins)

              the vm is live.

              according to the stats almost nothing is happening on this server or the 2 storage

              1 Reply Last reply Reply Quote 1
              • D Offline
                Davidj 0 @olivierlambert
                last edited by

                @olivierlambert
                Is the CPU on the sending host or the receiving host the limiting factor for single disk migrations?

                M 1 Reply Last reply Reply Quote 0
                • olivierlambertO Offline
                  olivierlambert Vates 🪐 Co-Founder CEO
                  last edited by

                  I can't really tell, gut feeling is the sending host, but I have no numbers to confirm.

                  1 Reply Last reply Reply Quote 0
                  • M Offline
                    magicker @Davidj 0
                    last edited by magicker

                    @Davidj-0 in my case there CPU activity is minimal. I think something is wrong with the software raid 10 setup. On an identical setup warm migration between to the raid 10 array between hosts is showing horrible iowait similar to the sr to sr transfer on the other host

                    bb40264a-7678-4906-aca5-788bc217c7f4-image.png

                    1 Reply Last reply Reply Quote 0
                    • olivierlambertO Offline
                      olivierlambert Vates 🪐 Co-Founder CEO
                      last edited by

                      Maybe the IO scheduler is not the right one?

                      T 1 Reply Last reply Reply Quote 0
                      • T Online
                        tosh @olivierlambert
                        last edited by

                        @olivierlambert
                        I think I found the root cause of the slow Storage Migration performance I reported above.

                        After tracing sparse_dd, the issue appears to be an interaction between Nagle's algorithm and TCP delayed ACKs on the NBD connection.

                        With strace, I found that sparse_dd sends the NBD request header and payload using separate write() calls:

                        write(fd, <NBD header>, 28) = 28
                        write(fd, <data>, 2097152) = 2097152
                        read(fd, <NBD reply>, 16) = 16
                        

                        Normally this is fast, but there are also periodic 512-byte requests:

                        write(fd, <NBD header>, 28) = 28
                        write(fd, <data>, 512) = 512
                        
                        ~40 ms delay
                        
                        read(fd, <NBD reply>, 16) = 16
                        

                        During these stalls, ss -tinp showed:

                        ato:40
                        unacked:1
                        notsent:512
                        

                        This seems to produce the following sequence:

                        1. The 28-byte NBD header is sent.
                        2. It remains unacknowledged.
                        3. The following 512-byte payload is queued.
                        4. Nagle's algorithm prevents that small payload from being sent while the previous data is unacknowledged.
                        5. The peer's delayed ACK timer expires after approximately 40 ms.
                        6. The ACK arrives and the 512-byte payload is finally transmitted.

                        I also checked a packet capture. The destination iSCSI write only occurred after this delay, so the storage itself was not causing the 40 ms stall.

                        To verify this before modifying the package, I attached GDB to the running sparse_dd process and enabled TCP_NODELAY with setsockopt() on its existing TCP socket.

                        The effect was immediate: the notsent:512 stalls disappeared and migration throughput increased substantially.

                        I then patched ocaml/vhd-tool/src/impl.ml.

                        The current code is:

                        let socket sockaddr =
                          let family =
                            match sockaddr with
                            | Lwt_unix.ADDR_INET (addr, port) ->
                                Unix.domain_of_sockaddr (Lwt_unix.ADDR_INET (addr, port))
                            | Lwt_unix.ADDR_UNIX _ ->
                                Unix.PF_UNIX
                          in
                          Lwt_unix.socket family Unix.SOCK_STREAM 0
                        

                        I changed it to enable TCP_NODELAY for TCP sockets:

                        --- a/ocaml/vhd-tool/src/impl.ml
                        +++ b/ocaml/vhd-tool/src/impl.ml
                        @@
                        -  Lwt_unix.socket family Unix.SOCK_STREAM 0
                        +  let sock = Lwt_unix.socket family Unix.SOCK_STREAM 0 in
                        +  ( match sockaddr with
                        +  | Lwt_unix.ADDR_INET _ ->
                        +      Lwt_unix.setsockopt sock Unix.TCP_NODELAY true
                        +  | Lwt_unix.ADDR_UNIX _ ->
                        +      ()
                        +  ) ;
                        +  sock
                        

                        I rebuilt vhd-tool for XCP-ng 8.3 and tested Storage Migration again.

                        Before the patch, I was consistently seeing only around:

                        30-40 MB/s
                        

                        After enabling TCP_NODELAY, I am seeing roughly:

                        150-300 MB/s
                        

                        depending on storage activity.

                        For example, during one test:

                        eth2: ~174 MB/s
                        eth3: ~174 MB/s
                        lo:   ~320 MB/s
                        

                        and there were also physical-interface peaks around 300 MB/s.

                        After the patch, ss still shows ato:40, which is expected because delayed ACK is still enabled on the peer:

                        rtt:0.059/0.017 ato:40 ... unacked:1
                        

                        but the important difference is that the persistent:

                        notsent:512
                        

                        is gone, so the delayed ACK timer no longer stalls the NBD payload.

                        I also checked the upstream xen-api v26.1.16 source, and the socket creation code still does not enable TCP_NODELAY.

                        So I believe this explains the ~30-40 MB/s limitation I was seeing with sparse_dd NBD Storage Migration.

                        Would it make sense to enable TCP_NODELAY for the ADDR_INET socket in vhd-tool upstream?

                        1 Reply Last reply Reply Quote 0
                        • olivierlambertO Offline
                          olivierlambert Vates 🪐 Co-Founder CEO
                          last edited by

                          Worth mentioning @Team-Storage

                          1 Reply Last reply Reply Quote 0

                          Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                          Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                          With your input, this post could be even better 💗

                          Register Login
                          • First post
                            Last post