@olivierlambert
I think I found the root cause of the slow Storage Migration performance I reported above.
After tracing sparse_dd, the issue appears to be an interaction between Nagle's algorithm and TCP delayed ACKs on the NBD connection.
With strace, I found that sparse_dd sends the NBD request header and payload using separate write() calls:
write(fd, <NBD header>, 28) = 28
write(fd, <data>, 2097152) = 2097152
read(fd, <NBD reply>, 16) = 16
Normally this is fast, but there are also periodic 512-byte requests:
write(fd, <NBD header>, 28) = 28
write(fd, <data>, 512) = 512
~40 ms delay
read(fd, <NBD reply>, 16) = 16
During these stalls, ss -tinp showed:
ato:40
unacked:1
notsent:512
This seems to produce the following sequence:
- The 28-byte NBD header is sent.
- It remains unacknowledged.
- The following 512-byte payload is queued.
- Nagle's algorithm prevents that small payload from being sent while the previous data is unacknowledged.
- The peer's delayed ACK timer expires after approximately 40 ms.
- The ACK arrives and the 512-byte payload is finally transmitted.
I also checked a packet capture. The destination iSCSI write only occurred after this delay, so the storage itself was not causing the 40 ms stall.
To verify this before modifying the package, I attached GDB to the running sparse_dd process and enabled TCP_NODELAY with setsockopt() on its existing TCP socket.
The effect was immediate: the notsent:512 stalls disappeared and migration throughput increased substantially.
I then patched ocaml/vhd-tool/src/impl.ml.
The current code is:
let socket sockaddr =
let family =
match sockaddr with
| Lwt_unix.ADDR_INET (addr, port) ->
Unix.domain_of_sockaddr (Lwt_unix.ADDR_INET (addr, port))
| Lwt_unix.ADDR_UNIX _ ->
Unix.PF_UNIX
in
Lwt_unix.socket family Unix.SOCK_STREAM 0
I changed it to enable TCP_NODELAY for TCP sockets:
--- a/ocaml/vhd-tool/src/impl.ml
+++ b/ocaml/vhd-tool/src/impl.ml
@@
- Lwt_unix.socket family Unix.SOCK_STREAM 0
+ let sock = Lwt_unix.socket family Unix.SOCK_STREAM 0 in
+ ( match sockaddr with
+ | Lwt_unix.ADDR_INET _ ->
+ Lwt_unix.setsockopt sock Unix.TCP_NODELAY true
+ | Lwt_unix.ADDR_UNIX _ ->
+ ()
+ ) ;
+ sock
I rebuilt vhd-tool for XCP-ng 8.3 and tested Storage Migration again.
Before the patch, I was consistently seeing only around:
30-40 MB/s
After enabling TCP_NODELAY, I am seeing roughly:
150-300 MB/s
depending on storage activity.
For example, during one test:
eth2: ~174 MB/s
eth3: ~174 MB/s
lo: ~320 MB/s
and there were also physical-interface peaks around 300 MB/s.
After the patch, ss still shows ato:40, which is expected because delayed ACK is still enabled on the peer:
rtt:0.059/0.017 ato:40 ... unacked:1
but the important difference is that the persistent:
notsent:512
is gone, so the delayed ACK timer no longer stalls the NBD payload.
I also checked the upstream xen-api v26.1.16 source, and the socket creation code still does not enable TCP_NODELAY.
So I believe this explains the ~30-40 MB/s limitation I was seeing with sparse_dd NBD Storage Migration.
Would it make sense to enable TCP_NODELAY for the ADDR_INET socket in vhd-tool upstream?