XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    HA causes reboot of xcp-ng nodes

    Scheduled Pinned Locked Moved Unsolved Management
    22 Posts 4 Posters 472 Views 3 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • J
      john.c
      last edited by john.c

      @tjkreidl When you’re back online my posts above, may reveal something about behaviours you noticed in Citrix XenServer and Citrix Hypervisor, during your employment there as CTP.

      @poddingue This is a potential source for a joint patch to Xen Hypervisor kernel to fix this as part of upstreaming, something that will benefit from multiple eyes and hands working on it.

      tjkreidlT 1 Reply Last reply
      Reply Quote 0
      • tjkreidlT
        tjkreidl Ambassador @john.c
        last edited by

        @john.c Interesting information about the strong v. weak host model.
        Active-active and LACP bonds will split the network traffic between the NICs, while an active-passive bond only uses the primary. WHat type of bonds do you have set up?

        J 1 Reply Last reply
        Reply Quote 0
        • J
          john.c @tjkreidl
          last edited by john.c

          @tjkreidl said:

          @john.c Interesting information about the strong v. weak host model.
          Active-active and LACP bonds will split the network traffic between the NICs, while an active-passive bond only uses the primary. WHat type of bonds do you have set up?

          @tjkreidl It is either a balance-alb (Mode 6) or a balance-xor (Mode 2) bond topology on this specific deployment (as the upstream router has limited port capacity and does not support LACP).

          However, the brilliant part about this architectural issue is that the underlying Layer 3 routing vulnerability remains identical regardless of which of these two modes is active. Both modes open up multiple physical paths for outbound traffic under a single logical bond, giving the kernel's default weak host model the perfect opportunity to misroute packets under load.

          Here is how the weak host model breaks both configurations:

          • If it is balance-alb (Mode 6): The bonding driver actively performs Layer 2 ARP and MAC manipulation to balance paths without switch assistance. The weak host model completely undermines this logic because the kernel treats the IP address as globally accessible to the whole host. Under load, the kernel's Layer 3 routing engine completely ignores the bonding driver's intended pathing boundaries, picks a "cheaper" path via global metrics, and leaks the packet out of a completely separate infrastructure interface (like Management or Storage).
          • If it is balance-xor (Mode 2): The driver relies on a strict hash policy (like layer2 or layer2+3) to statically map traffic to a destination across a specific physical NIC slave. Yet, if a heavy background process (like a backup job or storage replication) alters local routing table costs or causes transient congestion, the kernel’s Layer 3 logic overrides that static Layer 2 pathing—spilling packets out of an unrelated physical port.

          Ultimately, OpenMetrics packet-flow telemetry caught this exact moment of divergence: the Layer 3 stack bypassed the logical bond boundary entirely. The packet exited on an unintended infrastructure port carrying the wrong source IP, where adjacent switches or firewalls dropped it as unroutable. To XAPI and the HA daemon, the heartbeat was instantly lost on that specific port, triggering the self-fencing reboot loop.

          This is why it's a universal vulnerability across multi-path bonding modes on a multi-homed system, and why an upstream patch enforcing a strong host model for the Infrastructure Plane is the cleanest solution.

          tjkreidlT 1 Reply Last reply
          Reply Quote 0
          • J
            john.c
            last edited by john.c

            @AtaxyaNetwork Depending on your current or past homelab topography, this behavior might ring a few bells for you as well.

            Homelab and prosumer environments are highly susceptible to this exact weak host routing leak. Because we often multiplex distinct infrastructure planes (Management, dedicated storage networks, and backup backplanes) across multi-port NIC bonds—frequently using balance-alb or balance-xor because the upstream switches lack stacked enterprise LACP support—the conditions are perfect for a routing collision.

            If a heavy data operation (like a massive VM migration or a backup sync) alters the local metric weightings or causes micro-congestion, the default weak host model can silently push Dom0 host-terminated packets onto the wrong physical interface segment. It's a classic hidden variable that can cause erratic connection drops or unexplainable HA timeouts on otherwise perfectly configured hardware.

            1 Reply Last reply
            Reply Quote 0
            • tjkreidlT
              tjkreidl Ambassador @john.c
              last edited by

              @john.c Most interesting, and I fully agree, that a strong host model enforcement policy is really the best and only recourse for such a topology, or so it would seem.
              The only other option that comes to mind would be to not put the heartbeat connection on any sort of bond or multipath. That's, of course, not ideal.

              J 1 Reply Last reply
              Reply Quote 0
              • J
                john.c @tjkreidl
                last edited by john.c

                @tjkreidl said:

                @john.c Most interesting, and I fully agree, that a strong host model enforcement policy is really the best and only recourse for such a topology, or so it would seem.
                The only other option that comes to mind would be to not put the heartbeat connection on any sort of bond or multipath. That's, of course, not ideal.

                @tjkreidl Exactly, and you have hit on the exact architectural compromise we've all been forced to make for years.

                Moving the heartbeat connection off a bond/multipath and onto a dedicated, single physical NIC does reduce the Layer 3 path alternatives that trigger the weak host leak. However, as you rightly pointed out, it's highly non-ideal. By removing the bond, we introduce a single point of failure (SFP/cable/switch port) directly into the critical HA backplane. We shouldn't have to sacrifice physical hardware redundancy just to keep the kernel's Layer 3 routing engine from misbehaving.

                This is precisely why enforcing a strong host model for the Infrastructure Plane is the true architectural solution. It allows administrators to safely use balance-alb, balance-xor, or multipathing for maximum hardware resilience, while ensuring that Dom0-terminated traffic strictly honors its designated interface boundaries regardless of global metrics.

                Since we are in full agreement on the root cause and the ideal fix, this looks like a prime candidate for a strategic architectural shift. Hopefully, @TeddyAstie and the @Team-Hypervisor-Kernel can look at how we can implement this structural protection by default—perhaps utilizing targeted sysctl overrides (arp_ignore, arp_announce, rp_filter) or interface socket-binding for host-terminated infrastructure networks.

                1 Reply Last reply
                Reply Quote 1
                • C
                  carloum70 @john.c
                  last edited by

                  @john.c now I remember, during the installation we used the eth4 as the management interface.

                  [16:54 dacshyp002 ~]# cat /etc/firstboot.d/data/management.conf 
                  LABEL='eth4'
                  MODE='static'
                  IP='172.28.4.11'
                  NETMASK='255.255.252.0'
                  GATEWAY='172.28.4.1'
                  MODEV6='none'
                  DNS='8.8.8.8'
                  

                  Afterwords we created a bond bond0 with eth2/eth3 and used this as the management interface.
                  The xe pif-list command shows the following:

                  device ( RO)                   : bond0
                                 management ( RO): true
                      IP-configuration-mode ( RO): Static
                                         IP ( RO): 172.28.4.11
                  
                  device ( RO)                   : eth4
                                 management ( RO): false
                      IP-configuration-mode ( RO): Static
                                         IP ( RO): 172.28.4.11
                  

                  Is it that easy to remove the IP address from the eth4 interface? Because it make no sense to have duplicate ip-addresses.
                  I think we've had duplicate IP addresses all along, but we only noticed the issue after enabling HA.
                  By the way what is the correct way to change the management interface? We used the xsconsole.

                  J 1 Reply Last reply
                  Reply Quote 0
                  • J
                    john.c @carloum70
                    last edited by john.c

                    @carloum70 said:

                    @john.c now I remember, during the installation we used the eth4 as the management interface.

                    [16:54 dacshyp002 ~]# cat /etc/firstboot.d/data/management.conf 
                    LABEL='eth4'
                    MODE='static'
                    IP='172.28.4.11'
                    NETMASK='255.255.252.0'
                    GATEWAY='172.28.4.1'
                    MODEV6='none'
                    DNS='8.8.8.8'
                    

                    Afterwords we created a bond bond0 with eth2/eth3 and used this as the management interface.
                    The xe pif-list command shows the following:

                    device ( RO)                   : bond0
                                   management ( RO): true
                        IP-configuration-mode ( RO): Static
                                           IP ( RO): 172.28.4.11
                    
                    device ( RO)                   : eth4
                                   management ( RO): false
                        IP-configuration-mode ( RO): Static
                                           IP ( RO): 172.28.4.11
                    

                    Is it that easy to remove the IP address from the eth4 interface? Because it make no sense to have duplicate ip-addresses.
                    I think we've had duplicate IP addresses all along, but we only noticed the issue after enabling HA.
                    By the way what is the correct way to change the management interface? We used the xsconsole.

                    If you’re using a bonded NIC what mode is used? If not bonding it will allow for multiple NICs, via bond0-12 etc to show a single NIC. But only use once my above fix has landed, to avoid a return to this issue, unless forcing strong host yourself before hand.

                    Remove the IPs with the following process:-

                    1. xe pif-list params=uuid,device,IP,management
                    2. xe pif-reconfigure-ip uuid=<PIF_UUID> mode=none
                    3. xe-toolstack-restart
                    1 Reply Last reply
                    Reply Quote 0
                    • C
                      carloum70
                      last edited by

                      @john.c We are using a LCAP bond.
                      For my understanding, if I remove the IP settings from eth4, the problem will still not be solved because of the known issue with the bond that you explained in “Weak Host Model & HA Heartbeat Failures.”
                      (At the moment we have HA disabled.)

                      J 1 Reply Last reply
                      Reply Quote 0
                      • J
                        john.c @carloum70
                        last edited by

                        @carloum70 said:

                        @john.c We are using a LCAP bond.
                        For my understanding, if I remove the IP settings from eth4, the problem will still not be solved because of the known issue with the bond that you explained in “Weak Host Model & HA Heartbeat Failures.”
                        (At the moment we have HA disabled.)

                        What use is eth4 put to is it still management, if so put it into bond0? In one of my earlier posts on this I also detailed how to in the meantime use sysctl to force strong host, but best to wait for the official update from Vates.

                        1 Reply Last reply
                        Reply Quote 0

                        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                        With your input, this post could be even better 💗

                        Register Login
                        • First post
                          Last post