XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    HA causes reboot of xcp-ng nodes

    Scheduled Pinned Locked Moved Unsolved Management
    27 Posts 6 Posters 612 Views 5 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • tjkreidlT
      tjkreidl Ambassador @john.c
      last edited by

      @john.c Most interesting, and I fully agree, that a strong host model enforcement policy is really the best and only recourse for such a topology, or so it would seem.
      The only other option that comes to mind would be to not put the heartbeat connection on any sort of bond or multipath. That's, of course, not ideal.

      J 1 Reply Last reply
      Reply Quote 0
      • J
        john.c @tjkreidl
        last edited by john.c

        @tjkreidl said:

        @john.c Most interesting, and I fully agree, that a strong host model enforcement policy is really the best and only recourse for such a topology, or so it would seem.
        The only other option that comes to mind would be to not put the heartbeat connection on any sort of bond or multipath. That's, of course, not ideal.

        @tjkreidl Exactly, and you have hit on the exact architectural compromise we've all been forced to make for years.

        Moving the heartbeat connection off a bond/multipath and onto a dedicated, single physical NIC does reduce the Layer 3 path alternatives that trigger the weak host leak. However, as you rightly pointed out, it's highly non-ideal. By removing the bond, we introduce a single point of failure (SFP/cable/switch port) directly into the critical HA backplane. We shouldn't have to sacrifice physical hardware redundancy just to keep the kernel's Layer 3 routing engine from misbehaving.

        This is precisely why enforcing a strong host model for the Infrastructure Plane is the true architectural solution. It allows administrators to safely use balance-alb, balance-xor, or multipathing for maximum hardware resilience, while ensuring that Dom0-terminated traffic strictly honors its designated interface boundaries regardless of global metrics.

        Since we are in full agreement on the root cause and the ideal fix, this looks like a prime candidate for a strategic architectural shift. Hopefully, @TeddyAstie and the @Team-Hypervisor-Kernel can look at how we can implement this structural protection by default—perhaps utilizing targeted sysctl overrides (arp_ignore, arp_announce, rp_filter) or interface socket-binding for host-terminated infrastructure networks.

        1 Reply Last reply
        Reply Quote 1
        • C
          carloum70 @john.c
          last edited by

          @john.c now I remember, during the installation we used the eth4 as the management interface.

          [16:54 dacshyp002 ~]# cat /etc/firstboot.d/data/management.conf 
          LABEL='eth4'
          MODE='static'
          IP='172.28.4.11'
          NETMASK='255.255.252.0'
          GATEWAY='172.28.4.1'
          MODEV6='none'
          DNS='8.8.8.8'
          

          Afterwords we created a bond bond0 with eth2/eth3 and used this as the management interface.
          The xe pif-list command shows the following:

          device ( RO)                   : bond0
                         management ( RO): true
              IP-configuration-mode ( RO): Static
                                 IP ( RO): 172.28.4.11
          
          device ( RO)                   : eth4
                         management ( RO): false
              IP-configuration-mode ( RO): Static
                                 IP ( RO): 172.28.4.11
          

          Is it that easy to remove the IP address from the eth4 interface? Because it make no sense to have duplicate ip-addresses.
          I think we've had duplicate IP addresses all along, but we only noticed the issue after enabling HA.
          By the way what is the correct way to change the management interface? We used the xsconsole.

          J 1 Reply Last reply
          Reply Quote 0
          • J
            john.c @carloum70
            last edited by john.c

            @carloum70 said:

            @john.c now I remember, during the installation we used the eth4 as the management interface.

            [16:54 dacshyp002 ~]# cat /etc/firstboot.d/data/management.conf 
            LABEL='eth4'
            MODE='static'
            IP='172.28.4.11'
            NETMASK='255.255.252.0'
            GATEWAY='172.28.4.1'
            MODEV6='none'
            DNS='8.8.8.8'
            

            Afterwords we created a bond bond0 with eth2/eth3 and used this as the management interface.
            The xe pif-list command shows the following:

            device ( RO)                   : bond0
                           management ( RO): true
                IP-configuration-mode ( RO): Static
                                   IP ( RO): 172.28.4.11
            
            device ( RO)                   : eth4
                           management ( RO): false
                IP-configuration-mode ( RO): Static
                                   IP ( RO): 172.28.4.11
            

            Is it that easy to remove the IP address from the eth4 interface? Because it make no sense to have duplicate ip-addresses.
            I think we've had duplicate IP addresses all along, but we only noticed the issue after enabling HA.
            By the way what is the correct way to change the management interface? We used the xsconsole.

            If you’re using a bonded NIC what mode is used? If not bonding it will allow for multiple NICs, via bond0-12 etc to show a single NIC. But only use once my above fix has landed, to avoid a return to this issue, unless forcing strong host yourself before hand.

            Remove the IPs with the following process:-

            1. xe pif-list params=uuid,device,IP,management
            2. xe pif-reconfigure-ip uuid=<PIF_UUID> mode=none
            3. xe-toolstack-restart
            1 Reply Last reply
            Reply Quote 0
            • C
              carloum70
              last edited by

              @john.c We are using a LCAP bond.
              For my understanding, if I remove the IP settings from eth4, the problem will still not be solved because of the known issue with the bond that you explained in “Weak Host Model & HA Heartbeat Failures.”
              (At the moment we have HA disabled.)

              J 1 Reply Last reply
              Reply Quote 0
              • J
                john.c @carloum70
                last edited by

                @carloum70 said:

                @john.c We are using a LCAP bond.
                For my understanding, if I remove the IP settings from eth4, the problem will still not be solved because of the known issue with the bond that you explained in “Weak Host Model & HA Heartbeat Failures.”
                (At the moment we have HA disabled.)

                What use is eth4 put to is it still management, if so put it into bond0? In one of my earlier posts on this I also detailed how to in the meantime use sysctl to force strong host, but best to wait for the official update from Vates.

                1 Reply Last reply
                Reply Quote 0
                • bvitnikB
                  bvitnik
                  last edited by bvitnik

                  @carloum70 I can't really help you with anything specific but I must make a note that when analyzing network setup you must also take into account the Open vSwitch. Looking at the classic Linux networking, bridging and routing facilities is not enough to get you the full picture.

                  For example, if I remember correctly, Active-Active bonding (non LACP) is implemented by Open vSwitch with so called balance-slb mode. This mode is not available with classic Linux kernel bonding options. What you are seeing (duplicated IPs) is maybe perfectly normal for this kind of setup with Open vSwitch.

                  I would search for the cause of the issue somewhere else like network congestion. Heart beats are very sensitive to network congestion and can easily fail if you do not dedicate network links for this kind of traffic. If you are sharing the links over which heart beats are sent with some high traffic stuff (storage maybe?), it can easily make problems.

                  UPDATE: unfortunately I don't have access to these kind of bonded setups any more so I can't check if the configuration differs in any way from yours. My home lab thingies are single network interface only.

                  C J 2 Replies Last reply
                  Reply Quote 0
                  • C
                    carloum70 @bvitnik
                    last edited by

                    @bvitnik I think there is some confusing.
                    During the initial installation we have have used the eth4 (172.28.4.11) as the management interface. Later on we created a bond with eth2 and eth3 and changed the management interface to this bond using the xsconsole, and also used the same ip-address (172.28.4.11). We thought that the IP address had also been removed from eth4 during the reconfiguration, but we didn't check it afterwards. It turns out that this was not the case.
                    So now I removed the ip from eth4 using the xe pif-reconfigure-ip command.

                    Next step is to use the sysctl workaround @john.c mentioned in an earlier post and enable HA again.

                    (I also think it would be wise to use dedicated bond for the heartbeat.)

                    1 Reply Last reply
                    Reply Quote 0
                    • J
                      john.c @bvitnik
                      last edited by john.c

                      @bvitnik said:

                      @carloum70 I can't really help you with anything specific but I must make a note that when analyzing network setup you must also take into account the Open vSwitch. Looking at the classic Linux networking, bridging and routing facilities is not enough to get you the full picture.

                      For example, if I remember correctly, Active-Active bonding (non LACP) is implemented by Open vSwitch with so called balance-slb mode. This mode is not available with classic Linux kernel bonding options. What you are seeing (duplicated IPs) is maybe perfectly normal for this kind of setup with Open vSwitch.

                      I would search for the cause of the issue somewhere else like network congestion. Heart beats are very sensitive to network congestion and can easily fail if you do not dedicate network links for this kind of traffic. If you are sharing the links over which heart beats are sent with some high traffic stuff (storage maybe?), it can easily make problems.

                      UPDATE: unfortunately I don't have access to these kind of bonded setups any more so I can't check if the configuration differs in any way from yours. My home lab thingies are single network interface only.

                      The Open vSwitch would have been already taken into account with packet data flow analysis. Though if balance-slb is Open vSwitch exclusive analysis will be needed, but if influenced by kernel then it will also likely be affected.

                      1 Reply Last reply
                      Reply Quote 0
                      • stormiS
                        stormi Vates 🪐 XCP-ng Team
                        last edited by

                        CCing @Team-XAPI-Network and @Team-Hypervisor-Kernel for them to have a look at this strong vs weak host model and the impact on HA.

                        1 Reply Last reply
                        Reply Quote 0
                        • C
                          carloum70
                          last edited by

                          Some update:

                          After removing the ip-address from eth4 the following messages disappeared

                           [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=30408.
                          [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
                          

                          By the way I forgot to mention that each time I enable HA the following warning appears in xha.log

                          Sep 29 13:31:56 CEST 2026 [warn] BM: cannot open bonding status file (/proc/net/bonding/bond0). (2)
                          Sep 29 13:31:56 CEST 2026 [info] BM: this is not bonded. Terminating bonding thread.
                          

                          bond0 (eth2 + eth3) is the management interface and also used for the HA.

                          1 Reply Last reply
                          Reply Quote 0

                          Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                          Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                          With your input, this post could be even better 💗

                          Register Login
                          • First post
                            Last post