XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    HA causes reboot of xcp-ng nodes

    Scheduled Pinned Locked Moved Unsolved Management
    31 Posts 6 Posters 750 Views 5 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • J
      john.c @carloum70
      last edited by

      @carloum70 said:

      @john.c We are using a LCAP bond.
      For my understanding, if I remove the IP settings from eth4, the problem will still not be solved because of the known issue with the bond that you explained in “Weak Host Model & HA Heartbeat Failures.”
      (At the moment we have HA disabled.)

      What use is eth4 put to is it still management, if so put it into bond0? In one of my earlier posts on this I also detailed how to in the meantime use sysctl to force strong host, but best to wait for the official update from Vates.

      1 Reply Last reply
      Reply Quote 0
      • bvitnikB
        bvitnik
        last edited by bvitnik

        @carloum70 I can't really help you with anything specific but I must make a note that when analyzing network setup you must also take into account the Open vSwitch. Looking at the classic Linux networking, bridging and routing facilities is not enough to get you the full picture.

        For example, if I remember correctly, Active-Active bonding (non LACP) is implemented by Open vSwitch with so called balance-slb mode. This mode is not available with classic Linux kernel bonding options. What you are seeing (duplicated IPs) is maybe perfectly normal for this kind of setup with Open vSwitch.

        I would search for the cause of the issue somewhere else like network congestion. Heart beats are very sensitive to network congestion and can easily fail if you do not dedicate network links for this kind of traffic. If you are sharing the links over which heart beats are sent with some high traffic stuff (storage maybe?), it can easily make problems.

        UPDATE: unfortunately I don't have access to these kind of bonded setups any more so I can't check if the configuration differs in any way from yours. My home lab thingies are single network interface only.

        C J 2 Replies Last reply
        Reply Quote 0
        • C
          carloum70 @bvitnik
          last edited by

          @bvitnik I think there is some confusing.
          During the initial installation we have have used the eth4 (172.28.4.11) as the management interface. Later on we created a bond with eth2 and eth3 and changed the management interface to this bond using the xsconsole, and also used the same ip-address (172.28.4.11). We thought that the IP address had also been removed from eth4 during the reconfiguration, but we didn't check it afterwards. It turns out that this was not the case.
          So now I removed the ip from eth4 using the xe pif-reconfigure-ip command.

          Next step is to use the sysctl workaround @john.c mentioned in an earlier post and enable HA again.

          (I also think it would be wise to use dedicated bond for the heartbeat.)

          1 Reply Last reply
          Reply Quote 0
          • J
            john.c @bvitnik
            last edited by john.c

            @bvitnik said:

            @carloum70 I can't really help you with anything specific but I must make a note that when analyzing network setup you must also take into account the Open vSwitch. Looking at the classic Linux networking, bridging and routing facilities is not enough to get you the full picture.

            For example, if I remember correctly, Active-Active bonding (non LACP) is implemented by Open vSwitch with so called balance-slb mode. This mode is not available with classic Linux kernel bonding options. What you are seeing (duplicated IPs) is maybe perfectly normal for this kind of setup with Open vSwitch.

            I would search for the cause of the issue somewhere else like network congestion. Heart beats are very sensitive to network congestion and can easily fail if you do not dedicate network links for this kind of traffic. If you are sharing the links over which heart beats are sent with some high traffic stuff (storage maybe?), it can easily make problems.

            UPDATE: unfortunately I don't have access to these kind of bonded setups any more so I can't check if the configuration differs in any way from yours. My home lab thingies are single network interface only.

            The Open vSwitch would have been already taken into account with packet data flow analysis. Though if balance-slb is Open vSwitch exclusive analysis will be needed, but if influenced by kernel then it will also likely be affected.

            1 Reply Last reply
            Reply Quote 0
            • stormiS
              stormi Vates 🪐 XCP-ng Team
              last edited by

              CCing @Team-XAPI-Network and @Team-Hypervisor-Kernel for them to have a look at this strong vs weak host model and the impact on HA.

              1 Reply Last reply
              Reply Quote 0
              • C
                carloum70
                last edited by

                Some update:

                After removing the ip-address from eth4 the following messages disappeared

                 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=30408.
                [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
                

                By the way I forgot to mention that each time I enable HA the following warning appears in xha.log

                Sep 29 13:31:56 CEST 2026 [warn] BM: cannot open bonding status file (/proc/net/bonding/bond0). (2)
                Sep 29 13:31:56 CEST 2026 [info] BM: this is not bonded. Terminating bonding thread.
                

                bond0 (eth2 + eth3) is the management interface and also used for the HA.

                tjkreidlT J 2 Replies Last reply
                Reply Quote 0
                • tjkreidlT
                  tjkreidl Ambassador @carloum70
                  last edited by

                  @carloum70 One option to address the HA error messages: Delete the HA configuration and re-configure it from scratch once the bonds are all configured OK, which it appears you think they now are.

                  1 Reply Last reply
                  Reply Quote 0
                  • J
                    john.c @carloum70
                    last edited by john.c

                    @carloum70 said:

                    Some update:

                    After removing the ip-address from eth4 the following messages disappeared

                     [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=30408.
                    [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
                    

                    By the way I forgot to mention that each time I enable HA the following warning appears in xha.log

                    Sep 29 13:31:56 CEST 2026 [warn] BM: cannot open bonding status file (/proc/net/bonding/bond0). (2)
                    Sep 29 13:31:56 CEST 2026 [info] BM: this is not bonded. Terminating bonding thread.
                    

                    bond0 (eth2 + eth3) is the management interface and also used for the HA.

                    Warning: Unconfigured Backup/Migration Networks Trigger HA Reboots

                    Watch out if a dedicated backup or migration network is not configured in Vates VMS, all of that traffic defaults to the management network.

                    When massive backup jobs or VM migrations saturate the management interface, it chokes the High Availability (HA) cluster heartbeats. This network congestion causes packet loss, triggers false split-brain conditions, and forces the host into a spontaneous reboot (self-fencing). This will especially occur on highly busy and congested instances of Vates VMS.

                    While sysctl ARP filtering patches the routing leak, it cannot fix physical network saturation.

                    The Long-Term Fix: Quad-Port Ethernet Upgrades

                    I highly recommend upgrading all server LAN cards (especially XCP-ng hosts) to multiple quad-port modules across multiple slots or daughter cards. Abundant physical interfaces allow you to build dedicated LACP or Active-Backup bonds, ensuring complete infrastructure isolation:

                    • Dedicated HA & Management: Keeps critical XAPI orchestration and heartbeat checks isolated from heavy data bursts.
                    • Dedicated Storage: Keeps iSCSI, NFS, or XOSTOR disk I/O running smoothly on its own low-latency pipe.
                    • Dedicated VM Traffic: Isolates production guest network traffic from infrastructure management tasks.
                    • Dedicated Backup & Migration: Physical ports can be carved out exclusively for VM motions and backup windows—or at least provide a dedicated, shared channel away from HA traffic.

                    tjkreidlT 1 Reply Last reply
                    Reply Quote 0
                    • tjkreidlT
                      tjkreidl Ambassador @john.c
                      last edited by

                      @john.c Keeping the various network traffic isolated according to specific usage (management, storage, VMs, etc.) is always a good idea. The last system I managed had 10GiB LACP bonds and using VLANs to isolate traffic and that worked fine with a four-node pool running around 80 or so XenDesktop instances per node. Never experienced any congestion issues. Each dom0 had a ton of memory and I believe it was either 8 or 16 VCPUs to make sure there were sufficient compute and memory allocations to allow for sometimes very heavy loads. It also helped that I eventually added GPUs to take on some of the computational load, in particular when some of the VMs were running applications employing heavy graphics.

                      J 1 Reply Last reply
                      Reply Quote 0
                      • J
                        john.c @tjkreidl
                        last edited by john.c

                        @tjkreidl said:

                        @john.c Keeping the various network traffic isolated according to specific usage (management, storage, VMs, etc.) is always a good idea. The last system I managed had 10GiB LACP bonds and using VLANs to isolate traffic and that worked fine with a four-node pool running around 80 or so XenDesktop instances per node. Never experienced any congestion issues. Each dom0 had a ton of memory and I believe it was either 8 or 16 VCPUs to make sure there were sufficient compute and memory allocations to allow for sometimes very heavy loads. It also helped that I eventually added GPUs to take on some of the computational load, in particular when some of the VMs were running applications employing heavy graphics.

                        @tjkreidl Thanks Tobias, that’s great validation! Running 10GiB LACP bonds with proper VLAN isolation is definitely the ultimate goal for production stability, especially when pushing 80+ VMs per node.

                        Your point about dom0 resource allocation is also huge—people often forget that saturated vCPUs and starved dom0 memory can bottleneck network processing just as fast as a saturated physical link under heavy loads. Giving dom0 the extra compute headroom ensures the orchestration layer doesn't drop packets when backups or migrations scale up across that many instances.

                        1 Reply Last reply
                        Reply Quote 0

                        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                        With your input, this post could be even better 💗

                        Register Login
                        • First post
                          Last post