XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    HA causes reboot of xcp-ng nodes

    Scheduled Pinned Locked Moved Unsolved Management
    22 Posts 4 Posters 475 Views 3 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • C
      carloum70
      last edited by carloum70

      Hi all,

      I need some help regarding HA.
      I have a 3 node cluster running xcp-ng 8.3 (002 is the master)

      [13:55 dacshyp002 ~]# xe host-list
      uuid ( RO)                : 6b99f1ab-6f4c-4a8d-b766-d16a8a942bdf
                name-label ( RW): dacshyp003
          name-description ( RW): Default install
      
      
      uuid ( RO)                : d99e150c-079a-4092-8909-ad1a36e07dec
                name-label ( RW): dacshyp001
          name-description ( RW): Default install
      
      
      uuid ( RO)                : 41a2a448-a5dc-44c6-be44-c07540d75c60
                name-label ( RW): dacshyp002
          name-description ( RW): Default install
      

      Last Thursday I enabled HA on the pool and also on some om the VM's. On Friday I also did an "rolling pool update"
      Today I had a "spontaneous" reboot of the dacshyp001 and dacshyp003.
      Boot time of dacshyp001:

      [Mon Sep 21 11:21:31 2026] Linux version 4.19.0+1 (mockbuild@1efe98ffea2245bcb4c11893890153ac) (gcc version 4.8.5 20150623 (Red Hat 4.8.5-28) (GCC)) #1 SMP Thu Aug 13 13:06:11 UTC 2026
      

      Around that time I also see the following message in the xensource.log:

      Sep 21 11:21:08 dacshyp002 xapi: [debug||20398 ha_monitor|HA monitor D:7b5068fa6bfc|xapi_ha_vm_failover] Setting host dacshyp001 to dead
      Sep 21 11:21:28 dacshyp002 xapi: [debug||20398 ha_monitor|HA monitor D:7b5068fa6bfc|xapi_ha_vm_failover] Setting host dacshyp001 to dead
      Sep 21 11:21:48 dacshyp002 xapi: [debug||20398 ha_monitor|HA monitor D:7b5068fa6bfc|xapi_ha_vm_failover] Setting host dacshyp001 to dead
      

      I will attached all logging around this time.
      I am trying to understand what triggered the reboot and how I can troubleshoot this.

      I am running XO community commit 6a441 . I am aware this is not the latest version. Before upgrading I want to know the cause of the reboot.

      Also ignore the error:

      /var/lib/xcp/xapi|dispatch:VDI.get_by_uuid D:1f5b4c8d329e|backtrace] VDI.get_by_uuid D:9df5c0255224 failed with exception Db_exn.Read_missing_uuid("VDI", "", "9219113f-65e4-4368-b42d-1b8bdc7614a9")
      

      I will create another ticket for this.
      HA-problem.txt
      Thanks in advance.
      Carlo

      poddingueP tjkreidlT 2 Replies Last reply
      Reply Quote 0
      • poddingueP
        poddingue Vates 🪐 @carloum70
        last edited by

        From what I read in the docs (HA isn't my strong suit), a host in an HA pool that loses its heartbeat in certain ways is designed to reboot itself, which they call self-fencing, so this may well be HA doing its job rather than something crashing. 🤔

        The log you attached is from the master and starts at 11:21:00, and dacshyp001 is already marked as not live in the very first liveset at 11:21:08, so I think whatever triggered it happened just before that and isn't in this file. 🤷

        The three Setting host dacshyp001 to dead lines look to me like one event being re-checked every 20 seconds while 001 was still coming back, though I could be misreading that.

        I also suspect 003 went down at a different moment, because it still shows as alive in that same liveset. There's a doc section for this case, https://docs.xcp-ng.org/troubleshooting/troubleshooting-ha#my-host-rebooted-why-did-it-reboot, which points at /var/log/xha.log on the host that rebooted.

        Could you post that file from dacshyp001 and dacshyp003 for the few minutes before each reboot, plus 003's boot time?

        1 Reply Last reply
        Reply Quote 0
        • C
          carloum70
          last edited by carloum70

          dacshyp003 booted at

          [Mon Sep 21 05:44:37 2026] Linux version 4.19.0+1 (mockbuild@1efe98ffea2245bcb4c11893890153ac) (gcc version 4.8.5 20150623 (Red Hat 4.8.5-28) (GCC)) #1 SMP Thu Aug 13 13:06:11 UTC 2026
          

          This is the part of the xensource logging of the dacshyp001 before the reboot:

          Sep 21 11:17:52 dacshyp001 xapi: [debug||17 ha_monitor|HA monitor D:58d1715e0726|xapi_ha] The node we think is the master is still alive and marked as master; this is OK
          ^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@
          

          Same behavior on dacshyp003.
          I also attached the xha.log from both nodes.

          xha_log_001.txt xha_log_003.txt

          1 Reply Last reply
          Reply Quote 1
          • C
            carloum70
            last edited by carloum70

            I also checked the xha.log on the dacshyp002:

            Sep 21 11:39:49 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=30408.
            Sep 21 11:40:09 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 11:40:29 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=31466.
            Sep 21 11:40:50 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 11:46:30 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=35178.
            Sep 21 11:46:50 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 11:47:10 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=33233.
            Sep 21 11:47:30 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 11:55:51 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=40853.
            Sep 21 11:56:11 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 11:56:31 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=32914.
            Sep 21 11:56:51 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 11:59:31 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=33377.
            Sep 21 11:59:51 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:01:51 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=38588.
            Sep 21 12:02:12 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:10:12 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=38125.
            Sep 21 12:10:12 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=38323.
            Sep 21 12:10:32 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:19:33 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=43934.
            Sep 21 12:19:53 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:20:13 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=32989.
            Sep 21 12:20:33 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:21:13 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=42077.
            Sep 21 12:21:33 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:21:53 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=31325.
            Sep 21 12:22:13 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:25:54 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=34497.
            Sep 21 12:26:14 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:26:34 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=32558.
            Sep 21 12:26:54 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:27:14 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=30616.
            Sep 21 12:27:34 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:31:14 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=33975.
            Sep 21 12:31:34 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:38:15 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=34797.
            Sep 21 12:38:35 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:53:56 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=37000.
            Sep 21 12:53:56 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=40176.
            Sep 21 12:54:16 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:56:16 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=36380.
            Sep 21 12:56:36 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            Sep 21 12:56:57 CEST 2026 [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[2].time_since_last_hb=31439.
            Sep 21 12:57:17 CEST 2026 [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
            

            I can conclude the heartbeat communication is generally unstable in my HA pool. Tomorrow I check the logging of our switch/router.

            1 Reply Last reply
            Reply Quote 0
            • poddingueP poddingue marked this topic as a question
            • tjkreidlT
              tjkreidl Ambassador @carloum70
              last edited by tjkreidl

              @carloum70 Are all your hosts properly time sychronized to NTP or chronyc? Check each host for offsets. They need to be really close in time with each other. And what HA heartbeat client setup are you using? As long as you have a quorum, HA should continue to work fine. WHen you get down to two hosts and one cannot communicate with the other is when things get critical (unless running HA-Lizard on a two-host pool).

              1 Reply Last reply
              Reply Quote 0
              • C
                carloum70
                last edited by

                I did some further investigation and I think there is something wrong with the ip-settings of the management interface.

                At the xsconsole I see the following:

                Management Network Parameters    
                
                Device        bond0         
                address       172.28.4.11    
                Netmask       255.255.252.0   
                Gateway       172.28.4.1
                

                So it's using the bond0 interface.

                # xe pif-list host-uuid=41a2a448-a5dc-44c6-be44-c07540d75c60
                
                uuid ( RO)                  : 0761d887-268e-5cf0-401b-a08ad7c419aa
                                device ( RO): bond0
                                   MAC ( RO): 30:3e:a7:1d:b0:90
                    currently-attached ( RO): true
                                  VLAN ( RO): -1
                          network-uuid ( RO): 28d51793-0226-cf3f-75d3-ed7c5a9d8b33
                             host-uuid ( RO): 41a2a448-a5dc-44c6-be44-c07540d75c60
                

                This is part of the output of xe bond-list

                # xe bond-list
                uuid ( RO)      : 79c525c3-22f9-04ce-4a80-21d18ef65ebc
                    master ( RO): 0761d887-268e-5cf0-401b-a08ad7c419aa
                    slaves ( RO): 7ac18886-e0a5-057b-bac0-36b909588dba ; 2f18fc9a-8c1c-dfce-9416-fa8c657c5f63 
                

                I already checked that:

                7ac18886-e0a5-057b-bac0-36b909588dba --> eth3
                2f18fc9a-8c1c-dfce-9416-fa8c657c5f63 --> eth2
                

                And now the confusing part. This is part of the output of the ip a command:

                2: eth2: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq master ovs-system state UP group default qlen 1000
                    link/ether 30:3e:a7:1d:b0:90 brd ff:ff:ff:ff:ff:ff
                3: eth3: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq master ovs-system state UP group default qlen 1000
                    link/ether 30:3e:a7:1d:b0:91 brd ff:ff:ff:ff:ff:ff
                4: eth4: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq master ovs-system state UP group default qlen 1000
                    link/ether 30:3e:a7:1d:b0:92 brd ff:ff:ff:ff:ff:ff
                
                
                13: xenbr4: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000
                    link/ether 30:3e:a7:1d:b0:92 brd ff:ff:ff:ff:ff:ff
                    inet 172.28.4.11/22 brd 172.28.7.255 scope global xenbr4
                       valid_lft forever preferred_lft forever
                

                As you can see the xenbr4 has also the management ip-address and has the same mac-address as the eth4 interface.
                OK check the following commands:

                [17:21 dacshyp002 ~]# ip addr show xapi1
                15: xapi1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000
                    link/ether 30:3e:a7:1d:b0:90 brd ff:ff:ff:ff:ff:ff
                    inet 172.28.4.11/22 brd 172.28.7.255 scope global xapi1
                       valid_lft forever preferred_lft forever
                [17:21 dacshyp002 ~]# ip addr show xenbr4
                13: xenbr4: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000
                    link/ether 30:3e:a7:1d:b0:92 brd ff:ff:ff:ff:ff:ff
                    inet 172.28.4.11/22 brd 172.28.7.255 scope global xenbr4
                       valid_lft forever preferred_lft forever
                

                And If I look at the routing table:

                # ip route
                default via 172.28.4.1 dev xapi1 
                172.18.8.0/23 dev xenbr0 proto kernel scope link src 172.18.8.11 
                172.18.10.0/23 dev xenbr7 proto kernel scope link src 172.18.10.11 
                172.28.4.0/22 dev xenbr4 proto kernel scope link src 172.28.4.11 
                172.28.4.0/22 dev xapi1 proto kernel scope link src 172.28.4.11 
                

                According to me the node has 2 interfaces xenbr4 / xapi1 with the same ip-address and also 2 different routes for the same ip-range 172.28.4.0/22.
                And that's why the heartbeat has some issues. Can someone confirm this ?
                By the way I really don't know how this happened.

                tjkreidlT 1 Reply Last reply
                Reply Quote 0
                • tjkreidlT
                  tjkreidl Ambassador @carloum70
                  last edited by

                  @carloum70 If you are willing to try this, you can reset the primary management interface (PMI) on a bond, WIthin the bash shell:
                  :
                  xe pif-list
                  This will show the PIFs and their current network associations.

                  Identify the bond master PIF
                  If you have a bonded network, the master PIF is the one that represents the bond. You can list it with:

                  xe pif-list network-uuid=<bond-uuid>
                  The master PIF is the one that will be used as the primary interface for the bond.

                  Reset the management interface to the bond master
                  Use the xe-reset-networking command with the --reset-primary option:

                  xe-reset-networking --reset-primary
                  This will move the management interface to the bond master PIF, which is the correct way to reassign it when the bond is created or reconfigured.

                  1 Reply Last reply
                  Reply Quote 0
                  • C
                    carloum70
                    last edited by carloum70

                    @tjkreidl

                    According to xsconsole the management interface is bond0

                    Current Management Interface          
                                                                  
                    Device           bond0                
                    MAC Address	 30:3e:a7:1d:b0:90    
                    DHCP/Static IP   Static               
                    IP address       172.28.4.11          
                    Netmask          255.255.252.0        
                    Gateway          172.28.4.1           
                    Hostname         dacshyp002  
                    

                    According the output of xe pif-list

                    uuid ( RO)                  : 0761d887-268e-5cf0-401b-a08ad7c419aa
                                    device ( RO): bond0
                                       MAC ( RO): 30:3e:a7:1d:b0:90
                        currently-attached ( RO): true
                                      VLAN ( RO): -1
                              network-uuid ( RO): 28d51793-0226-cf3f-75d3-ed7c5a9d8b33
                                 host-uuid ( RO): 41a2a448-a5dc-44c6-be44-c07540d75c60
                    

                    Let's check the network

                    # xe network-list uuid=28d51793-0226-cf3f-75d3-ed7c5a9d8b33
                    uuid ( RO)                : 28d51793-0226-cf3f-75d3-ed7c5a9d8b33
                              name-label ( RW): mgmt-bond
                        name-description ( RW): 
                                  bridge ( RO): xapi1
                    

                    At linux level:

                    # ip addr show xapi1
                    15: xapi1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000
                        link/ether 30:3e:a7:1d:b0:90 brd ff:ff:ff:ff:ff:ff
                        inet 172.28.4.11/22 brd 172.28.7.255 scope global xapi1
                           valid_lft forever preferred_lft forever
                    

                    xapi1 is a bond of et2 and eth3. So far so good.
                    But if I am running the xe-reset-networking command

                    Your network will be re-configured as follows:
                    
                    Management interface:   eth4
                    Reset interface name rules: Yes
                    IP configuration mode:  dhcp
                    IPv6 configuration mode:none
                    
                    If you want to change any of the above settings, type 'no' and re-run
                    the command with appropriate arguments (use --help for a list of options).
                    
                    Type 'yes' to continue.
                    Type 'no' to cancel.
                    no
                    

                    So for some reason xcp-ng thinks the eth4 is management interface, which also has the ip-address 172.28.4.11/22 .
                    By the way --reset-primary is not a valid option:

                    # xe-reset-networking --reset-primary
                    Usage: xe-reset-networking [options]
                    
                    xe-reset-networking: error: no such option: --reset-primary
                    
                    tjkreidlT poddingueP 2 Replies Last reply
                    Reply Quote 0
                    • tjkreidlT
                      tjkreidl Ambassador @carloum70
                      last edited by

                      @carloum70 Sorry, that option was deprecated.

                      You may need to be more specific in the reset, something like the following with is just an example:
                      xe-reset-networking -m IP_of_Master --device=eth0 --mode=static --ip=192.168.1.50 --netmask=255.255.255.0 --gateway=192.168.1.1 --dns=192.168.1.254

                      Check the options and see which ones you actually need in your case.
                      Worst case, you could possibly dissolve the bond, redo the PMI and then re-create the bond. WHy eth4 shows up is hard to guess. I have seen before that on some hosts in a pool that the NIC order was not the same, even though the hardware and OS versions were identical. In that case, you have to shut down the NICs and use interface-rename to reassign specific NIC names to the corresponding MAC addresses.

                      1 Reply Last reply
                      Reply Quote 0
                      • poddingueP
                        poddingue Vates 🪐 @carloum70
                        last edited by

                        @carloum70 the same address on both xenbr4 and xapi1 looks like the best lead so far. 🤷
                        I haven't tested it, but two interfaces answering for 172.28.4.11, with two routes to the same subnet, could send heartbeat traffic out the wrong NIC now and then, and that would fit the drops in your xha.log. bleader untangled something close to it in https://xcp-ng.org/forum/topic/12472 (two bridges on one subnet, replies leaving by the wrong one).

                        As for eth4: without --device, xe-reset-networking takes the NIC recorded at install time in /etc/firstboot.d/data/management.conf, so I think it's just remembering the installer's choice.

                        I'd hold off on @tjkreidl's worst-case reset for now, because the docs say it wipes all PIF, bond and VLAN config, force-stops VMs, and isn't supported while HA is on (Emergency Network Reset). xe pif-list host-name-label=dacshyp002 params=device,IP-configuration-mode,IP,management only lists, and it would show whether XAPI itself thinks eth4 has that IP; if it does, @Team-XAPI-Network will know the clean way to drop it. 🤞

                        J 1 Reply Last reply
                        Reply Quote 0
                        • J
                          john.c @poddingue
                          last edited by john.c

                          @poddingue said:

                          @carloum70 the same address on both xenbr4 and xapi1 looks like the best lead so far. 🤷
                          I haven't tested it, but two interfaces answering for 172.28.4.11, with two routes to the same subnet, could send heartbeat traffic out the wrong NIC now and then, and that would fit the drops in your xha.log. bleader untangled something close to it in https://xcp-ng.org/forum/topic/12472 (two bridges on one subnet, replies leaving by the wrong one).

                          As for eth4: without --device, xe-reset-networking takes the NIC recorded at install time in /etc/firstboot.d/data/management.conf, so I think it's just remembering the installer's choice.

                          I'd hold off on @tjkreidl's worst-case reset for now, because the docs say it wipes all PIF, bond and VLAN config, force-stops VMs, and isn't supported while HA is on (Emergency Network Reset). xe pif-list host-name-label=dacshyp002 params=device,IP-configuration-mode,IP,management only lists, and it would show whether XAPI itself thinks eth4 has that IP; if it does, @Team-XAPI-Network will know the clean way to drop it. 🤞

                          @Team-Hypervisor-Kernel May also wish to have a look, as part of the problem can stem from the kernel’s weak host model, as it’s not just networking in XAPI but also kernel, due to asymmetric routing. This is down to packets being routed on the lowest cost path.

                          C 1 Reply Last reply
                          Reply Quote 1
                          • J
                            john.c
                            last edited by john.c

                            Technical Summary: Weak Host Model & HA Heartbeat Failures

                            The Problem: Asymmetric Routing & The Weak Host Model

                            Linux defaults to a weak host model for IPv4 networking. The kernel treats IP addresses as belonging to the entire host, rather than being strictly bound to a specific physical or logical network interface (NIC).

                            When a multi-homed cluster node transmits High Availability (HA) heartbeats or storage synchronization packets, the kernel evaluates its local routing table. If multiple paths exist to the destination subnet, the kernel will choose the egress interface based on the lowest routing metric (cost).

                            This causes two major issues in HA clusters:

                            1. Asymmetric Outbound Traffic: The kernel may send HA packets out of a secondary interface (e.g., a storage or management NIC) instead of the dedicated HA/heartbeat interface.
                            2. ARP Flapping & Confusion: Because of the weak host model, a node might reply to ARP requests for its HA IP address on any active interface. This poisons the ARP tables of adjacent switches and cluster peers, misrouting traffic to the wrong physical ports.
                            Expected Path:  [Node A: Dedicated HA NIC] ---------> [Switch] ---------> [Node B: Dedicated HA NIC]
                            Actual Path:    [Node A: Management NIC]   -- (Lowest Metric Route) --> [Node B: Management NIC]
                            

                            The Impact on HA Clusters (XCP-ng / XAPI)

                            When heartbeats leak onto the wrong network or get dropped due to strict firewalling (e.g., a management network blocking cluster traffic), the cluster experiences a false split-brain condition.

                            • Peer nodes assume the host is dead because heartbeats stopped arriving on the designated HA network.
                            • The HA fencing mechanism is triggered, causing the host to spontaneously reboot or self-fence to protect shared storage from corruption.

                            Recommended Subsystem Remedies

                            To enforce a strong host model where traffic strictly respects interface boundaries, the kernel and network orchestration layer (XAPI) must implement specific controls:

                            • Strict ARP Filtering (sysctl)
                              Prevent the kernel from answering ARP requests for an IP on the wrong interface.
                              net.ipv4.conf.all.arp_ignore = 1
                              net.ipv4.conf.all.arp_announce = 2
                              
                            • Reverse Path Filtering (RP-Filter):
                              Ensure the kernel drops packets that arrive on an interface if the return path would not normally use that same interface.
                              net.ipv4.conf.all.rp_filter = 1
                              
                            • Policy-Based Routing (PBR):
                              Configure separate routing tables for the HA interface so that traffic sourced from the HA IP is explicitly forced out of the HA NIC, ignoring global metrics in the main routing table.

                            So outside of altering these settings, if not altered at kernel source code level with a patch (that’s upstreamed to Xen Project), it will default to weak host model, a patch is required for it to instead default to the strong host model. If necessary it can target management, storage, backup, migration and/or HA networks specifically.

                            Doing a kernel source code patch will especially help out TWINSTOR development, currently in technical preview!

                            1 Reply Last reply
                            Reply Quote 1
                            • J
                              john.c
                              last edited by john.c

                              @tjkreidl When you’re back online my posts above, may reveal something about behaviours you noticed in Citrix XenServer and Citrix Hypervisor, during your employment there as CTP.

                              @poddingue This is a potential source for a joint patch to Xen Hypervisor kernel to fix this as part of upstreaming, something that will benefit from multiple eyes and hands working on it.

                              tjkreidlT 1 Reply Last reply
                              Reply Quote 0
                              • tjkreidlT
                                tjkreidl Ambassador @john.c
                                last edited by

                                @john.c Interesting information about the strong v. weak host model.
                                Active-active and LACP bonds will split the network traffic between the NICs, while an active-passive bond only uses the primary. WHat type of bonds do you have set up?

                                J 1 Reply Last reply
                                Reply Quote 0
                                • J
                                  john.c @tjkreidl
                                  last edited by john.c

                                  @tjkreidl said:

                                  @john.c Interesting information about the strong v. weak host model.
                                  Active-active and LACP bonds will split the network traffic between the NICs, while an active-passive bond only uses the primary. WHat type of bonds do you have set up?

                                  @tjkreidl It is either a balance-alb (Mode 6) or a balance-xor (Mode 2) bond topology on this specific deployment (as the upstream router has limited port capacity and does not support LACP).

                                  However, the brilliant part about this architectural issue is that the underlying Layer 3 routing vulnerability remains identical regardless of which of these two modes is active. Both modes open up multiple physical paths for outbound traffic under a single logical bond, giving the kernel's default weak host model the perfect opportunity to misroute packets under load.

                                  Here is how the weak host model breaks both configurations:

                                  • If it is balance-alb (Mode 6): The bonding driver actively performs Layer 2 ARP and MAC manipulation to balance paths without switch assistance. The weak host model completely undermines this logic because the kernel treats the IP address as globally accessible to the whole host. Under load, the kernel's Layer 3 routing engine completely ignores the bonding driver's intended pathing boundaries, picks a "cheaper" path via global metrics, and leaks the packet out of a completely separate infrastructure interface (like Management or Storage).
                                  • If it is balance-xor (Mode 2): The driver relies on a strict hash policy (like layer2 or layer2+3) to statically map traffic to a destination across a specific physical NIC slave. Yet, if a heavy background process (like a backup job or storage replication) alters local routing table costs or causes transient congestion, the kernel’s Layer 3 logic overrides that static Layer 2 pathing—spilling packets out of an unrelated physical port.

                                  Ultimately, OpenMetrics packet-flow telemetry caught this exact moment of divergence: the Layer 3 stack bypassed the logical bond boundary entirely. The packet exited on an unintended infrastructure port carrying the wrong source IP, where adjacent switches or firewalls dropped it as unroutable. To XAPI and the HA daemon, the heartbeat was instantly lost on that specific port, triggering the self-fencing reboot loop.

                                  This is why it's a universal vulnerability across multi-path bonding modes on a multi-homed system, and why an upstream patch enforcing a strong host model for the Infrastructure Plane is the cleanest solution.

                                  tjkreidlT 1 Reply Last reply
                                  Reply Quote 0
                                  • J
                                    john.c
                                    last edited by john.c

                                    @AtaxyaNetwork Depending on your current or past homelab topography, this behavior might ring a few bells for you as well.

                                    Homelab and prosumer environments are highly susceptible to this exact weak host routing leak. Because we often multiplex distinct infrastructure planes (Management, dedicated storage networks, and backup backplanes) across multi-port NIC bonds—frequently using balance-alb or balance-xor because the upstream switches lack stacked enterprise LACP support—the conditions are perfect for a routing collision.

                                    If a heavy data operation (like a massive VM migration or a backup sync) alters the local metric weightings or causes micro-congestion, the default weak host model can silently push Dom0 host-terminated packets onto the wrong physical interface segment. It's a classic hidden variable that can cause erratic connection drops or unexplainable HA timeouts on otherwise perfectly configured hardware.

                                    1 Reply Last reply
                                    Reply Quote 0
                                    • tjkreidlT
                                      tjkreidl Ambassador @john.c
                                      last edited by

                                      @john.c Most interesting, and I fully agree, that a strong host model enforcement policy is really the best and only recourse for such a topology, or so it would seem.
                                      The only other option that comes to mind would be to not put the heartbeat connection on any sort of bond or multipath. That's, of course, not ideal.

                                      J 1 Reply Last reply
                                      Reply Quote 0
                                      • J
                                        john.c @tjkreidl
                                        last edited by john.c

                                        @tjkreidl said:

                                        @john.c Most interesting, and I fully agree, that a strong host model enforcement policy is really the best and only recourse for such a topology, or so it would seem.
                                        The only other option that comes to mind would be to not put the heartbeat connection on any sort of bond or multipath. That's, of course, not ideal.

                                        @tjkreidl Exactly, and you have hit on the exact architectural compromise we've all been forced to make for years.

                                        Moving the heartbeat connection off a bond/multipath and onto a dedicated, single physical NIC does reduce the Layer 3 path alternatives that trigger the weak host leak. However, as you rightly pointed out, it's highly non-ideal. By removing the bond, we introduce a single point of failure (SFP/cable/switch port) directly into the critical HA backplane. We shouldn't have to sacrifice physical hardware redundancy just to keep the kernel's Layer 3 routing engine from misbehaving.

                                        This is precisely why enforcing a strong host model for the Infrastructure Plane is the true architectural solution. It allows administrators to safely use balance-alb, balance-xor, or multipathing for maximum hardware resilience, while ensuring that Dom0-terminated traffic strictly honors its designated interface boundaries regardless of global metrics.

                                        Since we are in full agreement on the root cause and the ideal fix, this looks like a prime candidate for a strategic architectural shift. Hopefully, @TeddyAstie and the @Team-Hypervisor-Kernel can look at how we can implement this structural protection by default—perhaps utilizing targeted sysctl overrides (arp_ignore, arp_announce, rp_filter) or interface socket-binding for host-terminated infrastructure networks.

                                        1 Reply Last reply
                                        Reply Quote 1
                                        • C
                                          carloum70 @john.c
                                          last edited by

                                          @john.c now I remember, during the installation we used the eth4 as the management interface.

                                          [16:54 dacshyp002 ~]# cat /etc/firstboot.d/data/management.conf 
                                          LABEL='eth4'
                                          MODE='static'
                                          IP='172.28.4.11'
                                          NETMASK='255.255.252.0'
                                          GATEWAY='172.28.4.1'
                                          MODEV6='none'
                                          DNS='8.8.8.8'
                                          

                                          Afterwords we created a bond bond0 with eth2/eth3 and used this as the management interface.
                                          The xe pif-list command shows the following:

                                          device ( RO)                   : bond0
                                                         management ( RO): true
                                              IP-configuration-mode ( RO): Static
                                                                 IP ( RO): 172.28.4.11
                                          
                                          device ( RO)                   : eth4
                                                         management ( RO): false
                                              IP-configuration-mode ( RO): Static
                                                                 IP ( RO): 172.28.4.11
                                          

                                          Is it that easy to remove the IP address from the eth4 interface? Because it make no sense to have duplicate ip-addresses.
                                          I think we've had duplicate IP addresses all along, but we only noticed the issue after enabling HA.
                                          By the way what is the correct way to change the management interface? We used the xsconsole.

                                          J 1 Reply Last reply
                                          Reply Quote 0
                                          • J
                                            john.c @carloum70
                                            last edited by john.c

                                            @carloum70 said:

                                            @john.c now I remember, during the installation we used the eth4 as the management interface.

                                            [16:54 dacshyp002 ~]# cat /etc/firstboot.d/data/management.conf 
                                            LABEL='eth4'
                                            MODE='static'
                                            IP='172.28.4.11'
                                            NETMASK='255.255.252.0'
                                            GATEWAY='172.28.4.1'
                                            MODEV6='none'
                                            DNS='8.8.8.8'
                                            

                                            Afterwords we created a bond bond0 with eth2/eth3 and used this as the management interface.
                                            The xe pif-list command shows the following:

                                            device ( RO)                   : bond0
                                                           management ( RO): true
                                                IP-configuration-mode ( RO): Static
                                                                   IP ( RO): 172.28.4.11
                                            
                                            device ( RO)                   : eth4
                                                           management ( RO): false
                                                IP-configuration-mode ( RO): Static
                                                                   IP ( RO): 172.28.4.11
                                            

                                            Is it that easy to remove the IP address from the eth4 interface? Because it make no sense to have duplicate ip-addresses.
                                            I think we've had duplicate IP addresses all along, but we only noticed the issue after enabling HA.
                                            By the way what is the correct way to change the management interface? We used the xsconsole.

                                            If you’re using a bonded NIC what mode is used? If not bonding it will allow for multiple NICs, via bond0-12 etc to show a single NIC. But only use once my above fix has landed, to avoid a return to this issue, unless forcing strong host yourself before hand.

                                            Remove the IPs with the following process:-

                                            1. xe pif-list params=uuid,device,IP,management
                                            2. xe pif-reconfigure-ip uuid=<PIF_UUID> mode=none
                                            3. xe-toolstack-restart
                                            1 Reply Last reply
                                            Reply Quote 0

                                            Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                                            Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                                            With your input, this post could be even better 💗

                                            Register Login
                                            • First post
                                              Last post