HA causes reboot of xcp-ng nodes
-
I did some further investigation and I think there is something wrong with the ip-settings of the management interface.
At the xsconsole I see the following:
Management Network Parameters Device bond0 address 172.28.4.11 Netmask 255.255.252.0 Gateway 172.28.4.1So it's using the bond0 interface.
# xe pif-list host-uuid=41a2a448-a5dc-44c6-be44-c07540d75c60 uuid ( RO) : 0761d887-268e-5cf0-401b-a08ad7c419aa device ( RO): bond0 MAC ( RO): 30:3e:a7:1d:b0:90 currently-attached ( RO): true VLAN ( RO): -1 network-uuid ( RO): 28d51793-0226-cf3f-75d3-ed7c5a9d8b33 host-uuid ( RO): 41a2a448-a5dc-44c6-be44-c07540d75c60This is part of the output of xe bond-list
# xe bond-list uuid ( RO) : 79c525c3-22f9-04ce-4a80-21d18ef65ebc master ( RO): 0761d887-268e-5cf0-401b-a08ad7c419aa slaves ( RO): 7ac18886-e0a5-057b-bac0-36b909588dba ; 2f18fc9a-8c1c-dfce-9416-fa8c657c5f63I already checked that:
7ac18886-e0a5-057b-bac0-36b909588dba --> eth3 2f18fc9a-8c1c-dfce-9416-fa8c657c5f63 --> eth2And now the confusing part. This is part of the output of the ip a command:
2: eth2: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq master ovs-system state UP group default qlen 1000 link/ether 30:3e:a7:1d:b0:90 brd ff:ff:ff:ff:ff:ff 3: eth3: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq master ovs-system state UP group default qlen 1000 link/ether 30:3e:a7:1d:b0:91 brd ff:ff:ff:ff:ff:ff 4: eth4: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq master ovs-system state UP group default qlen 1000 link/ether 30:3e:a7:1d:b0:92 brd ff:ff:ff:ff:ff:ff 13: xenbr4: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000 link/ether 30:3e:a7:1d:b0:92 brd ff:ff:ff:ff:ff:ff inet 172.28.4.11/22 brd 172.28.7.255 scope global xenbr4 valid_lft forever preferred_lft foreverAs you can see the xenbr4 has also the management ip-address and has the same mac-address as the eth4 interface.
OK check the following commands:[17:21 dacshyp002 ~]# ip addr show xapi1 15: xapi1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000 link/ether 30:3e:a7:1d:b0:90 brd ff:ff:ff:ff:ff:ff inet 172.28.4.11/22 brd 172.28.7.255 scope global xapi1 valid_lft forever preferred_lft forever [17:21 dacshyp002 ~]# ip addr show xenbr4 13: xenbr4: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000 link/ether 30:3e:a7:1d:b0:92 brd ff:ff:ff:ff:ff:ff inet 172.28.4.11/22 brd 172.28.7.255 scope global xenbr4 valid_lft forever preferred_lft foreverAnd If I look at the routing table:
# ip route default via 172.28.4.1 dev xapi1 172.18.8.0/23 dev xenbr0 proto kernel scope link src 172.18.8.11 172.18.10.0/23 dev xenbr7 proto kernel scope link src 172.18.10.11 172.28.4.0/22 dev xenbr4 proto kernel scope link src 172.28.4.11 172.28.4.0/22 dev xapi1 proto kernel scope link src 172.28.4.11According to me the node has 2 interfaces xenbr4 / xapi1 with the same ip-address and also 2 different routes for the same ip-range 172.28.4.0/22.
And that's why the heartbeat has some issues. Can someone confirm this ?
By the way I really don't know how this happened. -
@carloum70 If you are willing to try this, you can reset the primary management interface (PMI) on a bond, WIthin the bash shell:
:
xe pif-list
This will show the PIFs and their current network associations.Identify the bond master PIF
If you have a bonded network, the master PIF is the one that represents the bond. You can list it with:xe pif-list network-uuid=<bond-uuid>
The master PIF is the one that will be used as the primary interface for the bond.Reset the management interface to the bond master
Use the xe-reset-networking command with the --reset-primary option:xe-reset-networking --reset-primary
This will move the management interface to the bond master PIF, which is the correct way to reassign it when the bond is created or reconfigured. -
According to xsconsole the management interface is bond0
Current Management Interface Device bond0 MAC Address 30:3e:a7:1d:b0:90 DHCP/Static IP Static IP address 172.28.4.11 Netmask 255.255.252.0 Gateway 172.28.4.1 Hostname dacshyp002According the output of xe pif-list
uuid ( RO) : 0761d887-268e-5cf0-401b-a08ad7c419aa device ( RO): bond0 MAC ( RO): 30:3e:a7:1d:b0:90 currently-attached ( RO): true VLAN ( RO): -1 network-uuid ( RO): 28d51793-0226-cf3f-75d3-ed7c5a9d8b33 host-uuid ( RO): 41a2a448-a5dc-44c6-be44-c07540d75c60Let's check the network
# xe network-list uuid=28d51793-0226-cf3f-75d3-ed7c5a9d8b33 uuid ( RO) : 28d51793-0226-cf3f-75d3-ed7c5a9d8b33 name-label ( RW): mgmt-bond name-description ( RW): bridge ( RO): xapi1At linux level:
# ip addr show xapi1 15: xapi1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000 link/ether 30:3e:a7:1d:b0:90 brd ff:ff:ff:ff:ff:ff inet 172.28.4.11/22 brd 172.28.7.255 scope global xapi1 valid_lft forever preferred_lft foreverxapi1 is a bond of et2 and eth3. So far so good.
But if I am running the xe-reset-networking commandYour network will be re-configured as follows: Management interface: eth4 Reset interface name rules: Yes IP configuration mode: dhcp IPv6 configuration mode:none If you want to change any of the above settings, type 'no' and re-run the command with appropriate arguments (use --help for a list of options). Type 'yes' to continue. Type 'no' to cancel. noSo for some reason xcp-ng thinks the eth4 is management interface, which also has the ip-address 172.28.4.11/22 .
By the way --reset-primary is not a valid option:# xe-reset-networking --reset-primary Usage: xe-reset-networking [options] xe-reset-networking: error: no such option: --reset-primary -
@carloum70 Sorry, that option was deprecated.
You may need to be more specific in the reset, something like the following with is just an example:
xe-reset-networking -m IP_of_Master --device=eth0 --mode=static --ip=192.168.1.50 --netmask=255.255.255.0 --gateway=192.168.1.1 --dns=192.168.1.254Check the options and see which ones you actually need in your case.
Worst case, you could possibly dissolve the bond, redo the PMI and then re-create the bond. WHy eth4 shows up is hard to guess. I have seen before that on some hosts in a pool that the NIC order was not the same, even though the hardware and OS versions were identical. In that case, you have to shut down the NICs and use interface-rename to reassign specific NIC names to the corresponding MAC addresses. -
@carloum70 the same address on both xenbr4 and xapi1 looks like the best lead so far.

I haven't tested it, but two interfaces answering for 172.28.4.11, with two routes to the same subnet, could send heartbeat traffic out the wrong NIC now and then, and that would fit the drops in your xha.log. bleader untangled something close to it in https://xcp-ng.org/forum/topic/12472 (two bridges on one subnet, replies leaving by the wrong one).As for eth4: without
--device, xe-reset-networking takes the NIC recorded at install time in/etc/firstboot.d/data/management.conf, so I think it's just remembering the installer's choice.I'd hold off on @tjkreidl's worst-case reset for now, because the docs say it wipes all PIF, bond and VLAN config, force-stops VMs, and isn't supported while HA is on (Emergency Network Reset).
xe pif-list host-name-label=dacshyp002 params=device,IP-configuration-mode,IP,managementonly lists, and it would show whether XAPI itself thinks eth4 has that IP; if it does, @Team-XAPI-Network will know the clean way to drop it.
-
@carloum70 the same address on both xenbr4 and xapi1 looks like the best lead so far.

I haven't tested it, but two interfaces answering for 172.28.4.11, with two routes to the same subnet, could send heartbeat traffic out the wrong NIC now and then, and that would fit the drops in your xha.log. bleader untangled something close to it in https://xcp-ng.org/forum/topic/12472 (two bridges on one subnet, replies leaving by the wrong one).As for eth4: without
--device, xe-reset-networking takes the NIC recorded at install time in/etc/firstboot.d/data/management.conf, so I think it's just remembering the installer's choice.I'd hold off on @tjkreidl's worst-case reset for now, because the docs say it wipes all PIF, bond and VLAN config, force-stops VMs, and isn't supported while HA is on (Emergency Network Reset).
xe pif-list host-name-label=dacshyp002 params=device,IP-configuration-mode,IP,managementonly lists, and it would show whether XAPI itself thinks eth4 has that IP; if it does, @Team-XAPI-Network will know the clean way to drop it.
@Team-Hypervisor-Kernel May also wish to have a look, as part of the problem can stem from the kernel’s weak host model, as it’s not just networking in XAPI but also kernel, due to asymmetric routing. This is down to packets being routed on the lowest cost path.
-
Technical Summary: Weak Host Model & HA Heartbeat Failures
The Problem: Asymmetric Routing & The Weak Host Model
Linux defaults to a weak host model for IPv4 networking. The kernel treats IP addresses as belonging to the entire host, rather than being strictly bound to a specific physical or logical network interface (NIC).
When a multi-homed cluster node transmits High Availability (HA) heartbeats or storage synchronization packets, the kernel evaluates its local routing table. If multiple paths exist to the destination subnet, the kernel will choose the egress interface based on the lowest routing metric (cost).
This causes two major issues in HA clusters:
- Asymmetric Outbound Traffic: The kernel may send HA packets out of a secondary interface (e.g., a storage or management NIC) instead of the dedicated HA/heartbeat interface.
- ARP Flapping & Confusion: Because of the weak host model, a node might reply to ARP requests for its HA IP address on any active interface. This poisons the ARP tables of adjacent switches and cluster peers, misrouting traffic to the wrong physical ports.
Expected Path: [Node A: Dedicated HA NIC] ---------> [Switch] ---------> [Node B: Dedicated HA NIC] Actual Path: [Node A: Management NIC] -- (Lowest Metric Route) --> [Node B: Management NIC]The Impact on HA Clusters (XCP-ng / XAPI)
When heartbeats leak onto the wrong network or get dropped due to strict firewalling (e.g., a management network blocking cluster traffic), the cluster experiences a false split-brain condition.
- Peer nodes assume the host is dead because heartbeats stopped arriving on the designated HA network.
- The HA fencing mechanism is triggered, causing the host to spontaneously reboot or self-fence to protect shared storage from corruption.
Recommended Subsystem Remedies
To enforce a strong host model where traffic strictly respects interface boundaries, the kernel and network orchestration layer (XAPI) must implement specific controls:
- Strict ARP Filtering (
sysctl)
Prevent the kernel from answering ARP requests for an IP on the wrong interface.net.ipv4.conf.all.arp_ignore = 1 net.ipv4.conf.all.arp_announce = 2 - Reverse Path Filtering (RP-Filter):
Ensure the kernel drops packets that arrive on an interface if the return path would not normally use that same interface.net.ipv4.conf.all.rp_filter = 1 - Policy-Based Routing (PBR):
Configure separate routing tables for the HA interface so that traffic sourced from the HA IP is explicitly forced out of the HA NIC, ignoring global metrics in the main routing table.
So outside of altering these settings, if not altered at kernel source code level with a patch (that’s upstreamed to Xen Project), it will default to weak host model, a patch is required for it to instead default to the strong host model. If necessary it can target management, storage, backup, migration and/or HA networks specifically.
Doing a kernel source code patch will especially help out TWINSTOR development, currently in technical preview!
-
@tjkreidl When you’re back online my posts above, may reveal something about behaviours you noticed in Citrix XenServer and Citrix Hypervisor, during your employment there as CTP.
@poddingue This is a potential source for a joint patch to Xen Hypervisor kernel to fix this as part of upstreaming, something that will benefit from multiple eyes and hands working on it.
-
@john.c Interesting information about the strong v. weak host model.
Active-active and LACP bonds will split the network traffic between the NICs, while an active-passive bond only uses the primary. WHat type of bonds do you have set up? -
@john.c Interesting information about the strong v. weak host model.
Active-active and LACP bonds will split the network traffic between the NICs, while an active-passive bond only uses the primary. WHat type of bonds do you have set up?@tjkreidl It is either a
balance-alb(Mode 6) or abalance-xor(Mode 2) bond topology on this specific deployment (as the upstream router has limited port capacity and does not support LACP).However, the brilliant part about this architectural issue is that the underlying Layer 3 routing vulnerability remains identical regardless of which of these two modes is active. Both modes open up multiple physical paths for outbound traffic under a single logical bond, giving the kernel's default weak host model the perfect opportunity to misroute packets under load.
Here is how the weak host model breaks both configurations:
- If it is
balance-alb(Mode 6): The bonding driver actively performs Layer 2 ARP and MAC manipulation to balance paths without switch assistance. The weak host model completely undermines this logic because the kernel treats the IP address as globally accessible to the whole host. Under load, the kernel's Layer 3 routing engine completely ignores the bonding driver's intended pathing boundaries, picks a "cheaper" path via global metrics, and leaks the packet out of a completely separate infrastructure interface (like Management or Storage). - If it is
balance-xor(Mode 2): The driver relies on a strict hash policy (likelayer2orlayer2+3) to statically map traffic to a destination across a specific physical NIC slave. Yet, if a heavy background process (like a backup job or storage replication) alters local routing table costs or causes transient congestion, the kernel’s Layer 3 logic overrides that static Layer 2 pathing—spilling packets out of an unrelated physical port.
Ultimately, OpenMetrics packet-flow telemetry caught this exact moment of divergence: the Layer 3 stack bypassed the logical bond boundary entirely. The packet exited on an unintended infrastructure port carrying the wrong source IP, where adjacent switches or firewalls dropped it as unroutable. To XAPI and the HA daemon, the heartbeat was instantly lost on that specific port, triggering the self-fencing reboot loop.
This is why it's a universal vulnerability across multi-path bonding modes on a multi-homed system, and why an upstream patch enforcing a strong host model for the Infrastructure Plane is the cleanest solution.
- If it is
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login