XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    • Profile
    • Following 0
    • Followers 1
    • Topics 18
    • Posts 435
    • Groups 0
    J Offline
    1. Home
    2. john.c
    3. Posts

    Posts

    Recent Best Controversial
    • RE: Future Architecture?

      @herhin2017 said:

      @john.c
      Your absolute wright, i do my best and train about 25 young engineers per year on linux (we use debian). I also see more and more startups or small companies building their infrastructure on linux.
      Best greetings from austria

      It may be worth getting in line and noting the mention of EFI based servers for official Vates support with XCP-ng version 9.0 and above, in government reports. That way when the school’s hardware is refreshed it can be ensured that your provided with EFI capable servers, in time for XCP-ng version 8.3 EOL.

      I personally are already ready for XCP-ng version 9.0 due to my servers being Dell PowerEdge R620 for XCP-ng hosts. Along with Debian version 13.6 on the VMs. While using a Dell Precision 3590 to manage those systems.

      The “wright” makes your above now sound like a wheel wright, and its profession instead of “right” in for when correct or agreeing with someone or something.

      posted in News
      J
      john.c
    • RE: Future Architecture?

      @herhin2017 said:

      Hello John C.
      In Austria, the Government pays Microsoft Licenses for everything. So every school can use their Software (you know what i mean). To use Linux or something else is not really well received. I am not able to use corporations with 3rd party comp. and i am not allowed to take money from somewere else!

      So that may be about to change, depending on how far an on going transition in Austrian government goes. As with your country’s military now on Linux and LibreOffice your academic system is going to need to be training, Austrian children on those systems to prepare them, for if they join the military.

      https://linuxsecurity.com/news/government/linux-security-defense

      posted in News
      J
      john.c
    • RE: Future Architecture?

      @herhin2017 said:

      Yes, your wright! I know about that. As i mentioned, our machines are really working perfect (they can easily be a productive system) I am asking because our school does not have much money (in fact nothing) and i will go on with my lessons in hypervisors with xcp-ng.
      THANKS TO EVERYONE who works on this project - i like it!!
      P.S. If you need user tests in grater scale, we can do it (10 machines) and we have time.

      Welcome to the forums!

      To answer your question while expanding on @Poddingue's excellent point: it all comes down to the AlmaLinux base chosen for the dom0 control domain. While AlmaLinux 10 does offer an alternative x86_64-v2 compilation option for older processors, relying on a sub-baseline build for a modern hypervisor can limit your performance and future third-party repository support.

      To truly get the most out of XCP-ng 9.0 when it reaches stable release—and to avoid the hard limitations of the modern EFI and Secure Boot requirements—you will want to target at least generation 9 of HPE servers (Gen9) at a minimum.

      Since you are dealing with a tight school budget but need to acquire a large number of machines for your network engineering students, here are the most effective ways to source that volume:

      • Corporate Hardware Donations: Depending on your school's official legal status (such as Gemeinnützigkeit), you might be able to get these machines donated for free. Many large companies in Austria cycle out their enterprise hardware every few years and actively look to donate working Gen9 or Gen10 servers to schools for tax benefits or community goodwill. It is highly worth reaching out to local IT departments or tech companies!

      • Professionally Refurbished Bulk Batches: Sourcing former enterprise hardware in large, uniform batches is your next best path. Gen9 servers offer excellent performance for a very low price right now. Keeping the hardware identical also makes managing your pool templates significantly easier.

      • Academic Programs & Vates Discounts: If your rollout scales up and you eventually need official commercial backing, it is well worth reaching out directly to Vates. Let them know you are an educational institution; they are often very supportive of academia and may offer steep volume or academic purchasing discounts for schools.

      Good luck setting up the lab for your hypervisor lessons!

      P.S. A quick, friendly tip for your English writing: drop the "w" when you want to agree with someone. If you use "write", it sounds like the action of writing words down on paper (schreiben). To agree with someone, the correct phrases are "You're right" or "You are right" (Du hast recht / Sie haben recht).

      posted in News
      J
      john.c
    • RE: Future Architecture?

      It depends on whether the Alma Linux core chosen for dom0 is the one for v3 and v4 or the one for v2. As AlmaLinux also has a variant in vanilla not just for baseline v3 and v4 but also for one baseline below in other words v2.

      posted in News
      J
      john.c
    • RE: Pool metadata backup failed after xoa upgrade

      @poddingue said:

      There's a similar report on 6.8 in https://xcp-ng.org/forum/topic/12453, where @flakpyro said it got fixed through a support ticket and that @florent would know the exact fix.

      I read through the 6.9.0 changelog, but I couldn't find an entry about metadata backups or Body Timeout Error, so I can't confirm that moving to the latest channel fixes it, though it may have gone in without a changelog line. 🤷

      If you have a support contract, a ticket that points at that thread might be the quickest way to the patch Florent made.

      It is in the 6.9.0 change log but not where you might expect, even as a backup bug fix. In this case it’s filed under the miscellaneous (misc) section as “Fixed: BodyTimeoutError during long transfers”. Given where it is in the change log you may have missed it.

      If it isn’t in the change log then why is it in the release announcement?

      posted in Backup
      J
      john.c
    • RE: XOA Unable to connect xo server every 30s

      @GregBinSD said:

      Here are two more notes regarding the XO6 "Unable to connect to XO server. Retry" message, which pops up after 30 seconds.

      It occurs when either the Chrome or the Microsoft Edge browsers are used on my Windows 11 PC.

      However, I often use a Samsung Tab-A9 (tablet), and it does not have this issue with XO6. It uses the Chrome browser.

      To enlighten you the Microsoft Edge your talking about is not the original release (from Windows 10). It’s the Chromium based release from during Windows 10 and has been that one ever since. The original release of Microsoft Edge had its own rendering engine called MSHTML. The current modern Edge effectively shares a common upstream code base with Google Chrome, namely Chromium.

      The Samsung Tab-A9 doesn’t have the issue even though it, uses the same browser namely Google Chrome. This is the case because the tablet uses a version of Google Android, which has its own kernel, which is a fork or variation of the Linux Kernel.

      posted in Xen Orchestra
      J
      john.c
    • RE: Pool metadata backup failed after xoa upgrade

      @jacob.becker said:

      Hi!
      After the XOA update to 6.8.2 (Stable), the Backup of the pool metadata fail with

      Error: Body Timeout Error
      

      after aprox. 5 minutes.
      Other Backup Jobs run without issues.

      @jacob.becker On the latest channel for XOA updates version 6.9 (6.9.0) has a fix for Body Timeout Error issue, no sign of a fix in 6.8.0 to 6.8.2 version series. Change channel and you’ll have the fix or open for Vates staff via PM a support tunnel access so they can patch your 6.8.2 with the one specific for this version.

      This is an SERIOUSLY urgent action to fix the backup issue, but the Critical part is that it also fixes an important security vulnerability (VSA-2026-044).

      posted in Backup
      J
      john.c
    • RE: Internet connectivity - Check XOA failed.

      @acebmxer said:

      @john.c

      We currently dont use ipv6 internally. No network changes were made at this location other then moving the proxy to the correct network for nbd connections. With that move some how made the ipv6 issue appear. So it was just easy to disable ipv6 in the proxy. I guess if we ever switch to ipv6 (no plans too) then i guess i will have to look back into it.

      @acebmxer @zorro If any of your VMs are facing the public internet, completing your IPv6 Readiness compliance is vital. With regional internet registries completely exhausted of free-pool IPv4 space, modern endpoints and cloud architectures are increasingly deploying IPv6-only infrastructure.

      If you are referring to an internet access proxy, disabling IPv6 introduces significant architectural risk. If any upstream transit provider, carrier, or edge CDN in the path to your target FQDN transitions to an IPv6-only topology, your access path will break, causing a hard outage. If you are specifically utilizing a transit provider like XO Proxy, transitioning to dual-stack or IPv6-only transport is even more critical to ensure deterministic routing across the wider internet footprint.

      From an architecture and security standpoint, IPv6 introduces critical enterprise enhancements:

      • SLAAC Privacy Extensions: Enables temporary, rotating addresses to mitigate device fingerprinting and endpoint tracking.
      • Native IPSec Integration: While RFC 8200 technically shifted IPSec from a hard protocol requirement to an optional component, it remains a native architectural element of the IPv6 stack. Unlike IPv4—where IPSec must be bolted on as an awkward overlay—IPv6 accommodates encryption headers natively, simplifying the deployment of secure end-to-end transport encryption across enterprise and government domains.

      If you need to pitch this network-wide transition to leadership for project approval, I highly recommend framing it around business continuity and risk mitigation. Pointing out the looming vulnerability of upstream IPv6-only transit paths—combined with the compliance advantages of native architectural security—should give you the exact leverage needed to get this budgeted, planned, and implemented.

      posted in Management
      J
      john.c
    • RE: Internet connectivity - Check XOA failed.

      @acebmxer said:

      @poddingue

      I had issue with a remote proxy.. showed error in xoa untill i disabled ipv6. But only with one of two remote proxies. This was in a support ticket.

      Did you test your full stack IP v6 readiness with one of the online testers? The reason being your local LAN maybe ready even at your router level, but if your ISP doesn’t have full stack (or even dual full stack - v4 and v6) then FQDN addresses which are only on IP v6 only may not work). Also v4 and v6 IP address has to also be associated with the FQDN being contacted.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @tjkreidl said:

      @john.c Keeping the various network traffic isolated according to specific usage (management, storage, VMs, etc.) is always a good idea. The last system I managed had 10GiB LACP bonds and using VLANs to isolate traffic and that worked fine with a four-node pool running around 80 or so XenDesktop instances per node. Never experienced any congestion issues. Each dom0 had a ton of memory and I believe it was either 8 or 16 VCPUs to make sure there were sufficient compute and memory allocations to allow for sometimes very heavy loads. It also helped that I eventually added GPUs to take on some of the computational load, in particular when some of the VMs were running applications employing heavy graphics.

      @tjkreidl Thanks Tobias, that’s great validation! Running 10GiB LACP bonds with proper VLAN isolation is definitely the ultimate goal for production stability, especially when pushing 80+ VMs per node.

      Your point about dom0 resource allocation is also huge—people often forget that saturated vCPUs and starved dom0 memory can bottleneck network processing just as fast as a saturated physical link under heavy loads. Giving dom0 the extra compute headroom ensures the orchestration layer doesn't drop packets when backups or migrations scale up across that many instances.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @carloum70 said:

      Some update:

      After removing the ip-address from eth4 the following messages disappeared

       [warn] SC: (script_service_do_query_liveset) reporting "heartbeat approaching timeout". host[1].time_since_last_hb=30408.
      [info] SC: (script_service_do_query_liveset) "Heartbeat approaching timeout" turned FALSE 
      

      By the way I forgot to mention that each time I enable HA the following warning appears in xha.log

      Sep 29 13:31:56 CEST 2026 [warn] BM: cannot open bonding status file (/proc/net/bonding/bond0). (2)
      Sep 29 13:31:56 CEST 2026 [info] BM: this is not bonded. Terminating bonding thread.
      

      bond0 (eth2 + eth3) is the management interface and also used for the HA.

      Warning: Unconfigured Backup/Migration Networks Trigger HA Reboots

      Watch out if a dedicated backup or migration network is not configured in Vates VMS, all of that traffic defaults to the management network.

      When massive backup jobs or VM migrations saturate the management interface, it chokes the High Availability (HA) cluster heartbeats. This network congestion causes packet loss, triggers false split-brain conditions, and forces the host into a spontaneous reboot (self-fencing). This will especially occur on highly busy and congested instances of Vates VMS.

      While sysctl ARP filtering patches the routing leak, it cannot fix physical network saturation.

      The Long-Term Fix: Quad-Port Ethernet Upgrades

      I highly recommend upgrading all server LAN cards (especially XCP-ng hosts) to multiple quad-port modules across multiple slots or daughter cards. Abundant physical interfaces allow you to build dedicated LACP or Active-Backup bonds, ensuring complete infrastructure isolation:

      • Dedicated HA & Management: Keeps critical XAPI orchestration and heartbeat checks isolated from heavy data bursts.
      • Dedicated Storage: Keeps iSCSI, NFS, or XOSTOR disk I/O running smoothly on its own low-latency pipe.
      • Dedicated VM Traffic: Isolates production guest network traffic from infrastructure management tasks.
      • Dedicated Backup & Migration: Physical ports can be carved out exclusively for VM motions and backup windows—or at least provide a dedicated, shared channel away from HA traffic.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @bvitnik said:

      @carloum70 I can't really help you with anything specific but I must make a note that when analyzing network setup you must also take into account the Open vSwitch. Looking at the classic Linux networking, bridging and routing facilities is not enough to get you the full picture.

      For example, if I remember correctly, Active-Active bonding (non LACP) is implemented by Open vSwitch with so called balance-slb mode. This mode is not available with classic Linux kernel bonding options. What you are seeing (duplicated IPs) is maybe perfectly normal for this kind of setup with Open vSwitch.

      I would search for the cause of the issue somewhere else like network congestion. Heart beats are very sensitive to network congestion and can easily fail if you do not dedicate network links for this kind of traffic. If you are sharing the links over which heart beats are sent with some high traffic stuff (storage maybe?), it can easily make problems.

      UPDATE: unfortunately I don't have access to these kind of bonded setups any more so I can't check if the configuration differs in any way from yours. My home lab thingies are single network interface only.

      The Open vSwitch would have been already taken into account with packet data flow analysis. Though if balance-slb is Open vSwitch exclusive analysis will be needed, but if influenced by kernel then it will also likely be affected.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @carloum70 said:

      @john.c We are using a LCAP bond.
      For my understanding, if I remove the IP settings from eth4, the problem will still not be solved because of the known issue with the bond that you explained in “Weak Host Model & HA Heartbeat Failures.”
      (At the moment we have HA disabled.)

      What use is eth4 put to is it still management, if so put it into bond0? In one of my earlier posts on this I also detailed how to in the meantime use sysctl to force strong host, but best to wait for the official update from Vates.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @carloum70 said:

      @john.c now I remember, during the installation we used the eth4 as the management interface.

      [16:54 dacshyp002 ~]# cat /etc/firstboot.d/data/management.conf 
      LABEL='eth4'
      MODE='static'
      IP='172.28.4.11'
      NETMASK='255.255.252.0'
      GATEWAY='172.28.4.1'
      MODEV6='none'
      DNS='8.8.8.8'
      

      Afterwords we created a bond bond0 with eth2/eth3 and used this as the management interface.
      The xe pif-list command shows the following:

      device ( RO)                   : bond0
                     management ( RO): true
          IP-configuration-mode ( RO): Static
                             IP ( RO): 172.28.4.11
      
      device ( RO)                   : eth4
                     management ( RO): false
          IP-configuration-mode ( RO): Static
                             IP ( RO): 172.28.4.11
      

      Is it that easy to remove the IP address from the eth4 interface? Because it make no sense to have duplicate ip-addresses.
      I think we've had duplicate IP addresses all along, but we only noticed the issue after enabling HA.
      By the way what is the correct way to change the management interface? We used the xsconsole.

      If you’re using a bonded NIC what mode is used? If not bonding it will allow for multiple NICs, via bond0-12 etc to show a single NIC. But only use once my above fix has landed, to avoid a return to this issue, unless forcing strong host yourself before hand.

      Remove the IPs with the following process:-

      1. xe pif-list params=uuid,device,IP,management
      2. xe pif-reconfigure-ip uuid=<PIF_UUID> mode=none
      3. xe-toolstack-restart
      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @tjkreidl said:

      @john.c Most interesting, and I fully agree, that a strong host model enforcement policy is really the best and only recourse for such a topology, or so it would seem.
      The only other option that comes to mind would be to not put the heartbeat connection on any sort of bond or multipath. That's, of course, not ideal.

      @tjkreidl Exactly, and you have hit on the exact architectural compromise we've all been forced to make for years.

      Moving the heartbeat connection off a bond/multipath and onto a dedicated, single physical NIC does reduce the Layer 3 path alternatives that trigger the weak host leak. However, as you rightly pointed out, it's highly non-ideal. By removing the bond, we introduce a single point of failure (SFP/cable/switch port) directly into the critical HA backplane. We shouldn't have to sacrifice physical hardware redundancy just to keep the kernel's Layer 3 routing engine from misbehaving.

      This is precisely why enforcing a strong host model for the Infrastructure Plane is the true architectural solution. It allows administrators to safely use balance-alb, balance-xor, or multipathing for maximum hardware resilience, while ensuring that Dom0-terminated traffic strictly honors its designated interface boundaries regardless of global metrics.

      Since we are in full agreement on the root cause and the ideal fix, this looks like a prime candidate for a strategic architectural shift. Hopefully, @TeddyAstie and the @Team-Hypervisor-Kernel can look at how we can implement this structural protection by default—perhaps utilizing targeted sysctl overrides (arp_ignore, arp_announce, rp_filter) or interface socket-binding for host-terminated infrastructure networks.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @AtaxyaNetwork Depending on your current or past homelab topography, this behavior might ring a few bells for you as well.

      Homelab and prosumer environments are highly susceptible to this exact weak host routing leak. Because we often multiplex distinct infrastructure planes (Management, dedicated storage networks, and backup backplanes) across multi-port NIC bonds—frequently using balance-alb or balance-xor because the upstream switches lack stacked enterprise LACP support—the conditions are perfect for a routing collision.

      If a heavy data operation (like a massive VM migration or a backup sync) alters the local metric weightings or causes micro-congestion, the default weak host model can silently push Dom0 host-terminated packets onto the wrong physical interface segment. It's a classic hidden variable that can cause erratic connection drops or unexplainable HA timeouts on otherwise perfectly configured hardware.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @tjkreidl said:

      @john.c Interesting information about the strong v. weak host model.
      Active-active and LACP bonds will split the network traffic between the NICs, while an active-passive bond only uses the primary. WHat type of bonds do you have set up?

      @tjkreidl It is either a balance-alb (Mode 6) or a balance-xor (Mode 2) bond topology on this specific deployment (as the upstream router has limited port capacity and does not support LACP).

      However, the brilliant part about this architectural issue is that the underlying Layer 3 routing vulnerability remains identical regardless of which of these two modes is active. Both modes open up multiple physical paths for outbound traffic under a single logical bond, giving the kernel's default weak host model the perfect opportunity to misroute packets under load.

      Here is how the weak host model breaks both configurations:

      • If it is balance-alb (Mode 6): The bonding driver actively performs Layer 2 ARP and MAC manipulation to balance paths without switch assistance. The weak host model completely undermines this logic because the kernel treats the IP address as globally accessible to the whole host. Under load, the kernel's Layer 3 routing engine completely ignores the bonding driver's intended pathing boundaries, picks a "cheaper" path via global metrics, and leaks the packet out of a completely separate infrastructure interface (like Management or Storage).
      • If it is balance-xor (Mode 2): The driver relies on a strict hash policy (like layer2 or layer2+3) to statically map traffic to a destination across a specific physical NIC slave. Yet, if a heavy background process (like a backup job or storage replication) alters local routing table costs or causes transient congestion, the kernel’s Layer 3 logic overrides that static Layer 2 pathing—spilling packets out of an unrelated physical port.

      Ultimately, OpenMetrics packet-flow telemetry caught this exact moment of divergence: the Layer 3 stack bypassed the logical bond boundary entirely. The packet exited on an unintended infrastructure port carrying the wrong source IP, where adjacent switches or firewalls dropped it as unroutable. To XAPI and the HA daemon, the heartbeat was instantly lost on that specific port, triggering the self-fencing reboot loop.

      This is why it's a universal vulnerability across multi-path bonding modes on a multi-homed system, and why an upstream patch enforcing a strong host model for the Infrastructure Plane is the cleanest solution.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @tjkreidl When you’re back online my posts above, may reveal something about behaviours you noticed in Citrix XenServer and Citrix Hypervisor, during your employment there as CTP.

      @poddingue This is a potential source for a joint patch to Xen Hypervisor kernel to fix this as part of upstreaming, something that will benefit from multiple eyes and hands working on it.

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      Technical Summary: Weak Host Model & HA Heartbeat Failures

      The Problem: Asymmetric Routing & The Weak Host Model

      Linux defaults to a weak host model for IPv4 networking. The kernel treats IP addresses as belonging to the entire host, rather than being strictly bound to a specific physical or logical network interface (NIC).

      When a multi-homed cluster node transmits High Availability (HA) heartbeats or storage synchronization packets, the kernel evaluates its local routing table. If multiple paths exist to the destination subnet, the kernel will choose the egress interface based on the lowest routing metric (cost).

      This causes two major issues in HA clusters:

      1. Asymmetric Outbound Traffic: The kernel may send HA packets out of a secondary interface (e.g., a storage or management NIC) instead of the dedicated HA/heartbeat interface.
      2. ARP Flapping & Confusion: Because of the weak host model, a node might reply to ARP requests for its HA IP address on any active interface. This poisons the ARP tables of adjacent switches and cluster peers, misrouting traffic to the wrong physical ports.
      Expected Path:  [Node A: Dedicated HA NIC] ---------> [Switch] ---------> [Node B: Dedicated HA NIC]
      Actual Path:    [Node A: Management NIC]   -- (Lowest Metric Route) --> [Node B: Management NIC]
      

      The Impact on HA Clusters (XCP-ng / XAPI)

      When heartbeats leak onto the wrong network or get dropped due to strict firewalling (e.g., a management network blocking cluster traffic), the cluster experiences a false split-brain condition.

      • Peer nodes assume the host is dead because heartbeats stopped arriving on the designated HA network.
      • The HA fencing mechanism is triggered, causing the host to spontaneously reboot or self-fence to protect shared storage from corruption.

      Recommended Subsystem Remedies

      To enforce a strong host model where traffic strictly respects interface boundaries, the kernel and network orchestration layer (XAPI) must implement specific controls:

      • Strict ARP Filtering (sysctl)
        Prevent the kernel from answering ARP requests for an IP on the wrong interface.
        net.ipv4.conf.all.arp_ignore = 1
        net.ipv4.conf.all.arp_announce = 2
        
      • Reverse Path Filtering (RP-Filter):
        Ensure the kernel drops packets that arrive on an interface if the return path would not normally use that same interface.
        net.ipv4.conf.all.rp_filter = 1
        
      • Policy-Based Routing (PBR):
        Configure separate routing tables for the HA interface so that traffic sourced from the HA IP is explicitly forced out of the HA NIC, ignoring global metrics in the main routing table.

      So outside of altering these settings, if not altered at kernel source code level with a patch (that’s upstreamed to Xen Project), it will default to weak host model, a patch is required for it to instead default to the strong host model. If necessary it can target management, storage, backup, migration and/or HA networks specifically.

      Doing a kernel source code patch will especially help out TWINSTOR development, currently in technical preview!

      posted in Management
      J
      john.c
    • RE: HA causes reboot of xcp-ng nodes

      @poddingue said:

      @carloum70 the same address on both xenbr4 and xapi1 looks like the best lead so far. 🤷
      I haven't tested it, but two interfaces answering for 172.28.4.11, with two routes to the same subnet, could send heartbeat traffic out the wrong NIC now and then, and that would fit the drops in your xha.log. bleader untangled something close to it in https://xcp-ng.org/forum/topic/12472 (two bridges on one subnet, replies leaving by the wrong one).

      As for eth4: without --device, xe-reset-networking takes the NIC recorded at install time in /etc/firstboot.d/data/management.conf, so I think it's just remembering the installer's choice.

      I'd hold off on @tjkreidl's worst-case reset for now, because the docs say it wipes all PIF, bond and VLAN config, force-stops VMs, and isn't supported while HA is on (Emergency Network Reset). xe pif-list host-name-label=dacshyp002 params=device,IP-configuration-mode,IP,management only lists, and it would show whether XAPI itself thinks eth4 has that IP; if it does, @Team-XAPI-Network will know the clean way to drop it. 🤞

      @Team-Hypervisor-Kernel May also wish to have a look, as part of the problem can stem from the kernel’s weak host model, as it’s not just networking in XAPI but also kernel, due to asymmetric routing. This is down to packets being routed on the lowest cost path.

      posted in Management
      J
      john.c