XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    • Profile
    • Following 0
    • Followers 7
    • Topics 0
    • Posts 257
    • Groups 1
    tjkreidlT Offline
    1. Home
    2. tjkreidl

    tjkreidl

    @tjkreidl

    Ambassador

    Originally an astronomer for 15 years and later, an NAU employee in IT for 25+ years, most of which as a Team Lead. I was a Citrix CTP and NVIDIA NGCA for four years prior to retirement. Over 10 years' experience with XenServer/Citrix Hypervisor and close to that with NVIDIA GRID products. I was also a Red Hat Linux administrator and system programmer. Still trying to contribute what knowledge I have for the benefit of the IT community.

    125
    Reputation
    552
    Profile views
    257
    Posts
    7
    Followers
    0
    Following
    Joined
    Last Online
    Location Somewhere, USA

    tjkreidl Unfollow Follow
    Ambassador
    • RE: Introduce yourself!

      Hi, everyone. Nice to see this project turning into reality. I will try to spend time here as possible, which is hard with already being spread thinly. I've been a XenServer user for around a decade and am as interesting in learning as well as contributing whatever knowledge might be helpful to the community.

      Best regards,
      -=Tobias

      posted in Off topic
      tjkreidlT
      tjkreidl
    • RE: Remove a host from a pool

      And from the CLI:

      1. xe host-list (to get the UUID of the host)
      2. xe pool-eject host-uuid=<host_UUID>
      posted in Management
      tjkreidlT
      tjkreidl
    • RE: Server Admin Guide: A Tale of Two Servers: BIOS, GPU, and NUMA Tuning for XCP-ng: Preserving the valuable work done by Tobias Kreidl (@tjkreidl)

      @poddingue said:

      This is great to see, thank you for taking the time to rescue this; and thanks to @john.c for the recovery work and to @tjkreidl for writing it in the first place.
      I went looking, and there is a small XCP-ng-specific piece on this in the official docs under NUMA affinity (https://docs.xcp-ng.org/compute#numa-affinity), but it's nothing like the depth of the Tale of Two Servers series, so having the originals archived is genuinely useful. 👏
      I won't pretend to judge how much of the 2019 BIOS and GPU-scheduler guidance still maps cleanly onto current hardware and XCP-ng versions; others here will know where it's aged and where it hasn't.
      I'll make sure this is on our radar on the docs side, because it keeps coming up. Really appreciate you keeping this from disappearing. 👍

      Thank you kindly for your positive comments and appreciation. To me, it's amazing how some information can stay relevant for long periods of time, even given the rapid state of evolution in the technology sectors. I will try to get the full HTML docs uploaded soon, as well. There are a number of other XenServer articles I discovered a while back on a Polish server, and will see what else I can retrieve.
      My original avocation for 15 years was that of an astronomer, so research is in my blood and diving into specific issues and doing extensive testing have always been a big part of my motivation to better understand as well as share knowledge.

      posted in Hardware
      tjkreidlT
      tjkreidl
    • RE: Socket/core configuration in VM

      @robyt It depends on (1) licensing, if any, as some licenses go by cores vs. sockets, and (2) NUMA/VNUMA depending on how critical the performance is depending on how the VCPUs get allocated between sockets or on a single socket. Best way IMO is to try all and test with benchmarks. See, for example, this article and the previous two articles, as well as articles by Frank Denneman and others: https://blogs.mycugc.org/2019/04/30/a-tale-of-two-servers-part-3-the-influence-of-numa-cpus-and-sockets-cores-persocket-plus-other-vm-settings-on-apps-and-gpu-performance/

      posted in Compute
      tjkreidlT
      tjkreidl
    • RE: NUMA-impact - Xeon/Epyc - 1P vs 2P

      @olivierlambert said in NUMA-impact - Xeon/Epyc - 1P vs 2P:

      There is no universal answer (because it's mostly depending on your VM load and what do you expect). As usual, my advice is to keep it simple if you don't have a problem with it (ie: you are satisfied by the perf.). Even a default EPYC configuration will be likely always better than a Xeon one.

      After that, if you want to go deeper and learn the details, it's OK, let me just ping @tjkreidl who did a remarkable job (if I remember correctly) on this very topic.

      Thanks for the mention, @olivierlambert ! Here's a link to part 3, which contains links back to parts 1 and 2. Note that NUMA will affect EPYC processors differently as they changed the die configuration at one point with the number of cores. I'm open for any questions on this topic. 🙂 https://blogs.mycugc.org/2019/04/30/a-tale-of-two-servers-part-3-the-influence-of-numa-cpus-and-sockets-cores-persocket-plus-other-vm-settings-on-apps-and-gpu-performance/

      posted in Compute
      tjkreidlT
      tjkreidl
    • RE: vCPU Over-Subscription...

      @epretorious I would add that you have to be careful about overprovisioning when NUMA/vNUMA kicks in, that is when you allocate more VCPUs to exceed the number of physical CPUs of a bank of them as well as the associated physical memory (assume, for the sake of argument, you have two banks of physical CPUs and each has directly accessible to it one of two banks of memory) then things get inefficient because a CPU may need to go across to a different bank of memory to access data and there is additional overhead involved. See for example this article and the two preceding it:
      https://blogs.mycugc.org/2019/04/30/a-tale-of-two-servers-part-3-the-influence-of-numa-cpus-and-sockets-cores-persocket-plus-other-vm-settings-on-apps-and-gpu-performance/

      -=Tobias

      posted in Compute
      tjkreidlT
      tjkreidl
    • RE: Overprovisioning CPU + RAM?

      @MichaelCropper CPUs can be over-provisoned, but not memory. You can use DMC (dynamic memory control) to regulate how much memory a VM will actually use, but in total, you still cannot exceed the total amount of physical memory available on a server.

      CPU over-provisioning is very common, especially if loads change significantly over time (day/night weekday/weekend, special event and holidays/regular days, etc.).

      Watching the load with top and xentop will give you an idea about overall performance of dom0 and all VMs, respectively.

      As to a VM powred off, it will use up neither memory nor CPU resources.

      There are a lot of subtleties involved that would entail a much longer discussion, but hopefully this will help for starters. You can google a lot of information about memory and VCPU allocation; there is a lot of information out there.

      posted in Xen Orchestra
      tjkreidlT
      tjkreidl
    • RE: How to Re-attach an SR

      @olivierlambert Agreed. The Citrix forum used to be very active, but especially since Citrix was taken over, https://community.citrix.com has had way less activity, sadly.
      It's still gratifying that a lot of the functionality still is common to both platforms, although as XCP-ng evolves, there will be continually less commonality.

      posted in XCP-ng
      tjkreidlT
      tjkreidl
    • RE: How to Re-attach an SR

      @Chrome Cheers -- always glad to help out. I put in many thousands of posts on the old Citrix XenServer site, and am happy to share whatever knowledge I still have, as long as it's still relevant! In a few years, it probably won't be, so carpe diem!

      posted in XCP-ng
      tjkreidlT
      tjkreidl
    • RE: How to Re-attach an SR

      @Chrome Fantastic! Please mark my post as helpful if you found it as such. Was traveling much of today, hence the late response.

      BTW, it's always good to make a backup and/or archive of your LVM configuration anytime you change it, as the restore option is the cleanest way to deal with connectivity issues if there is some sort of corruption. It's saved my rear end before, I can assure you!

      Yeah, if the SSD drive got wiped, there's no option to get those back unless you made a backup somewhere of all that before you installed XCP-ng onto it.

      BTW, another very useful command for LVM is "vgchange -ay" which will attempt to renew VG information if a VG seems missing or the like.

      posted in XCP-ng
      tjkreidlT
      tjkreidl
    • RE: HA causes reboot of xcp-ng nodes

      @john.c Indeed, John, and I almost forgot that the backup network on each host was actually on a separate, isolated NIC, and not at all on the VLAN.

      posted in Management
      tjkreidlT
      tjkreidl
    • RE: HA causes reboot of xcp-ng nodes

      @john.c Keeping the various network traffic isolated according to specific usage (management, storage, VMs, etc.) is always a good idea. The last system I managed had 10GiB LACP bonds and using VLANs to isolate traffic and that worked fine with a four-node pool running around 80 or so XenDesktop instances per node. Never experienced any congestion issues. Each dom0 had a ton of memory and I believe it was either 8 or 16 VCPUs to make sure there were sufficient compute and memory allocations to allow for sometimes very heavy loads. It also helped that I eventually added GPUs to take on some of the computational load, in particular when some of the VMs were running applications employing heavy graphics.

      posted in Management
      tjkreidlT
      tjkreidl
    • RE: HA causes reboot of xcp-ng nodes

      @carloum70 One option to address the HA error messages: Delete the HA configuration and re-configure it from scratch once the bonds are all configured OK, which it appears you think they now are.

      posted in Management
      tjkreidlT
      tjkreidl
    • RE: HA causes reboot of xcp-ng nodes

      @john.c Most interesting, and I fully agree, that a strong host model enforcement policy is really the best and only recourse for such a topology, or so it would seem.
      The only other option that comes to mind would be to not put the heartbeat connection on any sort of bond or multipath. That's, of course, not ideal.

      posted in Management
      tjkreidlT
      tjkreidl
    • RE: HA causes reboot of xcp-ng nodes

      @john.c Interesting information about the strong v. weak host model.
      Active-active and LACP bonds will split the network traffic between the NICs, while an active-passive bond only uses the primary. WHat type of bonds do you have set up?

      posted in Management
      tjkreidlT
      tjkreidl
    • RE: V2V migrated VMs don't autostart

      @jr-m4 Hi there... A warm VMware V2V migration should indeed start up by itself, so I'm not sure what you are experiencing is normal. This is documented, as you've probably seen, at: https://xcp-ng.org/blog/2022/10/19/migrate-from-vmware-to-xcp-ng/

      That said, in addition, if you want your imported VMs to start up automatically every time your XCP-ng server or pool boots up, you have to enable the Auto Power On / Auto Start configuration.

      For this to work, it must be enabled at both the Pool level and the individual VM level.

      Method using Xen Orchestra (Recommended):
      Go to the VM tab, select your imported VM, and navigate to the Advanced view.
      Toggle the Auto power on switch to ON. (XO will handle configuring both the host pool and the VM flags for you).

      Or alternatively, via the bash shell:
      --> Get your Pool UUID
      xe pool-list

      --> Set the auto_poweron parameter to true
      xe pool-param-set uuid=<POOL_UUID> other-config:auto_poweron=true

      Enable Autostart on the Imported VM:
      --> Get your VM UUID
      xe vm-list name-label="Your_Imported_VM"

      --> Set the VM param to true
      xe vm-param-set uuid=<VM_UUID> other-config:auto_poweron=true

      posted in Migrate to XCP-ng
      tjkreidlT
      tjkreidl
    • RE: HA causes reboot of xcp-ng nodes

      @carloum70 Sorry, that option was deprecated.

      You may need to be more specific in the reset, something like the following with is just an example:
      xe-reset-networking -m IP_of_Master --device=eth0 --mode=static --ip=192.168.1.50 --netmask=255.255.255.0 --gateway=192.168.1.1 --dns=192.168.1.254

      Check the options and see which ones you actually need in your case.
      Worst case, you could possibly dissolve the bond, redo the PMI and then re-create the bond. WHy eth4 shows up is hard to guess. I have seen before that on some hosts in a pool that the NIC order was not the same, even though the hardware and OS versions were identical. In that case, you have to shut down the NICs and use interface-rename to reassign specific NIC names to the corresponding MAC addresses.

      posted in Management
      tjkreidlT
      tjkreidl
    • RE: HA causes reboot of xcp-ng nodes

      @carloum70 If you are willing to try this, you can reset the primary management interface (PMI) on a bond, WIthin the bash shell:
      :
      xe pif-list
      This will show the PIFs and their current network associations.

      Identify the bond master PIF
      If you have a bonded network, the master PIF is the one that represents the bond. You can list it with:

      xe pif-list network-uuid=<bond-uuid>
      The master PIF is the one that will be used as the primary interface for the bond.

      Reset the management interface to the bond master
      Use the xe-reset-networking command with the --reset-primary option:

      xe-reset-networking --reset-primary
      This will move the management interface to the bond master PIF, which is the correct way to reassign it when the bond is created or reconfigured.

      posted in Management
      tjkreidlT
      tjkreidl
    • RE: HA causes reboot of xcp-ng nodes

      @carloum70 Are all your hosts properly time sychronized to NTP or chronyc? Check each host for offsets. They need to be really close in time with each other. And what HA heartbeat client setup are you using? As long as you have a quorum, HA should continue to work fine. WHen you get down to two hosts and one cannot communicate with the other is when things get critical (unless running HA-Lizard on a two-host pool).

      posted in Management
      tjkreidlT
      tjkreidl
    • RE: GPU Passthrough

      @coolsport00 Sorry about the VMW need for the Cisco product. Sounds like you have a number of constraints, finances being I'm sure one of them! At least you have time on your hands and the means to experiment. You may, als, end up with a number of different platforms to meet your needs. We ran both Sun Microsystems and Red Hat Linux and Microsoft Windows servers, each taking on specific duties. It's far from ideal and probably not very cost-effective, but you do what you have to to get stuff to work.

      posted in Management
      tjkreidlT
      tjkreidl