XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    XCP-ng 8.3 updates announcements and testing

    Scheduled Pinned Locked Moved News
    683 Posts 57 Posters 694.8k Views 76 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • X
      XCP-ng-JustGreat
      last edited by

      I deferred applying the preview patches this time in order to try my luck again with RPU. Unfortunately, it did not work for me again as it has not in the past. My results are similar to others here: primary host patch application went fine including the reboot. However, once it began doing a secondary host, it started throwing errors e.g. CANNOT_EVACUATE_HOST and VM_REQUIRES_SR etc. However, manually putting the host in maintenance mode from the GUI evacuated each host just fine and I was able to apply the patches and reboot each subsequent host from the XO GUI. Also, as with others here, the RPU task hung in the task list and even a reboot of the XO VM would not clear it forcing me to delete the task using xo-cli e.g. xo-cli rest del tasks.
      ENVIRONMENT: Home lab consisting of 4 x Dell OptiPlex 7040 i7-6700 SFF hosts, 48GB RAM each, 10 Gbps storage connections to a TrueNAS home-built NAS via NFS and XO from source (XOS) using @ronivay build script on AlmaLinux 10.2 minimal install VM with XO commit 6a441 compiled on 2026-08-28 from master branch. FWIW, RPU functionality remains unavailable to me, though obviously, this is not a showstopper for a 4 x host home lab pool. If there is anything I can do to help resolve this, please let me know as I remain an enthusiastic proponent of the Vates virtualization stack.

      gduperreyG 1 Reply Last reply
      Reply Quote 0
      • gduperreyG
        gduperrey Vates πŸͺ XCP-ng Team @acebmxer
        last edited by

        Hello @acebmxer,

        I can see from the support ticket and the summary that the issue appears to be resolved and was linked to the NFS update on the Synology. It’s great that you were able to find the solution.

        1 Reply Last reply
        Reply Quote 0
        • gduperreyG
          gduperrey Vates πŸͺ XCP-ng Team @XCP-ng-JustGreat
          last edited by

          Hello @XCP-ng-JustGreat,

          I cannot comment on RPU issues via XO. I would recommend creating a new topic to report the problem directly to the XO team so they can analyze it.

          As you can see, we recommend validating updates via the command line directly on XCP-ng before releasing them. That is a different scenario.

          However, if several of you are experiencing RPU issues with XO, a dedicated topic will allow that team to investigate the situation and consult other teams if necessary. This would ensure the issue is better addressed and analyzed πŸ™‚

          M 1 Reply Last reply
          Reply Quote 0
          • M
            manilx @gduperrey
            last edited by

            @gduperrey https://xcp-ng.org/forum/topic/12439/rpu-issue

            1 Reply Last reply
            Reply Quote 1
            • stormiS
              stormi Vates πŸͺ XCP-ng Team @mthird
              last edited by

              @mthird said:

              Spoke to soon. While the updates succeeded, one of the nodes is rebooting every few minutes due to an HA self-fence.

              Could you open a dedicated thread and ping me there?

              1 Reply Last reply
              Reply Quote 0
              • gduperreyG
                gduperrey Vates πŸͺ XCP-ng Team
                last edited by

                We have just released security updates for xen and blktap. Full details are available on the blog: https://xcp-ng.org/blog/2026/09/08/september-2026-security-updates-1-for-xcp-ng-8-3-lts/

                B 1 Reply Last reply
                Reply Quote 2
                • B
                  bufanda @gduperrey
                  last edited by

                  @gduperrey Installed on all pools. No issues so far.

                  1 Reply Last reply
                  Reply Quote 1
                  • leroyjL
                    leroyj @TeddyAstie
                    last edited by

                    Hi @TeddyAstie,
                    I use a Ryzen 7 5875u CPU (Zen3) shipped in 2022.
                    The sensor/K10temp approach doesn't work (nothing detected).
                    AFAIK Zen3 temp support was introduced around the. kernel 5.14 late 2021.
                    If it's not supported yet, I guess that it won't, right?
                    Is there another solution to monitor temp except to wait for a future xcp-ng v9 version?

                    I was wondering if the amd cpu temp be probed the way you did for the intel cpu.
                    (Disclamer the following patch is a first draft AI generated because I don't have the skills. I share it because even I can spot it's not perfect as is, it doesn't look bloated)

                    diff --git a/xen/arch/x86/include/asm/msr-index.h b/xen/arch/x86/include/asm/msr-index.h
                    index df52587c85..9d8f7bd4d1 100644
                    --- a/xen/arch/x86/include/asm/msr-index.h
                    +++ b/xen/arch/x86/include/asm/msr-index.h
                    @@ -115,6 +115,9 @@
                     #define MCU_OPT_CTRL_GDS_MIT_DIS       (_AC(1, ULL) << 4)
                     #define MCU_OPT_CTRL_GDS_MIT_LOCK      (_AC(1, ULL) << 5)
                     
                    +/* AMD Hardware Thermal Control */
                    +#define MSR_AMD_HARDWARE_THERMAL_CONTROL 0xc0010292
                    +
                     #define MSR_FRED_RSP_SL0               0x000001cc
                     #define MSR_FRED_RSP_SL1               0x000001cd
                     #define MSR_FRED_RSP_SL2               0x000001ce
                    diff --git a/xen/arch/x86/platform_hypercall.c b/xen/arch/x86/platform_hypercall.c
                    index 79bb99e0b6..e84a2ddd3e 100644
                    --- a/xen/arch/x86/platform_hypercall.c
                    +++ b/xen/arch/x86/platform_hypercall.c
                    @@ -86,6 +86,13 @@ static bool msr_read_allowed(unsigned int msr)
                     
                         case MSR_MCU_OPT_CTRL:
                             return cpu_has_srbds_ctrl;
                    +
                    +    /* Intel MSRs (From Vates Patch) */
                    +    case MSR_IA32_THERM_STATUS:
                    +    case MSR_IA32_TEMPERATURE_TARGET:
                    +        return boot_cpu_data.x86_vendor == X86_VENDOR_INTEL;
                    +        
                    +    /* AMD MSR */
                    +    case MSR_AMD_HARDWARE_THERMAL_CONTROL:
                    +        return boot_cpu_data.x86_vendor == X86_VENDOR_AMD;
                         }
                     
                         if ( ppin_msr && msr == ppin_msr )
                    diff --git a/tools/misc/xenpm.c b/tools/misc/xenpm.c
                    index 3a228c2..f9b8c71 100644
                    --- a/tools/misc/xenpm.c
                    +++ b/tools/misc/xenpm.c
                    @@ -1553,6 +1554,81 @@ void get_core_temp(int argc, char *argv[])
                     {
                         xc_physinfo_t physinfo = { 0 };
                         int i, max_cpu_nr;
                    +    bool is_amd = false;
                    +    bool is_intel = false;
                         
                         if ( argc > 0 )
                         {
                    -        fprintf(stderr, "Usage: xenpm get-core-temp\n");
                    +        fprintf(stderr, "Usage: xenpm get-core-temp\n"
                    +                        "       (Supports Intel via TjMax and AMD via HTC)\n");
                             exit(EINVAL);
                         }
                     
                         if ( xc_physinfo(xc_handle, &physinfo) )
                         {
                             fprintf(stderr, "Failed to get xc_physinfo. Err: %s\n", strerror(errno));
                             exit(EINVAL);
                         }
                         max_cpu_nr = physinfo.max_cpu_id + 1;
                     
                    +    /* Probe CPU 0 to determine vendor based on allowed MSRs */
                    +    xc_resource_op_t probe_intel = {
                    +        .cpu = 0,
                    +        .cmd = XEN_RESOURCE_OP_read_msr,
                    +        .u.msr = { .msr = 0x19c } /* MSR_IA32_THERM_STATUS */
                    +    };
                    +    xc_resource_op_t probe_amd = {
                    +        .cpu = 0,
                    +        .cmd = XEN_RESOURCE_OP_read_msr,
                    +        .u.msr = { .msr = 0xc0010292 } /* MSR_AMD_HARDWARE_THERMAL_CONTROL */
                    +    };
                    +
                    +    if ( xc_resource_op(xc_handle, 1, &probe_intel) == 0 )
                    +        is_intel = true;
                    +    else if ( xc_resource_op(xc_handle, 1, &probe_amd) == 0 )
                    +        is_amd = true;
                    +    else
                    +    {
                    +        fprintf(stderr, "Error: Current CPU architecture is not supported for temperature reading.\n");
                    +        exit(ENOTSUP);
                    +    }
                    +
                         printf("CPU\tCurrent Temperature\n");
                         for ( i = 0; i < max_cpu_nr; i++ )
                         {
                    -        /* 
                    -         * ... Existing Vates Intel logic ... 
                    -         * (Reading MSR_IA32_TEMPERATURE_TARGET and MSR_IA32_THERM_STATUS)
                    -         */
                    +        if ( is_intel )
                    +        {
                    +            /* --- INTEL LOGIC (Vates implementation) --- */
                    +            xc_resource_op_t op_tjmax = { .cpu = i, .cmd = XEN_RESOURCE_OP_read_msr, .u.msr = { .msr = 0x1a2 } };
                    +            xc_resource_op_t op_therm = { .cpu = i, .cmd = XEN_RESOURCE_OP_read_msr, .u.msr = { .msr = 0x19c } };
                    +            
                    +            if ( xc_resource_op(xc_handle, 1, &op_tjmax) || xc_resource_op(xc_handle, 1, &op_therm) )
                    +                continue;
                    +                
                    +            uint32_t tjmax = (op_tjmax.u.msr.value >> 16) & 0xFF;
                    +            uint32_t readout = (op_therm.u.msr.value >> 16) & 0x7F;
                    +            
                    +            printf("%d\t%d C\n", i, tjmax - readout);
                    +        }
                    +        else if ( is_amd )
                    +        {
                    +            /* --- AMD LOGIC --- */
                    +            xc_resource_op_t op_amd = {
                    +                .cpu = i,
                    +                .cmd = XEN_RESOURCE_OP_read_msr,
                    +                .u.msr = { .msr = 0xc0010292 }
                    +            };
                    +
                    +            if ( xc_resource_op(xc_handle, 1, &op_amd) )
                    +                continue; /* Skip offline/unavailable CPUs */
                    +
                    +            uint64_t msr_val = op_amd.u.msr.value;
                    +            
                    +            /* Extract CurTmp (bits [31:21]) and multiply by 0.125Β°C */
                    +            uint32_t cur_tmp_raw = (msr_val >> 21) & 0x7FF;
                    +            float temp_c = cur_tmp_raw * 0.125;
                    +
                    +            printf("%d\t%.2f C\n", i, temp_c);
                    +        }
                         }
                     }
                    
                    TeddyAstieT 1 Reply Last reply
                    Reply Quote 0
                    • TeddyAstieT
                      TeddyAstie Vates πŸͺ XCP-ng Team Xen Guru @leroyj
                      last edited by TeddyAstie

                      @leroyj
                      We already know that our k10temp module doesn't have support beyond Zen 2, and needs to be updated. That's not incredibly complex, but still needs to be done.

                      I was wondering if the amd cpu temp be probed the way you did for the intel cpu.

                      No, AMD temperature infos are not exposed through MSR but through PCIe/MMIO (through various subsystems, like "SMU" and other ones). Like what does https://github.com/torvalds/linux/blob/master/drivers/hwmon/k10temp.c or https://github.com/ocerman/zenpower.
                      (regarding the AI draft, there is no documented MSR 0xc0010292 in neither APM nor public Zen3 PPM)

                      That doesn't require any specific Xen support aside the right Linux drivers in Dom0.

                      1 Reply Last reply
                      Reply Quote 0
                      • glehG
                        gleh Vates πŸͺ XCP-ng Team
                        last edited by

                        New maintenance update candidates for XCP-ng 8.3 LTS

                        This batch of updates contains performance improvements, fixes, a guest tools update, and other improvements.

                        What changed

                        Virtualization & System

                        • kernel: Fix potential deadlocks in the NFS subsystem.

                        • xen:

                          • Enable Viridian enlightenments to prepare for support of Windows VMs with >64 vCPUs
                          • Sync with XenServer release 4.17.7-2. It fixes XSA-511, which didn't affect XCP-ng.
                        • systemd: Fix detection of NVMe controllers with multiple namespaces

                        Control Plane

                        • xapi:

                          • Improve RPU stability by only migrating VMs to already-updated hosts
                          • Improve VDI migration performance by pipelining NBD writes
                        • guest-templates-json:

                          • Sync with XenServer's guest-templates-json-2.0.16-1. It removes the Kylin Linux 7 Guest template.
                          • Disable Viridian in the "Other install media" template by default. As a reminder, Windows VMs must use Windows templates.

                        Storage

                        • blktap:

                          1. A fix for a tapdisk crash happening when tapdisk pauses a VDI when a coalesce is in progress. This issue was caused by a use-after-free bug. We took time to thoroughly audit the code around pause/commit/cancel operations, and thus robustify the pause and commit code in this release.
                          2. Customers could get tapdisk with infinite stalled IO on QCOW2 VDI. A race-condition was identified that could allow to miss a notification from the front-end and therefore never check the blkif ring for new requests.
                          3. On a tapdisk running on a supporter node, CBT metadata does not flush before a pause operation. During the pause operation, the master host modifies the CBT metadata, then after the unpause the old metadata are flushed and discard master's modifications. Finally the CBT logs are detected as inconsistent and discarded.
                          4. Grow up the QCOW2 metadata cache to help on big disks.
                        • sm:

                          • Robustify iSCSI calls against race condition and bad access using lock context manager.

                          • Retry host key tag removal when XAPI is unreachable or restarted to prevent errors and finish action.

                          • Explicitly notify a missing LINSTOR resource definition instead of displaying a human-unreadable trace.

                          • Repair LINSTOR RAW volume creation. Regression caused by QCOW2 code refactoring.

                          • Fix LINSTOR SR creation when a VG is missing on a host. Changes in behavior brought by an update of LINSTOR RPMs broke this feature, which prevented the creation of an SR in this situation. We're expanding the test suite to prevent this in the future.

                          • To help recovering from an eventual LINSTOR database corruption, it's now backed up regularly and after every major operation, locally on the master, and on the DRBD LINSTOR database volume.

                          • Fixed an issue on shared LVM-based SR (LVMoISCSI, LVMoHBA) where a failed live leaf coalesce of a QCOW2 image on a slave host caused the SRs to remain in an SR_FAILURE_1200 error state until the VM using the VDI was stopped or tapdisk was manually paused.

                          • Reduce qemu-img log verbosity in SMlog.

                          • Fix LVM metadata corruption on slave when CBT log deletion is triggered on VDI activation.

                          • Prevent an incorrect online coalesce to be reported as successful and causing issues with the QCOW2 chain.

                          • Add QCOW2 support during interrupted leaf-coalesce operations on file-based SRs.

                        Network

                        • openvswitch: The openvswitch package no longer installs the unused openvswitch-cfg-update XAPI plugin file, which has been superseded by openvswitch-config-update from xapi-core. This is only meant to avoid confusion and brings no functional change.

                        UI

                        • xo-lite:
                          • [Treeview] Add VM tree actions (PR #10304)
                          • [Pool/networks] Add the possibility to create new network or bonded network (PR #10145)
                          • [VM] Add VDIs page with table and side panel (PR #10269)
                          • [Host/Network] Add ability to rescan physical network interfaces (PIFs) (PR #10147)
                          • Fix inconsistent spacing in side panel cards (PR #10279)
                          • Fix missing collapse buttons on pools and hosts in the tree sidebar (added hasChildren prop) (PR #10244)
                          • [Pool/networks] Add the possibility to create new internal network (PR #10235)

                        Drivers and Middleware

                        • vendor-drivers: The out of tree microsemi-aacraid driver from the RPM has proven to be unreliable. It has several issues, is missing patches from upstream and hasn't been updated for a while. So, we switch back to the in-tree driver just like XCP-ng 8.2 was doing.

                        • xcp-ng-pv-tools:

                          • Update to XCP-ng Windows Guest Tools 9.2.385.
                          • Include the XenTools-fix-9.2.351.msp hotfix tool for users running 9.2.350.

                        Others

                        • openssl: Update to 3.5.5

                        Versions

                        • blktap: 3.55.5-9.5.xcpng8.3 -> 3.55.5-11.1.xcpng8.3
                        • gpumon: 24.1.0-96.1.xcpng8.3 -> 24.1.0-98.1.xcpng8.3
                        • guest-templates-json: 2.0.15-1.1.xcpng8.3 -> 2.0.16-1.1.xcpng8.3
                        • kernel: 4.19.19-8.0.46.10.xcpng8.3 -> 4.19.19-8.0.46.12.xcpng8.3
                        • openssl: 3.0.9-2.0.1.3.xcpng8.3 -> 3.5.5-1.3.xcpng8.3
                        • openvswitch: 2.17.7-4.1.xcpng8.3 -> 2.17.7-4.2.xcpng8.3
                        • sm: 3.2.12-23.5.xcpng8.3 -> 3.2.12-25.1.xcpng8.3
                        • systemd: 219-57.5.xcpng8.3 -> 219-57.5.1.xcpng8.3
                        • vendor-drivers: 2.0.3-1.1.xcpng8.3 -> 2.0.3-1.2.xcpng8.3
                        • xapi: 26.1.16-1.2.xcpng8.3 -> 26.1.19-1.1.xcpng8.3
                        • xcp-featured: 1.2.1-4.xcpng8.3 -> 1.2.1-5.xcpng8.3
                        • xcp-ng-pv-tools: 8.3-18.xcpng8.3 -> 8.3-19.xcpng8.3
                        • xen: 4.17.6-12.3.xcpng8.3 -> 4.17.7-2.3.xcpng8.3
                        • xo-lite: 0.24.0-1.xcpng8.3 -> 0.25.0-1.xcpng8.3

                        Test on XCP-ng 8.3

                        yum clean metadata --enablerepo=xcp-ng-testing,xcp-ng-candidates
                        yum update --enablerepo=xcp-ng-testing,xcp-ng-candidates
                        reboot
                        

                        The usual update rules apply: pool coordinator first, etc.

                        What to test

                        As usual, normal use and anything else you want to test.

                        Test window before official release of the updates

                        4 days

                        We would like to thank users who shared feedback since our last call for testing:
                        @bufanda, @acebmxer, @MajorP93, @flakpyro, @Andrew, @mthird, @manilx, @XCP-ng-JustGreat

                        1 Reply Last reply
                        Reply Quote 1

                        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                        With your input, this post could be even better πŸ’—

                        Register Login
                        • First post
                          Last post