• XCP-ng 8.3 updates announcements and testing

    Pinned News
    699
    1 Votes
    699 Posts
    714k Views
    bvitnikB
    @gleh said: Disable Viridian in the "Other install media" template by default. As a reminder, Windows VMs must use Windows templates. What is the rationale for disabling Viridian? Does it have any negative effects for non Windows VMs?
  • 1 Votes
    6 Posts
    153 Views
    dthenotD
    @racom.jiristerba Sorry, I missed answering your questions the first time Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage? Yes, it's a known issue, most of them will be fixed with the latest update, you might need to activate the VDI again but a migration during RPU will be enough Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open? It's because in the case of VHD, it's the case but with QCOW2 we don't stop the tapdisk process accessing the VDI, it's a change that was missed Is the qcow2OLD_<UUID> naming seen in rollback expected, compared with QCOW2-OLD_<UUID> seen during successful cleanup? It's a pre-existing bug that I already have on my TODO list Is there a supported update, hotfix, or workaround for this configuration? Installing the latest release sm-3.2.12-25.1 during updates (and maybe launching a xe sr-scan uuid=<SR UUI> after updating it so it auto-resolve the undo) should be enough What additional logs are required to identify the original failure of the second QCOW2 VM? I don't think we need any more logs since the errors I'm seeing should already be fixed. If you have any more issues after installing the newest packages, I will take another look What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups? The sr-scan after updating should do it automatically, it shouldn't need any manipulation. Looking at the storage logs in /var/log/SMlog for any irregularities could help to see problems.
  • Migrating an offline VM disk between two local SRs is slow

    Unsolved Xen Orchestra
    35
    1
    0 Votes
    35 Posts
    9k Views
    A
    Data point from a 2-host XCP-ng 8.3 pool that supports @tosh's diagnosis, plus a no-patch workaround that gave us ~2.7x until the TCP_NODELAY fix lands. Setup XCP-ng 8.3, xapi-core / vhd-tool 26.1.16-1.2.xcpng8.3, dom0 kernel 4.19.0+1 Two hosts, local ext SRs on NVMe (Samsung 970 EVO Plus / WD Red SN700) Dedicated migration network: 2.5 GbE (Intel i226-V, igc), MTU 9000, no errors Offline storage migration (halted VM, VM.migrate_send) of 2 VDIs (disk + snapshot), 16 GiB virtual / 7773 MiB allocated each; sparse_dd ... -prezeroed ... -dest-proto nbd over TLS Baseline (stock) 132 s per VDI, i.e. ~59 MiB/s on the wire, flat line, same in both directions Link at ~20 % of line rate, source SR read latency 0.3 ms, dom0 total <= 0.5 vCPU, no single dom0 vCPU above 0.08 (RRD, 60 s averages) Sending host with turbo or capped at 2.5 GHz gave the same 132 s, so not CPU-bound here Workaround: disable delayed ACKs on the receiver, only for the migration-network route # on the destination host; adapt prefix/dev/src to the output of: ip route show <migration-net> ip route change 10.10.10.0/24 dev xenbr1 proto kernel scope link src 10.10.10.1 quickack 1 With the receiver ACKing immediately, Nagle on the sender only waits one RTT instead of the ~40 ms delayed-ACK timer. Result (same VM, same direction, same data) stock quickack 1 on receiver per VDI 132 s / 132 s 50 s / 46 s throughput ~59 MiB/s ~155-169 MiB/s whole migration 4:50 2:00 dom0 on the receiver went up to ~1.1 vCPU total, still nothing saturated. Caveats One run per arm, one pool. I did not capture ss -ti during the runs, so the notsent:512 signature is inferred, not observed here. It does not address the reply gating @TeddyAstie described, it only removes the delayed-ACK half. Consistent with that, we land at ~2.7x, below the ~4.3x reported above for TCP_NODELAY (different rig, so not directly comparable). Not persistent: lost on reboot and when xapi re-plugs the PIF. We re-apply it from a small systemd unit at boot as a stopgap and will remove it once vhd-tool ships TCP_NODELAY. Scope is only connections routed via the migration network (storage and live migration). More ACK packets on that link, nothing else changes. Is there a PR or issue for the TCP_NODELAY patch in xapi-project/xen-api that we can follow? I could not find one.
  • 0 Votes
    1 Posts
    16 Views
    No one has replied
  • Synchronize snapshots

    Backup
    5
    2
    0 Votes
    5 Posts
    210 Views
    acebmxerA
    Latest commit fixed the synchronized snapshots showing up as a vm... Commit 8c2f3 [image: Screenshot_20261005_060854.png]
  • Change Pool Master does not update connection string in XOA

    Unsolved Management
    3
    2
    0 Votes
    3 Posts
    48 Views
    ForzaF
    @poddingue I'm using XOA 6.2.2. Unfortunately not able to update at the moment. Which reminds me of another thing. Where is there release information for the XOA XVA itself? When I login to my account, there is only the download button but no information on what version it is. [image: image.jpeg] I would like to know when a new version is published and a changelog for it. When I import the current XVA it is tagged with 2025.12 [image: image.jpeg]
  • Pool metadata backup failed after xoa upgrade

    Unsolved Backup
    4
    0 Votes
    4 Posts
    100 Views
    J
    @poddingue said: There's a similar report on 6.8 in https://xcp-ng.org/forum/topic/12453, where @flakpyro said it got fixed through a support ticket and that @florent would know the exact fix. I read through the 6.9.0 changelog, but I couldn't find an entry about metadata backups or Body Timeout Error, so I can't confirm that moving to the latest channel fixes it, though it may have gone in without a changelog line. If you have a support contract, a ticket that points at that thread might be the quickest way to the patch Florent made. It is in the 6.9.0 change log but not where you might expect, even as a backup bug fix. In this case it’s filed under the miscellaneous (misc) section as “Fixed: BodyTimeoutError during long transfers”. Given where it is in the change log you may have missed it. If it isn’t in the change log then why is it in the release announcement?
  • XCP-ng Shutdown Hangs When an SR Is Unavailable

    Unsolved XCP-ng
    3
    0 Votes
    3 Posts
    62 Views
    poddingueP
    This came up in https://xcp-ng.org/forum/topic/11728 last winter, with the same kind of setup, an ISO SR served by a VM on the host. Olivier's answer there was to eject the CDs from the VMs, unplug the ISO SR's PBD, and only then shut the host down, using xe vm-cd-eject --multiple and xe pbd-unplug uuid=<PBD UUID> for the first two steps (the PBD one is in the CLI reference: https://docs.xcp-ng.org/appendix/cli_reference#pbd-unplug). @mickwilli then added the unplug to his UPS shutdown script, and his host shut down cleanly after that. I haven't tried it with an SMB share myself, so I can't promise it behaves like his NFS one, but your NUT script on the Pi looks like the natural place for those two steps, just before the shutdown command.
  • Cleanup orphans backups

    Unsolved Backup
    10
    0 Votes
    10 Posts
    487 Views
    P
    @johnnezero You can use the docs directly in /rest/v0/docs. You just have to connect and fill in the uuid you need. It will work
  • is Xo Proxy available in community version

    Unsolved Xen Orchestra
    14
    0 Votes
    14 Posts
    3k Views
    B
    Just pinging this thread to keep it alive as an answer is still very much of interest.
  • XOA Unable to connect xo server every 30s

    Unsolved Xen Orchestra
    7
    0 Votes
    7 Posts
    1k Views
    J
    @GregBinSD said: Here are two more notes regarding the XO6 "Unable to connect to XO server. Retry" message, which pops up after 30 seconds. It occurs when either the Chrome or the Microsoft Edge browsers are used on my Windows 11 PC. However, I often use a Samsung Tab-A9 (tablet), and it does not have this issue with XO6. It uses the Chrome browser. To enlighten you the Microsoft Edge your talking about is not the original release (from Windows 10). It’s the Chromium based release from during Windows 10 and has been that one ever since. The original release of Microsoft Edge had its own rendering engine called MSHTML. The current modern Edge effectively shares a common upstream code base with Google Chrome, namely Chromium. The Samsung Tab-A9 doesn’t have the issue even though it, uses the same browser namely Google Chrome. This is the case because the tablet uses a version of Google Android, which has its own kernel, which is a fork or variation of the Linux Kernel.
  • Remote desktop on Gnome hangs randomly

    Unsolved Hardware
    24
    0 Votes
    24 Posts
    2k Views
    O
    @dinhngtu i've manage to fixed it. Your build adds some extra entries in /var/lib/xcp/state.db that aren't in the version from testing, that broke things when i downgraded the packages. ipv4_dns VIF.ipv6_dns
  • Kioxia CM7 PCIe pass-through crash

    Unsolved Compute
    12
    0 Votes
    12 Posts
    2k Views
    D
    Hello, this issue has been fixed in XAPI 26.1.19-1.1 available in the testing updates. For now you have to configure your VMs manually (https://docs.xcp-ng.org/troubleshooting/common-problems/#pci-passthrough-error-unsupported-msi-delivery-mode-7) but the fix will be activated by default in the future.
  • [SOLVED] Just FYI: current update seams to break NUT dependancies

    Solved XCP-ng
    38
    0 Votes
    38 Posts
    9k Views
    K
    @FritzGerald Thanks! Running systemctl enable --now ups-driver.service got ups-driver.service going, then following the instructions you provided earlier got the rest of the two services running, even after a reboot. I'm now able to run upsc apc@localhost and get the correct output listing all the properties of the UPS even after a reboot.
  • Internet connectivity - Check XOA failed.

    Unsolved Management
    8
    1 Votes
    8 Posts
    195 Views
    Z
    @poddingue said: Thanks for posting the fix. The same thing came up in an older thread, https://xcp-ng.org/forum/topic/9957, where disabling IPv6 also made xoa check go green. @HamiltonWDS explained a likely reason in https://xcp-ng.org/forum/post/87831: Node tries the addresses with a short timeout and can end up on the IPv6 one. Your curl -6 test helps a lot here, because it shows IPv6 doesn't connect at all on that network, so it's more than a slow path. Turning it off in XOA seems reasonable if you don't use IPv6 there; if you do, I'd guess the router side is where it really needs fixing, though I could be wrong (still haven't migrated to IPv6 myself ). Curious whether it sorts out the Cloud Backup too. XOA is using our management network where ipv6 is not needed and not enabled. Looks like our cloud backup finally ran for the first time in while too.[image: cloudbackup.jpg]
  • [V2V] Without VDDK problems

    Solved Migrate to XCP-ng
    4
    0 Votes
    4 Posts
    90 Views
    J
    @mpiton Thanks for a fantasticly quick answer, and detailed as well. I will test this immediately! Update: I can confirm that this was indeed the problem!
  • XO Tasks - backups just pilling up

    Unsolved Management
    5
    1
    0 Votes
    5 Posts
    200 Views
    poddingueP
    I didn't know, so I went and looked at the code. XO prunes that list in a cleanup step: backup runs are kept for 31 days, and other tasks are kept by count, the newest 1000 (https://github.com/vatesfr/xen-orchestra/blob/b59c8f3d20954aa1d72e92c1e8f2db635d74fcba/@xen-orchestra/mixins/Tasks.mjs#L119-L127). You can change both with tasks.gc.keep and tasks.gc.backupKeepDuration in the xo-server config. As far as I can tell the cleanup runs when xo-server starts, not on a timer, so a long-running XO can collect quite a few before the next restart. My guess for the Cloud Backup entries @acebmxer still sees after 30 days is that they're kept by count rather than by age.
  • Unclear errors

    Unsolved Management
    4
    1
    0 Votes
    4 Posts
    126 Views
    poddingueP
    I tried it on a lab host (XCP-ng 8.3, xapi 26.1.16): calling VM.pool_migrate on a halted VM gives me your error, VM_BAD_POWER_STATE with running, halted, so the two values are the state XAPI wanted and the one it found. On the XO side, a migration inside the same pool with no SR or network mapping goes straight to VM.pool_migrate without checking the power state first: https://github.com/vatesfr/xen-orchestra/blob/b59c8f3d20954aa1d72e92c1e8f2db635d74fcba/packages/xo-server/src/xapi/index.mjs#L942-L981. So a stopped VM in that batch would give this, which fits what you remember. I haven't tried the batch migrate from the UI itself though, so I don't know if the UI is supposed to filter halted VMs out before that. It's a good concrete example for the feedback post anyway, XO could just say "this VM is halted" instead of passing XAPI's raw code through.
  • PCIe Pass-through lanes and lane performance

    Unsolved Compute
    57
    0 Votes
    57 Posts
    9k Views
    pandusenP
    @stormi I am glad to hear there is progress, and I never doubted that there was interest and will from the Dev team, it was just stated, by Oliver himself, that the focus was mostly on the classic hypervisor features and that resources for edge cases were scarce. The numbers might not be spot on, but it was something similar.. Anyways I know you care and that the team always take interest in home-lab scenarios. And that, is why I keep recommending XCP-ng. You are good and motivated people, you care and interact with the community and you have a product that is simple in structure and free of most of the technical bureaucracy that the other ecosystems have. And yes, this particular Intel Arc issue, is not only on XCP-ng, The intel toolbox and drivers for this platform is currently a mess for Linux. And you are right to take home-labbers under your wing and prioritize it. As I'm sure you know, although it might not be of direct importance for revenue, they do have a huge impact on popularity. I dont think Proxmox would have had the traction it has, without them Thank you for replying, and thank you to the team for their hard work!
  • 0 Votes
    20 Posts
    2k Views
    M
    @anthoineb, Thanks for the fix! We plan to install blktap-3.55.5-11.1.xcpng8.3 on all three hypervisors tomorrow. We will fully shut down and start the data VMs one at a time, then verify that all running tapdisk processes use the updated binary. As the stalls are intermittent, we will monitor for recurrence and report back with the results.