• Future Architecture?

    News
    10
    1 Votes
    10 Posts
    361 Views
    J
    @herhin2017 said: @john.c Your absolute wright, i do my best and train about 25 young engineers per year on linux (we use debian). I also see more and more startups or small companies building their infrastructure on linux. Best greetings from austria It may be worth getting in line and noting the mention of EFI based servers for official Vates support with XCP-ng version 9.0 and above, in government reports. That way when the school’s hardware is refreshed it can be ensured that your provided with EFI capable servers, in time for XCP-ng version 8.3 EOL. I personally are already ready for XCP-ng version 9.0 due to my servers being Dell PowerEdge R620 for XCP-ng hosts. Along with Debian version 13.6 on the VMs. While using a Dell Precision 3590 to manage those systems. The “wright” makes your above now sound like a wheel wright, and its profession instead of “right” in for when correct or agreeing with someone or something.
  • 1 Votes
    6 Posts
    263 Views
    dthenotD
    @racom.jiristerba Sorry, I missed answering your questions the first time Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage? Yes, it's a known issue, most of them will be fixed with the latest update, you might need to activate the VDI again but a migration during RPU will be enough Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open? It's because in the case of VHD, it's the case but with QCOW2 we don't stop the tapdisk process accessing the VDI, it's a change that was missed Is the qcow2OLD_<UUID> naming seen in rollback expected, compared with QCOW2-OLD_<UUID> seen during successful cleanup? It's a pre-existing bug that I already have on my TODO list Is there a supported update, hotfix, or workaround for this configuration? Installing the latest release sm-3.2.12-25.1 during updates (and maybe launching a xe sr-scan uuid=<SR UUI> after updating it so it auto-resolve the undo) should be enough What additional logs are required to identify the original failure of the second QCOW2 VM? I don't think we need any more logs since the errors I'm seeing should already be fixed. If you have any more issues after installing the newest packages, I will take another look What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups? The sr-scan after updating should do it automatically, it shouldn't need any manipulation. Looking at the storage logs in /var/log/SMlog for any irregularities could help to see problems.
  • Migrating an offline VM disk between two local SRs is slow

    Unsolved Xen Orchestra
    35
    1
    0 Votes
    35 Posts
    10k Views
    A
    Data point from a 2-host XCP-ng 8.3 pool that supports @tosh's diagnosis, plus a no-patch workaround that gave us ~2.7x until the TCP_NODELAY fix lands. Setup XCP-ng 8.3, xapi-core / vhd-tool 26.1.16-1.2.xcpng8.3, dom0 kernel 4.19.0+1 Two hosts, local ext SRs on NVMe (Samsung 970 EVO Plus / WD Red SN700) Dedicated migration network: 2.5 GbE (Intel i226-V, igc), MTU 9000, no errors Offline storage migration (halted VM, VM.migrate_send) of 2 VDIs (disk + snapshot), 16 GiB virtual / 7773 MiB allocated each; sparse_dd ... -prezeroed ... -dest-proto nbd over TLS Baseline (stock) 132 s per VDI, i.e. ~59 MiB/s on the wire, flat line, same in both directions Link at ~20 % of line rate, source SR read latency 0.3 ms, dom0 total <= 0.5 vCPU, no single dom0 vCPU above 0.08 (RRD, 60 s averages) Sending host with turbo or capped at 2.5 GHz gave the same 132 s, so not CPU-bound here Workaround: disable delayed ACKs on the receiver, only for the migration-network route # on the destination host; adapt prefix/dev/src to the output of: ip route show <migration-net> ip route change 10.10.10.0/24 dev xenbr1 proto kernel scope link src 10.10.10.1 quickack 1 With the receiver ACKing immediately, Nagle on the sender only waits one RTT instead of the ~40 ms delayed-ACK timer. Result (same VM, same direction, same data) stock quickack 1 on receiver per VDI 132 s / 132 s 50 s / 46 s throughput ~59 MiB/s ~155-169 MiB/s whole migration 4:50 2:00 dom0 on the receiver went up to ~1.1 vCPU total, still nothing saturated. Caveats One run per arm, one pool. I did not capture ss -ti during the runs, so the notsent:512 signature is inferred, not observed here. It does not address the reply gating @TeddyAstie described, it only removes the delayed-ACK half. Consistent with that, we land at ~2.7x, below the ~4.3x reported above for TCP_NODELAY (different rig, so not directly comparable). Not persistent: lost on reboot and when xapi re-plugs the PIF. We re-apply it from a small systemd unit at boot as a stopgap and will remove it once vhd-tool ships TCP_NODELAY. Scope is only connections routed via the migration network (storage and live migration). More ACK packets on that link, nothing else changes. Is there a PR or issue for the TCP_NODELAY patch in xapi-project/xen-api that we can follow? I could not find one.
  • Cleanup orphans backups

    Unsolved Backup
    10
    0 Votes
    10 Posts
    609 Views
    P
    @johnnezero You can use the docs directly in /rest/v0/docs. You just have to connect and fill in the uuid you need. It will work
  • Remote desktop on Gnome hangs randomly

    Unsolved Hardware
    24
    0 Votes
    24 Posts
    3k Views
    O
    @dinhngtu i've manage to fixed it. Your build adds some extra entries in /var/lib/xcp/state.db that aren't in the version from testing, that broke things when i downgraded the packages. ipv4_dns VIF.ipv6_dns
  • Kioxia CM7 PCIe pass-through crash

    Unsolved Compute
    12
    0 Votes
    12 Posts
    2k Views
    D
    Hello, this issue has been fixed in XAPI 26.1.19-1.1 available in the testing updates. For now you have to configure your VMs manually (https://docs.xcp-ng.org/troubleshooting/common-problems/#pci-passthrough-error-unsupported-msi-delivery-mode-7) but the fix will be activated by default in the future.
  • Internet connectivity - Check XOA failed.

    Unsolved Management
    8
    1 Votes
    8 Posts
    308 Views
    Z
    @poddingue said: Thanks for posting the fix. The same thing came up in an older thread, https://xcp-ng.org/forum/topic/9957, where disabling IPv6 also made xoa check go green. @HamiltonWDS explained a likely reason in https://xcp-ng.org/forum/post/87831: Node tries the addresses with a short timeout and can end up on the IPv6 one. Your curl -6 test helps a lot here, because it shows IPv6 doesn't connect at all on that network, so it's more than a slow path. Turning it off in XOA seems reasonable if you don't use IPv6 there; if you do, I'd guess the router side is where it really needs fixing, though I could be wrong (still haven't migrated to IPv6 myself ). Curious whether it sorts out the Cloud Backup too. XOA is using our management network where ipv6 is not needed and not enabled. Looks like our cloud backup finally ran for the first time in while too.[image: cloudbackup.jpg]
  • [V2V] Without VDDK problems

    Solved Migrate to XCP-ng
    4
    0 Votes
    4 Posts
    158 Views
    J
    @mpiton Thanks for a fantasticly quick answer, and detailed as well. I will test this immediately! Update: I can confirm that this was indeed the problem!
  • XO Tasks - backups just pilling up

    Unsolved Management
    5
    1
    0 Votes
    5 Posts
    263 Views
    poddingueP
    I didn't know, so I went and looked at the code. XO prunes that list in a cleanup step: backup runs are kept for 31 days, and other tasks are kept by count, the newest 1000 (https://github.com/vatesfr/xen-orchestra/blob/b59c8f3d20954aa1d72e92c1e8f2db635d74fcba/@xen-orchestra/mixins/Tasks.mjs#L119-L127). You can change both with tasks.gc.keep and tasks.gc.backupKeepDuration in the xo-server config. As far as I can tell the cleanup runs when xo-server starts, not on a timer, so a long-running XO can collect quite a few before the next restart. My guess for the Cloud Backup entries @acebmxer still sees after 30 days is that they're kept by count rather than by age.
  • Unclear errors

    Unsolved Management
    4
    1
    0 Votes
    4 Posts
    160 Views
    poddingueP
    I tried it on a lab host (XCP-ng 8.3, xapi 26.1.16): calling VM.pool_migrate on a halted VM gives me your error, VM_BAD_POWER_STATE with running, halted, so the two values are the state XAPI wanted and the one it found. On the XO side, a migration inside the same pool with no SR or network mapping goes straight to VM.pool_migrate without checking the power state first: https://github.com/vatesfr/xen-orchestra/blob/b59c8f3d20954aa1d72e92c1e8f2db635d74fcba/packages/xo-server/src/xapi/index.mjs#L942-L981. So a stopped VM in that batch would give this, which fits what you remember. I haven't tried the batch migrate from the UI itself though, so I don't know if the UI is supposed to filter halted VMs out before that. It's a good concrete example for the feedback post anyway, XO could just say "this VM is halted" instead of passing XAPI's raw code through.
  • PCIe Pass-through lanes and lane performance

    Unsolved Compute
    57
    0 Votes
    57 Posts
    10k Views
    pandusenP
    @stormi I am glad to hear there is progress, and I never doubted that there was interest and will from the Dev team, it was just stated, by Oliver himself, that the focus was mostly on the classic hypervisor features and that resources for edge cases were scarce. The numbers might not be spot on, but it was something similar.. Anyways I know you care and that the team always take interest in home-lab scenarios. And that, is why I keep recommending XCP-ng. You are good and motivated people, you care and interact with the community and you have a product that is simple in structure and free of most of the technical bureaucracy that the other ecosystems have. And yes, this particular Intel Arc issue, is not only on XCP-ng, The intel toolbox and drivers for this platform is currently a mess for Linux. And you are right to take home-labbers under your wing and prioritize it. As I'm sure you know, although it might not be of direct importance for revenue, they do have a huge impact on popularity. I dont think Proxmox would have had the traction it has, without them Thank you for replying, and thank you to the team for their hard work!
  • 0 Votes
    20 Posts
    2k Views
    M
    @anthoineb, Thanks for the fix! We plan to install blktap-3.55.5-11.1.xcpng8.3 on all three hypervisors tomorrow. We will fully shut down and start the data VMs one at a time, then verify that all running tapdisk processes use the updated binary. As the stalls are intermittent, we will monitor for recurrence and report back with the results.
  • XOA 6.9 Update

    Xen Orchestra
    9
    1 Votes
    9 Posts
    573 Views
    S
    @MathieuRA We reverted back to 6.8.2 for now. We can wait for patch, otherwise, I have opened a tunnel #38081.
  • 1 Votes
    3 Posts
    150 Views
    D
    @trelane Hello, the bug affects Xen drivers of any 9.0 series (considering the upstream code, it goes back to circa September 2019). 9.2.385 contains the fix, which was merged upstream just this month. The bug also impacts network connections, however to a lesser extent due to its nature. I believe that the bug is also present in XenServer VM Tools 9.6.0 and by extension, older versions of it. We recommend that you switch to the XCP-ng tools whenever possible, also due to its other benefits (fully open source, validated for XCP-ng and optimized for Xen Orchestra).
  • XCP-ng Windows PV tools announcements

    Moved News
    116
    0 Votes
    116 Posts
    45k Views
    olivierlambertO
    Great, thanks for the feedback @yomeyo !
  • HA causes reboot of xcp-ng nodes

    Unsolved Management
    32
    0 Votes
    32 Posts
    2k Views
    tjkreidlT
    @john.c Indeed, John, and I almost forgot that the backup network on each host was actually on a separate, isolated NIC, and not at all on the VLAN.
  • Enable Maintenance Mode = Host Not Enough Memory

    Unsolved XCP-ng
    5
    0 Votes
    5 Posts
    225 Views
    J
    @poddingue Our HA-pool is a 3 host system, yes. I'll add a note about our plans to go paid in that feedback-item. Thanks for the tip. In the meantime, I'm testing and reporting as much as I can. In order to hopefully help the product be better for all. Cheers!
  • [dedicated thread] Dell Open Manage Appliance (OME)

    Solved Compute
    104
    1
    0 Votes
    104 Posts
    65k Views
    S
    @ataxyanetwork hey, thanks for sharing this. I am new to the forum so maybe i am not looking in the right place. Your Blog seems to be offline, so did you link the Appliance here in the forum or is it only available on the blog thx
  • DRBD reactor metrics in k8s

    XOSTOR
    7
    3
    0 Votes
    7 Posts
    457 Views
    J
    @poddingue Turns out it is all secondary-to-secondary out of sync. So that makes it less intense, though if primary dies I wonder how it will resolve this, or if it will become split brained. Still unsure of how it happened, but the ones I manually cleaned up have not come back. Going to continue to manually clear them up. If it happens again I will have an alert setup to notify me, and I have all the xcp-ng logs and everything to be able to see what happened. If that happens I will post here with details and logs for the resource so we can see how it occurs. [image: image.jpeg]
  • Slow VM migration on Linstor SR

    Unsolved XOSTOR
    4
    0 Votes
    4 Posts
    261 Views
    K
    @ronan-a For testing purposes, I did the following: I created a VM on the local SR of a host, installed Linux on it, and configured an ISCSI target. I then added that ISCSI target as an ISCSI SR within xcp-ng pool, and created a small VM on that repository, using XO. While the machine was running, I performed a migration without losing a single ping.I assume this confirms my previous assertion regarding the use of a standard storage system.