XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    1. Home
    2. Popular
    Log in to post
    • All Time
    • Day
    • Week
    • Month
    • All Topics
    • New Topics
    • Watched Topics
    • Unreplied Topics

    • All categories
    • A

      Backup fails with "Body Timeout Error", "all targets have failed, step: writer.run()"

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Backup
      105
      0 Votes
      105 Posts
      13k Views
      J
      @christopher-petzel Ok! Thanks!
    • stormiS

      XCP-ng 8.3 updates announcements and testing

      Watching Ignoring Scheduled Pinned Locked Moved News
      674
      1 Votes
      674 Posts
      580k Views
      X
      I deferred applying the preview patches this time in order to try my luck again with RPU. Unfortunately, it did not work for me again as it has not in the past. My results are similar to others here: primary host patch application went fine including the reboot. However, once it began doing a secondary host, it started throwing errors e.g. CANNOT_EVACUATE_HOST and VM_REQUIRES_SR etc. However, manually putting the host in maintenance mode from the GUI evacuated each host just fine and I was able to apply the patches and reboot each subsequent host from the XO GUI. Also, as with others here, the RPU task hung in the task list and even a reboot of the XO VM would not clear it forcing me to delete the task using xo-cli e.g. xo-cli rest del tasks. ENVIRONMENT: Home lab consisting of 4 x Dell OptiPlex 7040 i7-6700 SFF hosts, 48GB RAM each, 10 Gbps storage connections to a TrueNAS home-built NAS via NFS and XO from source (XOS) using @ronivay build script on AlmaLinux 10.2 minimal install VM with XO commit 6a441 compiled on 2026-08-28 from master branch. FWIW, RPU functionality remains unavailable to me, though obviously, this is not a showstopper for a 4 x host home lab pool. If there is anything I can do to help resolve this, please let me know as I remain an enthusiastic proponent of the Vates virtualization stack.
    • D

      XCP-ng Windows PV tools announcements

      Watching Ignoring Scheduled Pinned Locked Moved News
      110
      0 Votes
      110 Posts
      36k Views
      A
      @dinhngtu Something called "Elpha Secure" ...none of our other antivirus shows it being bad but I wanted to ask around before I unflagged it.
    • J

      [PACKER] soucis avec cd_files

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved French (Français)
      19
      1 Votes
      19 Posts
      882 Views
      J
      @AtaxyaNetwork Merci pour tes recherches ! Oui "cd_label" serait cool comme ajout au plugin ce qui permet sur les distro type Fedora/Redhat de ne pas avoir de boot_command à gérer
    • acebmxerA

      Veeam for Xen Orchestra has been release today 13.1

      Watching Ignoring Scheduled Pinned Locked Moved Backup
      25
      0 Votes
      25 Posts
      2k Views
      acebmxerA
      So I ended up going with the Windows B&R as i could net setup the windows mount server when setting up the repo for the backups. It would setup the linux one but not the windows one. This was with using the veeam console from a windows client. So i ended up with the Windows one. Now after a few backups i am see this warning / error.... 8/5/2026 5:34:50 PM Warning : Failed to use CBT: [Task e75c4247-2ee1-7b51-088c-b970ace60f3d (Async.VDI.list_changed_blocks) failed: . SR_BACKEND_FAILURE_460. . Failed to calculate changed blocks for given VDIs. [opterr=Source and target VDI are unrelated] First i thought it was because i maped the new job to older backups from Beta v1. So i purged all old backups and started fresh. The full backups were successful. Now on first delta 4 out of 7 vms have that warrning. This a Veeam issue or xcp-ng?
    • msupportM

      Veeam 13.1 Rocky9 Linux Appliance: Potential Data Loss with CBT and Workers with Expired Tokens

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Backup
      14
      2 Votes
      14 Posts
      955 Views
      acebmxerA
      Veeam scheduled a remote call with me and pulled more log files. Of coarse when we ran the backup job twice in a row both times al vms were successful. Veeam needs to baby sit our backups :). The call was cut short do to internet going down. I have uploaded the logs and waiting to hear back. Update - Veeam took alot more logs from Veeam and from xcp-ng pool. Their response back - I've got someone else getting similiar results, so I'm providing both of your logs to get some insights. Basically when you see the error, it's because something happened to the bitmap we left behind on the previous run and so next run, we re-read the entire disk. I've not found anything super clear to what's going wrong with the bitmap and why its gone, even from the Xen server logs, so I'm hoping from QA's eyes might see what I might be missing. I will keep you posted if they have any details. Update 8.26.26 - I just wanted to provide an update, the QA team is still checking stuff, but they did advised the following. They noticed for the disks, they show there are configured XO native backups: ie. xo:backup:deltaChainLength: 4; xo:backup:contentKey: 4bce47c7-04bb-44ae-a687-c35c9cfbb2b2; xo:backup:job: e3616a64-6b83-4bd0-80b1-00523678e909; xo:backup:includeNonNbdQcow2Fix: true; xo:backup:schedule: 3488baee-fe48-4acb-a5d0-e728fe28efa1; xo:backup:vm: 5703adef-d804-6b15-ba2f-7b3357a711bb; xo:backup:datetime: 20260819T01:00:33Z I believe you said the native XO Backup was disabled, can you re-confirm if that is accurate, and provide a screenshot for XO's Job and backup list to verify. QA did confirm that native XO Backups can cause issues as it treats the objects as two different chains and so that's why a CBT comparison can fail. QA says they are looking into ways to change how the comparison works in a future update, and expects it to help if in situations if there are native backups, but asked if the Native backup can be verified as disabled for now. I have respnded stating that Veeam backup is only used for Windows based vms. Backup in Xen Orchestra is only used for linux base vms.
    • CyrilleC

      Xen Orchestra Container Storage Interface (CSI) for Kubernetes

      Watching Ignoring Scheduled Pinned Locked Moved Infrastructure as Code
      29
      5 Votes
      29 Posts
      5k Views
      K
      @Cyrille We can't disable the embedded CCM. Disabling the embedded CCM in RKE2 impacts core cluster bootstrap behavior because it is a bootstrap-critical component responsible for core node lifecycle management.
    • olivierlambertO

      🛰️ XO 6: dedicated thread for all your feedback!

      Watching Ignoring Scheduled Pinned Locked Moved Xen Orchestra
      254
      7 Votes
      254 Posts
      113k Views
      poddingueP
      Thanks!
    • B

      Native Ceph RBD SM driver for XCP-ng

      Watching Ignoring Scheduled Pinned Locked Moved Development
      29
      3 Votes
      29 Posts
      6k Views
      dicode-nlD
      @benapetr @olivierlambert I've made a new release which includes SMAPIv1 improvements and a proper SMAPIv3 volume + datapath plugin. https://github.com/dicode-nl/xcp-ng-ceph-rbd/releases#release-v20260827 Let me know your thoughts and if there is anything you'll like to see added / changed / tested. Next step for me is CBT and SXM.
    • D

      Smart Reboot blocked in XO, and no Rolling Pool Update

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved XCP-ng
      9
      0 Votes
      9 Posts
      419 Views
      D
      @poddingue said: What I can't tell you is what set that particular combination on your VM in the first place. Does it ring a bell? I have no Idea. I had it on "Protect from accidental shutdown" but turned that off again, later. Doing this again (on, off) helped, as you said. Thank you so much!
    • D

      Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Compute uefi cloud-init slow-boot ubuntu
      8
      0 Votes
      8 Posts
      377 Views
      D
      @poddingue Thanks for running the -31 numbers — good to have it confirmed that the ttyS0 removal stays worth ~3-4s even with the clock fixed. Agreed on not rushing -proposed to production; we'll pick up -31 when it promotes and keep the cloud-init tweak permanently.
    • J

      PCIe Pass-through lanes and lane performance

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Compute
      44
      0 Votes
      44 Posts
      6k Views
      pandusenP
      @andriy.sultanov @andriy.sultanov said: @pandusen As Teddy said above, you can't passthrough a PCI bridge, so there's no PCI devices xapi shouldn't omit here. I am not trying to pass through the bridge only the end points. The Intel arc's have 2 end points: The GPU and the Sound device. "xe pci-list" only reveals the GPU, not the sound device. (this works for nvidia and AMD) But "going the xen-cmdline way" shouldn't break anything, that's what xe pci-disable-dom0-access does behind the scenes. What issues did you see? Which steps did you follow? the sound device is available in the lspci list and can be passed through using CLI. But doing so, (using CLI for passtrough) undoes everything done using xe or the passthrough gui in XO. and results in this: https://xcp-ng.org/forum/topic/10609/xcp-ng-8.3-pci-passthrough-issue so yes, its does break something.
    • ForzaF

      Migrating an offline VM disk between two local SRs is slow

      Watching Ignoring Scheduled Pinned Locked Moved Xen Orchestra
      29
      1
      0 Votes
      29 Posts
      8k Views
      olivierlambertO
      @TeddyAstie you were right, and it's the bigger win. I prototyped the pipelining you described and measured it. Same rig, same 107 GiB disk, same direction as the earlier runs. Arm Wall clock vs control Peak rate Stalls A : stock behaviour 2080.8 s 1.00x 63 MiB/s 62.2% B : TCP_NODELAY only 743.8 s 2.80x 210 MiB/s 0.6% C : + pipelined writes, depth 8 377.4 s 5.51x 432 MiB/s 1.6% Pipelining is worth a further 1.97x on top of the Nagle fix. Your diagnosis was correct: the per-request reply gating, not TCP, is the dominant limit. Verified byte for byte, source and destination md5 of the 107 GiB disk both ce647d9436b48401cd4b489c955ef0f7. That mattered more than the stopwatch here, for reasons below. It's cheaper than you thought: the multiplexer already exists No NBD redesign is needed. nbd/lib/client.ml:78 is already module Rpc = Mux.Make (NbdRpc), and that multiplexer: assigns every request a unique handle (get_handle) registers a waiter in id_to_wakeup keyed by that handle serialises only the send under outgoing_mutex, then returns a promise runs a background dispatcher thread that reads replies and wakes the matching waiter So concurrent Client.write calls already interleave correctly. The whole request/reply machinery is there and unused. The serialisation is one fold_left in stream_nbd that awaits each write before pulling the next element. The rest of the chain was already fine too: Layer Verdict xapi nbdproxy Unixext.proxy, raw bidirectional byte copy, never parses NBD, cannot serialise tapdisk NBD server NBD_SERVER_NUM_REQS 8, per-client request pool destination storage fio at 2 MiB blocks: 399 MiB/s at qd=1, ~1700 MiB/s at qd=2..8, collapses at qd=16 Depth 8 matches tapdisk's pool. The fio sweep says deeper is not better. The actual trap: buffer ownership This is the part worth writing down, because it is invisible from the protocol level and it bites silently. Vhd_format.F.expand_copy allocates one 2 MiB buffer and hands out slices of it: let buffer = Memory.alloc twomib_bytes in ... let data = Cstruct.sub buffer 0 (this * 512) in really_read h (sector_start ** 512L) data >>= fun () -> return (Cons (`Sectors data, next)) It refills that same buffer on every step. The current sequential code is safe only as a side effect of awaiting each write before pulling the next element. Pipeline it naively and you get: launch write N, pull element N+1, really_read overwrites the buffer, write N puts block N+1's bytes at block N's offset. The migration completes, reports success, and the disk is corrupt. Nothing in the stack flags it. So any implementation of this needs buffer ownership solved alongside the concurrency. My prototype takes the cheap local route: a pool of depth buffers in stream_nbd, one memcpy per 2 MiB block, buffer returned only once its write completes. The proper fix is a buffer pool inside expand_copy itself, but f.ml is a shared library with other consumers, so that is a wider change than I wanted for a measurement. Prototype patch Against xapi-project/xen-api, ocaml/vhd-tool/src/impl.ml, on top of the TCP_NODELAY patch from the previous post. This is a measurement prototype, not mergeable as-is. Known gaps: progress reporting counts issued rather than completed work a failed write leaves its siblings unawaited rather than cancelled the per-block memcpy is a workaround for the shared buffer, not the right fix --- a/ocaml/vhd-tool/src/impl.ml +++ b/ocaml/vhd-tool/src/impl.ml @@ stream_nbd (if not prezeroed then expand_empty s else return s) >>= fun s -> expand_copy s >>= fun s -> + (* Pipelined writes. The NBD client already multiplexes: every request gets a + unique handle and a background dispatcher matches replies back to waiters, + so several writes may be outstanding at once. Issuing them one at a time + makes every request pay a full round trip. + + Depth 8 matches NBD_SERVER_NUM_REQS in tapdisk's NBD server. Deeper just + queues. + + Buffer ownership matters here. [expand_copy] hands out slices of a single + shared 2MiB buffer that it refills on every step, so an in-flight write + cannot keep pointing at it: pulling the next element would overwrite the + bytes before they reach the wire. Each outstanding write therefore gets a + private buffer from a pool sized to the pipeline depth, returned only once + the write has completed. *) + let depth = 8 in + let twomib = 2 * 1024 * 1024 in + let free = ref (List.init depth (fun _ -> IO.alloc twomib)) in + let inflight = ref [] in + let reap () = + match !inflight with + | [] -> + return () + | l -> + Lwt.nchoose_split (List.map fst l) >>= fun (_, pending) -> + let still, done_ = + List.partition (fun (t, _) -> List.memq t pending) l + in + inflight := still ; + free := List.map snd done_ @ !free ; + return () + in + let rec drain () = + if !inflight = [] then return () else reap () >>= fun () -> drain () + in fold_left (fun (sector, work_done) x -> ( match x with - | `Sectors data -> ( - Client.write server (Int64.mul sector 512L) [data] >>= function - | Ok () -> - return Int64.(of_int (Cstruct.length data)) - | Error _e -> - fail (Failure "Got error from NBD library") - ) + | `Sectors data -> + (* Block only when the pipeline is full. *) + (if !free = [] then reap () else return ()) >>= fun () -> + let buf = List.hd !free in + free := List.tl !free ; + let len = Cstruct.length data in + let mine = Cstruct.sub buf 0 len in + Cstruct.blit data 0 mine 0 len ; + let t = + Client.write server (Int64.mul sector 512L) [mine] >>= function + | Ok () -> + return () + | Error _e -> + fail (Failure "Got error from NBD library") + in + inflight := (t, buf) :: !inflight ; + return Int64.(of_int len) | `Empty _n -> (* must be prezeroed *) assert prezeroed ; return 0L @@ (0L, 0L) s.elements >>= fun _ -> + (* Every write must land before the stream is declared complete. *) + drain () >>= fun () -> p total_work ; return (Some total_work) Caveats on the numbers The destination SSD degrades partway through a large transfer (DRAM-less Lexar NM790, host with 3.9 GB RAM), which is why arm C starts at ~432 MiB/s and settles around 275. Arm C runs closest to that ceiling so it feels it most. On better destination storage the gap should widen, not narrow. Also worth saying: after pipelining, TCP_NODELAY matters much less, since with 8 requests in flight there is almost always at least an MSS queued. It is still correct to set it, and it is what every other NBD client in the stack does, but the 2.8x from the previous post should be read as "what you get today with a 12 line change", not as something that stacks cleanly onto the 5.5x.
    • K

      Intermittent Xen blkfront I/O stalls: all guest tags busy while tapdisk reports zero outstanding requests

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Compute
      16
      0 Votes
      16 Posts
      994 Views
      A
      @mike.potapov Can you upgrade to the latest blktap-3.55.5-9.3.xcpng8.3 to check is the issue is still there?
    • O

      Remote desktop on Gnome hangs randomly

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Hardware
      14
      0 Votes
      14 Posts
      2k Views
      O
      @dinhngtu Hello. I've updated to the latest commit. There is no need to give me credit. I just want to help so others and myself included can benefit from this. Anyway thank you for your work.
    • C

      Backup failures with odd connection refused errors

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Backup
      7
      0 Votes
      7 Posts
      409 Views
      poddingueP
      Thanks for the feedback.
    • F

      is Xo Proxy available in community version

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Xen Orchestra
      13
      0 Votes
      13 Posts
      3k Views
      B
      @poddingue Fistst of all I appreciate your answer and your position. The thing is that, even though the proxy code itself is opensource, the functionality of the plugin is basicaly behind a paywall. We are not talking about support. Actual functioning of the plugin after compiling from sources depends on license availability and there is no option to select no support or something along the lines "I built it myself from sources". Without patching the code even though the proxy is otherwise functional the backups won't work because of missing license. Hopefully the powers that can will provide an acceptable albeit community supported way to use the proxy cleanly, without touching license checks. Best regards!
    • C

      Bringing container visibility back to XO

      Watching Ignoring Scheduled Pinned Locked Moved Xen Orchestra
      6
      1
      0 Votes
      6 Posts
      274 Views
      poddingueP
      Nice, thanks for the feeder entry and the explanation, @CAPS!
    • P

      Error mirroring full backups to backblaze b2

      Watching Ignoring Scheduled Pinned Locked Moved Solved Backup
      33
      2
      0 Votes
      33 Posts
      4k Views
      poddingueP
      Thanks a lot for this feedback, @pedro!
    • henri9813H

      Slow boot on rocky linux 10 latest kernel

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Compute
      31
      2
      0 Votes
      31 Posts
      3k Views
      poddingueP
      Thanks for actually booting one, that's the bit I skipped. -84s versus -10s without console=ttyS0 matches the Ubuntu ratio, and it's the first EL10 number anyone has measured rather than read from the source. That settles the question I left open. A real-world measurement is vastly better than a source-code read, right? Thanks for the backport request, too. Since CentOS Stream sits upstream of RHEL and Rocky, if the backport lands there, it should be the earliest signal that the rest of the family will follow.