XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    1. Home
    2. poddingue
    3. Posts
    poddingueP Online
    • Profile
    • Following 0
    • Followers 2
    • Topics 1
    • Posts 92
    • Groups 2

    Posts

    Recent Best Controversial
    • RE: PCIe Passthrough of Radeon iGPU fails

      steff22, the answer you want is one post above yours. 😉
      Yann confirmed at post #4 that AMD iGPU passthrough doesn't work on XCP-ng today and that he's working on it, so your Invalid PCI ROM header signature looks like the known symptom rather than something specific to your board.
      He also linked https://xcp-ng.org/forum/topic/8909 , same error string, if you want the longer history. I don't know enough about the ROM BAR side to add anything useful to what he wrote, and I'd rather not guess in front of him. 😊
      Worth following this thread, he said he'd keep at it.

      posted in Hardware
      poddingueP
      poddingue
    • RE: xcp-ng server crashed/rebooted due to issues with drbd/linstor?

      That timeline is the clearest version of this yet. xen04 going down the same way as xen05 makes the race look real to me rather than coincidental, unless I am reading the ordering wrong.
      The fault lands inside drbd_bm_count_bits, so in the DRBD module itself, which makes me think this is an upstream LINBIT question as much as an XCP-ng one. I do not know enough about how that module gets packaged here to say who owns the fix though. Might be worth another mention to @Team-Storage so someone can open your bundles with your post #10 next to them.
      On the padding idea I am out of my depth: no idea whether LINSTOR can be told to pre-round a volume at creation, but asking LINBIT directly seems worth a try since it would sidestep drbd_bm_resize entirely. 🤷

      posted in XOSTOR
      poddingueP
      poddingue
    • RE: pure-plugin for Xen Orchestra

      There is a public alliance, announced back in April: Vates joined Everpure's Technology Alliance Program, and Everpure came in as an associate member of the Vates Alliance Network (https://vates.tech/blog/vates-joins-the-everpure-technology-alliance-program/).
      What that post actually commits to is validated reference architectures and compatibility documentation, with backup, disaster recovery and automation workflows named as planned expansion, so the area you're asking about is at least in scope.
      Whether any of that becomes an equivalent of the VMware pure-plugin, with per-VM restore from SafeMode snapshots, I genuinely don't know, and I'd rather not guess while you're deciding whether to spend weeks on it. 🤷
      Before you go far it might be worth a mention to @Team-Storage, since they'd know whether the array-side work is heading that way.
      Selfishly I'd rather you asked first than quietly built something that ends up overlapping. I'm not close enough to the storage roadmap to give you a straight yes or no.

      posted in Management
      poddingueP
      poddingue
    • RE: xcp-ng server crashed/rebooted due to issues with drbd/linstor?

      On the bundle, the docs warn it can carry sensitive details about your setup, so better not to attach it publicly: https://docs.xcp-ng.org/troubleshooting/log-files#produce-a-status-report. As far as I can tell there is no self-serve upload spot, it is usually someone from Vates dropping you a private Nextcloud link when you ask, so flagging here that you have the bundle ready is probably the quickest route.

      Since this was a kernel panic, is there anything in /var/crash on the host that went down, and which host was it this time? The earlier ones were xen04 and xen03.

      posted in XOSTOR
      poddingueP
      poddingue
    • RE: 🛰️ XO 6: dedicated thread for all your feedback!

      @john.c your PR landed on the 15th and the live page now reads "XCP-ng 8.3 LTS", so the table matches the advice sitting right above it again: https://github.com/vatesfr/xen-orchestra/pull/10092.
      You spotted it and then wrote the fix yourself, which is the part that saved the next person the trouble. 🙏

      posted in Xen Orchestra
      poddingueP
      poddingue
    • RE: Potential bug with Windows VM backup: "Body Timeout Error"

      Thanks, that complicates the large-VM theory in a good way. 👍
      If these are long-established VMs that backed up fine for years, "too much free space on a new VM" probably isn't the whole story here. 🤷
      Your deltas run from a completely different XO instance, on a different host, to different remotes, at non-overlapping times, so "deltas never fail, fulls do" might not be purely about job type, it could be tangled up with which instance or remote is doing the work.
      If you ever get the chance to run a full from the XO instance that normally handles your deltas, that would help tell whether the timeout follows the job type or the instance.
      I could easily be wrong, though, but that split feels worth isolating before we lean too hard on the free-space angle in https://github.com/vatesfr/xen-orchestra/issues/9181 .

      posted in Backup
      poddingueP
      poddingue
    • RE: xcp-ng server crashed/rebooted due to issues with drbd/linstor?

      That it's stayed quiet since you throttled Velero is a good sign, and it lines up with the concurrency theory rather than anything about the specific volumes. 🤔
      Bumping dom0 RAM makes sense too, since linstor-satellite spiking to roughly 926 log lines around the panic suggests dom0 was under real pressure right then. 👍
      If it does come back, a xen-bugtool --yestoall bundle from the host that panicked, plus the Loki window you already have, would give @Team-Storage something concrete to line up against the DRBD side.
      I'm not sure whether the real fix sits in Velero's pacing or in how LINSTOR handles concurrent create/delete, but the reproduction you've narrowed down is genuinely useful.
      Fingers crossed it stays boring from here. 🤞

      posted in XOSTOR
      poddingueP
      poddingue
    • RE: Unable to live migrate VM between 2 local storages SR

      I converted the topic to a question, then marked it solved. 😉
      Thanks!

      posted in XCP-ng
      poddingueP
      poddingue
    • RE: Error mirroring full backups to backblaze b2

      Quick one to close this off: florent's fix for the encrypted-remote mirror alignment landed in a numbered release now, XO 6.6.2 (2026-07-09).
      The PR was https://github.com/vatesfr/xen-orchestra/pull/10061, so it's out of master and into something you can just update to.
      Did your retest come good on it?
      I'm also curious whether you still need the minPartSize=100000000 setting with the fix in place, or whether that's redundant now, since it'd help us work out whether it's worth writing into the B2 docs.

      posted in Backup
      poddingueP
      poddingue
    • RE: xcp-ng server crashed/rebooted due to issues with drbd/linstor?

      The timing evidence here looks strong to me. Your second Grafana window is doing a lot of work: linstor-satellite.service jumps to roughly 8,700 log lines in the three minutes around the panic, well above linstor-controller at 926 and xapi at 460, and the priority chart goes red at the same moment. A backup that fans out volume creates and deletes across three replicas, landing on a DRBD race, fits what you are seeing.

      Throttling Velero should tell you a lot. If the crashes stop with clientQPS: 3 and itemBlockWorkerCount: 1, that narrows it to concurrency rather than anything about those particular volumes.

      For the logs question, a xen-bugtool --yestoall bundle from a host that has panicked is usually the thing people ask for first (https://docs.xcp-ng.org/troubleshooting/log-files), since it sweeps up the kernel side alongside the storage logs you already have in Loki. I don't know DRBD internals well enough to say which trace matters most, so it might be worth a mention to @Team-Storage. They can say what they actually need rather than have you guess, and a panic that reproduces on a schedule is a good deal easier for them to chase than most.

      posted in XOSTOR
      poddingueP
      poddingue
    • RE: PCIe Pass-through lanes and lane performance

      Before you go further down the bridge tree, redakula called this one early in the thread and it's worth rereading.

      Intel Arc cards put an internal PCIe bridge between the slot and the GPU, and lspci on Linux only reports the top of that chain. His own A310 shows PCIe x8 4.0 at x4 4.0 in GPU-Z under Windows while Linux insists it is x1. That is the same bridge topology you're now staring at on 81, 82 and 83.

      TeddyAstie said much the same when he called it mostly a display issue, and he later checked that resizable BAR does work on UEFI guests since the OVMF build supports it. So the Gen1x1 reading may just be lspci describing the bridge rather than the card. I could easily be wrong, and the people already in this thread know this hardware far better than I do. 🤷

      posted in Compute
      poddingueP
      poddingue
    • RE: Dual video adapters - what should I see, and where?

      Your write-up is the kind of thing that saves the next person a weekend. 👏

      One thing worth knowing before you commit to the whole-controller route. The passthrough page has a "Passing through Keyboards and Mice" section further down, and it says XCP-ng ships /etc/xensource/usb-policy.conf with DENY rules for mice and keyboards by default. You edit those to ALLOW, then refresh with /opt/xensource/libexec/usb_scan.py -d followed by xe pusb-scan host-uuid=<host_uuid>. It's at https://docs.xcp-ng.org/compute/ under USB Passthrough.

      I have no idea whether that covers your USB-to-serial adapter, which is a different device class, and passing the whole controller may still be the cleaner setup for two discrete workstations anyway. Might be worth a mention to @Team-Hypervisor-Kernel on the display question, because that one still puzzles me. 🤔

      posted in Hardware
      poddingueP
      poddingue
    • RE: ACL Permissions to CPU Topology on Self-Service Resource Set

      I had a look at your screenshots. The Topology dropdown is greyed out with a tooltip saying Requires admin permissions, so this looks deliberate rather than broken: the field seems gated on being a full XO admin, not on being admin of your own resource set.

      I couldn't find an existing report asking for it to respect resource-set admin instead, so feedback.vates.tech is probably the right place to raise it. It would carry more weight coming from you, with those screenshots, than from me. I don't know whether the ACL rework changes any of this, so I wouldn't count on it until someone who works on it says so.

      Might be worth a mention to @Team-XO-Backend, since where that permission gate lives is really their call.

      posted in Management
      poddingueP
      poddingue
    • RE: Unable to live migrate VM between 2 local storages SR

      Thanks for coming back and writing up what fixed it.

      I don't know why a disconnected SR blocks a migration between two other SRs. My guess, and it's only a guess, is that VM.migrate_send runs an SR.scan across every SR in the pool rather than just the source and the target, so one leftover SR with no plugin attached is enough to kill the whole prepare step. 🤔
      There is a troubleshooting entry with your symptom at https://docs.xcp-ng.org/troubleshooting/common-problems#unable-to-live-migrate-vdi-between-srs, but it only lists the empty xapi-pool-ca-bundle.pem cause, not orphaned SRs left behind after a host-forget. So it wouldn't have helped you anyway.

      Might be worth a mention to @Team-Storage, both to confirm the mechanism and maybe to get your case added to that page.

      posted in XCP-ng
      poddingueP
      poddingue
    • RE: VDI_IO_ERROR(Device I/O errors) Immediate HELP needed Please.

      You may have worked it out yourself already. 🤷
      A consistency check reporting inconsistent parity on Virtual Disk 1, plus Buffer I/O error on several dm- devices, is the storage layer underneath XCP-ng telling you something is wrong down there. The VDI_IO_ERROR is mostly XCP-ng saying it could not read the disk, not the cause itself.

      I would be careful about anything that writes to that array until someone who knows hardware RAID recovery better than I do has looked at it. I honestly don't know whether a rebuild helps or makes things worse from this state, and I'd rather say that than guess with your data.
      Might be worth a mention to @Team-Storage.

      posted in Xen Orchestra
      poddingueP
      poddingue
    • RE: Potential bug with Windows VM backup: "Body Timeout Error"

      Thanks Greg, that's a useful data point. 👍
      If your clocks are within a second across all three hosts and you're still seeing it, that makes me doubt the timezone angle as the root cause, even if the way dom0 displays the time is confusing.
      The thing I keep coming back to is the split you and I both see: full backups fail while the delta jobs on the same VMs never do. That's the same pattern in https://github.com/vatesfr/xen-orchestra/issues/9181, which points at large VMs or VMs with a lot of free disk space rather than anything clock-related.
      I'm not sure that's your case, but it might be worth checking whether the VMs that fail are the ones carrying the most free space inside the guest. 🤔

      posted in Backup
      poddingueP
      poddingue
    • RE: DUPLICATE_MAC_SEED

      I don't fully follow the mac-seed side of this, but a couple of things in the thread stand out. Tristis Oris's workaround looks like the practical unblock for now: removing the halted CR copy on the target host lets the migration go through, presumably because that replica VM is what collides on the mac-seed.
      Since you, KPS and Tristis Oris are all hitting the same DUPLICATE_MAC_SEED migrating into a replica target, this feels like something worth a GitHub issue on xen-orchestra with your XO commit, the exact steps, and whether a halted CR copy is present each time.
      It might also be worth a mention to @Team-XAPI-Network, since they'd know whether a CR replica is supposed to share its source's mac-seed.
      I could be wrong on the mechanism, so take that with a pinch of salt. 🤷

      posted in Xen Orchestra
      poddingueP
      poddingue
    • RE: XenOrchestra not showing VM Disks on Pool (on single Server working) - XCP-ng Center is showing them

      @kagbasi-wgsdac , thank you so much for the issue creation and the details, that will help for sure! 👍

      posted in Xen Orchestra
      poddingueP
      poddingue
    • RE: Potential bug with Windows VM backup: "Body Timeout Error"

      Thanks for checking. That's useful, even if it points away from where I was looking. 😥
      Clocks within a second of each other means drift probably isn't your problem, and I'd guess the MST/UTC difference is just how dom0 displays it, though I'm not sure. 🤔
      What I keep coming back to is that your full backups fail while the delta jobs on the same hosts never do. That's the same split in https://github.com/vatesfr/xen-orchestra/issues/9181, where full backups hit BodyTimeoutError on VMs with big disks or a lot of free space and the deltas are fine.
      If your failing VMs look like that, your dates and the MST/UTC detail would do more good on that issue than buried in here.

      posted in Backup
      poddingueP
      poddingue
    • RE: Backup fails with "Body Timeout Error", "all targets have failed, step: writer.run()"

      The 2026-05-28 build still being good narrows the window a lot more than "sometime since November". 👍
      What I keep coming back to is the shape of the failure: 14 out of 15, then 5 out of 6, so it is always exactly one that falls over and never the whole run. I don't know whether that points at a per-VM timeout or at something the last task in a run does differently, and someone on the XO team will read that better than me. 🤔
      @pierrebrunet still needs /var/log/xensource.log from the pool master covering one failed run's window, so even a slice from a job where only one VM failed should be enough.

      posted in Backup
      poddingueP
      poddingue