XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    1. Home
    2. Popular
    Log in to post
    • All Time
    • Day
    • Week
    • Month
    • All Topics
    • New Topics
    • Watched Topics
    • Unreplied Topics

    • All categories
    • D

      Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Compute uefi cloud-init slow-boot ubuntu
      8
      0 Votes
      8 Posts
      151 Views
      D
      @poddingue Thanks for running the -31 numbers — good to have it confirmed that the ttyS0 removal stays worth ~3-4s even with the clock fixed. Agreed on not rushing -proposed to production; we'll pick up -31 when it promotes and keep the cloud-init tweak permanently.
    • msupportM

      Veeam 13.1 Rocky9 Linux Appliance: Potential Data Loss with CBT and Workers with Expired Tokens

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Backup
      13
      2 Votes
      13 Posts
      546 Views
      msupportM
      our solution to the problem we want to share a post-incident analysis of a data-loss event on XCP-ng 8.3 LTS (shared block storage over FC, Veeam B&R 13.1 with CBT enabled on the pool) — consolidated from our own incident, topic 12402, the CBT feedback thread 9268, and the still-open blktap PR #17. The pieces fit together into one coherent failure chain, and we believe it may be worth a sticky/KB article. A note on the trigger, to be fair and complete: in our case the interrupted jobs (step 2) were caused by expired/invalid worker tokens on the XCP-ng hosts, so the backups died mid-run with the locks in place. We consider that our availability problem. However, an interrupted backup must never be able to damage production VM data — cleaning up VDI locks on job abort is the hypervisor's job, and that is the part that turned a backup hiccup into guest data loss. Failure chain (as we understand it) Veeam backup job with CBT runs and sets locks on VDIs — paused: true + host_OpaqueRef:xxx: RW entries in the VDI sm_config (XAPI state.db). The job is interrupted (timeout / crash / restart). The locks are never cleaned up → stale paused: true remains in sm_config. These entries are MRO, so xe vdi-param-remove can't clear them. A later leaf coalesce / commit runs into a cbtlog disk: per PR #17, tapdisk_vbd_first_image returns the cbtlog disk on td_commit, and the cbtlog driver has no commit action → commit fails early. Result: broken VDI chains, CBT metadata VDIs without a vhd parent, hundreds of orphaned VDIs, and .cbtlog files hanging coalesces (as reported in thread 9268). SR rescan believes a GC is already running and aborts; a host reboot was the only way to force the coalesce through (also reported in 9268). Storage cleaning freezes the VM disk briefly, but the un-freeze fails on the stale lock → failed to unpause tapdisk ... VMs using this tapdisk have lost access to the corresponding disk(s). The guest keeps writing on a frozen/lost disk → NTFS corruption inside the guest and, in our case, actual SQL Server data loss. Step 5 matches exactly the theory Veeam R&D is currently investigating ("storage cleaning freezes VM disks briefly during backup and sometimes fails to un-freeze them"). The stale paused:true lock appears to be the missing "why" behind the failed unpause. What helped us recover Patching the stale lock out of XAPI state.db (stop xapi, backup state.db, remove paused + host_OpaqueRef entries from the affected VDI's sm_config, start xapi). Then: reset CBT on the affected VDIs and trigger a full backup so CBT re-initializes cleanly — otherwise the next interrupted job re-creates the same situation. For the coalesce backlog: with the affected VMs powered off and CBT disabled, snapshot-create-then-delete to kick the GC, watch SMlog, iterate. (Same recipe a user documented in 9268.)
    • henri9813H

      Slow boot on rocky linux 10 latest kernel

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Compute
      29
      2
      0 Votes
      29 Posts
      3k Views
      poddingueP
      Rocky 10 is affected and has no fix in it. I pulled the source RPM for kernel-6.12.0-211.16.1.el10_2.0.1, which is your -211: jiffies.c still ends on core_initcall(init_jiffies_clocksource) with no cs_jiffies_registered, so f24df84cbe05 hasn't landed, and max_raw_delta sits in clocksource.h with nothing setting it early, so the regression is still there. So tsc_mode=2 and nomigrate stay the answer on Rocky until Red Hat picks it up. One more thing, because this thread reads like an AMD problem if you skim it. @dvinni measured the same stalled clock on Intel Xeon Gold in the sibling thread, and the upstream fix came from @teddyastie bisecting it on a Xen HVM guest. I read source rather than booting a Rocky VM, so if you have one behaving differently I'd like to hear it.
    • F

      i915 pass-through and Linux Mint - xcp-ng 8.3

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Compute
      2
      0 Votes
      2 Posts
      40 Views
      olivierlambertO
      Question for @teddyastie but I'm not 100% sure about Intel Coffee Lake iGPU passthrough (iGPU is always far more difficult to passthrough than a discrete GPU)
    • stormiS

      XCP-ng 8.3 updates announcements and testing

      Watching Ignoring Scheduled Pinned Locked Moved News
      653
      1 Votes
      653 Posts
      521k Views
      A
      @gduperrey Rolling pool update worked with released production patches.
    • olivierlambertO

      🛰️ XO 6: dedicated thread for all your feedback!

      Watching Ignoring Scheduled Pinned Locked Moved Xen Orchestra
      249
      7 Votes
      249 Posts
      102k Views
      poddingueP
      Depends which one you mean, because there are two different problems tangled together in this stretch of the thread and they have opposite answers. If you mean @escape222's report at #177, where a VM cloned from XO 6 sits at the TianoCore screen for a couple of minutes, then yes, that was an XO bug and it is fixed. PR #9867, merged 26 May, shipped in XO 6.5.0 on 28 May. The boot order was being rewritten whenever no new disk needed provisioning, so an HVM VM created from a template that already had a disk got network pushed to the front whether or not anyone asked for a network install, and the VM burned the PXE timeout before falling through to the disk. It now follows the install method only. Since you are on sources, anything past 6.5.0 has it. If you mean the slow UEFI boot @MajorP93 described at #182, with the installing Xen timer and spinlock lines, that one is not an XO bug and no XO patch will touch it. It is a regression in the Linux guest kernel introduced in 6.12.5. The cost lands per secondary vCPU, so the wider the VM, the worse it looks. There is a separate thread with the per-vCPU numbers: Ubuntu cloud images on XCP-ng 8.3 UEFI. Worth noting its title blames console=ttyS0, which we now think amplifies the same bug rather than being a second one. I measured that one here this week on a single host, changing only the guest kernel between runs and leaving everything else alone. Ubuntu 7.0.0-30 came in at 50 and 59 seconds across two runs. 7.0.0-31 came in at 0.4. Wall clock reboot to sshd went from 78 seconds to 31. The awkward part is the timing. The upstream fix is f24df84cbe05, in stable 6.12.97 and later, 6.18.y, 7.1.4 and later, and 7.2, but no default channel carries it yet. I re-checked the archives this evening: Debian trixie still ships 6.12.94-1, with 6.12.100-1 sitting in proposed-updates for the next point release, and Ubuntu 26.04 still ships 7.0.0-30 in updates, published today, while 7.0.0-31 has been in proposed since 10 August. Rocky and el10 I could not confirm either way. So keep whatever workaround you are on until a named version lands for your distro. There is arguably a third one at #183, where @Greg_E had Debian 13 and Windows Server 2022 refusing to boot at all when created through XO-lite with UEFI. As far as I know nobody has retested that since. If it is the kernel one you are hitting, this says which version you are waiting for: uname -r dmesg -T | grep -iE "installing Xen timer|spinlock event"
    • olivierlambertO

      Feedback on immutability

      Watching Ignoring Scheduled Pinned Locked Moved Backup
      58
      2 Votes
      58 Posts
      28k Views
      P
      @gsszuber Hi, Yes indeed, you need to preserve the root of the bucket from Lifecycle. We just had a customer with a similar issue. Can you help us by giving a small screenshot of the field to filter out the root (or filter in the three folders) please?