XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    1. Home
    2. dvinni
    D Online
    • Profile
    • Following 0
    • Followers 0
    • Topics 1
    • Posts 3
    • Groups 0

    dvinni

    @dvinni

    0
    Reputation
    2
    Profile views
    3
    Posts
    0
    Followers
    0
    Following
    Joined
    Last Online

    dvinni Unfollow Follow
    • RE: Slow boot on rocky linux 10 latest kernel

      Possibly related observation from an Intel pool (Xeon Gold, XCP-ng 8.3): Ubuntu 26.04 cloud image (kernel 7.0, UEFI) shows a similar-looking freeze at "installing Xen timer for CPU N". In my case console=ttyS0 from the cloud image's default cmdline amplified it ~7x — removing it dropped the sched_clock correction from 143s to 16s on 8 vCPUs, and unlike tsc_mode=2 it keeps live migration. Not sure it's the same root cause, but might be worth checking cmdline for those hitting this with cloud images.

      posted in Compute
      D
      dvinni
    • RE: Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      Update: possibly related to the TSC regression tracked in
      this thread
      (kernel 6.12.5+, upstream fix pending) — the symptom signature matches (slow “installing Xen timer” on UEFI only), though that thread’s repros are all AMD while my pool is Intel Ice Lake. The remaining ~2s/vCPU after the ttyS0 fix may be that underlying issue. Unlike the tsc_mode=2 workaround, removing ttyS0 keeps live migration intact.

      posted in Compute
      D
      dvinni
    • Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      Sharing a debugging result that may explain some of the “slow UEFI VM boot” reports (e.g. in the XO 6 feedback thread).

      Symptom: Ubuntu 26.04 cloud image (cloud-init template, UEFI, 8 vCPU) takes ~1.5 min of black screen before any kernel output. OVMF phase is fast (7s per qemu-dm log, ExitBootServices OK), and systemd-analyze claims ~8s total — yet wall clock says otherwise.

      Root cause found via sched_clock: early printk timestamps freeze during smp: Bringing up secondary CPUs / installing Xen timer for CPU N, and the correction shows up later:

      [    1.671429] smp: Bringing up secondary CPUs ...
      [    1.671429] installing Xen timer for CPU 1
      [    1.671429] smpboot: x86: Booting SMP configuration:
      [    1.671429] .... node  #0, CPUs:      #1
      [    1.671429] installing Xen timer for CPU 2
      [    1.671429]  #2
      [    1.671429] installing Xen timer for CPU 3
      [    1.671429]  #3
      [    1.671429] installing Xen timer for CPU 4
      [    1.671429]  #4
      [    1.671429] installing Xen timer for CPU 5
      [    1.671429]  #5
      [    1.671429] installing Xen timer for CPU 6
      [    1.671429]  #6
      [    1.671429] installing Xen timer for CPU 7
      [    1.671429]  #7
      [    1.671429] cpu 1 spinlock event irq 81
      [    1.671429] cpu 2 spinlock event irq 82
      [    1.671429] cpu 3 spinlock event irq 83
      [    1.671429] cpu 4 spinlock event irq 84
      [    1.671429] cpu 5 spinlock event irq 85
      [    1.697554] cpu 6 spinlock event irq 86
      [    1.757519] cpu 7 spinlock event irq 87
      [    1.758405] smp: Brought up 1 node, 8 CPUs
      ...
      [    2.549234] sched_clock: Marking stable (1444006278, 1105086856)->(145865443398, -143316350264)
      

      i.e. ~143 seconds of real time hidden at the SMP bringup stage (~15-20s per secondary vCPU). Scales linearly: with 2 vCPUs the correction is ~15s.

      Culprit: console=ttyS0 in the cloud image’s default kernel cmdline (/etc/default/grub.d/50-cloudimg-settings.cfg). Early boot printk output is written synchronously to the emulated 16550 UART; every byte is an I/O port access = VM exit. The verbose early boot output serializes around AP bringup with frozen clocks. An ISO-installed 26.04 on an identical VM config (same platform flags, same kernel 7.0.0-29) doesn’t have ttyS0 in cmdline and boots ~7x faster through this phase.

      Fix / proof:

      sed -i 's/console=tty1 console=ttyS0/console=tty1/' /etc/default/grub /etc/default/grub.d/50-cloudimg-settings.cfg
      update-grub
      

      After reboot the same VM shows sched_clock ... -16146235182 — down from 143s to 16s. Tested on both amd64 and amd64v3 builds, identical results.

      Remaining ~2s per vCPU seems to be the baseline UEFI/Xen-timer overhead others have reported — still there, but tolerable.

      Environment: XCP-ng 8.3 (fully patched), Xen 4.17, pool of Xeon Gold 6342/6348, guests: Ubuntu 26.04 kernel 7.0.0-29-generic, device-model qemu-upstream-uefi.

      For cloud-init templates the workaround is a runcmd in the cloud config applying the sed above.

      posted in Compute uefi cloud-init slow-boot ubuntu
      D
      dvinni