XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    • Profile
    • Following 0
    • Followers 0
    • Topics 1
    • Posts 7
    • Groups 0
    D Offline
    1. Home
    2. dvinni
    3. Posts

    Posts

    Recent Best Controversial
    • RE: Slow boot on rocky linux 10 latest kernel

      @poddingue Runtime confirmation for your source read: booted CentOS Stream 10 itself — GenericCloud x86_64, kernel 6.12.0-260.el10, 8 vCPUs, XCP-ng 8.3 UEFI. sched_clock correction is -84s with the image's default cmdline (it carries console=ttyS0,115200n8 too), -10s with it removed — same ~8x amplification as on Ubuntu.

      Backport request sent to devel@lists.centos.org: https://lists.centos.org/hyperkitty/list/devel@lists.centos.org/thread/OLHN4WZR3RYHHXDPPTJR2UDNAFFLCEJF/

      posted in Compute
      D
      dvinni
    • RE: Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      @poddingue Thanks for running the -31 numbers — good to have it confirmed that the ttyS0 removal stays worth ~3-4s even with the clock fixed. Agreed on not rushing -proposed to production; we'll pick up -31 when it promotes and keep the cloud-init tweak permanently.

      posted in Compute
      D
      dvinni
    • RE: Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      @poddingue Good news on that front — Canonical already picked it up: the fix ("time/jiffies: Register jiffies clocksource before usage") is in kernel 7.0.0-31, currently sitting in resolute-proposed as part of the 2026.08.03 SRU cycle. So it should reach -updates around early September without needing an upstream 7.0.y backport.

      Your reading of the commit matches what I see too — one stalled clock, two consumers (AP bringup waits + printk/UART delays), not two separate bugs.

      Removing console=ttyS0 from cloud images is still worth keeping in cloud-init templates though: verbose early printk to an emulated 16550 is pointless overhead on any kernel, and boot logs remain available via dmesg anyway.

      posted in Compute
      D
      dvinni
    • RE: Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      @poddingue Tested as requested: on a VM with the ttyS0 fix already applied (leftover was ~16s / ~2s per vCPU), setting tsc_mode=2 + nomigrate=true brings the sched_clock correction down to ~0.3s:

      [    0.281587] sched_clock: Marking stable (281005908, 562301)->(625590665, -344022456)
      

      vs ~-16s without it (reverted after the test). So it's the same regression, now also confirmed on Intel (Ice Lake, Xeon Gold 6342/6348) — with console=ttyS0 acting as a ~7x amplifier on top, presumably because every early printk to the emulated UART goes through the affected clock path.

      posted in Compute
      D
      dvinni
    • RE: Slow boot on rocky linux 10 latest kernel

      Possibly related observation from an Intel pool (Xeon Gold, XCP-ng 8.3): Ubuntu 26.04 cloud image (kernel 7.0, UEFI) shows a similar-looking freeze at "installing Xen timer for CPU N". In my case console=ttyS0 from the cloud image's default cmdline amplified it ~7x — removing it dropped the sched_clock correction from 143s to 16s on 8 vCPUs, and unlike tsc_mode=2 it keeps live migration. Not sure it's the same root cause, but might be worth checking cmdline for those hitting this with cloud images.

      posted in Compute
      D
      dvinni
    • RE: Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      Update: possibly related to the TSC regression tracked in
      this thread
      (kernel 6.12.5+, upstream fix pending) — the symptom signature matches (slow “installing Xen timer” on UEFI only), though that thread’s repros are all AMD while my pool is Intel Ice Lake. The remaining ~2s/vCPU after the ttyS0 fix may be that underlying issue. Unlike the tsc_mode=2 workaround, removing ttyS0 keeps live migration intact.

      posted in Compute
      D
      dvinni
    • Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

      Sharing a debugging result that may explain some of the “slow UEFI VM boot” reports (e.g. in the XO 6 feedback thread).

      Symptom: Ubuntu 26.04 cloud image (cloud-init template, UEFI, 8 vCPU) takes ~1.5 min of black screen before any kernel output. OVMF phase is fast (7s per qemu-dm log, ExitBootServices OK), and systemd-analyze claims ~8s total — yet wall clock says otherwise.

      Root cause found via sched_clock: early printk timestamps freeze during smp: Bringing up secondary CPUs / installing Xen timer for CPU N, and the correction shows up later:

      [    1.671429] smp: Bringing up secondary CPUs ...
      [    1.671429] installing Xen timer for CPU 1
      [    1.671429] smpboot: x86: Booting SMP configuration:
      [    1.671429] .... node  #0, CPUs:      #1
      [    1.671429] installing Xen timer for CPU 2
      [    1.671429]  #2
      [    1.671429] installing Xen timer for CPU 3
      [    1.671429]  #3
      [    1.671429] installing Xen timer for CPU 4
      [    1.671429]  #4
      [    1.671429] installing Xen timer for CPU 5
      [    1.671429]  #5
      [    1.671429] installing Xen timer for CPU 6
      [    1.671429]  #6
      [    1.671429] installing Xen timer for CPU 7
      [    1.671429]  #7
      [    1.671429] cpu 1 spinlock event irq 81
      [    1.671429] cpu 2 spinlock event irq 82
      [    1.671429] cpu 3 spinlock event irq 83
      [    1.671429] cpu 4 spinlock event irq 84
      [    1.671429] cpu 5 spinlock event irq 85
      [    1.697554] cpu 6 spinlock event irq 86
      [    1.757519] cpu 7 spinlock event irq 87
      [    1.758405] smp: Brought up 1 node, 8 CPUs
      ...
      [    2.549234] sched_clock: Marking stable (1444006278, 1105086856)->(145865443398, -143316350264)
      

      i.e. ~143 seconds of real time hidden at the SMP bringup stage (~15-20s per secondary vCPU). Scales linearly: with 2 vCPUs the correction is ~15s.

      Culprit: console=ttyS0 in the cloud image’s default kernel cmdline (/etc/default/grub.d/50-cloudimg-settings.cfg). Early boot printk output is written synchronously to the emulated 16550 UART; every byte is an I/O port access = VM exit. The verbose early boot output serializes around AP bringup with frozen clocks. An ISO-installed 26.04 on an identical VM config (same platform flags, same kernel 7.0.0-29) doesn’t have ttyS0 in cmdline and boots ~7x faster through this phase.

      Fix / proof:

      sed -i 's/console=tty1 console=ttyS0/console=tty1/' /etc/default/grub /etc/default/grub.d/50-cloudimg-settings.cfg
      update-grub
      

      After reboot the same VM shows sched_clock ... -16146235182 — down from 143s to 16s. Tested on both amd64 and amd64v3 builds, identical results.

      Remaining ~2s per vCPU seems to be the baseline UEFI/Xen-timer overhead others have reported — still there, but tolerable.

      Environment: XCP-ng 8.3 (fully patched), Xen 4.17, pool of Xeon Gold 6342/6348, guests: Ubuntu 26.04 kernel 7.0.0-29-generic, device-model qemu-upstream-uefi.

      For cloud-init templates the workaround is a runcmd in the cloud config applying the sed above.

      posted in Compute uefi cloud-init slow-boot ubuntu
      D
      dvinni