XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    Ubuntu cloud images on XCP-ng 8.3 UEFI: ~15s per secondary vCPU at boot, caused by console=ttyS0

    Scheduled Pinned Locked Moved Compute
    ueficloud-initslow-bootubuntu
    2 Posts 1 Posters 12 Views 1 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • D Online
      dvinni
      last edited by

      Sharing a debugging result that may explain some of the “slow UEFI VM boot” reports (e.g. in the XO 6 feedback thread).

      Symptom: Ubuntu 26.04 cloud image (cloud-init template, UEFI, 8 vCPU) takes ~1.5 min of black screen before any kernel output. OVMF phase is fast (7s per qemu-dm log, ExitBootServices OK), and systemd-analyze claims ~8s total — yet wall clock says otherwise.

      Root cause found via sched_clock: early printk timestamps freeze during smp: Bringing up secondary CPUs / installing Xen timer for CPU N, and the correction shows up later:

      [    1.671429] smp: Bringing up secondary CPUs ...
      [    1.671429] installing Xen timer for CPU 1
      [    1.671429] smpboot: x86: Booting SMP configuration:
      [    1.671429] .... node  #0, CPUs:      #1
      [    1.671429] installing Xen timer for CPU 2
      [    1.671429]  #2
      [    1.671429] installing Xen timer for CPU 3
      [    1.671429]  #3
      [    1.671429] installing Xen timer for CPU 4
      [    1.671429]  #4
      [    1.671429] installing Xen timer for CPU 5
      [    1.671429]  #5
      [    1.671429] installing Xen timer for CPU 6
      [    1.671429]  #6
      [    1.671429] installing Xen timer for CPU 7
      [    1.671429]  #7
      [    1.671429] cpu 1 spinlock event irq 81
      [    1.671429] cpu 2 spinlock event irq 82
      [    1.671429] cpu 3 spinlock event irq 83
      [    1.671429] cpu 4 spinlock event irq 84
      [    1.671429] cpu 5 spinlock event irq 85
      [    1.697554] cpu 6 spinlock event irq 86
      [    1.757519] cpu 7 spinlock event irq 87
      [    1.758405] smp: Brought up 1 node, 8 CPUs
      ...
      [    2.549234] sched_clock: Marking stable (1444006278, 1105086856)->(145865443398, -143316350264)
      

      i.e. ~143 seconds of real time hidden at the SMP bringup stage (~15-20s per secondary vCPU). Scales linearly: with 2 vCPUs the correction is ~15s.

      Culprit: console=ttyS0 in the cloud image’s default kernel cmdline (/etc/default/grub.d/50-cloudimg-settings.cfg). Early boot printk output is written synchronously to the emulated 16550 UART; every byte is an I/O port access = VM exit. The verbose early boot output serializes around AP bringup with frozen clocks. An ISO-installed 26.04 on an identical VM config (same platform flags, same kernel 7.0.0-29) doesn’t have ttyS0 in cmdline and boots ~7x faster through this phase.

      Fix / proof:

      sed -i 's/console=tty1 console=ttyS0/console=tty1/' /etc/default/grub /etc/default/grub.d/50-cloudimg-settings.cfg
      update-grub
      

      After reboot the same VM shows sched_clock ... -16146235182 — down from 143s to 16s. Tested on both amd64 and amd64v3 builds, identical results.

      Remaining ~2s per vCPU seems to be the baseline UEFI/Xen-timer overhead others have reported — still there, but tolerable.

      Environment: XCP-ng 8.3 (fully patched), Xen 4.17, pool of Xeon Gold 6342/6348, guests: Ubuntu 26.04 kernel 7.0.0-29-generic, device-model qemu-upstream-uefi.

      For cloud-init templates the workaround is a runcmd in the cloud config applying the sed above.

      D 1 Reply Last reply Reply Quote 0
      • D Online
        dvinni @dvinni
        last edited by

        Update: possibly related to the TSC regression tracked in
        this thread
        (kernel 6.12.5+, upstream fix pending) — the symptom signature matches (slow “installing Xen timer” on UEFI only), though that thread’s repros are all AMD while my pool is Intel Ice Lake. The remaining ~2s/vCPU after the ttyS0 fix may be that underlying issue. Unlike the tsc_mode=2 workaround, removing ttyS0 keeps live migration intact.

        1 Reply Last reply Reply Quote 0

        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

        With your input, this post could be even better 💗

        Register Login
        • First post
          Last post