XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    Slow boot on rocky linux 10 latest kernel

    Scheduled Pinned Locked Moved Unsolved Compute
    29 Posts 8 Posters 2.8k Views 7 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • olivierlambertO Offline
      olivierlambert Vates 🪐 Co-Founder CEO
      last edited by olivierlambert

      • 6.12.90 -> bad
      • 6.12.45 -> bad
      • 6.12.22 -> bad
      • 6.12.11 -> bad
      • 6.12.5 -> bad

      BUT 6.12.2 is GOOD! 😓 Almost there!

      edit: now 6.12.3 is also good. Getting close…

      edit: since 6.12.4 is good, then the issue is within 6.12.5. Investigating now.

      1 Reply Last reply Reply Quote 4
      • olivierlambertO Offline
        olivierlambert Vates 🪐 Co-Founder CEO
        last edited by olivierlambert

        Found the culprit, made & tested a patch that works.

        Recap at https://notes.vates.tech/share/v5jtq0iytw/p/slow-hvm-boot-on-linux-6-12-5wrvOvZKJ7

        I will let the rest of my team to try to get it fixed in upstream.

        M 1 Reply Last reply Reply Quote 1
        • M Offline
          MajorP93 @olivierlambert
          last edited by

          @olivierlambert Awesome work on tracking down the issue!
          Very nice, detailed, technical writeup.

          1 Reply Last reply Reply Quote 1
          • olivierlambertO Offline
            olivierlambert Vates 🪐 Co-Founder CEO
            last edited by

            Ping @Team-Hypervisor-Kernel for reference.

            1 Reply Last reply Reply Quote 0
            • TeddyAstieT Offline
              TeddyAstie Vates 🪐 XCP-ng Team Xen Guru
              last edited by

              @majorp93 @henri9813 @acebmxer
              Do you observe the same behavior after setting this for the VM ?

              xe vm-param-add uuid=$UUID param-name=platform tsc_mode=2
              xe vm-param-add uuid=$UUID param-name=platform nomigrate=true
              

              (beware you lose live migration support doing this, you can cancel these changes with matching vm-param-remove like xe vm-param-remove uuid=$UUID param-name=platform param-key=nomigrate)

              M 1 Reply Last reply Reply Quote 0
              • M Offline
                MajorP93 @TeddyAstie
                last edited by MajorP93

                @TeddyAstie said:

                param-name=platform nomigrate=true

                Hi @teddyastie , thanks for working on this.

                As per policy I am not allowed to test these parameters in production which is why I had to create a small test setup for being able to try your settings.

                I deployed a Debian 13 VM via Cloud-Init on a XCP-ng test host using the official Debian 13 cloud image.

                After deploying the VM I had the issue of slow boot.

                After shutting the VM down, applying the settings that you just sent and starting it again I can say that you are on the right track!

                In my case the boot time is completely normal now and on par with Debian 13 VMs that use BIOS instead of UEFI (for booting).

                As this is a workaround and disables live migration this is not an option for production environments but good to have a workaround available anyways for sure!

                Do you think it is possible to fix this on hypervisor level while still having live migration etc. enabled or do we have to wait for an upstream fix within Linux kernel tree?

                TeddyAstieT 1 Reply Last reply Reply Quote 0
                • TeddyAstieT Offline
                  TeddyAstie Vates 🪐 XCP-ng Team Xen Guru @MajorP93
                  last edited by

                  @MajorP93 said:
                  Do you think it is possible to fix this on hypervisor level while still having live migration etc. enabled or do we have to wait for an upstream fix within Linux kernel tree?

                  Yes it's possible to fix it on the hypervisor level (Invariant TSC in guest), but it's quite a bit of work that still needs to be done. A Linux upstream fix for the underlying bug should come at some point hopefully.

                  1 Reply Last reply Reply Quote 1
                  • olivierlambertO Offline
                    olivierlambert Vates 🪐 Co-Founder CEO
                    last edited by

                    2 paths we are doing in parallel:

                    1. We are doing our best to make it upstream in Linux, it's a regression after all. We know how to fix it, so hopefully this will be fixed quickly. Then, we'll have to wait for a Linux kernel update in main distros.
                    2. Invariant TSC in Xen is also a way to fix it, because we want to improve that anyway. But as Teddy said, it's more work and it will take more time.
                    1 Reply Last reply Reply Quote 1
                    • TeddyAstieT Offline
                      TeddyAstie Vates 🪐 XCP-ng Team Xen Guru
                      last edited by

                      Regarding upstream Linux, it should be addressed with https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git/commit/?id=f24df84cbe05e4471c04ac4b921fc0340bbc7752

                      Although, I have no ETA on when it will land to distros.

                      henri9813H 1 Reply Last reply Reply Quote 3
                      • henri9813H Offline
                        henri9813 @TeddyAstie
                        last edited by

                        Hello,

                        Thanks for all !

                        D 1 Reply Last reply Reply Quote 0
                        • D dvinni referenced this topic
                        • D Offline
                          dvinni @henri9813
                          last edited by

                          Possibly related observation from an Intel pool (Xeon Gold, XCP-ng 8.3): Ubuntu 26.04 cloud image (kernel 7.0, UEFI) shows a similar-looking freeze at "installing Xen timer for CPU N". In my case console=ttyS0 from the cloud image's default cmdline amplified it ~7x — removing it dropped the sched_clock correction from 143s to 16s on 8 vCPUs, and unlike tsc_mode=2 it keeps live migration. Not sure it's the same root cause, but might be worth checking cmdline for those hitting this with cloud images.

                          1 Reply Last reply Reply Quote 0
                          • poddingueP poddingue marked this topic as a question
                          • poddingueP Offline
                            poddingue Vates 🪐
                            last edited by poddingue

                            Coming back to this with what the distros actually ship, because "the fix is upstream" turned out not to mean much on its own. 🤷

                            The commit is f24df84cbe05. I checked each distro by grepping kernel/time/jiffies.c at the version they ship, rather than comparing version numbers, since some of them cherry-pick.

                            Already fixed, nothing to do:

                            • Fedora 44: 7.1.8
                            • Alpine 3.22: linux-virt 6.12.103. Alpine 3.23 and edge: 6.18.44

                            Not fixed in what you get today:

                            • Debian 13: trixie ships 6.12.94, which doesn't have it. 6.12.100 and 6.12.101 do, and they're in trixie-proposed-updates, so it should land with the next point release.
                            • Ubuntu 26.04: -updates moved to 7.0.0-30 this morning and that doesn't have it either. 7.0.0-31 does, sitting in resolute-proposed, expected early September.

                            Rocky 10 I couldn't settle. I don't see it in the CentOS Stream 10 kernel changelog, which does list per-commit subjects, and Stream is at 6.12.0-260 while @henri9813 is on the 10.2 branch at -211. So probably not yet. But I might just be failing to find it, so if someone can check properly I'd rather be corrected.

                            @acebmxer @MajorP93 on Debian, and anyone on Ubuntu: keep tsc_mode=2 and nomigrate for now. One thing that won't help you there, the console=ttyS0 removal from the other thread is an Ubuntu cloud image thing. Debian's cloud image recipe only sets a serial console for Azure, EC2 arm64 and ppc64el, so I don't think the generic amd64 image carries it. Worth a look at /proc/cmdline on yours though.

                            I did check that the fix works rather than assuming it: same VM, Ubuntu 7.0.0-30 against 7.0.0-31, and the sched_clock correction went from about -59s to -0.4s on 6 vCPUs.

                            acebmxerA 1 Reply Last reply Reply Quote 1
                            • acebmxerA Online
                              acebmxer @poddingue
                              last edited by

                              @poddingue

                              I haven't seen this issue since we last discussed it back in June. Also looking back at the older comments its seems those where having issues were on AMD systems. I have migrated off AMD in my home lab. Work was always Intel.

                              poddingueP 1 Reply Last reply Reply Quote 1
                              • poddingueP Offline
                                poddingue Vates 🪐 @acebmxer
                                last edited by poddingue

                                Rocky 10 is affected and has no fix in it. I pulled the source RPM for kernel-6.12.0-211.16.1.el10_2.0.1, which is your -211: jiffies.c still ends on core_initcall(init_jiffies_clocksource) with no cs_jiffies_registered, so f24df84cbe05 hasn't landed, and max_raw_delta sits in clocksource.h with nothing setting it early, so the regression is still there.
                                So tsc_mode=2 and nomigrate stay the answer on Rocky until Red Hat picks it up. 🤷
                                One more thing, because this thread reads like an AMD problem if you skim it. @dvinni measured the same stalled clock on Intel Xeon Gold in the sibling thread, and the upstream fix came from @teddyastie bisecting it on a Xen HVM guest. I read source rather than booting a Rocky VM, so if you have one behaving differently I'd like to hear it.

                                1 Reply Last reply Reply Quote 0

                                Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                                Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                                With your input, this post could be even better 💗

                                Register Login
                                • First post
                                  Last post