XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    HA causes reboot of xcp-ng nodes

    Scheduled Pinned Locked Moved Management
    1 Posts 1 Posters 12 Views 1 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • C
      carloum70
      last edited by carloum70

      Hi all,

      I need some help regarding HA.
      I have a 3 node cluster running xcp-ng 8.3 (002 is the master)

      [13:55 dacshyp002 ~]# xe host-list
      uuid ( RO)                : 6b99f1ab-6f4c-4a8d-b766-d16a8a942bdf
                name-label ( RW): dacshyp003
          name-description ( RW): Default install
      
      
      uuid ( RO)                : d99e150c-079a-4092-8909-ad1a36e07dec
                name-label ( RW): dacshyp001
          name-description ( RW): Default install
      
      
      uuid ( RO)                : 41a2a448-a5dc-44c6-be44-c07540d75c60
                name-label ( RW): dacshyp002
          name-description ( RW): Default install
      

      Last Thursday I enabled HA on the pool and also on some om the VM's. On Friday I also did an "rolling pool update"
      Today I had a "spontaneous" reboot of the dacshyp001 and dacshyp003.
      Boot time of dacshyp001:

      [Mon Sep 21 11:21:31 2026] Linux version 4.19.0+1 (mockbuild@1efe98ffea2245bcb4c11893890153ac) (gcc version 4.8.5 20150623 (Red Hat 4.8.5-28) (GCC)) #1 SMP Thu Aug 13 13:06:11 UTC 2026
      

      Around that time I also see the following message in the xensource.log:

      Sep 21 11:21:08 dacshyp002 xapi: [debug||20398 ha_monitor|HA monitor D:7b5068fa6bfc|xapi_ha_vm_failover] Setting host dacshyp001 to dead
      Sep 21 11:21:28 dacshyp002 xapi: [debug||20398 ha_monitor|HA monitor D:7b5068fa6bfc|xapi_ha_vm_failover] Setting host dacshyp001 to dead
      Sep 21 11:21:48 dacshyp002 xapi: [debug||20398 ha_monitor|HA monitor D:7b5068fa6bfc|xapi_ha_vm_failover] Setting host dacshyp001 to dead
      

      I will attached all logging around this time.
      I am trying to understand what triggered the reboot and how I can troubleshoot this.

      I am running XO community commit 6a441 . I am aware this is not the latest version. Before upgrading I want to know the cause of the reboot.

      Also ignore the error:

      /var/lib/xcp/xapi|dispatch:VDI.get_by_uuid D:1f5b4c8d329e|backtrace] VDI.get_by_uuid D:9df5c0255224 failed with exception Db_exn.Read_missing_uuid("VDI", "", "9219113f-65e4-4368-b42d-1b8bdc7614a9")
      

      I will create another ticket for this.
      HA-problem.txt
      Thanks in advance.
      Carlo

      1 Reply Last reply
      Reply Quote 0

      Hello! It looks like you're interested in this conversation, but you don't have an account yet.

      Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

      With your input, this post could be even better 💗

      Register Login
      • First post
        Last post