XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    Enable Maintenance Mode = Host Not Enough Memory

    Scheduled Pinned Locked Moved XCP-ng
    2 Posts 2 Posters 9 Views 2 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • J
      jr-m4
      last edited by jr-m4

      Doing some more testing, where I wished to place one Host in maintenance mode. I was then met with HOST_NOT_ENOUGH_FREE_MEMORY.
      The pool is a 3 host pool in HA, with load balancing enabled.

      My guess is that when trying to enable Maintenance mode. It tries to migrate ALL VMs on that host, to ONE single Host on the receiving end. And then it exhausts the amount of available memory.

      What I had to do, was manually disable load balancing. And manually migrate VMs to the two other hosts, to distribute them. This worked.

      What I expected to have happened:
      The evacuation of the host going into Maintenance mode, would be distributing the VMs in a smarter fashion. So that their workload would fit in the pooled resources available.

      Commit: 2f846
      XCP-NG: Fully updated as of writing

      host.setMaintenanceMode
      {
        "id": "<obfuscated>",
        "maintenance": true
      }
      {
        "code": "HOST_NOT_ENOUGH_FREE_MEMORY",
        "params": [
          "OpaqueRef:<obfuscated>"
        ],
        "task": {
          "uuid": "c3a83a2d-003d-ac0f-48b8-79a905c9e557",
          "name_label": "Async.host.evacuate",
          "name_description": "",
          "allowed_operations": [],
          "current_operations": {},
          "created": "20260929T07:46:40Z",
          "finished": "20260929T07:46:40Z",
          "status": "failure",
          "resident_on": "OpaqueRef:<obfuscated>",
          "progress": 1,
          "type": "<none/>",
          "result": "",
          "error_info": [
            "HOST_NOT_ENOUGH_FREE_MEMORY",
            "OpaqueRef:<obfuscated>"
          ],
          "other_config": {},
          "subtask_of": "OpaqueRef:NULL",
          "subtasks": [],
          "backtrace": "(((process xapi)(filename ocaml/xapi/xapi_host.ml)(line 629))((process xapi)(filename hashtbl.ml)(line 159))((process xapi)(filename hashtbl.ml)(line 165))((process xapi)(filename hashtbl.ml)(line 170))((process xapi)(filename ocaml/xapi/xapi_host.ml)(line 625))((process xapi)(filename ocaml/libs/xapi-stdext/lib/xapi-stdext-pervasives/pervasiveext.ml)(line 24))((process xapi)(filename ocaml/libs/xapi-stdext/lib/xapi-stdext-pervasives/pervasiveext.ml)(line 39))((process xapi)(filename ocaml/xapi/rbac.ml)(line 228))((process xapi)(filename ocaml/xapi/rbac.ml)(line 238))((process xapi)(filename ocaml/xapi/server_helpers.ml)(line 78)))"
        },
        "message": "HOST_NOT_ENOUGH_FREE_MEMORY(OpaqueRef:<obfuscated>)",
        "name": "XapiError",
        "stack": "XapiError: HOST_NOT_ENOUGH_FREE_MEMORY(OpaqueRef:<obfuscated>)
          at XapiError.wrap (file:///opt/xen-orchestra/packages/xen-api/_XapiError.mjs:16:12)
          at default (file:///opt/xen-orchestra/packages/xen-api/_getTaskResult.mjs:13:29)
          at Xapi._addRecordToCache (file:///opt/xen-orchestra/packages/xen-api/index.mjs:1358:24)
          at file:///opt/xen-orchestra/packages/xen-api/index.mjs:1392:14
          at Array.forEach (<anonymous>)
          at Xapi._processEvents (file:///opt/xen-orchestra/packages/xen-api/index.mjs:1382:12)
          at Xapi._watchEvents (file:///opt/xen-orchestra/packages/xen-api/index.mjs:1589:14)"
      }
      

      Update: Edited title, since it would suggest Load Balancing was the problem. It isn't. It's the assignment of migration target(s) that is the issue (imho)

      poddingueP 1 Reply Last reply
      Reply Quote 0
      • poddingueP
        poddingue Vates 🪐 @jr-m4
        last edited by

        Your thread looks a lot like https://xcp-ng.org/forum/topic/12321. Same error, same call, HA enabled there too. Olivier's answer on that one was that VMs whose HA restart priority isn't Restart aren't protected, so they never get a real evacuation plan.
        His two workarounds were setting those VMs to Restart, or turning HA off for the maintenance window.

        It doesn't land for everyone, though. On https://xcp-ng.org/forum/topic/12348 @MajorP93 fixed his by moving every VM from best-effort to restart, and @acebmxer tried the same thing and got HA_OPERATION_WOULD_BREAK_FAILOVER_PLAN instead. Same change, opposite outcome. 🤷

        So, what are your VMs set to, restart or best-effort? That would tell us whether you're looking at the same thing or something else entirely.

        Upstream there are two PRs off the back of that discussion. 7145 merged on 8 July and only makes the error message clearer. 7146 is the one that changes the evacuation behaviour itself, and it's still open. So if it is the same problem, I don't think there's anything you can pull down yet.

        1 Reply Last reply
        Reply Quote 0

        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

        With your input, this post could be even better 💗

        Register Login
        • First post
          Last post