XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    Tag-Based Automation Plugin: Tag-Based VM Performance & Permission Management via assigned tag(s)

    Scheduled Pinned Locked Moved Management
    15 Posts 6 Posters 1.6k Views 6 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • MathieuRAM Offline
      MathieuRA Vates πŸͺ XO Team @johnnezero
      last edited by MathieuRA

      Hi @johnnezero and thanks for the post!

      Just wanted to talk about PERMISSION SYNC.

      In the REST API the "permission sync" pattern is actually handled natively by the RBAC system using selectors.

      For example, if you want a role that allows a user to manage VM power state only for VMs tagged dev:

      • Start from the built-in role template β€œVMs power state manager” (just to speed up role creation, but totally optional)
      • Create or customize a role with the required VM power privileges (read, start, stop, reboot, etc.)
      • Scope each privileges using a selector like:
        tags:dev
      • Then assign the role to your user or group

      Once done, the access is fully dynamic:

      • Any VM with the dev tag is included in the scope
      • Removing the tag immediately revokes access
      • Adding the tag grants access instantly
      • No need to maintain per-VM ACL entries

      The key point is that RBAC evaluates privileges at request time based on selectors.
      You can also base selectors on other VM properties, not only tags (for example power state, name patterns ...).

      You can find the doc here
      and a dedicated forum thread here

      PS: For the moment the XO6 UI does not support the RBAC system, but we are working on it πŸ™‚

      julienXOvatesJ johnnezeroJ 2 Replies Last reply Reply Quote 1
      • julienXOvatesJ Offline
        julienXOvates Vates πŸͺ XO Team @MathieuRA
        last edited by

        @johnnezero
        To add one more info on top of Mathieu's post : ACL (XO5 only) and RBAC (in REST API + XO6 soon) are completely separate, so better use RBAC as it gives much more flexibility !
        Interesting plugin by the way !!! πŸ‘

        johnnezeroJ 1 Reply Last reply Reply Quote 0
        • johnnezeroJ Offline
          johnnezero @MathieuRA
          last edited by

          @MathieuRA Wow, and wow! Thanks for the detail on RBAC, I'll have give it a look. Unfortunately the current plugin doesn't get any where near the "deep-weeds" of all that, and basically just uses the built-in "Admin, Operator and Viewer" roles to auto-provision (via Autopilot) new VMs as they show up (i.e. during a migration project, etc. ). We needed something for our current VMware-to-XCP-ng migration project, and '"Where's there's a will, there's a way" kicked in and the plugin came to be.
          Thanks again for your input, and let me know if anything else comes to mind...
          Happy Day πŸ™‚

          1 Reply Last reply Reply Quote 0
          • johnnezeroJ Offline
            johnnezero @julienXOvates
            last edited by

            @julienXOvates Noted, thank you!

            fohdeeshaF 1 Reply Last reply Reply Quote 0
            • johnnezeroJ johnnezero referenced this topic
            • fohdeeshaF Offline
              fohdeesha Vates πŸͺ Pro Support Team @johnnezero
              last edited by fohdeesha

              @johnnezero Hi! Thank you for your contribution to the vates / XCP ecosystem! It's certainly appreciated. However there seems to be a few "vibe-code-isms" in the plugin source that has caused us some support tickets on Vates side - the main issue is that you're doing two synchronous NFS operations (fs.existsSync & fs.appendFileSync) for every single log entry with no buffering or batching. Being synchronous, the entire node event loop locks up dead until the NFS call completes. So your plugin writing ~500 log lines whenever it runs at the top of the hour (or whenever else) is doing 2000+ NFS sync calls, locking up the node event loop for nearly 60 full seconds - meaning XOA loses the ability to contact hosts, perform operations, etc. This is made exponentially worse by the logger writing to XOA backup mounts, as when an XOA backup is running, the NFS share is typically quite loaded with backup traffic, making the logging sync operations take many times longer.

              We've seen a couple customers whose backups started failing due to this which was interesting to track down - it appears as just the XOA appliance losing network connectivity to the pool, when in fact it's because xo-server is completely locked up waiting on logging activity from this plugin while the backup run is trying to call XAPI on the customer pool.

              there's some other stuff like the enforcePerformance performing four writes over xapi on every single VM unconditionally every single call, instead of checking if the value is already set correctly and doesn't need a write. Stuff like vm.VCPUs_params and vm.other_config which XOA already intelligently maintains a cache of, so 99% of these routine calls should theoretically require zero xapi writes, or even reads. Here's a full analysis from Claude (I am not a dev, and our current XO devs are swamped so I didn't want to bother them with this πŸ˜› )

              The hourly enforcement cycle blocks the Node.js event loop for tens of seconds.
              xo-server runs every pool connection in that same single threaded process, so
              the block starves xen-api's event watcher. Its event.from long poll has a
              60.1 second client side deadline, the timer cannot fire on time, and when it
              does xen-api treats the connection as dead and reconnects. Reconnecting
              flushes the XO object cache. Backup jobs starting in that window fail with
              no such object <pool-uuid>.

              The block is synchronous NFS I/O, one call per log line, inside loops over
              hundreds of rows.

              Evidence

              Four _watchEvents TimeoutError events fired within the same second, one second
              after the cycle's preload phase ended:

              call deadline actual
              1 60.1 s 99.4 s
              2 60.1 s 67.1 s
              3 60.1 s 66.4 s
              4 60.1 s 66.2 s

              Timers with deadlines spread across 33 seconds do not batch like that unless the
              loop was starved and then released. The pool master logged no XAPI errors during
              the window, and its session.login_with_password from the XOA address matches
              the reconnect exactly.

              Over a longer window, 18 timeout events since 1 August, 9 of them within 90
              seconds of the top of an hour.

              Findings

              1. Synchronous NFS I/O in the logging path (critical)

              writeLog (line 279) makes two blocking filesystem calls per log line, both
              against the NFS mount:

              if (!fs.existsSync(dir)) fs.mkdirSync(dir, { recursive: true });  // stat RPC
              fs.appendFileSync(logPath, message + "\n", "utf8");               // open+write+close
              

              logInfo (290) and logWarn (298) both call it, twice when summary is set.
              About 270 log lines per cycle at roughly four round trips each is about 1100
              blocking calls. Measured block was 39 seconds, implying about 35 ms per call,
              consistent with the share being under backup write load at the time.

              Other synchronous calls on the same mount: rotateLogs (231, 234, 244, 245,
              258, 261), readLogTail (310, 311) which reads up to 10 MB, CSV reads (587,
              619, 662, 750), and full CSV writes (645, 849).

              console.log is not a contributor. Node writes to a pipe asynchronously.

              2. Unconditional XAPI writes on every VM, every hour

              enforcePerformance, lines 450 to 453:

              await xapi.call("VM.remove_from_VCPUs_params", vm._xapiRef, "weight");
              await xapi.call("VM.add_to_VCPUs_params",      vm._xapiRef, "weight", String(matched.weight));
              await xapi.call("VM.remove_from_other_config", vm._xapiRef, "sched-pri");
              await xapi.call("VM.add_to_other_config",      vm._xapiRef, "sched-pri", String(matched.ioPri));
              

              No check for whether the value is already correct, though the cache already
              holds vm.VCPUs_params and vm.other_config. Four writes per tagged VM per
              hour, each journaled by XAPI and broadcast through event.from to every client,
              including the process that is currently frozen. The plugin generates the event
              storm it then cannot drain.

              The remove-then-add pair is also non-atomic. Between the calls the VM has
              neither weight nor sched-pri.

              Same pattern for tags at 688 and 794, where VM.add_tags is called blind and a
              duplicate key error is caught afterwards. The cached vm.tags already answers
              that.

              3. xo.getAllGroups() inside per-VM and per-tag loops

              Line 501 in enforcePermissions, line 552 in applyPermissionTag. The latter is
              called inline from runCsvSync (696) and processPreloadVms (831), multiplying
              the cost.

              4. xo.getXapi(vm) unguarded

              Lines 681 and 786. It throws no connection found for object <uuid> when the
              pool is disconnected, which is the state the plugin creates for itself. Neither
              is wrapped, so one disconnected pool aborts the whole cycle through
              runEnforcementCycle (891, which rethrows). enforcePerformance handles this
              correctly at 446; the other two do not.

              5. The preload queue never drains

              processPreloadVms (742) removes a row only on a successful match (839). A row
              naming a VM that does not exist, was decommissioned, lives on an unconnected
              pool, or was mistyped is retried forever. The observed instance held 216 such
              rows, producing 216 synchronous NFS writes per hour to report that nothing
              happened. That is roughly 80 percent of the plugin's blocking I/O.

              The message "not found or filtered" also conflates a missing VM with one
              rejected by isRealVm, and the lookup at 774 searches across every connected
              pool, taking whichever the index returns first on a name collision.

              6. Repeated full object scans and quadratic lookups

              Object.values(xo.getObjects({ type: "VM" })) is rebuilt at 430, 487, 583, 653,
              749, 860, 996, and 1018, several in the same cycle. VM resolution is a linear
              Array.find at 601, 678, and 774. The CSV is regenerated from the VM list, so
              row count tracks VM count and runCsvSync is quadratic in pool size.

              7. No guard against concurrent cycles

              runEnforcementCycle is reachable from the cron job (905, 941), the test
              action (1000), and xo-server-tag-automation.runSync (1005). None check whether
              a cycle is running. A manual "Run Now" during the hourly tick doubles everything.

              8. Scheduling

              getCron (182) maps hourly to 0 * * * *, and createSchedule (905, 914,
              941, 950) is called with no timezone, so it fires on the hour in appliance local
              time. That collides with every other scheduled job, including backups.

              9. Dead code in isRealVm

              Lines 346 to 349. All three exact matches contain the substring tested on 346,
              so the last three branches are unreachable.

              10. Unverified object model

              location code issue
              337, 338 vm.$type and vm.type both tested only one exists
              404 to 413 four property fallback chain for notes only one is the real field
              601, 678 (v.uuid \|\| v.id) both exist, with defined meanings

              vm._xapiRef (450 to 453, 688, 720, 794, 813) is an internal field that can go
              stale after a reconnect, and nothing revalidates it.

              11. Minor

              • rotateLogs (229) rotates only FILE_LOG; summary and daily logs grow without
                bound.
              • writeRefreshedCsv (640) never quotes fields while parseCsvLine (194)
                handles quotes, so a value containing a double quote round trips incorrectly.
              • The CSV is read twice per cycle (619, 662), three times with autopilot enabled
                (587).
              • configure() (925) reassigns _config, but a cycle already in flight holds
                the old reference.
              johnnezeroJ 1 Reply Last reply Reply Quote 1
              • johnnezeroJ Offline
                johnnezero @fohdeesha
                last edited by

                @fohdeesha Thank you very much for the corrections and details, I will be looking into addressing each issue directly....

                fohdeeshaF 1 Reply Last reply Reply Quote 1
                • fohdeeshaF Offline
                  fohdeesha Vates πŸͺ Pro Support Team @johnnezero
                  last edited by

                  @johnnezero if you copy/paste all those findings to a claude fable/opus or openAI Astra model, it's scarily good at fixing everything, just tell it to test test test πŸ™‚

                  J 1 Reply Last reply Reply Quote 0
                  • J Online
                    john.c @fohdeesha
                    last edited by

                    @fohdeesha said:

                    @johnnezero if you copy/paste all those findings to a claude fable/opus or openAI Astra model, it's scarily good at fixing everything, just tell it to test test test πŸ™‚

                    Not just that but also specify that you need to be professional grade code quality, so that it doesn’t bias itself towards the lower quality code, following it from its massive training dataset.

                    johnnezeroJ 1 Reply Last reply Reply Quote 0
                    • johnnezeroJ Offline
                      johnnezero @john.c
                      last edited by

                      @john.c Yes, will do! Thanks again for all the input, it is greatly appreciated.
                      Happy Day πŸ™‚

                      J 1 Reply Last reply Reply Quote 0
                      • J Online
                        john.c @johnnezero
                        last edited by john.c

                        @johnnezero said:

                        @john.c Yes, will do! Thanks again for all the input, it is greatly appreciated.
                        Happy Day πŸ™‚

                        Also tell it to additionally avoid regressions, as another bad habit of AI models, is it fixes one or two things though at times undoes another fix or fixes made earlier. Tests need to be present that prevent this from happening, in a way that prevents the model from undoing fixes, causing regression issues.

                        Also another bad habit is it can generate overly complex code, in a fashion that’s hard for humans to comprehend and understand. This leads to review and maintenance issues down the line.

                        Finally make sure as part of the professional code quality, it’s done in a secure method.

                        1 Reply Last reply Reply Quote 0

                        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                        With your input, this post could be even better πŸ’—

                        Register Login
                        • First post
                          Last post