XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login

    Veeam 13.1 Rocky9 Linux Appliance: Potential Data Loss with CBT and Workers with Expired Tokens

    Scheduled Pinned Locked Moved Unsolved Backup
    12 Posts 6 Posters 532 Views 8 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • acebmxerA Offline
      acebmxer @msupport
      last edited by acebmxer

      @msupport said:

      The CBT metadata VDIs remain as snapshots, but their VHD parent references point to nothing. The tapdisk crashes, the VM loses disk access, the filesystem becomes corrupt.

      For VM-A (SQL Server), both disks (60GB + 80GB) have broken CBT snapshots. virtual-size=0, physical-utilisation=0.

      For VM-B (File Server), the 80GB disk has a broken CBT snapshot. Interestingly: virtual-size=85899345920 (80GB, same as the base disk) and physical-utilisation=8388608 (8MB). This is different behavior from VM-A — possibly a different code path in the CBT lifecycle.

      In my testing i did see similar issue

      Edit - updated with actual error message from what I am getting. The first full backup were all successful. This first delta backup have 4 vms with this error. The first time I saw similar was only 2 vms.

      8/5/2026 5:34:50 PM Warning : Failed to use CBT: [Task e75c4247-2ee1-7b51-088c-b970ace60f3d (Async.VDI.list_changed_blocks) failed: . SR_BACKEND_FAILURE_460. . Failed to calculate changed blocks for given VDIs. [opterr=Source and target VDI are unrelated]
      

      For me i took it as because i original had backups from veeam 12.x the first xcp-ng beta. Pointed new backups from 13.1 to older backups. I then purged all older backups and started with a fresh backup chain.

      I ended up reverting to the windows base as i was not able to get the windows mount server setup. This was setting up from a windows client using the console to setup the repo.

      1 Reply Last reply Reply Quote 0
      • DanpD Offline
        Danp Pro Support Team @msupport
        last edited by

        @msupport Thanks for the detailed report. Please answer these questions to help us to better understand the source of the issue:

        • What OS were the affected VMs running?
        • Which guest tools were installed?
        • Were the affected VDIs in VHD or QCOW2 format?
        msupportM 2 Replies Last reply Reply Quote 0
        • msupportM Offline
          msupport @Danp
          last edited by msupport

          @Danp

          I have sent detailed log files from XCP-NG, Veeam, and Windows related to the incident to Veeam; the issue has already been escalated, and it appears to be a major problem.
          We're already in touch with Olga from the Veeam team

          Regards
          Micha


          Hello,
          Thank you for contacting Veeam Technical Support! My name is Olga and I will be assisting you with the case from now on.
          Thank you for sharing the logging and the analysis! I will now review the data provided and I will keep you posted on the findings and the next steps.
          Have a nice day,
          Olga Demidova,
          Customer Support Engineer,
          Veeam Technical Support.

          Previous Communication


          2026-08-06 14:11:22 UTC - system Generated Message
          servicenow_email_separator_e8dd76512b2e031caf56f5f7f891bfa7_c79268ad93ea47508575fcffb903d6e9
          ---Do not reply below this line---


          Hello,
          thank you for uploading!
          I provided the files to our R&D team. They will need some time to review everything and I will revert to you as soon as I have any updates.
          Best regards,
          Viktoria Nesmiyanova
          Technical Customer Support - EMEA
          Veeam Software

          Previous Communication


          Hello team,
          hope my email finds you well!
          I would like to inform you that we are actively researching the logs and details with R&D team. I will update you on the investigation status on a regular basis and will let you know as soon as I have any news.
          Thank you!
          Best regards,
          Viktoria Nesmiyanova
          Technical Customer Support - EMEA
          Veeam Software


          Unfortunately, we still haven't received a response from Veeam. We have now integrated the CBT & VM Health button, which checks the following:

          CBT & VM Health Check
          What is checked (4 checks):

          1. VDI Chain Depth — XOA API: Number of snapshots per VDI. Chain > 1 = Veeam has not cleaned up snapshots
          2. VHD Parent — XOA API: VDI.parent — if parent UUID is not in the VDI list = chain broken (critical)
          3. CBT Status — XOA API: VDI.cbt_enabled — indicates whether Changed Block Tracking is active
          4. Veeam Warnings — Veeam REST API: Sessions with a warning status (CBT reset indicator)

          fdab9dad-e37f-4216-92eb-43d71440bbea-image.jpeg


          thank you for your patience!
          The R&D team is still investigating the issue and to proceed further they need to have additional information. Here is the list:

          1. Logs from DC1-XCP-PROD-01 (10.71.1.11)
            Files: /var/log/SMlog*, /var/log/xensource.log*, /var/log/daemon.log*, for Aug 3-8 (including rotated/.gz files).

          2. The output of the commands, executed on any XEN host:
            xe message-list name="failed to unpause tapdisk" params=all
            yum list installed sm xapi

          3. Output of the below command executed for each affected VM disk:
            xe vdi-param-get uuid=<vdi> param-name=sm-config
            VDI UUID should be available in the the storage view. Here is the screenshot from my lab fro the reference:

          4. Details about NTFS corruption: how was it confirmed?
            Please let us know if any of the following info is available: chkdsk output, Windows error codes, disk going read-only, SQL errors, etc. If anything is saved to the files, please attach them to the case.

          5. Windows event logs and SQL Server ERRORLOG from affected VMs for the period of August 1st - August 8th.
            Here is the KB on exporting Windows Event logs (please make sure to include LocaleMetaData): https://www.veeam.com/kb1873
            Additionally, could you please clarify the following points about affected VMs?

          • Do the guests have XCP-ng drivers installed, or generic ones?
          • Timeline: when was each VM's problem first noticed, and was it force-restarted before restoring from backup?
            Also, we would like to ask for your cooperation and not to remove any orphaned disks or leftover snapshots, since this could greatly support our investigation.
            Currently our developers suspect that the issue might be due to the storage cleaning process freezes VM disks briefly during backup, and sometimes fails to un-freeze them. This theory is not fully confirmed as of now, but for the time being we'd recommend to consider those changes, that may decrease the probability of that happening:
          • Increase Windows disk timeout inside affected VMs (registry: HKLM\SYSTEM\CurrentControlSet\Services\Disk\TimeOutValue).
          • Configure the backup jobs to process less VMs located on the same storage simultaneously.
            Please let me know if you have any questions and I will be happy to address them.
            Thank you!
            Best regards,
            Viktoria Nesmiyanova
            Technical Customer Support - EMEA
            Veeam Software

          Previous Communication


          2026-08-17 15:26:13 UTC - Viktoria Nesmiyanova Additional Comments
          Hello Sacha,
          I would like to inform you that R&D team needs a bit more time for investigation. I will keep you informed on the updates from them.

          Thank you!

          Best regards,


          In the meantime, we have installed the Windows version of Veeam 13.1 (latest version 13.1.1.18), which has been working without any issues so far. It is possible that the problem only occurs with the Linux Rocky 9 version.

          1 Reply Last reply Reply Quote 1
          • acebmxerA Offline
            acebmxer
            last edited by acebmxer

            I have also opened a support ticket with Veeam. Well trying to... Issue with my account preventing me to and working to fix that issue.

            Veeam forum did make this statment about CBT and snapshots -

            • There is one known snapshot issue that I'm guessing could possibly have this side-effect (noted in the release notes BTW). We take standard Xen snapshots for CBT but then immediately rename them with our own naming convention. We've noticed that sometimes this rename operation returns success status even if it fails. In this case I can see how this might throw this sort of error but again best to confirm w/support

            Veeam support ticket - 08188639

            Update from Veeam -

            Thanks for the update

            I'm currently checking with my team and I will update as soon as possible, what I noticed in the debug logs is the following API failing

            \Backup\Plugins\XEN\Backup\Backup Job 1\Delilah ArcDC02 with debug enable

            2026-08-07 20:05:04.067 00006 DEBUG | [BackupService]:  <== Request localhost:19000, IsTaskFinished, body: {"taskId":"491371d9-4699-4087-b549-9862d2188177"}
            2026-08-07 20:05:04.067 00006 DEBUG | [BackupService]:  ==> Response localhost:19000, IsTaskFinished, success, duration: 0.1404 msec, body: false
            2026-08-07 20:05:04.843 00019 INFO  | [XenRpcClient]: Start ListChangedBlocks
            2026-08-07 20:05:04.843 00019 DEBUG | [XenRpcClient]: <== Request https://192.168.20.3/Async.VDI.list_changed_blocks, body: ["OpaqueRef:b77ca409-d438-d97e-d8c3-f212647ad469","OpaqueRef:8c83f698-8836-3b4c-1506-8e97315aae1f"]
            2026-08-07 20:05:04.845 00019 DEBUG | [XenRpcClient]: ==> Response https://192.168.20.3/Async.VDI.list_changed_blocks code: "OK", duration: 2 msec, body:{"opaque_ref":"OpaqueRef:c9389aa3-552a-0aa4-a5bc-f9b4f7a40e98"}
            2026-08-07 20:05:04.845 00019 DEBUG | [XenRpcClient]: <== Request https://192.168.20.3/task.get_record, body: ["OpaqueRef:c9389aa3-552a-0aa4-a5bc-f9b4f7a40e98"]
            2026-08-07 20:05:04.847 00019 DEBUG | [XenRpcClient]: ==> Response https://192.168.20.3/task.get_record code: "OK", duration: 2 msec, body:{"uuid":"72b5eb58-5053-d77a-9ee9-089ffaed4541","name_label":"Async.VDI.list_changed_blocks","name_description":"","allowed_operations":["cancel"],"current_operations":{},"created":"2026-08-08T00:05:04Z","finished":"1970-01-01T00:00:00Z","status":"pending","resident_on":"OpaqueRef:05e44734-86d0-c9f6-5079-25633df7bf06","progress":0.0,"type":"<none/>","result":"","error_info":[],"other_config":{},"subtask_of":"OpaqueRef:NULL","subtasks":[],"backtrace":"()","opaque_ref":null}
            2026-08-07 20:05:04.848 00019 DEBUG | [XenRpcClient]: Current task Async.VDI.list_changed_blocks:72b5eb58-5053-d77a-9ee9-089ffaed4541. Status: "pending" Progress: 0
            2026-08-07 20:05:06.848 00018 DEBUG | [XenRpcClient]: <== Request https://192.168.20.3/task.get_record, body: ["OpaqueRef:c9389aa3-552a-0aa4-a5bc-f9b4f7a40e98"]
            2026-08-07 20:05:06.850 00018 DEBUG | [XenRpcClient]: ==> Response https://192.168.20.3/task.get_record code: "OK", duration: 1 msec, body:{"uuid":"72b5eb58-5053-d77a-9ee9-089ffaed4541","name_label":"Async.VDI.list_changed_blocks","name_description":"","allowed_operations":[],"current_operations":{},"created":"2026-08-08T00:05:04Z","finished":"2026-08-08T00:05:05Z","status":"failure","resident_on":"OpaqueRef:05e44734-86d0-c9f6-5079-25633df7bf06","progress":1.0,"type":"<none/>","result":"","error_info":["SR_BACKEND_FAILURE_460","","Failed to calculate changed blocks for given VDIs. [opterr=Source and target VDI are unrelated]",""],"other_config":{},"subtask_of":"OpaqueRef:NULL","subtasks":[],"backtrace":"(((process xapi)(filename lib/backtrace.ml)(line 210))((process xapi)(filename ocaml/xapi/storage_utils.ml)(line 150))((process xapi)(filename ocaml/xapi/message_forwarding.ml)(line 141))((process xapi)(filename ocaml/libs/xapi-stdext/lib/xapi-stdext-pervasives/pervasiveext.ml)(line 24))((process xapi)(filename ocaml/libs/xapi-stdext/lib/xapi-stdext-pervasives/pervasiveext.ml)(line 39))((process xapi)(filename ocaml/xapi/rbac.ml)(line 228))((process xapi)(filename ocaml/xapi/rbac.ml)(line 238))((process xapi)(filename ocaml/xapi/server_helpers.ml)(line 78)))","opaque_ref":null}
            2026-08-07 20:05:06.850 00018 DEBUG | [XenRpcClient]: Current task Async.VDI.list_changed_blocks:72b5eb58-5053-d77a-9ee9-089ffaed4541. Status: "failure" Progress: 1
            2026-08-07 20:05:06.850 00018 DEBUG | [XenRpcClient]: <== Request https://192.168.20.3/task.destroy, body: ["OpaqueRef:c9389aa3-552a-0aa4-a5bc-f9b4f7a40e98"]
            2026-08-07 20:05:06.851 00018 DEBUG | [XenRpcClient]: ==> Response https://192.168.20.3/task.destroy code: "OK", duration: 0.735 msec, body:""
            2026-08-07 20:05:06.851 00018 ERROR | [XenRpcClient]: Failed ListChangedBlocks. Error: [Task 72b5eb58-5053-d77a-9ee9-089ffaed4541 (Async.VDI.list_changed_blocks) failed: . SR_BACKEND_FAILURE_460. . Failed to calculate changed blocks for given VDIs. [opterr=Source and target VDI are unrelated]. 
            Veeam.Vbf.Common.Exceptions.ExceptionWithDetail: [Task 72b5eb58-5053-d77a-9ee9-089ffaed4541 (Async.VDI.list_changed_blocks) failed: . SR_BACKEND_FAILURE_460. . Failed to calculate changed blocks for given VDIs. [opterr=Source and target VDI are unrelated]. 
             ---> Failed to calculate changed blocks for given VDIs.
               --- End of inner exception stack trace ---
               at Veeam.XenBackup.RestClient.XenRpcClient.GetTaskResult(XenRef`1 taskRef, CancellationToken cancellationToken)
               at Veeam.XenBackup.RestClient.XenRpcClient.<>c__DisplayClass94_0.<<ListChangedBlocksAsync>b__0>d.MoveNext()
            --- End of stack trace from previous location ---
               at Veeam.Vbf.Common.Helper.Retry.RetryHelper.ExecuteActionAsync[T](Func`2 asyncAction, String description, ILogger logger, LogLevel logLevel, CancellationToken cancellationToken)
            2026-08-07 20:05:06.851 00018 INFO  | [XenRpcClient]: Retry after 10 sec. Retry 6/10 for ListChangedBlocks 
            

            There is a XEN forum and guide from him, but I need to check internally as first

            https://docs.xenserver.com/en-us/xenserver/developer/changed-block-tracking-guide/troubleshoot.html#you-cant-list-changed-blocks-between-two-vdi-snapshots

            CBT: the thread to centralize your feedback | XCP-ng and XO forum
            https://xcp-ng.org/forum/topic/9268/cbt-the-thread-to-centralize-your-feedback/364
            Regards

            Update -

            Veeam came back and suggested I power off the vms with the warrning and power back on. This cleared the error for 2 out of the 4 vms. 1 vm still showed the error the other vm i was not able to power off at that time.

            @olivierlambert - From veeam...

            After checking with my team, is it possible if you can engage XEN support on this, I checked some of the XEN forum including the KB article about how to troubleshoot CBT errors, seems to be a clean metadata that can be done from their side Troubleshoot Changed Block Tracking | Develop for XenServer

            CBT: the thread to centralize your feedback | XCP-ng and XO forum

            Regards

            Looks like our issue may be different at the end even with involving CBT... Let me know if i should continue in a new post. Support ticket created - Ticket#7762393

            1 Reply Last reply Reply Quote 0
            • M Online
              MajorP93
              last edited by

              @msupport @acebmxer
              Hello guys.
              Thanks for sharing your experience with veeam and reporting these issues.
              My team is also planning on evaluating veeam and I was wondering: did you get a response from veeam?
              Did they give you a time line on when they plan to release a fix?

              msupportM 1 Reply Last reply Reply Quote 0
              • msupportM Offline
                msupport @MajorP93
                last edited by

                @MajorP93

                Last Message from Veeam:
                Veeam Support - Case # 08187386

                thank you for your email!
                We are currently waiting for R&D team conclusion, and a bit more time is required for the investigation. I am sorry for the possible inconveniences here!
                Best regards,
                Viktoria Nesmiyanova
                Technical Customer Support - EMEA
                Veeam Software

                D 1 Reply Last reply Reply Quote 1
                • poddingueP poddingue marked this topic as a question
                • D Offline
                  dsauce @msupport
                  last edited by

                  @msupport We're testing Veeam now. From what I read, I thought CBT was supposed to be disabled in XCP as Veeam uses it's own CBT engine, is that not correct?

                  Also, is it normal for the SR to show a bunch of veeamsnap files for all the VDI's it backed up? I was under the impression Veeam was supposed to remove those when the backup was complete, but perhaps one of them needs to stay for tracking?

                  acebmxerA 1 Reply Last reply Reply Quote 0
                  • acebmxerA Offline
                    acebmxer @dsauce
                    last edited by acebmxer

                    @dsauce

                    To test your theory out on my issue i just tried to disable CBT on XOA side and one vm gave me this error...

                    vdi.set
                    {
                      "id": "11286a97-b773-4ec4-a0b7-464ab87254ca",
                      "cbt": false
                    }
                    {
                      "code": "UUID_INVALID",
                      "params": [
                        "VDI",
                        "a52fe4ab-edc2-4975-a118-f6db84435379"
                      ],
                      "call": {
                        "duration": 1,
                        "method": "VDI.get_by_uuid",
                        "params": [
                          "* session id *",
                          "a52fe4ab-edc2-4975-a118-f6db84435379"
                        ]
                      },
                      "message": "UUID_INVALID(VDI, a52fe4ab-edc2-4975-a118-f6db84435379)",
                      "name": "XapiError",
                      "stack": "XapiError: UUID_INVALID(VDI, a52fe4ab-edc2-4975-a118-f6db84435379)
                        at Function.wrap (file:///usr/local/lib/node_modules/xo-server/node_modules/xen-api/_XapiError.mjs:16:12)
                        at file:///usr/local/lib/node_modules/xo-server/node_modules/xen-api/transports/json-rpc.mjs:38:21
                        at runNextTicks (node:internal/process/task_queues:64:5)
                        at processImmediate (node:internal/timers:452:9)
                        at process.callbackTrampoline (node:internal/async_hooks:130:17)"
                    }
                    

                    Another vm..

                    vdi.set
                    {
                      "id": "966fe072-4865-48a2-867a-0eb59b23dcc7",
                      "cbt": false
                    }
                    {
                      "code": "UUID_INVALID",
                      "params": [
                        "VDI",
                        "d6662a49-0d65-4c4b-a8fa-cddd42ab4cd5"
                      ],
                      "call": {
                        "duration": 1,
                        "method": "VDI.get_by_uuid",
                        "params": [
                          "* session id *",
                          "d6662a49-0d65-4c4b-a8fa-cddd42ab4cd5"
                        ]
                      },
                      "message": "UUID_INVALID(VDI, d6662a49-0d65-4c4b-a8fa-cddd42ab4cd5)",
                      "name": "XapiError",
                      "stack": "XapiError: UUID_INVALID(VDI, d6662a49-0d65-4c4b-a8fa-cddd42ab4cd5)
                        at Function.wrap (file:///usr/local/lib/node_modules/xo-server/node_modules/xen-api/_XapiError.mjs:16:12)
                        at file:///usr/local/lib/node_modules/xo-server/node_modules/xen-api/transports/json-rpc.mjs:38:21
                        at runNextTicks (node:internal/process/task_queues:64:5)
                        at processImmediate (node:internal/timers:452:9)
                        at process.callbackTrampoline (node:internal/async_hooks:130:17)"
                    }
                    

                    Update - after backup completed with warnings i looked back and CBT was re-enabled in XOA. Did not help with my Veeam issues.

                    1 Reply Last reply Reply Quote 0
                    • acebmxerA Offline
                      acebmxer
                      last edited by acebmxer

                      @msupport - Looks like veeam is still working with your on your issues. While veeam has pushed me off to vates / xen.

                      @poddingue - Any updates from Vates about these issues? Is it possible the least patches just pushed might help with either mine or @msupport's issue?

                      Update - Just got a reply back from veeam ...

                      As per internal testing, I’m escalation this to the next tier

                      Regards

                      poddingueP 1 Reply Last reply Reply Quote 0
                      • poddingueP Online
                        poddingue Vates 🪐 @acebmxer
                        last edited by

                        Nothing from our side that I can pass on, sorry.
                        Two public things I can point at, neither of which I've tested against your case: the updates that went live yesterday list tapdisk crash fixes among the storage changes (https://xcp-ng.org/blog/2026/08/18/august-2026-updates-1-for-xcp-ng-8-3-lts/), and there's an open PR on tapdisk picking up a cbtlog disk during commit and failing early because that driver has no commit action (https://github.com/xcp-ng/blktap/pull/17).
                        I'm reading a changelog and a PR body rather than reproducing anything, so treat both as leads.
                        @msupport, Danp's three questions from 6 August are still open (guest OS, guest tools, VHD or QCOW2), and filling those in is probably worth more than anything I can add here.
                        @dsauce, I don't know the answer to yours about CBT and the leftover veeamsnap VDIs, and I'd rather say so than guess on a thread about data loss.

                        1 Reply Last reply Reply Quote 0
                        • msupportM Offline
                          msupport @Danp
                          last edited by

                          @Danp
                          We encountered these issues with 4 VMs running different versions of Windows (2019–2022); one VM had the Rust Guest Agent, while the others had the Xen drivers. All VMs used VHD disks.

                          1 Reply Last reply Reply Quote 0

                          Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                          Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                          With your input, this post could be even better 💗

                          Register Login
                          • First post
                            Last post