@racom.jiristerba
Sorry, I missed answering your questions the first time
Is the tapdisk QCOW2 commit failure “The device is not writable: Permission denied” a known issue with these package versions on shared LVM/iSCSI storage?
Yes, it's a known issue, most of them will be fixed with the latest update, you might need to activate the VDI again but a migration during RPU will be enough
Why does rollback attempt to deactivate the active guest LV while tapdisk still holds it open?
It's because in the case of VHD, it's the case but with QCOW2 we don't stop the tapdisk process accessing the VDI, it's a change that was missed
Is the qcow2OLD_<UUID> naming seen in rollback expected, compared with QCOW2-OLD_<UUID> seen during successful cleanup?
It's a pre-existing bug that I already have on my TODO list
Is there a supported update, hotfix, or workaround for this configuration?
Installing the latest release sm-3.2.12-25.1 during updates (and maybe launching a xe sr-scan uuid=<SR UUI> after updating it so it auto-resolve the undo) should be enough
What additional logs are required to identify the original failure of the second QCOW2 VM?
I don't think we need any more logs since the errors I'm seeing should already be fixed.
If you have any more issues after installing the newest packages, I will take another look
What is the recommended recovery procedure without a guest outage, and how should we validate the disk chains before resuming snapshot backups?
The sr-scan after updating should do it automatically, it shouldn't need any manipulation. Looking at the storage logs in /var/log/SMlog for any irregularities could help to see problems.