We are running into this wired issue infra stack is vmware 7, vmware 8, 2 instances of Commvault one with vm proxies used for VMWARE 7 another one is HSX with HSX nodes for VMWARE 8. Datastore is on pure storage. We are in transit and evenrutally migrate to HSX
This is happening on both the environments V7 and V8 (separate cv instances)
The issue we run into is that when there are open snapshots; the Intellisnap backup randomly goes into hung state; where in it gets stuck during the phase(pre hardware snapshot phase) where VMWARE snapshots are taken and deleted and never proceeds into taking the storage snapshot; as a result the next morning there are open snapshots on all the VMs for that job during production hours which I then manually cull using a powershell script. What I don’t understand is how does manual deletion works (using a script) when Commvault support says they are getting timeout in the logs when trying to delete the snapshot and that could be the issue.
RemoveSnapshot() - Failed to remove Snapshot GX_BACKUP snapshot-200805 from VM: The operation has timed out
__WaitForTask() - Exceeded Maximum Wait for Task Attempt
I have been told that we never used to run into this issue when using nutanix; also this issue does not occur when we use traditional vmware streaming non-Intellisnap jobs.
The way we overcome this is by creating an exclusion (via tags) where we exclude all the VMs that have an open snapshots and the Intellisnap backups run fine.
Given that it is happening across environments and different instances of Commvault as well not sure how should I solve it. The work around works however I would rather resolve it.
I will be supplying fresh logs as the backups run when we have exclusions in place to see if we get the same timeout issue. I understand this is a vmware issue as CV is just waiting on VMWARE to create/delete the snapshot however not sure if changing these settings at CV end will change much; if it were a configuration issue at CV would it not happen irrespective of snapshots on the vm and impact all the jobs and the reason that it is happenning across diffrent vmware envrioments and CV instances.

