This is an archived copy of the Xen.org mailing list, which we have preserved to ensure that existing links to archives are not broken. The live archive, which contains the latest emails, can be found at http://lists.xen.org/
Home Products Support Community News


[Xen-changelog] [xen-unstable] remus: proper cleanup on checkpoint failu

To: xen-changelog@xxxxxxxxxxxxxxxxxxx
Subject: [Xen-changelog] [xen-unstable] remus: proper cleanup on checkpoint failure.
From: Xen patchbot-unstable <patchbot@xxxxxxx>
Date: Sat, 09 Apr 2011 09:20:30 +0100
Delivery-date: Sat, 09 Apr 2011 01:27:47 -0700
Envelope-to: www-data@xxxxxxxxxxxxxxxxxxx
List-help: <mailto:xen-changelog-request@lists.xensource.com?subject=help>
List-id: BK change log <xen-changelog.lists.xensource.com>
List-post: <mailto:xen-changelog@lists.xensource.com>
List-subscribe: <http://lists.xensource.com/mailman/listinfo/xen-changelog>, <mailto:xen-changelog-request@lists.xensource.com?subject=subscribe>
List-unsubscribe: <http://lists.xensource.com/mailman/listinfo/xen-changelog>, <mailto:xen-changelog-request@lists.xensource.com?subject=unsubscribe>
Reply-to: xen-devel@xxxxxxxxxxxxxxxxxxx
Sender: xen-changelog-bounces@xxxxxxxxxxxxxxxxxxx
# HG changeset patch
# User Shriram Rajagopalan <rshriram@xxxxxxxxx>
# Date 1302277744 -3600
# Node ID 13ec53a59a42d8a74a95e9439096d68e81ac2f32
# Parent  e917931a698b84293e94971209db62e37d7fcff8
remus: proper cleanup on checkpoint failure.

While running remus, when an error occurs during checkpointing
(e.g., timeouts on primary, failing to checkpoint network buffer
or disk or even communication failure) the domU is sometimes
left in suspended state on primary. Instead of blindly closing
the checkpoint file handle, attempt to resume the domain before
the close.

Signed-off-by: Shriram Rajagopalan <rshriram@xxxxxxxxx>
Committed-by: Ian Jackson <ian.jackson@xxxxxxxxxxxxx>

diff -r e917931a698b -r 13ec53a59a42 
--- a/tools/python/xen/lowlevel/checkpoint/checkpoint.c Fri Apr 08 16:40:58 
2011 +0100
+++ b/tools/python/xen/lowlevel/checkpoint/checkpoint.c Fri Apr 08 16:49:04 
2011 +0100
@@ -80,6 +80,9 @@
   CheckpointObject* self = (CheckpointObject*)obj;
+  if (checkpoint_resume(&self->cps) < 0)
+    fprintf(stderr, "%s\n", checkpoint_error(&self->cps));
diff -r e917931a698b -r 13ec53a59a42 tools/python/xen/remus/save.py
--- a/tools/python/xen/remus/save.py    Fri Apr 08 16:40:58 2011 +0100
+++ b/tools/python/xen/remus/save.py    Fri Apr 08 16:49:04 2011 +0100
@@ -158,9 +158,13 @@
             self.checkpointer.start(self.fd, self.suspendcb, self.resumecb,
                                     self.checkpointcb, self.interval)
-            self.checkpointer.close()
         except xen.lowlevel.checkpoint.error, e:
             raise CheckpointError(e)
+        finally:
+            try: #errors in checkpoint close are not critical atm.
+                self.checkpointer.close()
+            except:
+                pass
     def _resume(self):
         """low-overhead version of XendDomainInfo.resumeDomain"""

Xen-changelog mailing list

<Prev in Thread] Current Thread [Next in Thread>
  • [Xen-changelog] [xen-unstable] remus: proper cleanup on checkpoint failure., Xen patchbot-unstable <=