[9339] in athena10
Re: [Debathena] #1052: cluster reboots sometimes hang (1)
daemon@ATHENA.MIT.EDU (Debathena Trac)
Mon Jul 2 15:54:21 2012
MIME-Version: 1.0
Content-Type: text/plain; charset="utf-8"
From: "Debathena Trac" <debathena@MIT.EDU>
Cc: debathena@MIT.EDU
To: kaduk@MIT.EDU, jdreed@MIT.EDU
Date: Mon, 02 Jul 2012 19:54:17 -0000
Reply-To:
Message-ID: <056.0047add76b5fbc9a0c1023a504b49409@mit.edu>
In-Reply-To: <041.fa6b6c62bd09966d28310b32d01062b9@mit.edu>
Content-Transfer-Encoding: 8bit
#1052: cluster reboots sometimes hang (1)
---------------------+-----------------------------
Reporter: kaduk | Owner:
Type: defect | Status: new
Priority: blocker | Milestone: Precise Beta
Component: -- | Resolution:
Keywords: | Upstream bug:
---------------------+-----------------------------
Comment (by jdreed):
Some things are interesting about the qs-11-1 logs:
-modem-manager and wpa_supplicant were started inside the chroot, thanks
to D-Bus doing... something. They are also the only processes (other than
metrics, which hasn't yet been killed at this point, and probably should
be -- see the fact that debathena-reactivate-cleanup runs _first_ in the
postsession, not last) that stuck around inside the chroot. Even more
confusing is how they got started, because this is an Optiplex 760, which
lacks both a modem and a wireless card.
I didn't add the self-destruct call yet, it's easy enough to do later if
we think it's a good (or at least "not terrible") idea. I did add some
more logging, as well as the ability for snapshot-run to attempt to
recover and end the session if something fails early enough. I wonder if
it's worth having 16-killprocs-no-really go through its process killing
twice, so if something survived the kill -9, we at least know that was
attempted (as opposed to somehow missed, or worse, the process was
restarted after it was killed).
--
Ticket URL: <https://athena10.mit.edu/trac/ticket/1052#comment:19>
Debathena <http://debathena.mit.edu>
MIT Debathena Project