[9271] in athena10
Re: [Debathena] #1052: cluster reboots sometimes hang (1)
daemon@ATHENA.MIT.EDU (Debathena Trac)
Thu Jun 28 22:16:51 2012
MIME-Version: 1.0
Content-Type: text/plain; charset="utf-8"
From: "Debathena Trac" <debathena@MIT.EDU>
Cc: debathena@MIT.EDU
To: kaduk@MIT.EDU, jdreed@MIT.EDU
Date: Fri, 29 Jun 2012 02:16:25 -0000
Reply-To:
Message-ID: <056.1a62d8fcaac1ce1cb73e95ec7de61e7f@mit.edu>
In-Reply-To: <041.fa6b6c62bd09966d28310b32d01062b9@mit.edu>
Content-Transfer-Encoding: 8bit
#1052: cluster reboots sometimes hang (1)
---------------------+-----------------------------
Reporter: kaduk | Owner:
Type: defect | Status: new
Priority: blocker | Milestone: Precise Beta
Component: -- | Resolution:
Keywords: | Upstream bug:
---------------------+-----------------------------
Comment (by jdreed):
OK, I dug into the upstart code and looked at a variety of launchpad bugs.
I have a theory. This started with Natty, which coincided with when
Upstart grew support for chroots (see
https://wiki.ubuntu.com/NattyNarwhal/TechnicalOverviewUpstart). I kind of
wonder if the system bus is dying, getting respawned inside the chroot,
and then dying when the chroot exits, and there's no bus with which to
reboot. I wonder if fixing #775 will have any impact on this, and also I
haven't see this failure mode on Precise. But if we do see it, I'd like
to try disabling this support ("--no-sessions" on the command line) and
see if that helps.
--
Ticket URL: <https://athena10.mit.edu/trac/ticket/1052#comment:15>
Debathena <http://debathena.mit.edu>
MIT Debathena Project