[3227] in SIPB-AFS-requests
Re: migrating rosebud to a Sun
daemon@ATHENA.MIT.EDU (mhpower@MIT.EDU)
Mon Jan 11 14:46:34 1999
Date: Mon, 11 Jan 1999 14:46:27 -0500
From: mhpower@MIT.EDU
To: zacheiss@MIT.EDU
Cc: sipb-afsreq@MIT.EDU
In-Reply-To: "[3220] in SIPB-AFS-requests"
> ... One
>problem here is that the system disk in the SS5 might be big enough to
>hold one copy of /os and /srvd local (although it would be tight), but
>can't hold two, making OS upgrades very challenging.
I'm not sure I know what you mean by "OS upgrades". Is there some
expectation that the server would move from SunOS 5.6 to SunOS 5.7
while remaining on the same hardware? I don't believe that's
realistic. I think we would purchase replacement hardware in a
shorter time frame than that type of OS upgrade. This is, at least,
historically true of sipb, and also seems sensible given the cost of
obtaining faster Sun hardware relative to the sipb server budget.
Because of this, I don't think it's at all necessary to have enough
local disk space to hold the collections of files associated with both
SunOS 5.6 and SunOS 5.7.
> ... Tracking all of /os and
>/srvd local and then applying the Sun recommended patches takes quite a
>while.
> ... Even if one halves this time, it plus the
>time necessary to restore the AFS data on the disk adds up to a long
>time between the disk in question dying and the server being returned to
>service.
In the case of a server machine dying, I don't think that copying files
out of AFS via track and applying Sun patches is a reasonable way to
get the machine running again. Because of this, the amount of time
that that would take doesn't seem to me particularly relevant.
Recovering servers' system disks from backups should be faster and
simpler. We should have some type of tape drive that can be brought
into the machine room and hooked to servers when necessary for
recovery, and we should have tape backups of the local disks of server
machines. In the case of AFS servers, if we happen to be in a state
where multiple servers are supposed to be identical except for a few
files specifying the hostname and IP address, it should be reasonable
to have backups of only one such server. Also, in the case of AFS
servers where new content (except under the separately backed up
/vicep filesystems) is not added regularly, it should be reasonable to
have backups much less often than once a week.
Another advantage here is that getting the machine running again
requires less specialized knowledge. It should be fairly
straightforward for people familiar with maintaining Unix machines to
restore a server's filesystems from backups. Additional recovery
tasks, such as applying Sun patches (and dealing with cases where a
patch doesn't apply correctly), require a different set of skills
that I believe fewer people have. Unlike ASO, sipb has no guarantee
that the maintainers available at a time of a server failure will have
various useful skills such as understanding the Sun patch system.
Because of this, I think sipb's method of recovering from a failure of
a server system disk ought to be less complicated than ASO's.
Also, if the concern were only with bringing back dead AFS servers, it
might be somewhat more cost effective to have a spare external disk,
that was normally unused, on one AFS server with a copy of the various
system filesystems (/, /usr, /var, /os, /srvd) on it. This would allow
much quicker recovery from system-disk failures. Also, patches
could be regularly applied to content on the spare disk. Having a tape
drive is mainly of interest if we only wished to spend a few hundred
dollars and wanted that expense shared across all types of servers.
Matt