[3060] in SIPB-AFS-requests

home help back first fref pref prev next nref lref last post

Re: FY98 budgeting and AFS

daemon@ATHENA.MIT.EDU (mhpower@MIT.EDU)
Wed Jul 15 17:55:16 1998

Date: Wed, 15 Jul 1998 17:55:12 -0400
From: mhpower@MIT.EDU
To: ghudson@MIT.EDU
Cc: sipb-machine-room@MIT.EDU, sipb-afsreq@MIT.EDU
In-Reply-To: "[3059] in SIPB-AFS-requests",
 "[0020] in SIPB machine room issues"

>                                                             ... I
>think it would make our cell easier to manage if we started to
>transition to these units instead of maintaining a plethora of small,
>differently-sized disks.

I think it might make the cell easier to manage by people who
already have a great deal of familiarity with the details of the Box
Hill RAID implementation, perhaps due to working with the
Box Hill RAID implementation in other afs cells. 

>                      ... And, of course, RAID units would make the
>cell a lot more reliable.

I don't think we have much evidence that the particular hardware
configuration in question would, in practice, increase reliability. I
don't believe it follows that if use of RAID would increase the
reliability of the athena cell, then it would also increase the
reliability of the sipb cell. For example, if a RAID unit fails
catastrophically, the magnitude of the problem of restoring file
service is large enough that I believe it's almost mandatory that
there be people whose job it is to restore the service. I don't have
much confidence that a volunteer organization such as sipb would
regularly have the capability of handling this degree of service
restoral in an acceptable amount of time. Because of this, I believe
sipb would be better off with hardware configurations that tend to
have only more localized failures, even if these more localized
failures lead to a greater total amount of data-inaccessibility time
than with much less frequent failures of greater scale.

>        * Get opinions from AFS maintainers about buying RAID units.
>          Does anyone think it's a bad idea?

Yes. I think it's already difficult enough for new people to be able
to maintain sipb's file service, due to use of proprietary software
that isn't in common use at many places outside MIT (for definitions
of "many" that correspond to more than the number of Transarc customer
sites). Box Hill RAID, or for that matter any type of RAID, also most
likely won't have been encountered by an average new person who has
some typical previous Unix system-administration experience (for
definitions of "typical" that mostly correspond to personal Unix PCs).
If we add Box Hill RAID familiarity to the knowledge base required
before someone is able to to deal with sipb-cell disk-failure issues,
I think we'd soon head toward a situation where a very small number
of people were qualified to maintain the sipb cell. Very likely this
would consist of people who maintain the sipb cell and also happen to
be employed maintaining other Box Hill RAID installations. Admittedly,
it could also consist of people who maintain the sipb cell because
they wish to later be hired for positions that involve maintaining
other Box Hill RAID installations. In any case, I think it would
unnecessarily reduce the potential server-maintainer pool and would
not be a productive future direction for sipb to choose.

Matt

home help back first fref pref prev next nref lref last post