[3050] in SIPB-AFS-requests
Proceeding with new-rosebud
daemon@ATHENA.MIT.EDU (John Hawkinson)
Fri Jul 10 08:03:25 1998
Date: Fri, 10 Jul 1998 08:03:06 -0400
To: sipb-afsreq@MIT.EDU
From: John Hawkinson <jhawk@MIT.EDU>
Previously I had made noises about getting rosebud's replacement up over
IAP. Obviously this did not happen, but I'd like to make sure that it (and
ancillary happenings, see below) are complete before the start of
term. Hopefully substantially sooner than that.
This note tries to summarize discussions I've had with some maintainers about
how to procede; sorry if I omitted anyone's positions or opinions on the
relevent topics.
At this point, we have four disks on-hand, and are awaiting delivery on the
fifth (a replacement). While there are some possible issues with the quality of
the disks (c.f. [3047]), I'm working those in the background so we'll see what
happens, and would like to see us move forward with planning for the rosebud
upgrade.
A number of questions arise:
1. Should it run SunOS or Solaris?
2. If Solaris, should it run Solaris 2.6?
3. What version of AFS should it run?
4. Should it run an Athenized operating system or not?
I'll try to summarize both sides of each, and indicate my preference.
1. Should it run SunOS or Solaris?
A: No; SunOS is an operating system that SIPB has run AFS servers under before,
and nwhich we're reasonably confident of. It will run AFS 3.3a, so the rosebud
would not necessitate an AFS cell upgrade. It has not been unreliable for us in
the past, and has a high comfort level with current AFS cell maintainers.
B: Yes; Solaris is Sun's current "supported" operating system. While Sun claims
to support SunOS, in point of fact, they are slow to release patches and fixes
(security or otherwise), and actively recommend that customers use Solaris
instead. Solaris has comparable or better performance for almost all hardware
configurations (including ours), and more modern operating system
features. It's drastically easier to apply vendor patches and fixes. It's an
operating system that new SIPB service maintainers are more likely to be
familiar with than SunOS. Vendor patches and fixes are much more readily
available. Current Solaris is Y2K-compliant, and given SIPB's propensity to not
upgrade server operating systems, this has relevence. Solaris changes a number
of hardcoded manually-tuned parameters into adaptive and
automatically-configured parameters.
jhawk: I'd much rather run Solaris than SunOS. (B)
2. If Solaris, should it run Solaris 2.6?
A: No; SIPB currently has Solaris servers that run 2.5.1, and it's what we're
comfortable with. The current public Athena release is based upon 2.5.1.
Server maintainer's comfort level with 2.5.1 is greater than that of 2.6, which
is somewhat of an unknown. ASO is running 2.5.1 servers but not 2.6 servers.
B: Yes; Solaris 2.6 has been out for quite some time (over a year), and Sun is
getting ready to release 2.7. Solaris 2.6 is more readily supported than 2.5.1,
typically with security patches and fixes ahead of 2.5.1 by one to two
weeks. Solaris 2.6 is Y2K compliant, and Solaris 2.5.1 lags a fair bit in that
regard, though there may be some patches. Solaris 2.6 has more modern operating
system features than 2.5.1, but is not a light-year jump like SunOS ->
Solaris. The 8.2 beta to date has not uncovered any significant 2.5.1/2.6
failures-to-play-nice. We'll have to run 2.6 eventually if we deploy UltraSPARC
AFS servers.
jhawk: I'd like to leave the painful position of being in the dark
ages of operating system versions. I'd much rather run 2.6, I find it
a significantly friendlier OS. Athena 8.2 also has a number of excellent
benefits over Athena 8.1, and 2.6 instead of 2.5.1 brings Athena 8.2,
should we select Athenization. (B)
3. What version of AFS should it run?
3.3a: This is largely governed by the OS version installed. In the
event that SunOS is selected, 3.3a can be installed, maintaining the
cell at 3.3a.
3.4a: In the event that 3.4a is installed, the cell needs to be migrated to
3.4a. This involves some pain and careful machinations (See below, b) 3.4a
Migration). On the other hand, the Athena cell is already at 3.4a and it would
be very nice to be running both current AFS software, as well as running
software consistent with that of the athena.mit.edu.
jhawk: I'm already bound into 3.4a by my preference for Solaris 2.6,
but even independantly of that, I'd favor it. (3.4a)
4. Should it run an Athenized operating system or not?
A: Running a "vanilla" operating system means that we have the flexibility
to apply vendor patches and fixes. We are also freed of any security holes,
problems, or unfortunate baggage that accompanies Athena software.
B: Athena software is damned useful in our environment. With the relatively
open development model that is currently available, it's quite feasable to
get local features and requests into the Athena release where they make sense.
By using Athena software, we can leverage existing installation tools,
and avoid building lots of software on our own that is included with part
of the release.
jhawk: All of our current servers are Athenized to some degree, and
consistency is useful. A proof-of-concept test showed that I had no
problem applying all Sun recommended and security patches to 8.2.7
with a local /os and /srvd. (B)
Assuming there is some level of concensus that the "jhawk:" versions
carry:
1. Should it run SunOS or Solaris? Solaris
2. If Solaris, should it run Solaris 2.6? Yes, 2.6
3. What version of AFS should it run? 3.4a
4. Should it run an Athenized operating system or not? Yes
Some may argue that we're being too aggressive, and that we should
delay behind Ops in moving forward with this stuff. I think almost
the opposite, SIPB has the ability to be more aggressive, and I think
we should utilize that capability to our advantage.
Then there are some issues that are worthy of discussion:
a) Patches
b) 3.4a Migration
c) Data migration strategy
d) 2.6 server binaries
e) Machine configuration reproducibility
a) Patches
I think it's important that, at the time of installation, SIPB servers should
run current versions of all Sun recommended and security patches. It's
unfortunately often the case that Athena lags behind Sun by significant
amounts, even though it appears now that more effort is being payed to this
topic by Athena dev and releng.
While it is the case that applying patches potentially destabilizes the system
and introduces new bugs, Sun's patch release process does involve reasonable
tests of those patches before releasing them to customers, and in general
recommended or security patches have something of relative import associated
with them.
Additionally, as part of the version of Solaris on the install server
(currently Sun's 3/98 Solaris "update" release, I believe, for 2.6), a number
of non-security, non-recommended patches are installed. Roughly three
courses of action are possible:
i) Update all the currently-installed patches to current revisions. In
some cases, because a recent revision of patch #1 might depend on
patches #2 and #3, it may involve installing additional patches. It
may also be that patches #2 and #3 may be whoally unrelated to our
architecture or system, yet applying them would be necessary according
to this policy.
ii) Update none of the currently-installed but downrev patches. Potentially
something important may be missed. Additionally, when reporting or
analyzing problems which are affected by a patch, Sun requests that
current patches be installed.
iii) Evaluate the set of update-able patches on a case-by-case basis and
decide whether to upgrade one-by-one.
I'm not really sure how I feel about this issue. I nominally favor i), but
could be easily dissuaded. I think deciding might be best accomplished by
evaluating iii) and seeing if that suggests a particular course of action.
There also tend to be some number of non-recommended, non-security patches
that may be worth installing. I believe there are two of these at present
for Solaris 2.6, one of which is fixes for breakpoint problems in kadb. Not
something we're likely to frequently use on a server machine, but something
that should certainly work right if we ever need it. In any case,
any of these should get evaluated on a case-by-case basis.
The last issue with patches is the question of application of patches
to the OS as time goes on, post-installation. This is a topic frought
with peril. Installtion of patches on a running server may necessitate
reboots after those patches, or living under conditions that the vendor
may render inadvisable. This is not true for all patches, of course.
On the other hand, living with security holes and bugs is irritating.
Additionally, if machines are patched to current levels at installation
time, but patches are not maintained, machines installed at different times
will have different patchsets, and thus different operating environments.
I suspect it will be hard to come to resolution on this issue, so I would
suggest that again such patches be evaluated on a case-by-case basis.
b) 3.4a Migration
In the event that Solaris is installed, it's necessary to run 3.4a. Even if
not, it may be desirable.
It is inadvisable to mix-and-match 3.3a and 3.4a vlservers and ptservers,
so a migration strategy might be:
i) Remove rosebud's vlserver, ptserver, and buserver instances. Clients may
occasionally have to time out rosebud as a vlserver in the event
it remains in their CellServDBs, but either way this is minimal.
A plausible compromise between notifying the world and client slowdowns
is to have the Athena CellServDB updated, but not the worlds.
ii) Upgrade ronald-ann and reynelda's vlserver, ptserver, and buserver
to 3.4a. This is necessary as the 3.4a fileserver will not talk to
a 3.3a fileserver, but a 3.3a fileserver will talk to 3.4a database
servers.
iii) Replace the old rosebud with the new (see "c) Data migration
strategy", for more info), running 3.4a. The cell is now all Sun
SPARCstation-5 servers.
iv) Upgrade ronald-ann and reynelda's fileservers to 3.4a. The cell
is now all 3.4a.
v) Reinstate rosebud's vlserver, ptserver, and buserver instances.
The database servers are now redundant and can maintain quorum in
the event of a single-server failure. We're all done.
It would also be possible to ra and rc to Solaris in this process, but that
seems like a bad idea for a number of reasons; much better to let us gain
experience with one Solaris server and work out the bugs (additionally,
operating system upgrades are a pain in the neck and I think might constitute
"too many variables").
I think this is the way to go. I've confirmed my understanding of
3.4 to 3.3a interactions against the 3.4a release notes. See
/mit/afsdev/src/3.4a/src/doc/rel3_4a.ps.
An alternative option was to install new-rosebud running SunOS, upgrade the
cell as a whole to 3.4a, and then take it out of commission and upgrade it
to Solaris. That strikes me as more trouble, more work, and more outages
for user data.
With my preferred strategy, there is some question how much time
should happen in between each step. There's a desire to minimize the
overall length of the process, and to not run with only two database
servers for an extended period of time. On the other hand, it's desireable
to reduce the variables in each stage, and ensure that each step is stable.
I would suggest that waiting is important at these junctures:
ii) -> iii)
iii) -> iv)
I think that a wait of no less than half an hour, and no more than a week,
would be appropraite at each of those points. I'd be interested in opinions
as to what is best.
c) Data migration strategy
There are some choices as to how to migrate data off of old-rosebud
and onto new-rosebud. I believe the choies are:
i) Bring up new-rosebud as sipb-server-1, make it a fileserver,
and change it's address to rosebud afterwards. This allows us to
vos move volumes directly from old-rosebud to new-rosebud.
ii) Take down old-rosebud, remove it's disks, attach them to new-rosebud,
and bring it up serving files.
iii) Take one of the disks destined for new-rosebud, install it on
reynelda, vos move data from old-rosebud to reynelda, bring down
old-rosebud, bring up new-rosebud, and move data from reynelda to
rosebud. Optional whether the disk placed on reynelda stays there,
or whether reynelda goes down for an outage to move the disk back.
Note that 'vos changeaddr' is obsolete under 3.4a and that changing addresses
of servers is supposedly pretty easy.
Of all of these, I think i) makes the most sense. It gives us pretty easy
flexibility to continue to use rosebud as a fileserver and an easy backout
mechanism. It also means that there will be no outages of volumes visible
to users, except for read/write volumes during moves.
d) 2.6 server binaries
Jonathon has observed that, because of the fact that we (MIT) build our own
AFS server binaries, we do not currently have an AFS build for Solaris 2.6.
It's relatively easy to remedy this, though we need to queue a request to
Dev for it (possibly through an expedited path), or consider building our own.
Building our own is discouraged since I think everything is happiest and best
if the SIPB cell runs the same binaries as the Athena cell.
e) Machine configuration reproducibility
I think that it's important that we have a reasonable method of installing
the machine that we can reproduce in a reasonable fashion. While this doesn't
have to be (and probably cannot reasonbly be) as self-contained as
mkserv, I would like to avoid the situation where in order to duplicate
a machine's configuration, either a disk needs to be copied from
another machine, or tens or scores of discuss transactions must be found
and read through to track changes.
I recognize this nominal "requirement" may make some lives difficult,
particularly when checking out security issues. Nevertheless, I think
it's to our advantage on the whole in the future.
I'd like to see a timeline such as:
25 July Disk issues resolved
1 August Rest of cell to 3.4a
15 August New-Rosebud fully functional
those being relatively final dates -- I'd hope that we could alctually
take much less time.
That's all, folks.
--jhawk
(yawn!)