[8568] in Commercialization & Privatization of the Internet
Software archiving/searching
daemon@ATHENA.MIT.EDU (Peter Deutsch)
Mon Nov 22 20:29:44 1993
From: Peter Deutsch <peterd@bunyip.com>
Date: Mon, 22 Nov 1993 20:20:32 -0500
In-Reply-To: Barry Shein's message as of Nov 22, 19:27
To: bzs@world.std.com (Barry Shein), com-priv@psi.com
Cc: bajan@bunyip.com
g'day Barry,
[ You wrote: ]
>
> A MODEST PROPOSAL...
>
> I think the time is long overdue to sit down with some library science
> types and come up with a scheme for cataloguing software based on
> something similar to the Library of Congress catalogue schemes
> incorporating major application area, minor application areas, perhaps
> source language, etc into one tidy (as is practical) scheme (think:
> catalogue number or ISSN etc.)
As somebody said "There is nothing so powerful as an idea
whose time has come."
In fact, there has been a great deal of activity towards
this goal in the last little while. Here's just a quick
ten cent tour of some of the related work.
Over the past year of two there have been a number of
meetings between people in the library community, the
Internet services developer community, operators of
archive sites and services and so on. The gang at CNIDR
(Coalition for Network Information Discovery and
Retrieval), people from the Library of Congress and other
library groups, Internet service developers (including
people from Bunyip, Pandora Systems, WAIS Inc. and others)
and quite a few others have met a number of times to
discuss problems of common interest.
These meetings have occured at Educom and CNI meetings,
Usenix, the IETF and Interop (to name a few that I'm aware
of). The list of participants varies over time, but there
is a core group that seems to keep showing up, thus
ensuring continued representation from the various
communities with a stake in the outcome.
A number of things have come of all of these meetings. For
example, at the Washington IETF at the start of the year
an ad hoc group reviewed proposed changes to the US MARC
record format intended to allow cataloguing of online
information. There is work afoot to standardize
cataloguing and naming on the Internet, as well as work to
extend current search and retrieval protocols and develop
new ones.
Here's a pocket summary of the group's activities of late:
IAFA Records:
-------------
There is now a draft document (of which I am the coauthor
with my partner Alan Emtage) which represents a prelimiary
attempt to catalogue software, mailing lists, services and
a number of other informational objects on the net. This
document is the output of the Internet Anonymous FTP
Archives Working Group (IAFA) at the IETF and is currently
an Internet Draft, destined eventually we hope to be an
informational RFC. It describes each information record as
a set of attribute/value pairs and defines a mimimal
subset of these attributes which can be used to document
each object type.
We know this document is not perfect but it does represent
an initial attempt to get things started to allow us all
to gain some working experience in this medium. Feedback
and comments on the document are most welcome and should be
addressed to the working group's mailing list
"iafa@bunyip.com". The usual "-request" convention to
subscribe.
We recently sent out a first announcement to the anonFTP
server administrators we track in archie asking them to
consider filling out such cataloguing info so that
automated tools (such as ours) can collect and distribute
this information to others. Initial response has been quite
favourable, and I know of people now working on "quick and
dirty" docs explaining how to get started, simple software
to automate much of the tedious job of completing records
and so on.
Naming:
-------
There is also work afoot at the IETF to standardize naming
and references for the Internet. This work is being done
in the Uniform Resource Identifer Working Group (URI) and
is currently concentrated on producing standard pointers
(Uniform Resource Locators) and names (Uniform Resource
Names). The mailing list for this work is "uri@bunyip.com"
(again, add a "-request" when signing up).
URLs are intended to provide a mechanism for
non-persistent references (in effect, pointers to
information on the net). They are expected to find use in
tools such as archie, gopher, WWW and WAIS as transient
pointers to resources and information objects across the
net and in fact developers for all of these tools are
active in the working group.
URNs are intended to provide more long-lasting names, and
are intended to exist in an environment where users would
be able to dereference them into a URL for the
corresponding object, assuming a copy exists on the net.
They correspond much more closely with the standard
librarian's ISBN reference, although the implicit ability
to dereference to a URL, and even automate the entire
dereference and fetch cycle, obviously holds out the
promise of greater power once they're fully deployed.
An initial draft of the URL specification has been under
development for some time now, but the working group
believes that the worst on this one is now behind us and
we hope to see it advance to draft status soon. An initial
draft of the URN functional specifications and syntax
documents are now available, although there are still
disagreements among the group on some points and these are
probably farther back up the pike at this point.
Standardizing protocols:
------------------------
Efforts to standardize information protocols at the IETF
have taken a couple of different approaches. First, there
is a liason effort, typified by a push to get the various
Internet service protocol developers to publish their
current practice as Informational RFCs. The initial gopher
protocol RFC is now out in draft form, the initial WAIS
protocol doc was presented at the last IETF in Houston and
I understand that the authors of both the HTTP
(WorldWideWeb) and Prospero system protocols (also used as
the archie access protocol) intend to follow this route.
There is also some work afoot to define a new general
purpose directory search protocol known somewhat
whimsically as "WHOIS++". This protocol uses a generalized
data template model and simple query syntax and can be
used for white pages or yellow pages services, as well as
any other collection of information organized in general
attribute/value form. The protocol has been under
development for about a year now and is about to go into
the pipeline as an Internet Draft. Several initial server
implementations are already available and work continues
apace.
Other Activities:
-----------------
There's a lot more going on, but I trust this gives
everyone a flavour of the sort of work that is underway to
tame the Internet frontier. It will take some time for the
results of all of this work to find their way into our
tools, but if you want some flavour of what can be done,
take a look at Mosaic. it uses an initial draft form of
URLs now and provides a multiple protocol front end onto
the Internet. Once we have the added ability to name and
dereference objects, services such as archie will provide
such references in a standard format and things will
really cut loose. Watch this space...
> This would be extremely useful in numerous contexts.
>
> I realize some of the difficulties (e.g. a standardized method of
> assignation), but I don't believe any stand in the way of devising a
> scheme and fall mostly in the realm of what you do with it once it's
> been codified. Even so, a badly applied scheme would be far superior
> to the current chaos. Failings can be repaired incrementally over the
> next few centuries, better methods (e.g. Q&A programs) to help authors
> suggest catalogue numbers developed, etc.
Given the number of communities who will want to be
represented in the creation of these new cataloguing and
serving tools I'm not sure that there can be a single such
encoding scheme, but the above activities will I trust
provide us with a toolkit from which we can draw to help
move to more useful services.
Addressing your specific concern of cataloguing programs,
I'd point you at the IAFA Software Description templates
and ask anyone who produces software to consider making
the creation and maintenance of such a template a standard
part of the software release cycle. We're already
preparing to add these records to our existing archie
service, assuming they are completed and available.
- peterd
--
-----------------------------------------------------------------------------
"The Internet destroys the Greek tragedy of time and space..."
- Daniel Pimienta <pimienta!daniel@redid.org.do>
-----------------------------------------------------------------------------