[953] in Public-Access_Computer_Systems_Forum
Boolean & Probabilistic
daemon@ATHENA.MIT.EDU (Public-Access Computer Systems For)
Mon Aug 10 11:13:11 1992
Date: Mon, 10 Aug 1992 10:08:25 CDT
Reply-To: Public-Access Computer Systems Forum <PACS-L%UHUPVM1.BITNET@ricevm1.rice.edu>
From: Public-Access Computer Systems Forum <LIBPACS%UHUPVM1.BITNET@ricevm1.rice.edu>
To: Multiple recipients of list PACS-L <PACS-L%UHUPVM1.BITNET@ricevm1.rice.edu>
2 Messages, 68 Lines
*-----
Subject: Boolean & probabilistic
From: popwin@uu.psi.com
Reply-To: kwan@winthrop.org
In response to Charles Hildreth's & Michael Buckland's
messages, I would like to clarify what I meant by "not very good"
in my previous note.
>From my experience in teaching end-user searching, I know the
pain a library patron has to go through to convert his/her
question into a nice boolean query that will retrieve relevant
information for him/her. That is why I was so thrilled to learn
that there's a life beyond boolean. Users can put in queries in
natural language and based on that, system (including probalistic
model) can retrieve information.
I had a chance to do a little trial last year to "compare"
boolean and probabilistic searching. I did it in Aries's
Knowledge Finder because it was the only commercial product that
I knew which had the capability to do both. I chose 5 searches
from our patrons. I searched each one two times. One was
entered as the way patrons put it. The other time, I converted
the query into boolean format myself before searching. I used
the top 20 documents retrieved as the cut off point in the
probabilistic part. All the boolean searches retrieved less than
5 documents each and they were somehow relevant. The
probabilistic searches were not too bad. They did have some
relevant documents in the top 20 but a large percent of them were
way off. I also asked about 10 patrons of their opinion on the
probabilistic retrieval. (They all tried the system in both ways)
To my surprise, only one liked the probabilistic part. They said
they could not "trust" the system.
As you see, my trial was far from scientific. Perhaps, I was
working with a bad probabilistic retrieval system. I don't think
the result I got can judge which retrieval model is superior.
However, I do think users' reaction is very interesting.
Kathy Kwan
kwan@winthrop.org
*-----
From: "Jim Dwyer" <jim_dwyer@macgate.csuchico.edu>
Subject: RE: Boolean & Probabilistic
Just a reminder that whatever the search method--probabilistic, Boolean, simple
browse, etc.--current systems only search on the available data in the records.
Most catalog records have only a few subject headings or added entries. I urge
you to spend part of what you spend on systems in cataloging to add additional
access points including keyword indexed contents notes whenever possible.
Using the sort of capabilities discussed here to access the typical
bibliographic database is like buying an expensive food processor just to chop
carrots, or, to use Charles Hildreth's probably unintentional environmentally
irresponsible analogy, using a Volvo to drive to the corner store to buy a loaf
of bread. Shouldn't you walk or ride your bike? My point here (other than the
environmental one to "always think green") is that it's great to talk blue sky,
but many online catalog vendors still oversell rather crude and ineffective
systems. They promise the moon and deliver a meteor. Buy basic proven
functionality. Consider advanced software vaporware until the vendor can
deliver the real thing. If they tell you they do quality assurance in-house,
ask for a demonstration on something larger than a special test database after
you have stopped laughing or scowling. If you agree to be a beta site expect to
spend the next several months on the phone and demand an honest return and a
very special deal for your efforts.
Jim "I have been to the beta" Dwyer
Chico State University
Speaking only for myself, but I'm sure that not all the blood, bruises, and
infections from dull or rusty cutting edges are mine...