[992] in Public-Access_Computer_Systems_Forum

home help back first fref pref prev next nref lref last post

Re: Boolean & Probabilistic

daemon@ATHENA.MIT.EDU (news@Think.COM)
Fri Aug 14 10:14:05 1992

Date:         Fri, 14 Aug 1992 09:08:39 CDT
Reply-To: Public-Access Computer Systems Forum <PACS-L%UHUPVM1.BITNET@ricevm1.rice.edu>
From: news@Think.COM
To: Multiple recipients of list PACS-L <PACS-L%UHUPVM1.BITNET@ricevm1.rice.edu>

----------------------------Original message----------------------------
Path: news!ses
From: ses@cmns-sun.think.com (Simon Edward Spero)
Newsgroups: bit.listserv.pacs-l
Subject: Re: Boolean & Probabilistic
Date: 13 Aug 92 17:12:43
Organization: What's your clearance, citizen?
Lines: 29
Message-ID: <SES.92Aug13171243@cmns-sun.think.com>
References: <PACS-L%92081309441914@UHUPVM1.BITNET>
NNTP-Posting-Host: cmns-sun.think.com
In-reply-to: KINGH@SNYSYRV1.BITNET's message of Thu, 13 Aug 1992 09:41:59 CDT

In article <PACS-L%92081309441914@UHUPVM1.BITNET> KINGH@SNYSYRV1.BITNET writes:

   ----------------------------Original message----------------------------
   Who determines relevance?  The definition of relevance in ranked retrieval
   systems is the number of times the keyword occurs, right?  Is there a newer,
   more sophisticated, more valid and reliable method of assessing relevance?

There are a lot of other factors that go in to calculating the weight of a
document than just the number of times the keyword. For example, one very
common technique use the inverse of a terms frequency within the database
to calculate the relative weights of the search terms- rare words are more
important than common words. Also, the number of search terms that are matched
in a document can be used to filter out documents which match only  a few term
but match those terms very closely.

There are other techniques which show a lot of promise, such as using term
co-ocurence frequencies to built neural networks, but I don't know of any
production system that uses this yet.

Another technique which helps improve relevance calculations is called
relevance feedback; once the computer has run your search, you can mark
certain search items as being relevant toyour search; the query is then
modufied to try and find documents as similar as possible to the ones
you have selected.

Simon
--
-----
Guest account at TMC 	|	Just Another  WAIS Hacker | DOD# 0612

home help back first fref pref prev next nref lref last post