[Thanks to James Allen and Jimmy Lin]
Please Read Chapter 3 on Evaluation.
Run the following two queries on Google (http://google.com)
and on Teoma (http://www.teoma.com/) (do
not run the explanation; just run the query). You will judge 10 documents for
relevance to the query. The explanation will help you decide whether or not
something is relevant.
When judging pages for relevance, please note the following:
To decide which 10 documents to judge, use the last digit of your student ID
number:
Email your results in the following format to whomever we chose on Tuesday
[Ben Perry <pianoman@udel.edu>]:
G-or-T query-number doc’s-rank R-or-N doc’s-title
where G-or-T indicates whether you ran this on Google or Teoma, query-number
is 1 or 2 from above, doc’s-rank is the number from 1 to 20 of this document’s
rank, R-or-N is “R” if the page is relevant and “N”
otherwise, and doc’s-title is the title of the document (to help us verify
that results make sense).
As an example, evaluating the first document retrieved by Google on the query "gardening wet soil conditions" might result in the following (fictitious) judgment:
G 1 1 N Gardening for dummies
Since you are judging 10 pages from each of two queries on each of two search engines, the file you email should have 40 non-blank lines.
Whomever we choose on Tuesday [Ben Perry <pianoman@udel.edu>] should email the pooled results back to everyone in the class by the end of the day SEPTEMBER 9!
Class mailing list is: <CISC887-010-05F@udel.edu>
In the second part of this assignment, for one of the topics
[YOUR CHOICE--make this clear], you will
There should be judgments from N different people for each document (Web page).
The first question you'll answer is: How often do judges agree on relevance?
There are four possibilities:
For your chosen topic, figure out how often each case happens (both in terms of counts and in terms of percentage). Turn this information in. Pick two cases where judgments about a particular topic are not uniform, and briefly speculate why this may be so. Turn this in.
Adjudication is simply the process
of reconciling inconsistent judgments. Do this by simple majority voting. If
you have an equal split for a particular document then simply pick one randomly.
The result should be something like this:
G N Gardening for dummies
G R How to deal with wet soil conditions
...
Turn your adjudicated relevance judgments in.
Now, evaluate Teoma and Google using the adjudicated relevance judgments you just created (for the topic you chose). Issue the query to Google and Teoma again, and examine the top 20 hits. Turn in the following information for both search system:
In addition, answer the following questions:
3.A. The following list of R’s and N’s represents relevant (R) and non-relevant (N) documents in a ranked list of 50 documents. The “top” of the ranked list is on the left of the list, so that represents the most highly weighted document, the one that the system believes is most likely to be relevant. The list runs across the page to the right and then starts on the next line. This list shows 8 relevant documents. Assume that there are an additional two (2) relevant documents that were not retrieved by this system.
R R N N N R N N N N R N N N R N N N N R R N N N N N N N R N
N N N N N N N N N N N N N N N N N N N N
Based on that list, calculate the following measures:
3.B. Now, Imagine another system retrieves the following ranked list for the same query.
R R R R N N N N N N N N N N N N N N N N N N N N N N N N N N N
N N N N N N N N N N N N N N N N N N N R
Repeat parts (3.A.1), (3.A.2) and (3.A.3) for the above ranked list. Compare the two ranked lists on the basis of these 3 metrics that you have computed—i.e., if you were given only these 3 numbers (Mean Average Precision, Precision at 50% recall and Precision at 33% recall) what can you determine about the relative performance of the two systems in general?
3.C. Plot a recall/precision graph for the above two systems. Generate both
an uninterpolated and an interpolated graph (probably as two graphs to make
the four plots easier to see). What do the graphs tell you about the system
in A and the one in B?