Evaluation with informational and navigational intents

研究成果: Conference contribution

33 引用 (Scopus)

抄録

Given an ambiguous or underspecified query, search result diversification aims at accomodating different user intents within a single "entry- point" result page. However, some intents are informational, for which many relevant pages may help, while others are navigational, for which only one web page is required. We propose new evaluation metrics for search result diversification that considers this distinction, as well as a simple method for comparing the intuitiveness of a given pair of metrics quantitatively. Our main experimental findings are: (a) In terms of discriminative power which reflects statistical reliability, the proposed metrics, DIN#- nDCG and P+Q#, are comparable to intent recall and D-nDCG, and possibly superior to α-nDCG; (b) In terms of preference agreement with intent recall, P+Q# is superior to other diversity metrics and therefore may be the most intuitive as a metric that emphasises diversity; and (c) In terms of preference agreement with effective precision, DIN#-nDCG is superior to other diversity metrics and therefore may be the most intuitive as a metric that emphasises relevance. Moreover, DIN#-nDCG may be the most intuitive as a metric that considers both diversity and relevance. In addition, we demonstrate that the randomised Tukey's Honestly Significant Differences test that takes the entire set of available runs into account is substantially more conservative than the paired bootstrap test that only considers one run pair at a time, and therefore recommend the former approach for significance testing when a set of runs is available for evaluation.

元の言語English
ホスト出版物のタイトルWWW'12 - Proceedings of the 21st Annual Conference on World Wide Web
ページ499-508
ページ数10
DOI
出版物ステータスPublished - 2012
外部発表Yes
イベント21st Annual Conference on World Wide Web, WWW'12 - Lyon
継続期間: 2012 4 162012 4 20

Other

Other21st Annual Conference on World Wide Web, WWW'12
Lyon
期間12/4/1612/4/20

Fingerprint

Websites
Testing

ASJC Scopus subject areas

  • Computer Networks and Communications

これを引用

Sakai, T. (2012). Evaluation with informational and navigational intents. : WWW'12 - Proceedings of the 21st Annual Conference on World Wide Web (pp. 499-508) https://doi.org/10.1145/2187836.2187904

Evaluation with informational and navigational intents. / Sakai, Tetsuya.

WWW'12 - Proceedings of the 21st Annual Conference on World Wide Web. 2012. p. 499-508.

研究成果: Conference contribution

Sakai, T 2012, Evaluation with informational and navigational intents. : WWW'12 - Proceedings of the 21st Annual Conference on World Wide Web. pp. 499-508, 21st Annual Conference on World Wide Web, WWW'12, Lyon, 12/4/16. https://doi.org/10.1145/2187836.2187904
Sakai T. Evaluation with informational and navigational intents. : WWW'12 - Proceedings of the 21st Annual Conference on World Wide Web. 2012. p. 499-508 https://doi.org/10.1145/2187836.2187904
Sakai, Tetsuya. / Evaluation with informational and navigational intents. WWW'12 - Proceedings of the 21st Annual Conference on World Wide Web. 2012. pp. 499-508
@inproceedings{794071491d1f4621be85bb2aa85e5328,
title = "Evaluation with informational and navigational intents",
abstract = "Given an ambiguous or underspecified query, search result diversification aims at accomodating different user intents within a single {"}entry- point{"} result page. However, some intents are informational, for which many relevant pages may help, while others are navigational, for which only one web page is required. We propose new evaluation metrics for search result diversification that considers this distinction, as well as a simple method for comparing the intuitiveness of a given pair of metrics quantitatively. Our main experimental findings are: (a) In terms of discriminative power which reflects statistical reliability, the proposed metrics, DIN#- nDCG and P+Q#, are comparable to intent recall and D-nDCG, and possibly superior to α-nDCG; (b) In terms of preference agreement with intent recall, P+Q# is superior to other diversity metrics and therefore may be the most intuitive as a metric that emphasises diversity; and (c) In terms of preference agreement with effective precision, DIN#-nDCG is superior to other diversity metrics and therefore may be the most intuitive as a metric that emphasises relevance. Moreover, DIN#-nDCG may be the most intuitive as a metric that considers both diversity and relevance. In addition, we demonstrate that the randomised Tukey's Honestly Significant Differences test that takes the entire set of available runs into account is substantially more conservative than the paired bootstrap test that only considers one run pair at a time, and therefore recommend the former approach for significance testing when a set of runs is available for evaluation.",
keywords = "Diversification, Evaluation, Intents, Metrics, Novelty, Redundancy",
author = "Tetsuya Sakai",
year = "2012",
doi = "10.1145/2187836.2187904",
language = "English",
isbn = "9781450312295",
pages = "499--508",
booktitle = "WWW'12 - Proceedings of the 21st Annual Conference on World Wide Web",

}

TY - GEN

T1 - Evaluation with informational and navigational intents

AU - Sakai, Tetsuya

PY - 2012

Y1 - 2012

N2 - Given an ambiguous or underspecified query, search result diversification aims at accomodating different user intents within a single "entry- point" result page. However, some intents are informational, for which many relevant pages may help, while others are navigational, for which only one web page is required. We propose new evaluation metrics for search result diversification that considers this distinction, as well as a simple method for comparing the intuitiveness of a given pair of metrics quantitatively. Our main experimental findings are: (a) In terms of discriminative power which reflects statistical reliability, the proposed metrics, DIN#- nDCG and P+Q#, are comparable to intent recall and D-nDCG, and possibly superior to α-nDCG; (b) In terms of preference agreement with intent recall, P+Q# is superior to other diversity metrics and therefore may be the most intuitive as a metric that emphasises diversity; and (c) In terms of preference agreement with effective precision, DIN#-nDCG is superior to other diversity metrics and therefore may be the most intuitive as a metric that emphasises relevance. Moreover, DIN#-nDCG may be the most intuitive as a metric that considers both diversity and relevance. In addition, we demonstrate that the randomised Tukey's Honestly Significant Differences test that takes the entire set of available runs into account is substantially more conservative than the paired bootstrap test that only considers one run pair at a time, and therefore recommend the former approach for significance testing when a set of runs is available for evaluation.

AB - Given an ambiguous or underspecified query, search result diversification aims at accomodating different user intents within a single "entry- point" result page. However, some intents are informational, for which many relevant pages may help, while others are navigational, for which only one web page is required. We propose new evaluation metrics for search result diversification that considers this distinction, as well as a simple method for comparing the intuitiveness of a given pair of metrics quantitatively. Our main experimental findings are: (a) In terms of discriminative power which reflects statistical reliability, the proposed metrics, DIN#- nDCG and P+Q#, are comparable to intent recall and D-nDCG, and possibly superior to α-nDCG; (b) In terms of preference agreement with intent recall, P+Q# is superior to other diversity metrics and therefore may be the most intuitive as a metric that emphasises diversity; and (c) In terms of preference agreement with effective precision, DIN#-nDCG is superior to other diversity metrics and therefore may be the most intuitive as a metric that emphasises relevance. Moreover, DIN#-nDCG may be the most intuitive as a metric that considers both diversity and relevance. In addition, we demonstrate that the randomised Tukey's Honestly Significant Differences test that takes the entire set of available runs into account is substantially more conservative than the paired bootstrap test that only considers one run pair at a time, and therefore recommend the former approach for significance testing when a set of runs is available for evaluation.

KW - Diversification

KW - Evaluation

KW - Intents

KW - Metrics

KW - Novelty

KW - Redundancy

UR - http://www.scopus.com/inward/record.url?scp=84860858401&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=84860858401&partnerID=8YFLogxK

U2 - 10.1145/2187836.2187904

DO - 10.1145/2187836.2187904

M3 - Conference contribution

AN - SCOPUS:84860858401

SN - 9781450312295

SP - 499

EP - 508

BT - WWW'12 - Proceedings of the 21st Annual Conference on World Wide Web

ER -