Contributo in atti di convegno, 2011, ENG, 10.1145/1951365.1951379

Caching query-biased snippets for efficient retrieval

Lucchese C.; Perego R.; Ceccarelli D.; Silvestri F.; Orlando S.

CNR-ISTI, Pisa, Italy; CNR-ISTI, Pisa, Italy; Dipartimento di Informatica, Università di Pisa, Italy; Università di Venezia, Italy

Web Search Engines' result pages contain references to the top-k documents relevant for the query submitted by a user. Each document is represented by a title, a snippet and a URL. Snippets, i.e. short sentences showing the portions of the document being relevant to the query, help users to select the most interesting results. The snippet generation process is very expensive, since it may require to access a number of documents for each issued query. We assert that caching, a popular technique used to enhance performance at various levels of any computing systems, can be very e ective in this context. We design and experiment several cache organizations, and we introduce the concept of supersnippet, that is the set of sentences in a document that are more likely to answer future queries. We show that supersnippets can be built by exploiting query logs, and that in our experiments a supersnippet cache answers up to 62% of the requests, remarkably outperforming other caching approaches.

14th International Conference on Extending Database Technology, EDBT/ICDT '11, pp. 93–104, Uppsala, Sweden, March 21-24 2011

Keywords

Information retrieval

CNR authors

Lucchese Claudio, Ceccarelli Diego, Perego Raffaele

CNR institutes

ISTI – Istituto di scienza e tecnologie dell'informazione "Alessandro Faedo"

ID: 199719

Year: 2011

Type: Contributo in atti di convegno

Creation: 2013-01-22 14:10:00.000

Last update: 2017-10-13 19:37:14.000

External IDs

CNR OAI-PMH: oai:it.cnr:prodotti:199719

DOI: 10.1145/1951365.1951379

Scopus: 2-s2.0-79953863312