Bi-Phase Enhanced IVFPQ for Time-Efficient Ad-hoc Retrieval

10/11/2022
by   Peitian Zhang, et al.
2

IVFPQ is a popular index paradigm for time-efficient ad-hoc retrieval. Instead of traversing the entire database for relevant documents, it accelerates the retrieval operation by 1) accessing a fraction of the database guided the activation of latent topics in IVF (inverted file system), and 2) approximating the exact relevance measurement based on PQ (product quantization). However, the conventional IVFPQ is limited in retrieval performance due to the coarse granularity of its latent topics. On the one hand, it may result in severe loss of retrieval quality when visiting a small number of topics; on the other hand, it will lead to a huge retrieval cost when visiting a large number of topics. To mitigate the above problem, we propose a novel framework named Bi-Phase IVFPQ. It jointly uses two types of features: the latent topics and the explicit terms, to build the inverted file system. Both types of features are complementary to each other, which helps to achieve better coverage of the relevant documents. Besides, the documents' memberships to different IVF entries are learned by distilling knowledge from deep semantic models, which substantially improves the index quality and retrieval accuracy. We perform comprehensive empirical studies on popular ad-hoc retrieval benchmarks, whose results verify the effectiveness and efficiency of our proposed framework.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/16/2021

A Neural Passage Model for Ad-hoc Document Retrieval

Traditional statistical retrieval models often treat each document as a ...
research
07/21/2018

A Line in the Sand: Recommendation or Ad-hoc Retrieval?

The popular approaches to recommendation and ad-hoc retrieval tasks are ...
research
01/24/2022

HC4: A New Suite of Test Collections for Ad Hoc CLIR

HC4 is a new suite of test collections for ad hoc Cross-Language Informa...
research
02/22/2021

Graph-based Hierarchical Relevance Matching Signals for Ad-hoc Retrieval

The ad-hoc retrieval task is to rank related documents given a query and...
research
04/24/2023

Overview of the TREC 2022 NeuCLIR Track

This is the first year of the TREC Neural CLIR (NeuCLIR) track, which ai...
research
11/30/2022

CDSM: Cascaded Deep Semantic Matching on Textual Graphs Leveraging Ad-hoc Neighbor Selection

Deep semantic matching aims to discriminate the relationship between doc...
research
08/25/2021

Podcast Metadata and Content: Episode Relevance andAttractiveness in Ad Hoc Search

Rapidly growing online podcast archives contain diverse content on a wid...

Please sign up or login with your details

Forgot password? Click here to reset