DoSSIER@COLIEE 2021: Leveraging dense retrieval and summarization-based re-ranking for case law retrieval

08/09/2021
by   Sophia Althammer, et al.
0

In this paper, we present our approaches for the case law retrieval and the legal case entailment task in the Competition on Legal Information Extraction/Entailment (COLIEE) 2021. As first stage retrieval methods combined with neural re-ranking methods using contextualized language models like BERT achieved great performance improvements for information retrieval in the web and news domain, we evaluate these methods for the legal domain. A distinct characteristic of legal case retrieval is that the query case and case description in the corpus tend to be long documents and therefore exceed the input length of BERT. We address this challenge by combining lexical and dense retrieval methods on the paragraph-level of the cases for the first stage retrieval. Here we demonstrate that the retrieval on the paragraph-level outperforms the retrieval on the document-level. Furthermore the experiments suggest that dense retrieval methods outperform lexical retrieval. For re-ranking we address the problem of long documents by summarizing the cases and fine-tuning a BERT-based re-ranker with the summaries. Overall, our best results were obtained with a combination of BM25 and dense passage retrieval using domain-specific embeddings.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
01/05/2022

PARM: A Paragraph Aggregation Retrieval Model for Dense Document-to-Document Retrieval

Dense passage retrieval (DPR) models show great effectiveness gains in f...
research
05/26/2022

LeiBi@COLIEE 2022: Aggregating Tuned Lexical Models with a Cluster-driven BERT-based Model for Case Law Retrieval

This paper summarizes our approaches submitted to the case law retrieval...
research
07/11/2023

U-CREAT: Unsupervised Case Retrieval using Events extrAcTion

The task of Prior Case Retrieval (PCR) in the legal domain is about auto...
research
09/29/2020

Building Legal Case Retrieval Systems with Lexical Matching and Summarization using A Pre-Trained Phrase Scoring Model

We present our method for tackling the legal case retrieval task of the ...
research
12/21/2020

Cross-domain Retrieval in the Legal and Patent Domains: a Reproducibility Study

Domain specific search has always been a challenging information retriev...
research
04/17/2023

Statute-enhanced lexical retrieval of court cases for COLIEE 2022

We discuss our experiments for COLIEE Task 1, a court case retrieval com...
research
12/13/2022

Attentive Deep Neural Networks for Legal Document Retrieval

Legal text retrieval serves as a key component in a wide range of legal ...

Please sign up or login with your details

Forgot password? Click here to reset