Search results

Your Search

Sort Options

(1 - 8 of 8)

Abolghasemi, M.A.; Verberne, S.; Askari, A.; Azzopardi, L. 2023

Retrievability bias estimation using synthetically generated queries

Article in monograph or in proceedings

open access

Askari, A.; Abolghasemi, M.A.; Pasi, G.; Kraaij, W.; Verberne, S. 2023

Injecting the BM25 score as text improves BERT-based re-rankers

Article in monograph or in proceedings

open access

Askari, A.; Aliannejadi, M.; Abolghasemi, M.A.; Kanoulas, E.; Verberne, S. 2023

CLosER: conversational legal longformer with expertise-aware passage response ranker for long contexts

Article in monograph or in proceedings

open access

Askari, A.; Aliannejadi, M.; Meng, C.; Kanoulas, E.; Verberne, S 2023

Expand, highlight, generate: RL-driven document generation for passage reranking

Article in monograph or in proceedings

open access

Generating synthetic training data based on large language models (LLMs) for ranking models has gained attention recently. Prior studies use LLMs to build pseudo query-document pairs by generating... Show moreGenerating synthetic training data based on large language models (LLMs) for ranking models has gained attention recently. Prior studies use LLMs to build pseudo query-document pairs by generating synthetic queries from documents in a corpus. In this paper, we propose a new perspective of data augmentation: generating synthetic documents from queries. To achieve this, we propose DocGen, that consists of a three-step pipeline that utilizes the few-shot capabilities of LLMs. DocGen pipeline performs synthetic document generation by (i) expanding, (ii) highlighting the original query, and then (iii) generating a synthetic document that is likely to be relevant to the query. To further improve the relevance between generated synthetic documents and their corresponding queries, we propose DocGen-RL, which regards the estimated relevance of the document as a reward and leverages reinforcement learning (RL) to optimize DocGen pipeline. Extensive experiments demonstrate that DocGen pipeline and DocGen-RL significantly outperform existing state-of-theart data augmentation methods, such as InPars, indicating that our new perspective of generating documents leverages the capacity of LLMs in generating synthetic data more effectively. We release the code, generated data, and model checkpoints to foster research in this area. Show less

Askari, A.; Aliannejadi, M.; Kanoulas, E.; Verberne, S. 2023

A test collection of synthetic documents for training rankers: ChatGPT vs. human experts

Article in monograph or in proceedings

open access

Askari, A.; Verberne, S.; Pasi, G. 2022

Expert finding in legal community question answering

Article in monograph or in proceedings

open access

Abolghasemi, M.A.; Askari, A.; Verberne, S. 2022

On the interpolation of contextualized term-based ranking with BM25 for query-by-example retrieval

Article in monograph or in proceedings

open access

Althammer, S.; Askari, A.; Verberne, S.; Hanbury, A.H. 2021

DoSSIER@ COLIEE 2021: Leveraging dense retrieval and summarization-based re-ranking for case law retrieval

Article in monograph or in proceedings

open access

In this paper, we present our approaches for the case law retrieval and the legal case entailment task in the Competition on Legal Information Extraction/Entailment (COLIEE) 2021. As first stage... Show moreIn this paper, we present our approaches for the case law retrieval and the legal case entailment task in the Competition on Legal Information Extraction/Entailment (COLIEE) 2021. As first stage retrieval methods combined with neural re-ranking methods us- ing contextualized language models like BERT achieved great performance improvements for information retrieval in the web and news domain, we evaluate these methods for the legal domain. A distinct characteristic of legal case retrieval is that the query case and case description in the corpus tend to be long documents and therefore exceed the input length of BERT. We address this challenge by combining lexical and dense retrieval methods on the paragraph-level of the cases for the first stage retrieval. Here we demonstrate that the retrieval on the paragraph-level outperforms the retrieval on the document-level. Furthermore the experiments suggest that dense retrieval methods outperform lexical retrieval. For re-ranking we address the problem of long documents by sum- marizing the cases and fine-tuning a BERT-based re-ranker with the summaries. Overall, our best results were obtained with a com- bination of BM25 and dense passage retrieval using domain-specific embeddings.DoSSIER@ COLIEE 2021 Show less

Leiden University Scholarly Publications

Your Search

Enabled Filters

Sort Options

Refine Results

Resource Type

Availability

Creation Date

Faculty

Collection

Author

Language

Search results