Proceedings of the 26th Australasian Document Computing Symposium最新文献

The Task: Distinguishing Tasks and Sessions in Legal Information Retrieval 任务:区分法律信息检索中的任务和会话

Proceedings of the 26th Australasian Document Computing Symposium Pub Date : 2022-12-15 DOI: 10.1145/3572960.3572983

G. Wiggers, G. Zuccon

{"title":"The Task: Distinguishing Tasks and Sessions in Legal Information Retrieval","authors":"G. Wiggers, G. Zuccon","doi":"10.1145/3572960.3572983","DOIUrl":"https://doi.org/10.1145/3572960.3572983","url":null,"abstract":"Legal information retrieval (IR) is a form of professional search often associated with high recall. Information seeking in this context can consist of a single query with no clicks (known as updating behaviour), a literature review where a complex boolean query crafted over several iterations is performed and all documents returned are inspected, or a seeking task spanning days or weeks, consisting of multiple queries interleaved with other tasks. Analysis of query logs is paramount to the improvement of current legal IR systems, and in particular of the system we are associated with, the Dutch Legal Intelligence IR system. This analysis however requires the ability to automatically identify which queries of a user are related to the same search goal — or in other words, related to the same search task. The current practice of defining sessions — a set of user interactions with the IR system with no more than 30 minutes between user actions — and equating a session to representing a search task, might prove ineffective given the characteristics of this user group. In this paper we provide an initial analysis of a sub-set of the query log from the Dutch Legal Intelligence IR system, comprising of 970 queries issued by 10 users within the space of 1 year. From this query log, we used the 30-minutes heuristic to define sessions, and extract 126 sessions, ranging from 1 to 71 sessions per user. We then independently annotate the query log to manually identify search tasks: this activity leads to the identification of 55 tasks, ranging from 1 to 21 tasks per user. In doing this, we highlight how the currently employed heuristic is not adequate to extract search queries from a user that are related to the same search task. We also show why tasks are more informative than sessions with regards to legal information retrieval. We further describe the potential of using characteristics such as Levenshtein distance, common words and string matching for automated task classification.","PeriodicalId":106265,"journal":{"name":"Proceedings of the 26th Australasian Document Computing Symposium","volume":"51 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2022-12-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"132071204","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Robustness of Neural Rankers to Typos: A Comparative Study 神经排序器对错别字的鲁棒性比较研究

Proceedings of the 26th Australasian Document Computing Symposium Pub Date : 2022-12-15 DOI: 10.1145/3572960.3572981

Shengyao Zhuang, Xinyu Mao, G. Zuccon

引用次数: 1

Immediate-Access Indexing Using Space-Efficient Extensible Arrays 使用空间高效的可扩展数组的即时访问索引

Proceedings of the 26th Australasian Document Computing Symposium Pub Date : 2022-12-15 DOI: 10.1145/3572960.3572984

Alistair Moffat

引用次数: 2

Neural Rankers for Effective Screening Prioritisation in Medical Systematic Review Literature Search 医学系统评价文献检索中有效筛选优先级的神经排序方法

Proceedings of the 26th Australasian Document Computing Symposium Pub Date : 2022-12-15 DOI: 10.1145/3572960.3572980

Shuai Wang, Harrisen Scells, B. Koopman, G. Zuccon

{"title":"Neural Rankers for Effective Screening Prioritisation in Medical Systematic Review Literature Search","authors":"Shuai Wang, Harrisen Scells, B. Koopman, G. Zuccon","doi":"10.1145/3572960.3572980","DOIUrl":"https://doi.org/10.1145/3572960.3572980","url":null,"abstract":"Medical systematic reviews typically require assessing all the documents retrieved by a search. The reason is two-fold: the task aims for “total recall”; and documents retrieved using Boolean search are an unordered set, and thus it is unclear how an assessor could examine only a subset. Screening prioritisation is the process of ranking the (unordered) set of retrieved documents, allowing assessors to begin the downstream processes of the systematic review creation earlier, leading to earlier completion of the review, or even avoiding screening documents ranked least relevant. Screening prioritisation requires highly effective ranking methods. Pre-trained language models are state-of-the-art on many IR tasks but have yet to be applied to systematic review screening prioritisation. In this paper, we apply several pre-trained language models to the systematic review document ranking task, both directly and fine-tuned. An empirical analysis compares how effective neural methods compare to traditional methods for this task. We also investigate different types of document representations for neural methods and their impact on ranking performance. Our results show that BERT-based rankers outperform the current state-of-the-art screening prioritisation methods. However, BERT rankers and existing methods can actually be complementary, and thus, further improvements may be achieved if used in conjunction.","PeriodicalId":106265,"journal":{"name":"Proceedings of the 26th Australasian Document Computing Symposium","volume":"86 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2022-12-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"123599928","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 7

Investigating Language Use by Polarised Groups on Twitter: A Case Study of the Bushfires 调查推特上两极分化群体的语言使用:以丛林大火为例

Proceedings of the 26th Australasian Document Computing Symposium Pub Date : 2022-12-15 DOI: 10.1145/3572960.3572979

Mehwish Nasim, Naeha Sharif, Pranav Bhandari, Derek Weber, Martin Wood, L. Falzon, Y. Kashima

引用次数: 0

Pseudo-Relevance Feedback with Dense Retrievers in Pyserini Pyserini中密集检索器的伪相关反馈

Proceedings of the 26th Australasian Document Computing Symposium Pub Date : 2022-12-15 DOI: 10.1145/3572960.3572982

Hang Li

{"title":"Pseudo-Relevance Feedback with Dense Retrievers in Pyserini","authors":"Hang Li","doi":"10.1145/3572960.3572982","DOIUrl":"https://doi.org/10.1145/3572960.3572982","url":null,"abstract":"Transformer-based Dense Retrievers (DRs) are attracting extensive attention because of their effectiveness paired with high efficiency. In this context, few Pseudo-Relevance Feedback (PRF) methods applied to DRs have emerged. However, the absence of a general framework for performing PRF with DRs has made the empirical evaluation, comparison and reproduction of these methods challenging and time-consuming, especially across different DR models developed by different teams of researchers. To tackle this and speed up research into PRF methods for DRs, we showcase a new PRF framework that we implemented as a feature in Pyserini – an easy-to-use Python Information Retrieval toolkit. In particular, we leverage Pyserini’s DR framework and expand it with a PRF framework that abstracts the PRF process away from the specific DR model used. This new functionality in Pyserini allows to easily experiment with PRF methods across different DR models and datasets. Our framework comes with a number of recently proposed PRF methods built into it. Experiments within our framework show that this new PRF feature improves the effectiveness of the DR models currently available in Pyserini.","PeriodicalId":106265,"journal":{"name":"Proceedings of the 26th Australasian Document Computing Symposium","volume":"254 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2022-12-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"134525623","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 1

Proceedings of the 26th Australasian Document Computing Symposium 第26届澳洲文献计算研讨会论文集

Proceedings of the 26th Australasian Document Computing Symposium Pub Date : 1900-01-01 DOI: 10.1145/3572960

引用次数: 0