{"title":"Summarizing judicial documents: a hybrid extractive- abstractive model with legal domain knowledge","authors":"Yan Gao, Jie Wu, Zhengtao Liu, Juan Li","doi":"10.1007/s10506-025-09435-z","DOIUrl":"10.1007/s10506-025-09435-z","url":null,"abstract":"<div><p>The automatic summarization of judgment documents is a challenging task due to their length and the dispersed nature of the important information they contain. The prevailing approach to tackling the summarization of lengthy documents involves the integration of both extractive and abstractive summarization models. However, current extractive models face challenges in capturing all essential details due to the scattered distribution of pertinent information within judgment documents. Additionally, the existing abstractive models still grapple with the problem of \"hallucinations\" which leads to generating inaccurate information. In our work, we proposed a novel hybrid legal summarization method that incorporates legal domain knowledge into both the extractive model and abstractive model. The method consists of two parts: (1) The rhetorical role of sentences is identified by the sentence-level sequence labeling method, and the rhetorical information is integrated into the extractive model based on WoBERT through the conditional normalization to ensure that the identification of key sentences is both precise and complete. (2) The pre-trained model RoFormer is combined with Seq2Seq to construct a long text summarization model, and the prior knowledge in the external resources and the document itself is introduced into the decoding process to improve the faithfulness and coherence of the composed summary. In addition, the contrastive learning strategy is employed during the training process to enhance the robustness of the abstractive model. Experimental results on the CAIL2020 dataset show that the proposed model is superior to the baseline methods. Furthermore, our method outperforms GPT and other LLMs in processing judgment documents.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 2","pages":"493 - 521"},"PeriodicalIF":3.1,"publicationDate":"2025-04-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147968061","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Enhancing Indian legal judgment classification with embeddings, feature selection, and ensemble strategies","authors":"Priyanka Prabhakar, Peeta Basa Pati","doi":"10.1007/s10506-025-09438-w","DOIUrl":"10.1007/s10506-025-09438-w","url":null,"abstract":"<div><p>Legal document analysis presents significant challenges due to its complexity and domain-specific nature. This study introduces an innovative approach for classifying Indian court judgments into legal domains using diverse machine learning and deep learning techniques. The method incorporates feature engineering and deep learning algorithms for extracting meaningful features, supported by a wide range of classifiers, including voting classifiers, gradient boosting, and random forest. Embeddings are generated using models such as InLegalBERT, InCaseLawBERT, CustomInLawBERT, Mamba, T5, RoBERTa, CodeT5, SBERT, DistilBERT, XLM, XLM Large, LegalBERT, GPT2, ALBERT, Electra, DeBERTa, TFDeBERTa, FlanT5, FlanT-Large, BART, BigBird Pegasus, LongFormer, and LUKE. To address class imbalance, the SMOTE technique is employed, and dimensionality is reduced using PCA and forward feature selection. The T5+SMOTE+feature selection+voting classifier configuration achieves a notable accuracy of 98%, highlighting the effectiveness of the proposed approach. These advancements have significant implications for applications such as document retrieval, legal discovery, and case law analysis, enhancing the accuracy and efficiency of legal document classification.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 2","pages":"523 - 564"},"PeriodicalIF":3.1,"publicationDate":"2025-03-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147968148","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Understanding unnecessary stops and police use of force in NYPD Stop, Question, and Frisk with machine learning techniques","authors":"Passiri Bodhidatta, Daricha Sutivong","doi":"10.1007/s10506-025-09444-y","DOIUrl":"10.1007/s10506-025-09444-y","url":null,"abstract":"<div><p>Even though the New York Police Department (NYPD) reform in 2013 led to a substantial reduction in the total number of stops, unnecessary stops and weapon use against innocent citizens remain critical issues. This study analyzes stop-and-frisk records during 2014 – 2019 using tree-based machine learning approaches along with logistic regression and Multi-Layer Perceptron (MLP) models, in order to discover patterns and insights. By developing predictive models for both suspect convictions and the level of force applied by police, this study provides a basis for a discussion whether weapon usage aligns with indicators of guilt or conviction. Findings show that XGBoost outperforms other machine learning techniques in predicting both conviction and the level of force used. Key factors associated with a suspect’s conviction include weapon possession, carrying suspicious objects, and trespassing. However, an excessive number of unnecessary stops appear to be associated with inaccurate assumptions about suspects’ weapon possession, which are also linked to police gunfire against innocent citizens. Refining suspicion criteria for Criminal Possession of Weapon and suspect actions could help reduce unnecessary stops and excessive force.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 2","pages":"587 - 623"},"PeriodicalIF":3.1,"publicationDate":"2025-03-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147968005","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Topic classification of case law using a large language model and a new taxonomy for UK law: AI insights into summary judgment","authors":"Holli Sargeant, Ahmed Izzidien, Felix Steffek","doi":"10.1007/s10506-025-09434-0","DOIUrl":"10.1007/s10506-025-09434-0","url":null,"abstract":"<div><p>This paper addresses a critical gap in legal analytics by developing and applying a novel taxonomy for topic classification of summary judgment cases in the United Kingdom. Using a curated dataset of summary judgment cases, we use the Large Language Model Claude 3 Opus to explore functional topics and trends. We find that Claude 3 Opus correctly classified the topic with an accuracy of 87.13% and an F1 score of 0.87. The analysis reveals distinct patterns in the application of summary judgments across various legal domains. As case law in the United Kingdom is not originally labelled with keywords or a topic filtering option, the findings not only refine our understanding of the thematic underpinnings of summary judgments but also illustrate the potential of combining traditional and AI-driven approaches in legal classification. Therefore, this paper provides a new and general taxonomy for UK law. The implications of this work serve as a foundation for further research and policy discussions in the field of judicial administration and computational legal research methodologies.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 2","pages":"443 - 491"},"PeriodicalIF":3.1,"publicationDate":"2025-02-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://link.springer.com/content/pdf/10.1007/s10506-025-09434-0.pdf","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147968150","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"An interpretable approach to detect case law on housing and eviction issues within the HUDOC database","authors":"Mohammad Mohammadi, Martijn Wieling, Michel Vols","doi":"10.1007/s10506-025-09439-9","DOIUrl":"10.1007/s10506-025-09439-9","url":null,"abstract":"<div><p>Case law plays a critical role in shaping our understanding of human rights, including the right to adequate housing. However, analyzing large legal databases like HUDOC, which contains over 40,000 cases, is a challenging task that requires automated solutions. This study focuses on detecting cases related to housing—a topic encompassing issues such as eviction, access to adequate housing and etc.—from the HUDOC database. For this, we developed classifiers to identify cases related to both housing and eviction issues. We first constructed a dataset using an unsupervised process refined through manual corrections. Then, we trained the Adaptive Chordal Distance-based Subspace Learning Vector Quantization models. These models achieved classification accuracies of 93% for housing-related cases and 91.5% for eviction-specific cases, matching the performance of transformer-based models while requiring fewer computational resources. Furthermore, they provide interpretability by assigning word-level importance scores, helping legal scholars understand and verify the reasoning behind the model’s predictions. The models identified 2,305 potentially housing-related cases. Manual reviews confirmed that 278 of 340 reviewed cases were indeed relevant. By detecting overlooked cases and enriching legal datasets, this study highlights the utility of NLP methods in facilitating the analysis of human rights case law. This approach supports a deeper exploration of housing rights and eviction-related decisions under the European Court of Human Rights (ECtHR), offering transparency, efficiency, and scalability for legal research.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 2","pages":"565 - 586"},"PeriodicalIF":3.1,"publicationDate":"2025-02-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://link.springer.com/content/pdf/10.1007/s10506-025-09439-9.pdf","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147968104","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Irene Benedetto, Luca Cagliero, Michele Ferro, Francesco Tarasconi, Claudia Bernini, Giuseppe Giacalone
{"title":"Leveraging large language models for abstractive summarization of Italian legal news","authors":"Irene Benedetto, Luca Cagliero, Michele Ferro, Francesco Tarasconi, Claudia Bernini, Giuseppe Giacalone","doi":"10.1007/s10506-025-09431-3","DOIUrl":"10.1007/s10506-025-09431-3","url":null,"abstract":"<div><p>Condensing the key message conveyed by a long document into an informative summary is particularly helpful to lawyers and legal experts. State-of-the-art approaches to legal document summarization rely on Language Models (LMs) and are mostly trained on English documents. More limited research efforts have been devoted to summarizing legal documents in languages other than English. In this work, we investigate the applicability of Large Language Models (LLMs) to summarize Italian legal news documents. We benchmark state-of-the-art abstractive summarization techniques based on Language Models, Large and not, for headline and abstract generation from legal news documents. We run an extensive set of experiments on a proprietary legal dataset, evaluating the resulting summaries according to both quantitative metrics and human evaluation. As expected, latest LLMs outperform classical models such as BART, T5, particularly in terms of grammaticality and informativeness of the summary content. Fine-tuned LLMs also show a significant increase in performance, variable across law areas, compared to their zero-shot setting. Importantly, the level of specialization of the fine-tuned version already reaches a steady state after feeding the model with few hundreds of training data.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 2","pages":"421 - 441"},"PeriodicalIF":3.1,"publicationDate":"2025-02-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147967987","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Laura State, Alejandra Bringas Colmenarejo, Andrea Beretta, Salvatore Ruggieri, Franco Turini, Stephanie Law
{"title":"The explanation dialogues: an expert focus study to understand requirements towards explanations within the GDPR","authors":"Laura State, Alejandra Bringas Colmenarejo, Andrea Beretta, Salvatore Ruggieri, Franco Turini, Stephanie Law","doi":"10.1007/s10506-024-09430-w","DOIUrl":"10.1007/s10506-024-09430-w","url":null,"abstract":"<div><p>Explainable AI (XAI) provides methods to understand non-interpretable machine learning models. However, we have little knowledge about what legal experts expect from these explanations, including their legal compliance with, and value against European Union legislation. To close this gap, we present the <i>Explanation Dialogues</i>, an expert focus study to uncover the expectations, reasoning, and understanding of legal experts and practitioners towards XAI, with a specific focus on the European General Data Protection Regulation. The study consists of an online questionnaire and follow-up interviews, and is centered around a use-case in the credit domain. We extract both a set of hierarchical and interconnected codes using grounded theory, and present the standpoints of the participating experts towards XAI. We find that the presented explanations are hard to understand and lack information, and discuss issues that can arise from the different interests of the data controller and subject. Finally, we present a set of recommendations for developers of XAI methods, and indications of legal areas of discussion. Among others, recommendations address the presentation, choice, and content of an explanation, technical risks as well as the end-user, while we provide legal pointers to the contestability of explanations, transparency thresholds, intellectual property rights as well as the relationship between involved parties.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 2","pages":"361 - 420"},"PeriodicalIF":3.1,"publicationDate":"2025-01-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147968133","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Causality-inspired legal provision selection with large language model-based explanation","authors":"Zheng Wang, Yuanzhi Ding, Caiyuan Wu, Yuzhen Guo, Wei Zhou","doi":"10.1007/s10506-024-09429-3","DOIUrl":"10.1007/s10506-024-09429-3","url":null,"abstract":"<div><p>Accurate identification of legal provisions is crucial for adjudicating criminal cases, but the complexity and volume of legal texts pose significant challenges for legal professionals. This paper addresses these challenges by introducing a novel legal provision selection framework that transforms the task from a simple classification problem into a sophisticated system combining semantic matching with causal relationship learning. Leveraging large language models, our approach enhances the understanding and interpretation of legal language, by extracting nuanced features from legal texts for deeper contextual comprehension. Additionally, integrating causal learning aligns with the inherent causality in legal reasoning, improving model interpretability and mitigating data bias. Our method demonstrates superior accuracy and robustness through extensive experiments on the CAIL2018 dataset and its subsets. This research significantly advances legal AI applications, promoting efficiency and fairness in the criminal justice system by providing precise and reliable legal provision selection.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 2","pages":"335 - 359"},"PeriodicalIF":3.1,"publicationDate":"2024-12-23","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147968094","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Precedent-based reasoning with incomplete information for human-in-the-loop decision support","authors":"Daphne Odekerken, Floris Bex, Henry Prakken","doi":"10.1007/s10506-024-09421-x","DOIUrl":"10.1007/s10506-024-09421-x","url":null,"abstract":"<div><p>We define and study the notions of stability and relevance for precedent-based reasoning, focusing on Horty’s result model of precedential constraint. According to this model, precedents constrain the possible outcomes for a focus case, which is a yet undecided case, where precedents and the focus case are compared on their characteristics (called dimensions). In this paper, we refer to the enforced outcome for the focus case as its <i>justification</i> status. In contrast to earlier work, we do not assume that all dimension values of the focus case or the precedent cases have been established with certainty: rather, each dimension is assigned a set of possible values. We define a focus case as <i>stable</i> if its justification status is the same for every choice of the possible values. For focus cases that are not stable, we study the task of identifying <i>relevance</i>: which possible values should be excluded to make the focus case stable? In addition, we introduce the notion of <i>possibility</i> to verify if a user can assign an outcome to an unstable focus case without making the case base of precedents inconsistent. We show how the tasks of identifying justification, stability, relevance and possibility can be applied for human-in-the-loop decision support. Finally, we discuss the computational complexity of these tasks and provide efficient algorithms.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 1","pages":"107 - 152"},"PeriodicalIF":3.1,"publicationDate":"2024-12-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://link.springer.com/content/pdf/10.1007/s10506-024-09421-x.pdf","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147559198","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"It cannot be right if it was written by AI: on lawyers’ preferences of documents perceived as authored by an LLM vs a human","authors":"Jakub Harasta, Tereza Novotná, Jaromir Savelka","doi":"10.1007/s10506-024-09422-w","DOIUrl":"10.1007/s10506-024-09422-w","url":null,"abstract":"<div><p>Large Language Models (LLMs) enable a future in which certain types of legal documents may be generated automatically. This has a great potential to streamline legal processes, lower the cost of legal services, and dramatically increase access to justice. While many researchers focus on proposing and evaluating LLM-based applications supporting tasks in the legal domain, there is a notable lack of investigations into how legal professionals perceive content if they believe an LLM has generated it. Yet, this is a critical point as over-reliance or unfounded scepticism may influence whether such documents bring about appropriate legal consequences. This study is the necessary analysis of the ongoing transition towards mature generative AI systems. Specifically, we examined whether the perception of legal documents’ by lawyers and law students (n = 75) varies based on their assumed origin (human-crafted vs AI-generated). The participants evaluated the documents, focusing on their correctness and language quality. Our analysis revealed a clear preference for documents perceived as crafted by a human over those believed to be generated by AI. At the same time, most participants expect the future in which documents will be generated automatically. These findings could be leveraged by legal practitioners, policymakers, and legislators to implement and adopt legal document generation technology responsibly and to fuel the necessary discussions on how legal processes should be updated to reflect recent technological developments.</p></div>","PeriodicalId":51336,"journal":{"name":"Artificial Intelligence and Law","volume":"34 1","pages":"153 - 190"},"PeriodicalIF":3.1,"publicationDate":"2024-12-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147559069","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"社会学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}