The corpus of United States state statutes—design, construction and use

IF 2.1

Applied Corpus Linguistics Pub Date : 2023-08-01 DOI:10.1016/j.acorp.2023.100047

Jesse Egbert, Margaret Wood

引用次数: 0

Abstract

There is a need for more publicly available corpora of legal language. To help fill this gap, we have developed the Corpus of U.S. State Statutes, or CorUSSS, a new corpus comprising the statutory code from all 50 U.S. states. In total the corpus contains 1,785,742 texts, each of which represents the statutory text associated with a unique Universal Citation in one of the 50 U.S. states’ codes. This corpus provides us with the ability to explore language use in statutes within or across all 50 states. After motivating the need for this corpus, we describe its design and the methods we used to collect, clean and store the texts. We then report on a case study that illustrates the utility of this corpus for addressing important questions in statutory interpretation by investigating whether the word information can be used to refer to statements that are non-factual. We conclude with a call for researchers in law and corpus linguistics to rely on both legal and ordinary language when investigating questions of interpretation.

查看原文本刊更多论文

美国州法规文集——设计、建造和使用

有必要提供更多公开的法律语言语料库。为了帮助填补这一空白，我们开发了美国州法规语料库(CorUSSS)，这是一个包含美国所有50个州的法定代码的新语料库。该语料库总共包含1,785,742个文本，每个文本都代表与美国50个州法典之一的唯一通用引文相关的法定文本。这个语料库为我们提供了探索所有50个州内或跨州的法规中语言使用的能力。在激发了对这个语料库的需求之后，我们描述了它的设计以及我们用来收集、清理和存储文本的方法。然后，我们报告了一个案例研究，通过调查“信息”一词是否可以用来指非事实性陈述，说明了该语料库在解决法律解释中的重要问题方面的效用。最后，我们呼吁法律和语料库语言学的研究人员在调查解释问题时既依赖法律语言，也依赖普通语言。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊