Lost in Alignment: A Survey on Cross-Lingual Alignment Methods for Contextualized Representation

IF 28 1区计算机科学 Q1 COMPUTER SCIENCE, THEORY & METHODS

ACM Computing Surveys Pub Date : 2025-08-26 DOI:10.1145/3764112

Filippo Pallucchini, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica

引用次数: 0

Abstract

Cross-lingual word representations allow us to analyse word meanings across diverse language settings. It is crucial in aiding cross-lingual knowledge transfer when constructing natural language processing (NLP) models for languages with limited resources. This survey presents a comprehensive classification of cross-lingual contextual embedding models. We assess their data requirements and objective functions, and we introduce a taxonomy for categorising these approaches. Then, we present a comprehensive table containing a set of hierarchical criteria to compare them better, along with information regarding the availability of code and data to enable replication of the research. Furthermore, we delve into the evaluation methodologies employed for cross-lingual embeddings, exploring their practical applications and addressing their current associated challenges.

查看原文本刊更多论文

迷失在对齐中：语境化表征的跨语言对齐方法研究

跨语言的单词表示使我们能够分析不同语言背景下的单词含义。对于资源有限的语言，构建自然语言处理（NLP）模型对于帮助跨语言知识迁移至关重要。本研究提出了跨语言上下文嵌入模型的综合分类。我们评估了它们的数据需求和目标函数，并引入了对这些方法进行分类的分类法。然后，我们提出了一个综合表，其中包含一组层次标准，以便更好地比较它们，以及有关代码和数据的可用性的信息，以实现研究的复制。此外，我们深入研究了用于跨语言嵌入的评估方法，探索了它们的实际应用并解决了它们当前的相关挑战。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

ACM Computing Surveys 工程技术-计算机：理论方法

CiteScore

33.20

自引率

0.60%

发文量

372

审稿时长

12 months

期刊介绍： ACM Computing Surveys is an academic journal that focuses on publishing surveys and tutorials on various areas of computing research and practice. The journal aims to provide comprehensive and easily understandable articles that guide readers through the literature and help them understand topics outside their specialties. In terms of impact, CSUR has a high reputation with a 2022 Impact Factor of 16.6. It is ranked 3rd out of 111 journals in the field of Computer Science Theory & Methods. ACM Computing Surveys is indexed and abstracted in various services, including AI2 Semantic Scholar, Baidu, Clarivate/ISI: JCR, CNKI, DeepDyve, DTU, EBSCO: EDS/HOST, and IET Inspec, among others.