Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval最新文献_第7页

Passage retrieval vs. document retrieval for factoid question answering 段落检索与伪题回答的文档检索

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860534

C. Clarke, E. Terra

{"title":"Passage retrieval vs. document retrieval for factoid question answering","authors":"C. Clarke, E. Terra","doi":"10.1145/860435.860534","DOIUrl":"https://doi.org/10.1145/860435.860534","url":null,"abstract":"Question answering (QA) systems often contain an information retrieval subsystem that identifies documents or passages where the answer to a question might appear [1–3, 5, 6, 10]. The QA system generates queries from the questions and submits them to the IR subsystem. The IR subsystem returns the top-ranked documents or passages, and the QA system selects the answers from them. In many QA systems, the IR component retrieves entire documents. Then, in a post-retrieval step, the system scans the retrieved documents and locates groups of sentences that contain most or all of the question keywords [3,10, and others]. These sentences are subjected to further analysis to select the answer. In other QA systems, a passage-retrieval technique is employed to directly identify locations within the document collection where the answer might be found, avoiding the post-retrieval step [1, 2, 5, 6, and others]. In this context, a “relevant” document or passage is one that contains an answer. We utilize this notion of relevance to evaluate an IR subsystem in isolation from the rest of its QA system by applying standard measures of IR effectiveness. By restricting our evaluation to a single subsystem we hope to gain experience that is applicable to QA systems beyond our own. An assumption inherent in this approach is that improved precision in the IR subsystem will translate to improved performance of the QA system as a whole. This assumption holds for our own system, and should (at least) hold for any system that exploits redundancy—that takes advantage of the observation that answers tend to occur in more than one retrieved passage [1, 2, 5]. In this paper we compare a successful passage-retrieval method [1, 5] with a well-known and effective documentretrieval method: Okapi BM25 [7]. Our goal is to examine","PeriodicalId":209809,"journal":{"name":"Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval","volume":"1 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2003-07-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"121252637","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 54

Modeling annotated data 建模注释数据

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860460

D. Blei, Michael I. Jordan

引用次数: 1250

Word sense disambiguation in information retrieval revisited 再论信息检索中的词义消歧

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860466

Christopher Stokoe, M. Oakes, J. Tait

引用次数: 223

Query length in interactive information retrieval 交互式信息检索中的查询长度

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860474

N. Belkin, D. Kelly, G. Kim, Ja-Young Kim, Hyuk-Jin Lee, G. Muresan, Muh-Chyun Tang, Xiaojun Yuan, Colleen Cool

引用次数: 161

Transliteration of proper names in cross-language applications 跨语言应用中专有名称的音译

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860503

Paola Virga, S. Khudanpur

{"title":"Transliteration of proper names in cross-language applications","authors":"Paola Virga, S. Khudanpur","doi":"10.1145/860435.860503","DOIUrl":"https://doi.org/10.1145/860435.860503","url":null,"abstract":"Translation of proper names is generally recognized as a significant problem in many multi-lingual text and speech processing applications. Even when large bilingual lexicons used for machine translation (MT) and cross-lingual information retrieval (CLIR) provide significant coverage of the words encountered in the text, a significant portion of the tokens not covered by such lexicons are proper names (cf e.g. [3]). For CLIR applications in particular, proper names and technical terms are particularly important, as they carry some of the more distinctive information in a query. In IR systems where users provide very short queries (e.g. 2-3 words), their importance grows even further. Proper names are amenable to a speech-inspired translation approach. When writing a foreign name in ones native language, one tries to preserve the way it sounds. i.e. one uses an orthographic representation which, when “read aloud” by a native speaker of the language sounds as it would when spoken by a speaker of the foreign language — a process referred to as transliteration. If mechanisms were available (a) to render, say, an English name in its phonemic form, and (b) to convert this phonemic string into the orthography of, say, Mandarin Chinese, then one would have a mechanism for transliterating English names using Chinese characters. The first part has been addressed extensively in the automatic textto-speech synthesis literature. This paper describes a statistical approach for the second part. Several techniques have been proposed in the recent past for name transliteration. Finite state transducers that implement transformation rules for back-transliteration from Japanese to English are described in [2], and extended to Arabic in [5]. In both cases, the goal is to recognize words in Japanese or Arabic text which happen to be transliterations of English names. The strongly phonetic orthography of Korean is exploited in [1] to obtain good transliteration using relatively simple HMM-based models. A set of handcrafted rules for locally editing the phonemic spelling of an English name to conform to Mandarin syllabification is provided to a transformation-based learning algorithm in [4], which then learns how to convert an English phoneme sequence to a Mandarin syllable sequence. We describe here a fully data driven counterpart to the technique of [4] for English-to-Mandarin name transliteration. In addition to intrinsic evaluation, we test our transliteration system extrinsically for cross-lingual spoken document retrieval by usThis research was partially supported by DARPA via Grant No N66001-00-2-8910 and ONR via Grant No N00014-01-1-0685. Copyright is held by the author/owner. SIGIR’03, July 28–August 1, 2003, Toronto, Canada. ACM 1-58113-646-3/03/0007. ing English text queries to retrieve Mandarin audio from the Topic Detection and Tracking (TDT) corpus.","PeriodicalId":209809,"journal":{"name":"Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval","volume":"44 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2003-07-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"133179889","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 46

Text categorization by boosting automatically extracted concepts 文本分类通过提升自动提取的概念

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860470

Lijuan Cai, Thomas Hofmann

引用次数: 144

Document-self expansion for text categorization 用于文本分类的文档自展开

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860520

Yuen-Hsien Tseng, Da-Wei Juang

引用次数: 7

Structured use of external knowledge for event-based open domain question answering 结构化地使用外部知识进行基于事件的开放领域问答

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860444

G. Yang, Tat-Seng Chua, Shuguang Wang, Chun-Keat Koh

引用次数: 117

Topic hierarchy generation via linear discriminant projection 基于线性判别投影的主题层次生成

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860531

Tao Li, Shenghuo Zhu, M. Ogihara

{"title":"Topic hierarchy generation via linear discriminant projection","authors":"Tao Li, Shenghuo Zhu, M. Ogihara","doi":"10.1145/860435.860531","DOIUrl":"https://doi.org/10.1145/860435.860531","url":null,"abstract":"Text categorization has been receiving more and more attention with the ever-increasing growth of the on-line information. Automated text categorization is generally a supervised learning problem, defined as the problem of assigning pre-defined category labels to new documents based on the likelihood suggested by labeled documents. Most studies in the area have been focused on flat classification, where the predefined categories are treated individually and separately [5]. As the available information increases, when the number of categories grows significantly large, it will become much more difficult to browse and search categories. The most successful paradigm for organizing this mass of information and making it compressible is by categorizing documents according to their topics where the topics are organized in a hierarchy of increasing specificity [3]. Hierarchical structures identify the relationships of dependence between the categories and provides a valuable information source for many problems. Recently several researchers have investigated the use of hierarchies for text classification and obtained promising results [1, 4]. However, little has been done to explore the approaches to automatically generate topic hierarchies. Most of the reported techniques have been conducted on existential hierarchically structured corpora. The aim of automatic hierarchy generation has several motivations. First, manually building hierarchies is an expensive task since it requires domain experts to evaluate the documents’ relevance to the topics. Second, existing hierarchies are optimized for human use based on “human semantics”, but not necessarily for classifier use. Automatic generated hierarchies can be incorporated into various classification methods","PeriodicalId":209809,"journal":{"name":"Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval","volume":"118 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2003-07-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"127994597","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 18

Collaborative filtering via gaussian probabilistic latent semantic analysis 基于高斯概率潜在语义分析的协同过滤

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI: 10.1145/860435.860483

Thomas Hofmann

引用次数: 449