摘要
跨语言信息检索指以一种语言为检索词,检索出用另一种或几种语言描述的一种信息的检索技术,是信息检索领域重要的研究方向之一。近年来,跨语言词向量为跨语言信息检索提供了良好的词向量表示,受到很多学者的关注。该文首先利用跨语言词向量模型实现汉文查询词到蒙古文查询词的映射,其次提出串联式查询扩展、串联式查询扩展过滤、交叉验证筛选过滤三种查询扩展方法对候选蒙古文查询词进行筛选和排序,最后选取上下文相关的蒙古文查询词。实验结果表明:在蒙汉跨语言信息检索任务中引入交叉验证筛选方法对信息检索结果有很大的提升。
Cross-Language information retrieval is supposed to retrieve information in one language according to the queries of other languages.This paper uses cross language word vectors model to realize the mapping of Chinese query words to Mongolian query words.Three methods are proposed in this paper to perform mapping,namely Series,Series_opt and Cross_valid.These methods are used to map the Chinese queries as well as to select and sort the mapped words.Experimental conducted in the real environments show that our proposed algorithm can obtain improvement on Chinese-Mongolian information retrieval.
作者
马路佳
赖文
赵小兵
MA Lujia;LAI Wen;ZHAO Xiaobing(National Language Resource Monitoring & Research Center of Minority Languages,Minzu University of China,Beijing 100081,China)
出处
《中文信息学报》
CSCD
北大核心
2019年第6期27-34,共8页
Journal of Chinese Information Processing
基金
国家自然科学基金(61331013)
关键词
查询扩展
跨语言词向量
信息检索
query expansion
cross-language word vectors
information retrieval