期刊文献+
共找到2篇文章
< 1 >
每页显示 20 50 100
Benchmarks for Pirá2.0,a Reading Comprehension Dataset about the Ocean,the Brazilian Coast,and Climate Change
1
作者 Paulo Pirozelli Marcos M.José +5 位作者 Igor Silveira Flávio Nakasato Sarajane M.Peres Anarosa A.F.Brandão Anna H.R.Costa Fabio G.Cozman 《Data Intelligence》 EI 2024年第1期29-63,共35页
Piráis a reading comprehension dataset focused on the ocean,the Brazilian coast,and climate change,built from a collection of scientific abstracts and reports on these topics.This dataset represents a versatile l... Piráis a reading comprehension dataset focused on the ocean,the Brazilian coast,and climate change,built from a collection of scientific abstracts and reports on these topics.This dataset represents a versatile language resource,particularly useful for testing the ability of current machine learning models to acquire expert scientific knowledge.Despite its potential,a detailed set of baselines has not yet been developed for Pirá.By creating these baselines,researchers can more easily utilize Piráas a resource for testing machine learning models across a wide range of question answering tasks.In this paper,we define six benchmarks over the Pirádataset,covering closed generative question answering,machine reading comprehension,information retrieval,open question answering,answer triggering,and multiple choice question answering.As part of this effort,we have also produced a curated version of the original dataset,where we fixed a number of grammar issues,repetitions,and other shortcomings.Furthermore,the dataset has been extended in several new directions,so as to face the aforementioned benchmarks:translation of supporting texts from English into Portuguese,classification labels for answerability,automatic paraphrases of questions and answers,and multiple choice candidates.The results described in this paper provide several points of reference for researchers interested in exploring the challenges provided by the Pirádataset. 展开更多
关键词 Natural language processing Question answering Benchmarks Language resource DomainOriented dataset scientific knowledge text dataset
原文传递
Measuring Similarity of Academic Articles with Semantic Profile and Joint Word Embedding 被引量:11
2
作者 Ming Liu Bo Lang +1 位作者 Zepeng Gu Ahmed Zeeshan 《Tsinghua Science and Technology》 SCIE EI CAS CSCD 2017年第6期619-632,共14页
Long-document semantic measurement has great significance in many applications such as semantic searchs, plagiarism detection, and automatic technical surveys. However, research efforts have mainly focused on the sema... Long-document semantic measurement has great significance in many applications such as semantic searchs, plagiarism detection, and automatic technical surveys. However, research efforts have mainly focused on the semantic similarity of short texts. Document-level semantic measurement remains an open issue due to problems such as the omission of background knowledge and topic transition. In this paper, we propose a novel semantic matching method for long documents in the academic domain. To accurately represent the general meaning of an academic article, we construct a semantic profile in which key semantic elements such as the research purpose, methodology, and domain are included and enriched. As such, we can obtain the overall semantic similarity of two papers by computing the distance between their profiles. The distances between the concepts of two different semantic profiles are measured by word vectors. To improve the semantic representation quality of word vectors, we propose a joint word-embedding model for incorporating a domain-specific semantic relation constraint into the traditional context constraint. Our experimental results demonstrate that, in the measurement of document semantic similarity, our approach achieves substantial improvement over state-of-the-art methods, and our joint word-embedding model produces significantly better word representations than traditional word-embedding models. 展开更多
关键词 document semantic similarity text understanding semantic enrichment word embedding scientific literature analysis
原文传递
上一页 1 下一页 到第
使用帮助 返回顶部