摘要
为了研究拼音对检测和纠正语音识别文本错误的影响,提出了一种基于中文语义−音韵信息的文本校对模型。定义了5种拼音编码方法构建字符−音韵嵌入向量,以此作为基于GRU的Seq2Seq模型的输入,并应用注意力机制提取语句的语义−音韵信息来校对语音识别文本错误。针对标注语料不足的问题,提出了一种基于拼音声韵置换的数据增强方法。在AISHELL-3公开数据集的实验结果表明,拼音携带的音韵信息有利于校对语音识别文本错误,所提方法可提升模型的检错性能。
To study the influence of Chinese Pinyin on detecting and correcting text errors in speech recognition,a text proofreading model based on Chinese semantic and phonological information was proposed.Five Pinyin coding methods were designed to construct the character-Pinyin embedding vector that was employed as the input of the Seq2Seq model based on gated recurrent unit.At the same time,the attention mechanism was adopted to extract the Chinese semantic and phonological information of sentences to correct speech recognition errors.Aiming at the problem of insufficient labeled corpus,a data augmentation method was introduced,which could automatically obtain annotated corpora by exchanging the initials or finals of Chinese Pinyin.The experimental results on AISHELL-3’s public data show that phonological in-formation is conducive to the text proofreading model to detect and correct text errors after speech recognition,and the proposed data augmentation method can improve the error detection performance of the model.
作者
仲美玉
吴培良
窦燕
刘毅
孔令富
ZHONG Meiyu;WU Peiliang;DOU Yan;LIU Yi;KONG Lingfu(School of Information Science and Engineering,Yanshan University,Qinhuangdao 066004,China;The Key Laboratory for Computer Virtual Technology and System Integration of Hebei Province,Qinhuangdao 066004,China;The Key Laboratory of Software Engineering of Hebei Province,Qinhuangdao 066004,China)
出处
《通信学报》
EI
CSCD
北大核心
2022年第11期65-79,共15页
Journal on Communications
基金
国家重点研发计划基金资助项目(No.2018YFB1308300)
国家自然科学基金资助项目(No.62276028,No.U20A20167)
北京市自然科学基金资助项目(No.4202026)
河北省自然科学基金资助项目(No.F202103079)
河北省创新能力提升计划基金资助项目(No.22567626H)
河北省软件工程重点实验室基金资助项目(No.22567637H)。
关键词
文本校对
语音识别
拼音
注意力机制
text proofreading
speech recognition
Pinyin
attention mechanism