期刊文献+
共找到1篇文章
< 1 >
每页显示 20 50 100
Offline Urdu Nastaleeq Optical Character Recognition Based on Stacked Denoising Autoencoder 被引量:2
1
作者 Ibrar Ahmad Xiaojie Wang +1 位作者 Ruifan Li Shahid Rasheed 《China Communications》 SCIE CSCD 2017年第1期146-157,共12页
Offline Urdu Nastaleeq text recognition has long been a serious problem due to its very cursive nature. In order to get rid of the character segmentation problems, many researchers are shifting focus towards segmentat... Offline Urdu Nastaleeq text recognition has long been a serious problem due to its very cursive nature. In order to get rid of the character segmentation problems, many researchers are shifting focus towards segmentation free ligature based recognition approaches. Majority of the prevalent ligature based recognition systems heavily rely on hand-engineered feature extraction techniques. However, such techniques are more error prone and may often lead to a loss of useful information that might hardly be captured later by any manual features. Most of the prevalent Urdu Nastaleeq test recognition was trained and tested on small sets. This paper proposes the use of stacked denoising autoencoder for automatic feature extraction directly from raw pixel values of ligature images. Such deep learning networks have not been applied for the recognition of Urdu text thus far. Different stacked denoising autoencoders have been trained on 178573 ligatures with 3732 classes from un-degraded(noise free) UPTI(Urdu Printed Text Image) data set. Subsequently, trained networks are validated and tested on degraded versions of UPTI data set. The experimental results demonstrate accuracies in range of 93% to 96% which are better than the existing Urdu OCR systems for such large dataset of ligatures. 展开更多
关键词 offline printed ligature recognition urdu nastaleeq denoising autoencoder deep learning classification
下载PDF
上一页 1 下一页 到第
使用帮助 返回顶部