摘要
互联网上信息量的激增,迫切需要一些自动化的工具帮助人们在海量信息源中迅速找到真正需要的信息,如标题、链接e、mail和图片等,而HTML语言所表述的Web页面经浏览器分析后只适合浏览,不适合作为一种数据交换的方式由机器处理。介绍了HTMLParser的原理和java正则表达式相关知识,基于HTMLParser包和正则表达式。以提取网站内部email信息为例,提出了Web信息抽取系统设计方案,阐述了email信息抽取的工作原理和关键技术,给出了email抽取算法,并详细介绍了系统的抽取URL、email和存储模块,抽取结果保存于数据库中,供机器检索利用。
The rapid growth of the Web contents increasers the need for some automatic tools to help to find the exact information among the magnanimous information sources such as titles, links, emails, pictures etc. The Web pages expressed by HTML, after analyzed by Internet Explorer, are suitable for browse, but not for machine processing as the way of data exchange. The principle of HTMLParser and related knowledge of regular expression, package HTMLParser and regular expression were introduced. Taking extracting email information inside websites as an example, the scheme of design was proposed. The principle of email extraction and key technique were presented. The algorithm of email extraction was given. URL extraction module, email extraction module and storage module were described in detail. The result of extraction is stored in database for the use of data retrieval.
出处
《辽宁石油化工大学学报》
CAS
2006年第2期83-86,共4页
Journal of Liaoning Petrochemical University