摘要
针对目前通用搜索引擎搜索到的结果过多、与主题相关性不强的现状,提出一种基于网页分块技术的主题爬行器实现方法,并实现了一个原型系统Crawler1.实验结果表明,本系统性能较好,所爬网页的相关度在55%以上.
In the light of result returned currently by general-purpose search engines being excessive, and having no strong similarity with the topic, this paper covers a technique of dividing the web page to chunks to implement a focused crawler. With this method, Crawlerl, a prototype of a focused crawler has been realized. Experimental results indicate that Crawlerl has better performance. The number of topic web pages crawled by Crawlerl attains more than 55%.
出处
《吉林大学学报(理学版)》
CAS
CSCD
北大核心
2007年第6期959-965,共7页
Journal of Jilin University:Science Edition
基金
国家自然科学基金(批准号:60373099)
关键词
主题搜索
主题爬行
相关度分析
网页分块
topic-specific search
focused crawling
relevance analysis
page segmentation