The paper mainly discusses the application of dynamic relevance feedback technique in Web intelligent retrieval and stresses the dynamic similarity in the course of feedback.It discusses the application of dynamic rel...The paper mainly discusses the application of dynamic relevance feedback technique in Web intelligent retrieval and stresses the dynamic similarity in the course of feedback.It discusses the application of dynamic relevance feedback technique from the Web page text and image aspects,and proposes to use the user interest model and the neural network to raise the precision and feasibility of the search result in relevant feedback.展开更多
Content extraction of HTML pages is the basis of the web page clustering and information retrieval,so it is necessary to eliminate cluttered information and very important to extract content of pages accurately.A nove...Content extraction of HTML pages is the basis of the web page clustering and information retrieval,so it is necessary to eliminate cluttered information and very important to extract content of pages accurately.A novel and accurate solution for extracting content of HTML pages was proposed.First of all,the HTML page is parsed into DOM object and the IDs of all leaf nodes are generated.Secondly,the score of each leaf node is calculated and the score is adjusted according to the relationship with neighbors.Finally,the information blocks are found according to the definition,and a universal classification algorithm is used to identify the content blocks.The experimental results show that the algorithm can extract content effectively and accurately,and the recall rate and precision are 96.5% and 93.8%,respectively.展开更多
文摘The paper mainly discusses the application of dynamic relevance feedback technique in Web intelligent retrieval and stresses the dynamic similarity in the course of feedback.It discusses the application of dynamic relevance feedback technique from the Web page text and image aspects,and proposes to use the user interest model and the neural network to raise the precision and feasibility of the search result in relevant feedback.
基金Project(2012BAH18B05) supported by the Supporting Program of Ministry of Science and Technology of China
文摘Content extraction of HTML pages is the basis of the web page clustering and information retrieval,so it is necessary to eliminate cluttered information and very important to extract content of pages accurately.A novel and accurate solution for extracting content of HTML pages was proposed.First of all,the HTML page is parsed into DOM object and the IDs of all leaf nodes are generated.Secondly,the score of each leaf node is calculated and the score is adjusted according to the relationship with neighbors.Finally,the information blocks are found according to the definition,and a universal classification algorithm is used to identify the content blocks.The experimental results show that the algorithm can extract content effectively and accurately,and the recall rate and precision are 96.5% and 93.8%,respectively.