摘要
互联网上存在大量的免费、公开、有价值的非合约形式的对地观测数据源,这些数据源具有网页查询入口、海量数据隐藏在后台的大型数据库且数据共享平台多样、不同种类空间数据平台难以互联等特点,难以利用传统技术实现数据汇聚和共享。在阐述目前遇到的问题后,提出了一种基于暗网爬虫架构的非合约异构分布式数据源被动汇聚架构;设计出一套数据源识别标准、非合约式数据源发现机制、非合约式数据源搜索条件树构建模式、非合约式数据源索引机制以及数据源异步更新规则,成功汇聚了分布在国际上不同网络域的五个大型对地观测数据源,包括NASA、USGS、ASAR等三个国际上使用较为广泛的运行性数据源;形成了对地观测数据资源自动化汇聚和更新工具集,最终使用户可以通过统一查询界面获取非合约对地观测数据资源信息。
It is difficult to use the traditional technology to realize data aggregation and data sharing for the Internet, which contains a large number of free, open and valuable non-contractual earth obser- vation data sources. These data sources have the characteristics of webpage query entrance, massive data hidden in the network background database, data sharing platform diversity and different kinds of spatial data platform to interconnect etc. Considering these problems, a non-contractual heterogeneous distribu ted data sources passive aggregation architecture is proposed, which is based on deep web crawler tech- nology. Meanwhile, we design a data source identification standard, non-contractual data source discov- ery mechanism, non-contractual data source search tree building mode, non-contractual data source inde- xing mechanism and data source asynchronous update rules. Using this mechanism, we archive 5 data sources of large data sharing system including NASA, USGS, ASAR, these three widely used data re- sources and form earth observation data resouree automatic aggregation and update tool sets. Eventual- ly, through a unified query interface, users can obtain non-contractual earth observation data resouree information.
出处
《计算机工程与科学》
CSCD
北大核心
2013年第11期68-75,共8页
Computer Engineering & Science
基金
国家863计划资助项目(2012AA12A301)
关键词
对地观测数据搜索
非合约式数据源
暗网爬虫
增量爬虫
earth observation data
search non-contractual data sources
deep web crawler
incremental crawler