摘要
软件 DSM(distributed shared memory)系统在机群上构造了共享存储编程环境,结合了共享存储的易编程性和机群的可扩展性,引起了广泛的研究.由于软件 DSM 系统是一个分布式系统,系统失败风险大,需要实现容错技术以促进其实用化.利用用户级检查点技术,在支持域存储一致模型的软件 DSM 系统 JIAJIA 的基础上,设计并实现了一个可恢复的高可移植的软件 DSM 系统 JIACKPT(JIAjia with ChecKPoinTing).由于采用适合软件 DSM 系统的强全局一致状态以及多种优化措施,JIACKPT 易于实现且获得很好的性能.在一个 8 节点的 PC 机群上的应用测试表明,即使每分钟做一次检查点,大部分应用的检查点开销也小于 10%.此外,JIACKPT 还具有高可移植性.这些都表明 JIACKPT 已经成为一个比较实用的系统.
Software distributed shared memory (DSM) system has constructed a virtual shared memory abstract on cluster, which combines the programmability of shared memory and fine scalability of cluster. So it is widely studied. Software DSM system is easy to fail because it is a distributed system, some kinds of fault tolerance are necessary for it to be more practical. A recoverable and portable software DSM system, JIACKPT (JIAjia with ChecKPoinTing), has been designed and implemented to tolerate the fault of system. JIACKPT, based on JIAJIA, has adopted the checkpointing technology. By maintaining the strict global consistent state and using some optimization techniques, JIACKPT has gotten high performance. The experimental results on an 8-node PC cluster show that the checkpoint overhead is less than 10% of the whole execution time when checkpoint is done once per minute. JIACKPT also has good portability and can run on several operating systems, such as Linux, Solaris, etc. JIACKPT is a practical recoverable software DSM system.
出处
《软件学报》
EI
CSCD
北大核心
2005年第2期165-173,共9页
Journal of Software
基金
国家自然科学基金~~
关键词
软件DSM系统
检查点
全局一致状态
JIAJIA
Computer software
Computer workstations
Data communication systems
Fault tolerant computer systems