Distributed Virtual Diskless Checkpointing: A Highly Fault Tolerant Scheme for Virtualized Clusters

Ben Eckart,Xubin He,Chentao Wu,Ferrol Aderholdt,Fang Han,Stephen Scott

Parallel and Distributed Processing Symposium Workshops & PhD Forum（2012）

引用 5|浏览0

暂无评分

摘要

Today's high-end computing systems are facing a crisis of high failure rates due to increased numbers of components. Recent studies have shown that traditional fault tolerant techniques incur overheads that more than double execution times on these highly parallel machines. Thus, future high-end computing must be able to provide adequate fault tolerance at an acceptable cost or the burdens of fault management will severely affect the viability of such systems. Cluster virtualization offers a potentially unique solution for fault management, but brings significant overhead, especially for I/O. In this paper, we propose a novel diskless check pointing technique on clusters of virtual machines. Our technique splits Virtual Machines into sets of orthogonal RAID systems and distributes parity evenly across the cluster, similar to a RAID-5 configuration, but using VM images as data elements. Our theoretical analysis shows that our technique significantly reduces the overhead associated with check pointing by removing the disk I/O bottleneck.

查看译文

关键词

o bottleneck,highly fault tolerant scheme,high-end computing system,adequate fault tolerance,traditional fault tolerant technique,virtualized clusters,cluster virtualization,virtual diskless checkpointing,raid-5 configuration,significant overhead,future high-end computing,fault management,vm image,failure rate,data element,hardware,raid,virtual machine,virtual machines,distributed processing,fault tolerance,virtualisation

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要