Detecting Failures In Distributed Systems With The Falcon Spy Network

Joshua B. Leners,Hao Wu,Wei-Lun Hung,Marcos K. Aguilera,Michael Walfish

SOSP '11: ACM SIGOPS 23nd Symposium on Operating Systems Principles Cascais Portugal October, 2011（2011）

引用 150|浏览140

暂无评分

摘要

A common way for a distributed system to tolerate crashes is to explicitly detect them and then recover from them. Interestingly, detection can take much longer than recovery, as a result of many advances in recovery techniques, making failure detection the dominant factor in these systems' unavailability when a crash occurs.This paper presents the design, implementation, and evaluation of Falcon, a failure detector with several features. First, Falcon's common-case detection time is sub-second, which keeps unavailability low. Second, Falcon is reliable: it never reports a process as down when it is actually up. Third, Falcon sometimes kills to achieve reliable detection but aims to kill the smallest needed component. Falcon achieves these features by coordinating a network of spies, each monitoring a layer of the system. Falcon's main cost is a small amount of platform-specific logic. Falcon is thus the first failure detector that is fast, reliable, and viable. As such, it could change the way that a class of distributed systems is built.

查看译文

关键词

Failure detectors,high availability,reliable detection,layer-specific monitors,layer-specific probes,STONITH

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要