Apache hadoop goes realtime at Facebook.

Dhruba Borthakur, Jonathan Gray,Joydeep Sen Sarma, Kannan Muthukkaruppan,Nicolas Spiegelberg, Hairong Kuang,Karthik Ranganathan,Dmytro Molkov,Aravind Menon, Samuel Rash, Rodrigo Schmidt,Amitanand S. Aiyer

MOD(2011)

引用 698|浏览551
暂无评分
摘要
ABSTRACTFacebook recently deployed Facebook Messages, its first ever user-facing application built on the Apache Hadoop platform. Apache HBase is a database-like layer built on Hadoop designed to support billions of messages per day. This paper describes the reasons why Facebook chose Hadoop and HBase over other systems such as Apache Cassandra and Voldemort and discusses the application's requirements for consistency, availability, partition tolerance, data model and scalability. We explore the enhancements made to Hadoop to make it a more effective realtime system, the tradeoffs we made while configuring the system, and how this solution has significant advantages over the sharded MySQL database scheme used in other applications at Facebook and many other web-scale companies. We discuss the motivations behind our design choices, the challenges that we face in day-to-day operations, and future capabilities and improvements still under development. We offer these observations on the deployment as a model for other companies who are contemplating a Hadoop-based solution over traditional sharded RDBMS deployments.
更多
查看译文
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要