A Demonstration Of Geospark: A Cluster Computing Framework For Processing Big Spatial Data

2016 IEEE 32nd International Conference on Data Engineering (ICDE)(2016)

引用 51|浏览24
暂无评分
摘要
This paper demonstrates GEOSPARK a cluster computing framework for developing and processing large-scale spatial data analytics programs. GEOSPARK consists of three main layers: Apache Spark Layer, Spatial RDD Layer and Spatial Query Processing Layer. Apache Spark Layer provides basic Apache Spark functionalities as regular RDD operations. Spatial RDD Layer consists of three novel Spatial Resilient Distributed Datasets (SRDDs) which extend regular Apache Spark RDD to support geometrical and spatial objects with data partitioning and indexing. Spatial Query Processing Layer executes spatial queries (e.g., Spatial Join) on SRDDs. The dynamic status of SRDDs and spatial operations are visualized by GEOSPARK monitoring map interface. We demonstrate GEOSPARK using three spatial analytics applications (spatial aggregation, autocorrelation and co-location) to show how users can easily define their spatial analytics tasks and efficiently process such tasks on large-scale spatial data at interactive performance.
更多
查看译文
关键词
GeoSpark,cluster computing framework,large-scale spatial data analytics programs,Apache Spark layer,spatial RDD layer,spatial query processing layer,novel spatial resilient distributed datasets,SRDD
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要