Adaptive Data Depth via Multi-Armed Bandits

JOURNAL OF MACHINE LEARNING RESEARCH(2023)

引用 0|浏览10
暂无评分
摘要
Data depth, introduced by Tukey (1975), is an important tool in data science, robust statistics, and computational geometry. One chief barrier to its broader practical utility is that many common measures of depth are computationally intensive, requiring on the order of n(d) operations to exactly compute the depth of a single point within a data set of n points in d-dimensional space. Often however, we are not directly interested in the absolute depths of the points, but rather in their relative ordering. For example, we may want to find the most central point in a data set (a generalized median), or to identify and remove all outliers (points on the fringe of the data set with low depth). With this observation, we develop a novel instance-adaptive algorithm for adaptive data depth computation by reducing the problem of exactly computing n depths to an n-armed stochastic multi-armed bandit problem which we can efficiently solve. We focus our exposition on simplicial depth, developed by Liu (1990), which has emerged as a promising notion of depth due to its interpretability and asymptotic properties. We provide general data-dependent theoretical guarantees for our proposed algorithms, which readily extend to many other common measures of data depth including majority depth, Oja depth, and likelihood depth. When specialized to the case where the gaps in the data follow a power law distribution with parameter ff < 2, we reduce the complexity of identifying the deepest point in the data set (the simplicial median) from O(n(d)) to <(O)over tilde> (n(d-(d-1)a/2)), where (O) over tilde suppresses a logarithmic factor. We corroborate our theoretical results with numerical experiments on synthetic data, showing the practical utility of our proposed methods.
更多
查看译文
关键词
multi-armed bandits,data depth,adaptive computation,simplicial depth
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要