Low-latency, query-driven analytics over voluminous multidimensional, spatiotemporal datasets

Malensek, Matthew, author; Pallickara, Shrideep, advisor; Pallickara, Sangmi Lee, advisor; Bohm, A. P. Willem, committee member; Draper, Bruce, committee member; Breidt, F. Jay, committee member

Low-latency, query-driven analytics over voluminous multidimensional, spatiotemporal datasets

Files

Malensek_colostate_0053A_14370.pdf (5.63 MB)

Date

2017

Authors

Malensek, Matthew, author

Pallickara, Shrideep, advisor

Pallickara, Sangmi Lee, advisor

Bohm, A. P. Willem, committee member

Draper, Bruce, committee member

Breidt, F. Jay, committee member

Abstract

Ubiquitous data collection from sources such as remote sensing equipment, networked observational devices, location-based services, and sales tracking has led to the accumulation of voluminous datasets; IDC projects that by 2020 we will generate 40 zettabytes of data per year, while Gartner and ABI estimate 20-35 billion new devices will be connected to the Internet in the same time frame. The storage and processing requirements of these datasets far exceed the capabilities of modern computing hardware, which has led to the development of distributed storage frameworks that can scale out by assimilating more computing resources as necessary. While challenging in its own right, storing and managing voluminous datasets is only the precursor to a broader field of study: extracting knowledge, insights, and relationships from the underlying datasets. The basic building block of this knowledge discovery process is analytic queries, encompassing both query instrumentation and evaluation. This dissertation is centered around query-driven exploratory and predictive analytics over voluminous, multidimensional datasets. Both of these types of analysis represent a higher-level abstraction over classical query models; rather than indexing every discrete value for subsequent retrieval, our framework autonomously learns the relationships and interactions between dimensions in the dataset (including time series and geospatial aspects), and makes the information readily available to users. This functionality includes statistical synopses, correlation analysis, hypothesis testing, probabilistic structures, and predictive models that not only enable the discovery of nuanced relationships between dimensions, but also allow future events and trends to be predicted. This requires specialized data structures and partitioning algorithms, along with adaptive reductions in the search space and management of the inherent trade-off between timeliness and accuracy. The algorithms presented in this dissertation were evaluated empirically on real-world geospatial time-series datasets in a production environment, and are broadly applicable across other storage frameworks.

Subject

big data

analytic queries

distributed systems

URI

https://hdl.handle.net/10217/183998

Collections

2000-2019
Theses and Dissertations

Full item page

Low-latency, query-driven analytics over voluminous multidimensional, spatiotemporal datasets

Files

Date

Authors

Journal Title

Journal ISSN

Volume Title

Abstract

Description

Rights Access

Subject

Citation

URI

Associated Publications

Collections