Apache storm configuration platform for dynamic sampling and filtering of data streams

Citations

SCOPUS

5

초록

With the launch of the big data era, in recent years real-time data streams are being used in various fields. It is not practical to collect and process the whole of these big data streams. Thus, there is a need for a sampling method for extracting a good sample and/or a filtering method for extracting necessary data from the entire data stream. Apache Storm is a real-time distributed parallel processing framework for processing large data streams. However, Storm needs to modify the source code, redistribute it, and restart the process when changing the structure or algorithm of the input data. In this paper, we describe the problems of sampling and filtering methods in Storm environment, and define the requirements to solve them. In addition, we design a novel plan model consisting of the input, processing, and output modules in the data stream. Our proposed plan manager has features that can dynamically create, execute, and monitor the plan visually through the Web UI (Web User Interface). © 2019, ICIC International. All rights reserved.

키워드

Apache stormBig dataData streamDistributed processingFilteringSampling
제목
Apache storm configuration platform for dynamic sampling and filtering of data streams
저자
Kim, YoungkukSon, SiwoonMoon, Yang-sae
DOI
10.24507/icicelb.10.01.55
발행일
2019
유형
Article
저널명
ICIC Express Letters, Part B: Applications
10
1
페이지
55 ~ 61