Artificial neural networks based techniques for anomaly detection in Apache Spark
File(s)
Author(s)
Alnafessah, Ahmad
Casale, Giuliano
Type
Journal Article
Abstract
Late detection and manual resolutions of
performance anomalies in Cloud Computing and Big
Data systems may lead to performance violations and
financial penalties. Motivated by this issue, we propose an artificial neural network based methodology
for anomaly detection especially for the Apache Spark
in-memory processing platforms. Apache Spark has become widely adopted by industry because of its speed
and generality, however there is still a shortage of comprehensive performance anomaly detection methods applicable to this platform. We propose artificial neural
networks driven methodology to quickly sift through
Spark logs data and operating system monitoring metrics to accurately detect and classify anomalous behaviors based on the Spark resilient distributed dataset
(RDD) characteristics. The proposed method is evaluated against three popular machine learning algorithms, decision trees, nearest neighbor, and support
vector machine (SVM), as well as against four variants that consider different monitoring datasets. The
results prove that our proposed method outperforms
other methods, typically achieving 98%-99% F-scores,
and offering much greater accuracy than alternative
techniques to detect both the period in which anomalies
occurred and their type.
performance anomalies in Cloud Computing and Big
Data systems may lead to performance violations and
financial penalties. Motivated by this issue, we propose an artificial neural network based methodology
for anomaly detection especially for the Apache Spark
in-memory processing platforms. Apache Spark has become widely adopted by industry because of its speed
and generality, however there is still a shortage of comprehensive performance anomaly detection methods applicable to this platform. We propose artificial neural
networks driven methodology to quickly sift through
Spark logs data and operating system monitoring metrics to accurately detect and classify anomalous behaviors based on the Spark resilient distributed dataset
(RDD) characteristics. The proposed method is evaluated against three popular machine learning algorithms, decision trees, nearest neighbor, and support
vector machine (SVM), as well as against four variants that consider different monitoring datasets. The
results prove that our proposed method outperforms
other methods, typically achieving 98%-99% F-scores,
and offering much greater accuracy than alternative
techniques to detect both the period in which anomalies
occurred and their type.
Date Issued
2020-06-01
Date Acceptance
2019-10-03
Citation
Cluster Computing, 2020, 23, pp.1345-1360
ISSN
1386-7857
Publisher
Springer (part of Springer Nature)
Start Page
1345
End Page
1360
Journal / Book Title
Cluster Computing
Volume
23
Copyright Statement
© The Author(s) 2019. This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made.
License URL
Subjects
Science & Technology
Technology
Computer Science, Information Systems
Computer Science, Theory & Methods
Computer Science
Performance anomalies
Apache Spark
Neural network
Big data
Machine learning
Artificial intelligence
Resilient distributed dataset (RDD)
CONJUGATE-GRADIENT
0805 Distributed Computing
Distributed Computing
Publication Status
Published
Date Publish Online
2019-10-23
