Anomaly Detection in Cloud-Native Systems
Lomio, Francesco; Moreschini, Sergio; Li, Xiaozhou; Lenarduzzi, Valentina (2022)
Lomio, Francesco
Moreschini, Sergio
Li, Xiaozhou
Lenarduzzi, Valentina
Teoksen toimittaja(t)
Callico, Gustavo M.
Hebig, Regina
Wortmann, Andreas
IEEE
2022
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202402092223
https://urn.fi/URN:NBN:fi:tuni-202402092223
Kuvaus
Peer reviewed
Tiivistelmä
Companies develop cloud-native systems deployed on public and private clouds. Since private clouds have limited resources, the systems should run efficiently by keeping performance related anomalies under control. The goal of this work is to understand whether a set of five performance-related KPIs depends on the metrics collected at runtime by Kafka, Zookeeper, and other tools (168 different metrics). We considered four weeks worth of runtime data collected from a system running in production. We trained eight Machine Learning algorithms on three weeks worth of data and tested them on one week’s worth of data to compare their prediction accuracy and their training and testing time. It is possible to detect performance-related anomalies with a very high level of accuracy (higher than 95% AUC) and with very limited training time (between 8 and 17 minutes). Machine Learning algorithms can help to identify runtime anomalies and to detect them efficiently. Future work will include the identification of a proactive approach to recognize the root cause of the anomalies and to prevent them as early as possible.
Kokoelmat
- TUNICRIS-julkaisut [19020]