Automated learning of performance models for cloud-native applications
File(s)
Author(s)
Wang, Runan
Type
Thesis
Abstract
DevOps is widely used in software development, allowing engineers to release and update software products frequently. It is important to consider system performance to ensure the quality of cloud-native applications. However, working on system performance prediction for microservices and serverless-based applications can be costly and error-prone when models are parameterized on real data. In this thesis, we employ automated learning of performance models from coarse-grained to fine-grained levels. We first focus on accurate model parameterization for microservices. The challenge we tackle is to fit the monitoring data with higher-order moments of the distribution to precisely predict the performance. Next, we focus on automatically generating performance models for serverless applications. The main issue is the information gap between the structural modeling and monitoring granularity, as well as high-level automation. Lastly, we focus on performance diagnosis with code-level tracing strategies to provide inference for internal failures. Even though distributed tracing provides observations at multiple granularities, excessive code instrumentation increases the overheads in terms of computation and storage.
To address the first challenge, we propose a service demand distribution estimation method based on coarse-grained inter-departure times, for both single and multiple service classes. Our experiments yield better performance on fitting real traces of microservices compared to the baseline with Exponential distributions. For the gaps in the automatic performance model generation, we propose to build a compact queueing model based on static analysis and fine-grained profiling data. The experiments on serverless workflows deployed on a public cloud achieve accurate prediction, with a mean error under 7.3%. To mitigate the issues in performance diagnosis, we present a methodology combining code-level analysis to devise tracing data collection strategies and relate probe placements aided by novelty detection methods. Our experimental results show a higher recall score of 90% locating single and multiple faults for 15% latency anomalies.
To address the first challenge, we propose a service demand distribution estimation method based on coarse-grained inter-departure times, for both single and multiple service classes. Our experiments yield better performance on fitting real traces of microservices compared to the baseline with Exponential distributions. For the gaps in the automatic performance model generation, we propose to build a compact queueing model based on static analysis and fine-grained profiling data. The experiments on serverless workflows deployed on a public cloud achieve accurate prediction, with a mean error under 7.3%. To mitigate the issues in performance diagnosis, we present a methodology combining code-level analysis to devise tracing data collection strategies and relate probe placements aided by novelty detection methods. Our experimental results show a higher recall score of 90% locating single and multiple faults for 15% latency anomalies.
Version
Open Access
Date Issued
2024-01
Date Awarded
2024-05
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Casale, Giuliano
Filieri, Antonio
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
