Fine-grained tracing for performance anomaly diagnosis of serverless functions
File(s) 3785005.pdf (4.47 MB)
Published version
Author(s)
Wang, Runan
Yu, Guangba
Casale, Giuliano
Chen, Pengfei
Filieri, Antonio
Type
Journal Article
Abstract
Serverless function compositions subject to unpredictable faults are challenging to evaluate for root cause analysis. Even though distributed tracing provides observations at multiple levels of granularity for troubleshooting, excessive code instrumentation increases the tracing overheads in terms of both computation and storage. Therefore, developers face the challenge of where and how to instrument serverless functions to maximize the likelihood of locating faults based on tracing data while minimizing tracing overhead and costs.
In this article, we propose a methodology to instrument an application with code-level tracing to infer the location of faults, taking into account constraints in terms of the maximum cost of the instrumentation and testing. We encode the tracing probe placement based on the control flow graph of the application and devise heuristics-based tracing data collection strategies to relate possible probe placements with their ability to locate a fault. Then we train novelty detection models to identify the internal anomalies and present an enhanced global search algorithm that automatically computes a probe placement with optimal fault localization ability versus cost. Experimental results show high performance in locating single and multiple faults with over 90% recall score for up to 15% latency anomalies, with minimal instrumentation overhead.
In this article, we propose a methodology to instrument an application with code-level tracing to infer the location of faults, taking into account constraints in terms of the maximum cost of the instrumentation and testing. We encode the tracing probe placement based on the control flow graph of the application and devise heuristics-based tracing data collection strategies to relate possible probe placements with their ability to locate a fault. Then we train novelty detection models to identify the internal anomalies and present an enhanced global search algorithm that automatically computes a probe placement with optimal fault localization ability versus cost. Experimental results show high performance in locating single and multiple faults with over 90% recall score for up to 15% latency anomalies, with minimal instrumentation overhead.
Date Issued
2026-03-01
Date Acceptance
2025-11-28
Citation
ACM Transactions on Autonomous and Adaptive Systems, 2026, 21 (1)
ISSN
1556-4665
Publisher
Association for Computing Machinery (ACM)
Journal / Book Title
ACM Transactions on Autonomous and Adaptive Systems
Volume
21
Issue
1
Copyright Statement
Copyright © 2026 Copyright held by the owner/author(s). This work is licensed under Creative Commons Attribution International 4.0.
License URL
Identifier
10.1145/3785005
Subjects
Serverless Functions
Tracing
Code Instrumentation
Anomaly Detection
Optimization This work is licensed under Creative Commons Attribution International 4.0
Publication Status
Published
Article Number
5
Date Publish Online
2026-01-19
