DeepBAT: Performance and cost optimization of serverless inference using transformers
File(s) 2025_IPDPS_DeepBAT_CR.pdf (776.89 KB)
Accepted version
Author(s)
Sun, Bowen
Pinciroli, Riccardo
Casale, Giuliano
Smirni, Evgenia
Type
Conference Paper
Abstract
Serverless computing is an autoscaling pay-as-yougo paradigm that can efficiently support machine learning inference especially under bursty workload conditions. Within the serverless paradigm, batching ML inference requests before serving them is widely adopted. Thanks to its parallelism properties, batching can highly improve inference performance while reducing the monetary cost of serverless. Identifying the correct serverless parameterization to simultaneously meet conflicting targets (i.e., keep monetary cost at a minimum while meeting pre-defined service level objectives, SLO) may be cast as a resource allocation problem. In this paper, we illustrate that a deep surrogate model can quickly discover optimized serverless configurations by learning the relationship among the workload patterns and achieve performance measures. We develop DeepBAT, an SLO-aware framework that leverages the Transformer encoder and multi-head attention mechanism to optimize the performance of serverless inference subject to bursty and previously unobserved workloads. We illustrate the effectiveness of DeepBAT on a set of case studies and show that for the problem of inference serving on AWS Lambda, DeepBAT can speed up the solution time of state-of-the-art analytic solutions by over 55 times while generalizing remarkably well on unseen workloads.
Date Issued
2025-07-23
Date Acceptance
2025-06-01
Citation
2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 2025, pp.335-346
ISBN
979-8-3315-3238-3
ISSN
1530-2075
Publisher
IEEE Computer Society
Start Page
335
End Page
346
Journal / Book Title
2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
Copyright Statement
Copyright © 2025, IEEE. This is the author’s accepted manuscript made available under a CC-BY licence in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy)
License URL
Source
2025 International Parallel and Distributed Processing Symposium-IPDPS-Annual
Subjects
Batching Technique
Bursty Workloads
Computer Science
Computer Science, Theory & Methods
Multi-head Attention
Science & Technology
Serverless Computing
SLO-Aware
Technology
Transformer
Publication Status
Published
Start Date
2025-06-03
Finish Date
2025-06-07
Coverage Spatial
Milan, Italy
