ServiceRouter: hyperscale and minimal cost service mesh at Meta
File(s) osdi23-saokar.pdf (1.36 MB)
Published version
Author(s)
Type
Conference Paper
Abstract
Datacenter applications are often structured as many inter connected microservices, and the service mesh has become a
popular approach to route RPC traffic among services. This pa per presents ServiceRouter (SR), Meta’s global service mesh,
which has been in production since 2012. SR differs from
publicly known service meshes in several important ways.
First, SR is designed for hyperscale and currently uses mil lions of L7 routers to route tens of billions of requests per
second across tens of thousands of services. Second, while
SR adopts the common approach of using sidecar or remote
proxies to route 1% of RPC requests in our fleet, it employs a
routing library that is directly linked into service executables
to route the remaining 99% directly from clients to servers,
without the extra hop of going through a proxy. This approach
significantly reduces the hardware costs of our hyperscale ser vice mesh, saving hundreds of thousands of machines. Third,
SR provides built-in support for sharded services, which ac count for 68% of RPC requests in our fleet, whereas existing
general-purpose service meshes do not support sharded ser vices. Finally, SR introduces the concept of locality rings to
simultaneously minimize RPC latency and balance load across
geo-distributed datacenter regions, which, to our knowledge,
has not been attempted before.
popular approach to route RPC traffic among services. This pa per presents ServiceRouter (SR), Meta’s global service mesh,
which has been in production since 2012. SR differs from
publicly known service meshes in several important ways.
First, SR is designed for hyperscale and currently uses mil lions of L7 routers to route tens of billions of requests per
second across tens of thousands of services. Second, while
SR adopts the common approach of using sidecar or remote
proxies to route 1% of RPC requests in our fleet, it employs a
routing library that is directly linked into service executables
to route the remaining 99% directly from clients to servers,
without the extra hop of going through a proxy. This approach
significantly reduces the hardware costs of our hyperscale ser vice mesh, saving hundreds of thousands of machines. Third,
SR provides built-in support for sharded services, which ac count for 68% of RPC requests in our fleet, whereas existing
general-purpose service meshes do not support sharded ser vices. Finally, SR introduces the concept of locality rings to
simultaneously minimize RPC latency and balance load across
geo-distributed datacenter regions, which, to our knowledge,
has not been attempted before.
Date Issued
2023-07
Date Acceptance
2023-07-10
Citation
Proceedings of the 17th Usenix Symposium on Operating Systems Design and Implementation, OSDI 2023, 2023, pp.969-985
Publisher
USENIX ASSOC
Start Page
969
End Page
985
Journal / Book Title
Proceedings of the 17th Usenix Symposium on Operating Systems Design and Implementation, OSDI 2023
Copyright Statement
© 2023. The Author(s). The USENIX Association. Rights to individual papers remain with the author or the author’s employer. Permission is granted for the noncommercial reproduction of the complete work for educational or research purposes. Permission is granted to print, primarily for one person’s exclusive use, a single copy of these Proceedings.
Identifier
https://www.webofscience.com/api/gateway?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:001066453900053&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=a2bf6146997ec60c407a63945d4e92bb
Source
17th USENIX Symposium on Operating Systems Design and Implementation (OSDI)
Publication Status
Published
Start Date
2023-07-10
Finish Date
2023-07-12
Coverage Spatial
MA, Boston
