Hardware-efficient data imputation through DBMS extensibility
File(s)3681954.3682016.pdf (1.17 MB)
Published version
Author(s)
Mohr-Daurat, Hubert
Theodorakis, Georgios
Pirk, Holger
Type
Conference Paper
Abstract
The separation of data and code/queries has served Data Management Systems (DBMSs) well for decades. However, while the resulting soundness and rigidity is the basis for many performance-oriented optimizations, it lacks the flexibility to efficiently support modern data science applications: data cleansing, data ingestion/augmentation or generative models. To support such applications without sacrificing performance, we propose a new logical data model called Homoiconic Collection Processing (HCP). HCP is based on a well-known Meta-Programming concept called Homoiconicity (a unified representation for code and data).
In a DBMS, HCP supports the storage of “classic” relational data but also allows the storage and evaluation of code fragments we refer to as “Homoiconic Expressions”. Homoiconic Expressions enable applications such as data imputation directly in the database kernel. Implemented naïvely, such flexibility would come at a prohibitive cost in terms of performance. To make HCP performance-competitive with highly-tuned in-memory DBMSs, we develop a novel storage and processing model called Shape-Wise Micro-batching (SWM) and implement it in a system called BOSS. BOSS is performance-competitive with high-performance DBMSs while offering unprecedented extensibility. To demonstrate the extensibility, we implement an extension for impute-and-query workloads: BOSS outperforms state-of-the-art homoiconic runtimes and data imputation systems by two to five orders of magnitude.
In a DBMS, HCP supports the storage of “classic” relational data but also allows the storage and evaluation of code fragments we refer to as “Homoiconic Expressions”. Homoiconic Expressions enable applications such as data imputation directly in the database kernel. Implemented naïvely, such flexibility would come at a prohibitive cost in terms of performance. To make HCP performance-competitive with highly-tuned in-memory DBMSs, we develop a novel storage and processing model called Shape-Wise Micro-batching (SWM) and implement it in a system called BOSS. BOSS is performance-competitive with high-performance DBMSs while offering unprecedented extensibility. To demonstrate the extensibility, we implement an extension for impute-and-query workloads: BOSS outperforms state-of-the-art homoiconic runtimes and data imputation systems by two to five orders of magnitude.
Date Issued
2024-08-30
Date Acceptance
2024-06-29
Citation
Proceedings of the VLDB Endowment, Vol. 17, No. 11, 2024, 17, pp.3497-3510
ISSN
2150-8097
Publisher
VLDB Endowment
Start Page
3497
End Page
3510
Journal / Book Title
Proceedings of the VLDB Endowment, Vol. 17, No. 11
Volume
17
Copyright Statement
This work is licensed under the Creative Commons BY-NC-ND 4.0 International
License. Visit https://creativecommons.org/licenses/by-nc-nd/4.0/ to view a copy of
this license. For any use beyond those covered by this license, obtain permission by
emailing info@vldb.org. Copyright is held by the owner/author(s). Publication rights
licensed to the VLDB Endowment.
License. Visit https://creativecommons.org/licenses/by-nc-nd/4.0/ to view a copy of
this license. For any use beyond those covered by this license, obtain permission by
emailing info@vldb.org. Copyright is held by the owner/author(s). Publication rights
licensed to the VLDB Endowment.
Identifier
https://dl.acm.org/doi/10.14778/3681954.3682016
Source
VLDB 2024
Publication Status
Published
Start Date
2024-08-25
Finish Date
2024-08-29
Coverage Spatial
Guangzhou, China
Date Publish Online
2024-08-30