Rigorous assessment of model inference accuracy using language cardinality
File(s)3640332.pdf (1.76 MB)
Published version
Author(s)
Clun, Donato
Shin, Donghwan
Filieri, Antonio
Bianculli, Domenico
Type
Journal Article
Abstract
Models such as finite state automata are widely used to abstract the behavior of software systems by capturing the sequences of events observable during their execution. Nevertheless, models rarely exist in practice and, when they do, get easily outdated; moreover, manually building and maintaining models is costly and error-prone. As a result, a variety of model inference methods that automatically construct models from execution traces have been proposed to address these issues.
However, performing a systematic and reliable accuracy assessment of inferred models remains an open problem. Even when a reference model is given, most existing model accuracy assessment methods may return misleading and biased results. This is mainly due to their reliance on statistical estimators over a finite number of randomly generated traces, introducing avoidable uncertainty about the estimation and being sensitive to the parameters of the random trace generative process.
This article addresses this problem by developing a systematic approach based on analytic combinatorics that minimizes bias and uncertainty in model accuracy assessment by replacing statistical estimation with deterministic accuracy measures. We experimentally demonstrate the consistency and applicability of our approach by assessing the accuracy of models inferred by state-of-the-art inference tools against reference models from established specification mining benchmarks.
However, performing a systematic and reliable accuracy assessment of inferred models remains an open problem. Even when a reference model is given, most existing model accuracy assessment methods may return misleading and biased results. This is mainly due to their reliance on statistical estimators over a finite number of randomly generated traces, introducing avoidable uncertainty about the estimation and being sensitive to the parameters of the random trace generative process.
This article addresses this problem by developing a systematic approach based on analytic combinatorics that minimizes bias and uncertainty in model accuracy assessment by replacing statistical estimation with deterministic accuracy measures. We experimentally demonstrate the consistency and applicability of our approach by assessing the accuracy of models inferred by state-of-the-art inference tools against reference models from established specification mining benchmarks.
Date Issued
2024-05
Date Acceptance
2023-12-12
Citation
ACM Transactions on Software Engineering and Methodology, 2024, 33 (4), pp.1-39
ISSN
1049-331X
Publisher
Association for Computing Machinery (ACM)
Start Page
1
End Page
39
Journal / Book Title
ACM Transactions on Software Engineering and Methodology
Volume
33
Issue
4
Copyright Statement
Copyright © 2024 Copyright held by the owner/author(s).
This work is licensed under a Creative Commons Attribution International 4.0 License.
This work is licensed under a Creative Commons Attribution International 4.0 License.
License URL
Identifier
http://dx.doi.org/10.1145/3640332
Publication Status
Published
Article Number
95
Date Publish Online
2024-01-16