Quantitative Evaluation of Process Discovery Enhancement Using Clustering-Based Techniques

Article Preview

Abstract:

Business process discovery aims to reconstruct accurate models of organizational workflowsfrom event logs, yet classical algorithms often struggle with noisy and heterogeneous data,producing overly complex or imprecise models. This paper presents a systematic, quantitative evaluationof clustering-based preprocessing as a means to enhance process discovery. Using the BPIChallenge 2014 incident management log, we applied comprehensive feature engineering to extractstructural, temporal, and behavioral attributes, followed by multiple clustering techniques (K-Means,DBSCAN, HDBSCAN, GMM). Within each cluster, four established discovery algorithms (Alpha Miner,Heuristic Miner, Inductive Miner, ILP Miner) were executed from the PM4Py Python libraryand evaluated using fitness, precision, generalization, and simplicity metrics. Results show that clusteringconsistently improved model simplicity and precision, while fitness remained stable and generalizationtended to decrease. The extent of improvement depended on both the clustering methodand the discovery algorithm: K-Means with higher cluster numbers enhanced fitness and simplicityfor Alpha Miner; HDBSCAN improved precision and simplicity for Heuristic Miner; InductiveMiner benefited from K-Means even with small cluster numbers; whereas ILP Miner models remainedhighly complex regardless of clustering. These findings highlight the trade-off between interpretabilityand generalization, and demonstrate that clustering-based preprocessing can yield moretransparent and actionable process models.

You might also be interested in these eBooks

Info:

Periodical:

Pages:

435-450

Citation:

Online since:

August 2026

Export:

Price:

Permissions CCC:

Permissions PLS:

Сopyright:

© 2026 Trans Tech Publications Ltd. All Rights Reserved

Share:

Citation:

* - Corresponding Author

[1] W. M. P. van der Aalst. Process Mining: Data Science in Action. Springer, 2016.

Google Scholar

[2] M. Dumas, M. La Rosa, J. Mendling, and H. A. Reijers. Fundamentals of Business Process Management. Springer, 2018.

Google Scholar

[3] A. Augusto, R. Conforti, M. Dumas, M. La Rosa, F. M. Maggi, A. Marrella, and I. Weber. Automated discovery of process models from event logs: Review and benchmark. Transactions on Knowledge and Data Engineering, 31(4):686–705, 2019.

Google Scholar

[4] S.J.J. Leemans, D. Fahland, and W.M.P. van der Aalst. Discovering block-structured process models from event logs containing infrequent behaviour. In Song M. Wohed P. Lohmann, N., editor, Business Process Management Workshops. BPM 2013. Lecture Notes in Business Information Processing, volume 171. Springer, Cham., 2013.

DOI: 10.1007/978-3-319-06257-0_6

Google Scholar

[5] C. Di Francescomarino, M. Dumas, F. M. Maggi, and I. Teinemaa. Clustering-based predictive process monitoring. IEEE Transactions on Services Computing, 12(6):896–909, 2016.

DOI: 10.1109/tsc.2016.2645153

Google Scholar

[6] G. Greco, A. Guzzo, L. Pontieri, and D. Sacca. Discovering expressive process models by clustering log traces. IEEE Transactions on Knowledge and Data Engineering, 18(8):1010–1027, 2006.

DOI: 10.1109/tkde.2006.123

Google Scholar

[7] J. Evermann, J.R. Rehse, and P. Fettke. Predicting process behaviour using deep learning. Decision Support Systems, 100:129–140, 2017.

DOI: 10.1016/j.dss.2017.04.003

Google Scholar

[8] P. Pfeiffer, L. Abb, and P. Fettke. Learning from the data to predict the process. Business & Information Systems Engineering, 67:357–383, 2025.

DOI: 10.1007/s12599-025-00936-4

Google Scholar

[9] S. Weinzierl, S. Zilker, S. Dunzer, and M. Matzner. Machine learning in business process management: A systematic literature review. In Expert Systems with Applications, volume 253, 2024.

DOI: 10.1016/j.eswa.2024.124181

Google Scholar

[10] J.C.A.M. Buijs, B.F. van Dongen, and W.M.P. van der Aalst. Quality dimensions in process discovery: the importance of fitness, precision, generalization and simplicity. International Journal of Cooperative Information Systems, 23(1):1440001, 2014.

DOI: 10.1142/s0218843014400012

Google Scholar

[11] W. M. P. van der Aalst. Process mining manifesto. In Barkaoui K. Dustdar S. Daniel, F., editor, Business Process Management Workshops. BPM 2011, volume 99. Springer, Berlin, Heidelberg, 2011.

Google Scholar

[12] B. F. van Dongen. BPI Challenge 2014: Activity log for incidents. TU.ResearchData, Eindhoven University of Technology, 2014.

Google Scholar

[13] A. J. M. M. Weijters and J. Ribeiro. Flexible heuristics miner (fhm). In IEEE Symposium on Computational Intelligence and Data Mining (CIDM), pages 310–317. Paris, France, 2011.

DOI: 10.1109/cidm.2011.5949453

Google Scholar

[14] S. J. van Zelst, B. F. van Dongen, and W. M. P. van der Aalst. Ilp-based process discovery using hybrid regions. In CEUR Workshop Proceedings, volume 1371, pages 47–61, 2015.

Google Scholar

[15] J.-Y. Jung, J. Bae, and L. Liu. Hierarchical business process clustering. In IEEE International Conference on Services Computing, pages 613–616. Honolulu, HI, USA, 2008.

DOI: 10.1109/scc.2008.69

Google Scholar

[16] P. De Koninck, S. vanden Broucke, and J. De Weerdt. Act2vec, trace2vec, log2vec, and model2vec: Representation learning for business processes. In Montali M. Weber I. vom Brocke J. Weske, M., editor, Business Process Management. BPM 2018. Lecture Notes in Computer Science, volume 11080. Springer, Cham., 2018.

DOI: 10.1007/978-3-319-98648-7_18

Google Scholar

[17] A. Seeliger, S. Luettgen, T. Nolle, and M. Mühlhäuser. Learning of process representations using recurrent neural networks. In In Advanced Information Systems Engineering: 33rd International Conference, volume CAiSE 2021. Springer-Verlag, Berlin, Heidelberg, 2021.

Google Scholar

[18] J. Buijs, B. F. van Dongen, and W. M. P. van der Aalst. A genetic algorithm for discovering process trees. In IEEE Congress on Evolutionary Computation, pages 1–8. Brisbane, QLD, Australia, 2012.

DOI: 10.1109/cec.2012.6256458

Google Scholar

[19] van der Aalst W.M.P. Berti, A. A novel token-based replay technique to speed up conformance checking and process enhancement. In Kordon F. Pomello L. Koutny, M., editor, Transactions on Petri Nets and Other Models of Concurrency XV. Lecture Notes in Computer Science, volume 12530. Springer, Berlin, Heidelberg, 2021.

DOI: 10.1007/978-3-662-63079-2_1

Google Scholar

[20] A. Rozinat and W.M.P. van der Aalst. Conformance checking of processes based on monitoring real behavior. Information Systems, 33(1):64–95, 2008.

DOI: 10.1016/j.is.2007.07.001

Google Scholar

[21] W.M.P. van der Aalst, A. Adriansyah, and B. van Dongen. Replaying history on process models for conformance checking and performance analysis. In Rev. Data Min. and Knowl, volume Disc 2, pages 182–192. Wiley Int., 2012.

DOI: 10.1002/widm.1045

Google Scholar