Comparative Analysis of Data Normalization Effects on RFMBased Customer Segmentation Using K-Means and DBSCAN
DOI:
https://doi.org/10.52465/joiser.v4i2.86Keywords:
DBSCAN algorithm, K-Means algorithm, Data normalization, RFM analysis, Customer segmentationAbstract
Customer segmentation is widely used to analyze customer transaction patterns and support effective business strategies. However, previous studies have reported inconsistent findings regarding the impact of data normalization on clustering quality across different datasets and algorithms. This study investigates the effect of data normalization on RFM-based customer segmentation using K-Means and DBSCAN. Two transaction datasets, Online Retail II and TransJakarta, were analyzed under three preprocessing scenarios: no normalization, Min-Max normalization, and Z-Score normalization. Clustering performance was evaluated using the Silhouette Score and Davies–Bouldin Index (DBI). For the Online Retail II dataset, K-Means achieved the best performance without normalization (Silhouette Score = 0.9845), while DBSCAN produced valid clusters only after Z-Score normalization. For the TransJakarta dataset, both algorithms performed best without normalization, whereas DBSCAN identified up to 20 clusters and noise points. These findings highlight that the effectiveness of normalization depends on dataset characteristics and the clustering algorithm used.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information System Exploration and Research

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


