Journal of Transportation Research

Journal of Transportation Research

Clustering Intercity Road Axes Based on Traffic Volume, Speed, Fleet Composition, and Violation Patterns Using Road Traffic Count Data

Document Type : Original Article

Authors
1 M.Sc., Grad., Department of Civil Engineering, Faculty of Engineering, Ferdowsi University of Mashhad, Iran.
2 Associate Professor, Department of Civil Engineering, Faculty of Engineering, Ferdowsi University of Mashhad, Iran.
Abstract
Safety management and operation of intercity roads require understanding the behavioral differences among network corridors. Simple ranking of high-risk corridors, while useful for prioritization, fails to reveal similar behavioral patterns in terms of traffic volume, speed, fleet composition, and violation profiles. This study aims to cluster the intercity corridors of Razavi Khorasan Province based on traffic count data from the 141 system in the year 1404 (2025–2026). In this study, 136 unidirectional corridors were included in the analysis after data quality control and aggregation of corridor-level features. The features used included annual traffic volume, following-distance violation rate, speed violation rate, overtaking violation rate, weighted average speed, and the share of commercial/heavy vehicles. To reduce scaling effects, the data were standardized, and Principal Component Analysis was employed to visualize the data structure. Subsequently, two approaches—K-Means and Ward's hierarchical clustering—were evaluated for cluster counts ranging from 2 to 7. Based on the Silhouette index, k=4 with a coefficient of 0.269 provided the best K-Means configuration. The results revealed four behavioral types: "high-risk with high following-distance violations," "high-traffic with moderate behavioral risk," "commercial/heavy with controlled speed," and "lower-risk and more stable." The first cluster, comprising 47 corridors, had an average following-distance violation rate of 205.55 violations per 1,000 vehicles and a weighted rate of 244.92, placing it as the top priority for safety monitoring and intervention. The findings of this study indicate that data-driven clustering can serve as a suitable complement to risk indices and predictive models, assisting road authorities in managing similar corridors through more coordinated policies.
Keywords
Subjects

-Abduljabbar, R., et al, (2025). Machine learning traffic flow prediction models for smart cities: Review and applications. Infrastructures, , 7-10.
-Caliński, T., & Harabasz, J (2024).  A dendrite method for cluster analysis. Communications in Statistics 3(1), 1-27.
-Chai, A. B. Z., et al, (2024). Enhancing road safety with machine learning: Current state and future directions. Engineering Applications of Artificial Intelligence.
-Clustering Algorithms to Analyse Smart City Traffic Data, (2024). International Journal of Advanced Computer Science and Applications, , (15)8.
-Davies, D. L., & Bouldin, D. W. A.  (1979). cluster separation measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI 2(1), 224-227.
-Etemad, S., et al, (2023).Clustering of urban traffic patterns by K-Means and Dynamic Time Warping: Case study. arXiv: 230909830.
-Federal Highway Administration, Road Safety Fundamentals: Measuring Safety (2024).
-Federal Highway Administration (2008). Surrogate Safety Assessment Model and Validation Final Report .
 -Fredriksson, H, (2025). Exploring spatio-temporal traffic performance variation in road networks. Transportation  Research­\Preprint Report.
-Ikotun, A. M., Ezugwu, A. E., Abualigah, L., Abuhaija, B., & Heming, J, K (2023). means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big  data. Information Sciences  622, 178-220.
-International Transport Forum / OECD,  Road Safety Annual Report  (2024).
-Jain, A. K, Data clustering: 50 years beyondK-means. Pattern Recognition Letters, 31(18), 651-666.
-Jolliffe, I. TPrincipal Component Analysis. (2002). Springer,
-Kwon, Y., Lee, M., Lee, M. J., & Son, S.-W, Cluster formations of free and congested flows in urban road networks. arXiv, 240808122.
-Lloyd, S, Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2), 129-137.
-MacQueen, J, (1967).  Some methods for classification and analysis of multivariate observations. Proceedings of the Fifth Berkeley Symposium,
 -Morissette, L., & Chartier, S. . (2013). The k-means clustering technique: General considerations and implementation in Mathematica. Tutorials in Quantitative Methods for Psychology,, (9)1, 15-24.
-Pavlyshyn, V., et al, (2025). An adaptive machine learning approach to traffic pattern recognition. Smart Cities, (5)4, 151-152.
-Pedregosa, F., et al, (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, , (12), 2825-2830.
-Rousseeuw, P. J, Silhouettes, (1987).  A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics , 20-53-65.
-Sohail, A. M.  (2024). Unravelling traffic dynamics with K-means clustering and data preprocessing. Information Systems and Smart Cities.
-Ward, J. H., (2022). Hierarchical grouping to optimize an objective function. Journal of the American Statistical Association.
-World Health Organization, Global status report on road safety (2023).
-Xu, D., & Tian, Y (2015). A comprehensive survey of clustering algorithms. Annals of Data Science (2), 165-193.
-Yumak, A.  (2025). A machine learning approach to identify high-risk road segments. Applied Sciences (15)25, 1224-1225.