Environmental Sound Classification Using a Hybrid SVM-CNN Approach

Main Article Content

Nisreen Talib Abdulhusein
Basheera M. Mahammod

Abstract

Classifying environmental sounds is considered a challenging task because of the complex nature of acoustic signals. This classification involves identifying the acoustic flow associated with different sounds in the environment. Several factors affect it, such as the distances between the source and the destination, voice interference, and the huge number of sound sources. These problems have been addressed in this work, where a hybrid deep learning technique is proposed to improve the classification of sounds, focusing on fundamental audio features which are accurate in these tasks. The proposed hybrid classification framework combines a Support Vector Machine (SVM) and a Convolutional Neural Network (CNN). The SVM model uses feature extraction based on the orthogonal transformation (DCT), and the CNN model uses machine learning to represent data from Log-Mel spectral diagrams. Noisy sound signals are generated using the TIMIT and NOISEX-92 datasets by combining the clean signal with various types of environmental sounds, resulting in 17 categories, including the clean signal category. A soft voting strategy is then used to enhance overall decision-making. It calculates the weighted average probability for each class across all models and selects the class that has the highest average probability as the final output with more robust and reliable classification. The experimental results prove that the proposed model achieves high performance compared to current ways and individual models, reaching an accuracy of 99.28% with improvements in F1 score, accuracy and recall. The results obtained prove the efficiency of integrating ML and DL methods for classifying environmental sounds.

Downloads

Download data is not yet available.

Article Details

Section

Articles

How to Cite

“Environmental Sound Classification Using a Hybrid SVM-CNN Approach” (2026) Journal of Engineering, 32(8), pp. 232–255. doi:10.31026/j.eng.2026.08.11.

References

Abdoune, L., Fezari, M. and Dib, A. 2024. Indoor sound classification with support vector machines: State of the art and experimentation. International Journal of Computational Methods and Experimental Measurements, 12(3), pp. 269–279. https://doi.org/10.18280/ijcmem.120307.

Aji, NB., Kurnianingsih, K., Masuyama, N., Nojima, Y. 2024. CNN-LSTM for heartbeat sound classification. JOIV: International Journal on Informatics Visualization, 8(2), pp. 735–741. https://dx.doi.org/10.62527/joiv.8.2.2115.

Akter, R., Islam, MR., Debnath, SK., Sarker, PK., Uddin, MK. 2025. A hybrid CNN-LSTM model for environmental sound classification: Leveraging feature engineering and transfer learning. Digital Signal Processing, 163, P. 105234. https://doi.org/10.1016/j.dsp.2025.105234.

Albaji, A.O., Rashid, R.B.A. and Abdul Hamid, S.Z. 2023. Investigation on machine learning approaches for environmental noise classifications. Journal of Electrical and Computer Engineering, 2023(1), P. 3615137. https://doi.org/10.1155/2023/3615137.

Ali, Y.H., Rashid, R.A. and Hamid, S.Z.A. 2022. A machine learning for environmental noise classification in smart cities. Indonesian Journal of Electrical Engineering and Computer Science, 25(3), pp. 1777–1786. https://doi.org/10.11591/ijeecs.v25.i3.pp1777-1786.

Alsouda, Y., Pllana, S. and Kurti, A. 2019. IoT-based urban noise identification using machine learning: performance of SVM, KNN, bagging, and random forest. In Proceedings of the international conference on omni-layer intelligent systems, pp. 62–67. https://doi.org/10.1145/3312614.3312631.

Bansal, A. and Garg, N.K. 2022. Environmental sound classification: A descriptive review of the literature. Intelligent systems with applications, 16, P. 200115. https://doi.org/10.1016/j.iswa.2022.200115.

Bansal, A. and Garg, N.K. 2023. Environmental sound classification using hybrid ensemble model. Procedia Computer Science, 218, pp. 418–428. https://doi.org/10.1016/j.procs.2023.01.024.

Beritelli, L., Borzì, MG., Randieri, C., Avanzato, R., Beritelli, F. 2023. Techniques for recognising and classifying environmental noise using deep learning. In SYSTEM, pp. 62–67.

Blanco, V., Japón, A. and Puerto, J. 2022. A mathematical programming approach to SVM-based classification with label noise. Computers & Industrial Engineering, 172, P. 108611. https://doi.org/10.1016/j.cie.2022.108611.

Chan, C.-F. and Eric, W.M. 2010. An abnormal sound detection and classification system for surveillance applications. In 2010 18th European Signal Processing Conference, pp. 1851–1855.

Chen, F., Zhu, Z., C Sun, and Xia, L. 2025. Evaluating metric and contrastive learning in pretrained models for environmental sound classification. Applied Acoustics, 232, P. 110593. https://doi.org/10.1016/j.apacoust.2025.110593.

Chen, G., Zhang, B., Ding, Z., Xiao, K., Guan, P., Xiao, X., Wang, X., Yi, H., Hu, H., and Zhang, W. 2025. A lightweight dual branch masking network for environmental sound classification. Scientific Reports [Preprint]. https://doi.org/10.1038/s41598-025-33636-w.

Chu, H.-C., Zhang, Y.-L. and Chiang, H.-C. 2023. A CNN sound classification mechanism using data augmentation. Sensors, 23(15), P. 6972. https://doi.org/10.3390/s23156972.

Demir, F., Abdullah, D.A. and Sengur, A. 2020. A new deep CNN model for environmental sound classification. IEEE Access, 8, pp. 66529–66537. https://doi.org/10.1109/ACCESS.2020.2984903.

Goulão, M., Bandeira, L., Martins, B., and Oliveira, A L. 2024. Training environmental sound classification models for real-world deployment in edge devices. Discover Applied Sciences, 6(4), P. 166. https://doi.org/10.1007/s42452-024-05803-7.

Hasan, N.I. 2022. Bird species classification and acoustic features selection based on distributed neural network with two stage windowing of short-term features. arXiv preprint, arXiv:2201.00124 [Preprint]. https://doi.org/10.48550/arXiv.2201.00124.

Hasnain, FU., Iqbal, V., Javed, T., and Yasir, M. 2025. Adaptive Boosted Support Vector Machine-random Forest for Environmental Sound Classification. Journal of Computing & Biomedical Informatics, 9(02). https://doi.org/10.56979/902/2025.

Hu, R., Hu, K., Wang, L., Guan, Z., Zhou, X., Wang, N., and Ye, L. 2024. Using deep learning to classify environmental sounds in the habitat of western black-crested gibbons. Diversity, 16(8), P. 509. https://doi.org/10.3390/d16080509.

Hussein, HA., Hameed, SH., Mahmmod, B.M., Abdulhussain, S.H., Hussain, A.J. 2023. Dual stages of speech enhancement algorithm based on super gaussian speech models, Journal of Engineering, 29(09), pp. 1–13. https://doi.org/10.31026/j.eng.2023.09.01.

Ibraheem, Z.H. and Shihab, A.I. 2024. Speech isolation and recognition in crowded noise using a dual-path recurrent neural network. Iraqi Journal of Science, pp. 5784–5797. https://doi.org/10.24996/ijs.2024.65.10.37.

Ilahi, A.H.Z., Irwansyah, A. and Oktavianto, H. 2025. Comparative study of CNN architectures for real-time audio-based car accident detection on edge devices. JOIV: International Journal on Informatics Visualization, 9(3), pp. 1310–1318. https://dx.doi.org/10.62527/joiv.9.3.2985.

Iqbal, F., Abbasi, A., Almadhor, A., Alsubai, S., and Gregus, M. 2025. Real-time active-learning method for audio-based anomalous event identification and rare events classification for audio events detection. Frontiers in Computer Science, 7, P. 1517346. https://doi.org/10.3389/fcomp.2025.1517346.

Islam, M. and Ali, M.N.Y. 2024. Environmental sound classification using feature fusion of mfccs, mel-spectrogram, and chroma. In 2024 27th International Conference on Computer and Information Technology (ICCIT), pp. 3212–3217. https://doi.org/10.1109/ICCIT64611.2024.11021738.

Khayam, S.A. 2003 The discrete cosine transform (DCT): Theory and application. Michigan State University, 114(1), P. 31.

Malhotra, A. 2023. Noise identification in urban cities using performance of Random Forest, KNN, and SVM. In 2023 International Conference on Innovative Computing, Intelligent Communication and Smart Electrical Systems (ICSES), pp. 1–5. https://doi.org/10.1109/ICSES60034.2023.10465454.

Meedeniya, D., Ariyarathne, I., Bandara, M., Jayasundara, R., and Perera, C. 2023. A survey on deep learning based forest environment sound classification at the edge. ACM Computing Surveys, 56(3), pp. 1–36. https://doi.org/10.1145/3618104.

Nasir, R.J. and Abdulmohsin, H.A. 2024. Noise reduction techniques for enhancing speech. Iraqi Journal of Science, pp. 5798–5818. https://doi.org/10.24996/ijs.2024.65.10.38.

Nogueira, A.F.R., Oliveira, H.S., Machado, J.J.M., and Tavares, J.M.R.S. 2022. Sound classification and processing of urban environments: A systematic literature review. Sensors, 22(22), P. 8608. https://doi.org/10.3390/s22228608.

Piczak, K.J. 2015. Environmental sound classification with convolutional neural networks. In 2015 IEEE 25th international workshop on machine learning for signal processing (MLSP), pp. 1–6. https://doi.org/10.1109/MLSP.2015.7324337.

Sankavi, K., Jyothish Lal, G. and Premjith, B. 2024. Deep learning based automatic noisy speech classification for enhanced speech analysis. In 2024 IEEE 3rd World Conference on Applied Intelligence and Computing (AIC), pp. 1166–1171. https://doi.org/10.1109/AIC61668.2024.10730980.

Souadek, R. and Guellil, N., 2025. Environmental sound classification based on STFT features and Convolutional Neural Networks (CNNs). Przegląd Elektrotechniczny, 2025(7). https://doi.org/10.15199/48.2025.07.25.

Walden, F., Dasgupta, S., Rahman, M., and Islam, M. 2022. Improving the environmental perception of autonomous vehicles using deep learning-based audio classification. arXiv preprint arXiv:2209.04075 [Preprint]. https://doi.org/10.48550/arXiv.2209.04075.

Yassen, M.T. and Abdulhussain, S.H. 2025. Discrete cosine transform-driven hybrid orthogonal polynomials: Design and applications in signal processing.

Yousif, S.T., Mahmmod, B.M. 2026. A dual-stage perceptual-harmonic hybrid estimator for speech enhancement. Journal of Engineering, 32(3), pp. 173–192. https://doi.org/10.31026/j.eng.2026.03.10.

Zaman, K., Sah, M., Direkoglu, C., and Unoki, M. 2023. A survey of audio classification using deep learning. IEEE access, 11, pp. 106620–106649. https://doi.org/10.1109/ACCESS.2023.3318015.

Zhang, Y., Huang, Q., Sun, W., Chen, F., Lin, D., and Chen, F. 2024. Research on lung sound classification model based on dual-channel CNN-LSTM algorithm. Biomedical Signal Processing and Control, 94, P. 106257. https://doi.org/10.1016/j.bspc.2024.106257.

Similar Articles

You may also start an advanced similarity search for this article.