Journal of Information Technology Management

Journal of Information Technology Management

ConvNeXt-Tiny for High-Accuracy Natural Scene Image Classification with Test-Time Augmentation

Document Type : Research Paper

Authors
1 Department of Studies and Planning, University of Information Technology and Communications, Baghdad, Iraq.
2 Department of Scientific Affairs, University of Information Technology and Communications, Baghdad, Iraq.
3 Department of Business Information Technology, University of Information Technology and Communications, Baghdad, Iraq.
Abstract
The ConvNeXt-Tiny architecture combined with Test-Time Augmentation is used in this paper to present a strong deep learning system for multi-class natural scene image classification. Scene classification is a basic computer vision problem that has a wide range of applications, including autonomous navigation and environmental monitoring. We are using the Intel Image Classification dataset, which consists of 6 categories: buildings, forest, glacier, mountain, sea, and street. Thanks to this cutting-edge ConvNeXt-Tiny backbone, based on the transfer learning paradigm with a pre-trained ConvNeXt-Tiny backbone, a state-of-the-art CNN incorporating the design principles of Vision Transformers, we achieve very high feature extraction efficiency.TTA is also adopted so as to further enhance prediction robustness and reliability with visual changes during the inference stage. Experimental results have shown that the proposed method provides a Macro F1-score of 94.83% and an accuracy on test data of 94.70%. Based on the comparisons made herein, our method had an improved performance for all benchmarks from 2023–2025, including modified ResNet50 models and architectures for the Swin Transformers. The performance of our method is better than many available current benchmarks of 2023-2025 that include modified ResNet50 models and architectures of Swin Transformers, while maintaining the computational efficiency of the pure convolutional models.
Keywords

Adhikarla, E., Zhang, K., Yu, J., Sun, L., Nicholson, J., & Davison, B. D. (2023). Robust computer vision in an ever-changing world: A survey of techniques for tackling distribution shifts. arXiv preprint arXiv:2312.01540.
Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (2009, June). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255). Ieee.
Gupta, N., & Khobragade, P. (2023). Muti-class image classification using transfer learning. Int J Res Appl Sci Eng Technol11(1), 700-4.
Hekler, A., Brinker, T. J., & Buettner, F. (2023, June). Test time augmentation meets post-hoc calibration: uncertainty quantification under real-world conditions. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 37, No. 12, pp. 14856-14864).
Islam, N., Ahmed, M. R., Fahad, N. M., Islam, S., Islam, A. K. M., Mukta, S., & Shatabda, S. (2026). Remote Sensing Image Classification Using Deep Ensemble Learning. arXiv preprint arXiv:2603.05844.
Jiang, L., Yuan, B., Ma, W., & Wang, Y. (2023). JujubeNet: A high-precision lightweight jujube surface defect classification network with an attention mechanism. Frontiers in Plant Science, 13, 1108437.
Khan, W., & Vishwamitra, D. L. K. (2024). Image Classification using modified Convolutional Neural Network. Published in SCOPUS Journal of Electrical Systems, 20(3), 3465-3472.
Lin, C. H., Chen, T. Y., Chen, H. Y., & Chan, Y. K. (2024). Efficient and lightweight convolutional neural network architecture search methods for object classification. Pattern Recognition, 156, 110752.
Liu, S., Cao, S., Lu, X., Peng, J., Ping, L., Fan, X., ... & Liu, X. (2025). Lightweight deep learning model, ConvNeXt-U: an improved U-Net network for extracting cropland in complex landscapes from Gaofen-2 images. Sensors, 25(1), 261.
Liu, Z., Mao, H., Wu, C. Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022, June). A convnet for the 2020s. In 2022 IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp. 11966-11976). IEEE.
Maurício, J., Domingues, I., & Bernardino, J. (2023). Comparing vision transformers and convolutional neural networks for image classification: A literature review. Applied Sciences, 13(9), 5521.
Mei, S., Lian, J., Wang, X., Su, Y., Ma, M., & Chau, L. P. (2024). A comprehensive study on the robustness of deep learning-based image classification and object detection in remote sensing: Surveying and benchmarking. Journal of Remote Sensing, 4, 0219.
Musyaffa, M. S. I., Yudistira, N., Rahman, M. A., Basori, A. H., Mansur, A. B. F., & Batoro, J. (2024). IndoHerb: Indonesia medicinal plants recognition using transfer learning and deep learning. Heliyon, 10(23).
Nieradzik, L., Sieburg-Rockel, J., Helmling, S., Keuper, J., Weibel, T., Olbrich, A., & Stephani, H. (2024). Automating wood species detection and classification in microscopic images of fibrous materials with deep learning. Microscopy and Microanalysis, 30(3), 508-520.
Odbal, & Zhang, H. (2025, March). A Vision Transformer-Based Approach to Remote Sensing Scene Classification: Exploring the Impact of Dropout Regularization. In Proceedings of the 2025 6th International Conference on Computer Information and Big Data Applications (pp. 1488-1493).
Ozdemir, B., Sermet, F., & Pacal, I. (2025). Attention-enhanced ConvNeXt for accurate, efficient, and interpretable crack detection. Expert Systems with Applications, 129165.
Ozturk, E., Prabhushankar, M., & AlRegib, G. (2024, October). Intelligent multi-view test time augmentation. In 2024 IEEE International Conference on Image Processing (ICIP) (pp. 617-623). IEEE.
Rosy, N. A., Balasubadra, K., & Deepa, K. (2025). Are vision transformers replacing convolutional neural networks in scene interpretation?: A review. Discover Applied Sciences, 7(9), 932.
Sherkatghanad, Z., Abdar, M., Bakhtyari, M., Pławiak, P., & Makarenkov, V. (2025). BayTTA: Uncertainty-aware medical image classification with optimized test-time augmentation using Bayesian model averaging. Knowledge-Based Systems, 114123.
Teng, S., Zhu, L., Li, Y., Wang, X., & Jin, Q. (2023). ConvNeXt steel slag sand substitution rate detection method incorporating attention mechanism. Scientific Reports, 13(1), 10593.
Valverde, A., Coto, L. F. S., & Fernández, A. M. (2025). Back Home: A Computer Vision Solution to Seashell Identification for Ecological Restoration. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 5159-5168).
Wang, F., Li, H., Chao, W., Zhuo, Z., Ji, Y., Peng, C., & Sun, Y. (2025). E-convNeXt: A lightweight and efficient convNeXt variant with cross-stage partial connections. arXiv preprint arXiv:2508.20955.
Wu, L., & Wang, H. (2023). Global and pyramid convolutional neural network with hybrid attention mechanism for hyperspectral image classification. Geocarto International, 38(1), 2226112.
Xia, J., Yin, Y., & Li, X. (2025). An efficient medical image classification method based on a lightweight improved ConvNeXt-Tiny architecture. arXiv preprint arXiv:2508.11532.
Zhang, Z., Eli, E., Mamat, H., Aysa, A., & Ubul, K. (2023). EA-ConvNeXt: An approach to script identification in natural scenes based on edge flow and coordinate attention. Electronics, 12(13), 2837.
Zheng, F., Lin, S., Zhou, W., & Huang, H. (2023). A lightweight dual-branch Swin Transformer for remote sensing scene classification. Remote Sensing, 15(11), 2865.