Synergizing CNNs and ViTs: A Survey on Hybrid Models for Benchmark Image Classification
DOI:
https://doi.org/10.22456/2175-2745.147347Keywords:
Image Classification, Convolutional Neural Networks (CNNs), Vision Transformers (Vits), Hybrid Deep Learning Models, Explainable AI (XAI), Representation LearningAbstract
The rapid expansion of online fashion retail has led to a massive increase in product images, making accurate clothing classification a crucial task. Misclassification not only disrupts customer satisfaction but also contributes to higher return rates and operational inefficiencies. Effective apparel classification enhances search results, improves personalized recommendations, and optimizes inventory management. In response to these challenges, deep learning models have gained attention for visual recognition tasks. This review focuses on the capabilities of convolutional neural networks (CNNs), vision transformers (ViTs), and hybrid models combining both, with an emphasis on the Fashion MNIST benchmark. CNNs are widely used for extracting local spatial features, while ViTs excel at capturing global dependencies through self-attention mechanisms. Recent works have proposed hybrid architectures to leverage the strengths of both models, combining fine-grained feature extraction with broader contextual understanding. This paper offers a comprehensive survey of these hybrid approaches, evaluating their performance in terms of accuracy, scalability, and adaptability. Additionally, we highlight the importance of explainability in fashion classification models and discuss how hybrid solutions enhance model interpretability. Comparative findings suggest that hybrid models outperform individual architectures, making them highly effective for real-world fashion applications demanding both transparency and robust performance.Downloads
References
[1] SHEIKH, A.-S. et al. A deep learning system for predicting size and fit in fashion e-commerce. In: Proceedings of the 13th ACM conference on recommender systems. [S.l.: s.n.], 2019. p. 110–118.
[2] RANE, N. L. et al. Artificial intelligence, machine learning, and deep learning for advanced business strategies: a review. Partners Universal International Innovation Journal, v. 2, n. 3, p. 147–171, 2024.
[3] WANG, S.; QIU, J. A deep neural network model for fashion collocation recommendation using side information in e-commerce. Applied Soft Computing, Elsevier, v. 110, p. 107753, 2021.
[4] JIN, L.; CHEN, L. Exploring the impact of computer applications on cross-border e-commerce performance. IEEE Access, IEEE, 2024.
[5] DASH. E-commerce Statistics Report. 2025. ⟨https://dash.app/blog/ecommerce-statistics⟩. Accessed: 2023-06-18.
[6] WIZISHOP, S. Starting an E-commerce Business: Key Figures and Trends. 2025. ⟨https://www.statista.com/topics/871/online-shopping/#topicOverview⟩. Accessed: 2025-02-22.
[7] UNIFORMMARKET. Global Apparel Industry Statistics. 2025. ⟨https://www.uniformmarket.com/statistics/global-apparel-industry-statistics⟩. Accessed: 2025-08-17.
[8] DARGAN, S. et al. A survey of deep learning and its applications: a new paradigm to machine learning. Archives of computational methods in engineering, Springer, v. 27, p. 1071–1092, 2020.
[9] VIJAYARAJ, A. et al. Deep learning image classification for fashion design. Wireless Communications and Mobile Computing, Wiley Online Library, v. 2022, n. 1, p. 7549397, 2022.
[10] KIM, J.; FORSYTHE, S. Adoption of virtual try-on technology for online apparel shopping. Journal of interactive marketing, Elsevier, v. 22, n. 2, p. 45–59, 2008.
[11] BOUZIDI, S. et al. Convolutional neural networks and vision transformers for fashion mnist classification: A literature review. arXiv preprint arXiv:2406.03478, 2024.
[12] HCINI, G.; JDEY, I.; LTIFI, H. HSV-Net: a custom CNN for malaria detection with enhanced color representation. In: 2023 International Conference on Cyberworlds (CW). [S.l.]: IEEE, 2023. p. 337–340.
[13] HCINI, G. et al. Hyperparameter optimization in customized convolutional neural network for blood cells classification. Journal of Theoretical and Applied Information Technology, v. 99, p. 5425–5435, 2021.
[14] SLIMANI, N.; JDEY, I.; KHERALLAH, M. Performance comparison of machine learning methods based on CNN for satellite imagery classification. In: 2023 9th International Conference on Control, Decision and Information Technologies (CoDIT). [S.l.]: IEEE, 2023. p. 185–189.
[15] CONG, S.; ZHOU, Y. A review of convolutional neural network architectures and their optimizations. Artificial Intelligence Review, Springer, v. 56, n. 3, p. 1905–1969, 2023.
[16] BRAHMI, W.; JDEY, I. Automatic tooth instance segmentation and identification from panoramic x-ray images using deep CNN. Multimedia Tools and Applications, Springer, v. 83, n. 18, p. 55565–55585, 2024.
[17] JAMIL, S.; PIRAN, M. J.; KWON, O.-J. A comprehensive survey of transformers for computer vision. Drones, MDPI, v. 7, n. 5, p. 287, 2023.
[18] BOUZIDI, S.; JDEY, I.; ALIMI, A. A vision transformer approach with L2 regularization for sustainable fashion classification. 2024. Available at SSRN 4686032. ⟨https://ssrn.com/abstract=4686032⟩.
[19] KHANDAY, O. M.; DADVANDIPOUR, S.; LONE, M. A. Effect of filter sizes on image classification in CNN: A case study on CIFAR-10 and Fashion-MNIST datasets. International Journal of Artificial Intelligence, v. 2252, n. 8938, p. 8938, 2021.
[20] BOUZIDI, S.; JDEY, I.; DRIRA, F. XG-ViT: Explainable and generalizable vision transformer for benchmark image classification. In: International Conference on Verification and Evaluation of Computer and Communication Systems (VECoS). [S.l.: s.n.], 2025. p. 7–11.
[21] DOSOVITSKIY, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
[22] WANG, Y. et al. Vision transformers for image classification: A comparative survey. Technologies, MDPI, v. 13, n. 1, p. 32, 2025.
[23] KHAN, S. et al. Transformers in vision: A survey. ACM Computing Surveys (CSUR), New York, NY: ACM, v. 54, n. 10s, p. 1–41, 2022.
[24] ALADHADH, S. et al. An effective skin cancer classification mechanism via medical vision transformer. Sensors, MDPI, v. 22, n. 11, p. 4008, 2022.
[25] BANSAL, A. et al. Enhancing fashion cloth image classification through hybrid CNN-SVM modeling: A multi-class study. In: 2023 International Conference on Sustainable Computing and Smart Systems (ICSCSS). [S.l.]: IEEE, 2023. p. 484–489.
[26] LONG, H. Hybrid design of CNN and vision transformer: A review. In: Proceedings of the 2024 7th International Conference on Computer Information Science and Artificial Intelligence. [S.l.: s.n.], 2024. p. 121–127.
[27] PENG, Z. et al. Conformer: Local features coupling global representations for visual recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. [S.l.: s.n.], 2021. p. 367–376.
[28] DAI, Z. et al. Coatnet: Marrying convolution and attention for all data sizes. Advances in Neural Information Processing Systems, v. 34, p. 3965–3977, 2021.
[29] YUNUSA, H. et al. Exploring the synergies of hybrid CNNs and ViTs architectures for computer vision: A survey. arXiv preprint arXiv:2402.02941, 2024.
[30] HASSANI, A.; SHI, H. Dilated neighborhood attention transformer. arXiv preprint arXiv:2209.15001, 2022.
[31] XIN, J. et al. Convolutional neural network for fashion images classification (Fashion-MNIST). Journal of Applied Technology and Innovation, v. 7, n. 4, p. 11, 2023.
[32] AN, H. et al. Conceptual framework of hybrid style in fashion image datasets for machine learning. Fashion and Textiles, Springer, v. 10, n. 1, p. 18, 2023.
[33] RAO, J. et al. Watch and buy: A practical solution for real-time fashion product identification in live stream. In: Proceedings of the 1st Workshop on Multimodal Product Identification in Livestreaming and WAB Challenge. [S.l.: s.n.], 2021. p. 23–31.
[34] WU, H. et al. Fashion iq: A new dataset towards retrieving images by natural language feedback. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [S.l.: s.n.], 2021. p. 11307–11317.
[35] HU, S. X. et al. Pushing the limits of simple pipelines for few-shot learning: External data and fine-tuning make a difference. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [S.l.: s.n.], 2022. p. 9068–9077.
[36] TZIKAS, T.-R. et al. Towards fashion image annotation: A clothing category recognition procedure. In: Proceedings of the 11th Hellenic Conference on Artificial Intelligence (SETN 2020), Workshops. [S.l.]: CEUR Workshop Proceedings, 2020. p. 42–47.
[37] BOUZIDI, S.; JDEY, I.; DRIRA, F. Towards Explainable Skin Cancer Diagnosis: A Vision Transformer Approach with Grad-CAM Visualization. 2025. p. 1–8.
[38] PANG, S. et al. An efficient style virtual try on network for clothing business industry. arXiv preprint arXiv:2105.13183, 2021.
[39] MENG, X.; CHEN, W.; YANG, B. NEAT: Learning neural implicit surfaces with arbitrary topologies from multi-view images. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [S.l.: s.n.], 2023. p. 248–258.
[40] BHATT, D. et al. CNN variants for computer vision: History, architecture, application, challenges and future scope. Electronics, MDPI, v. 10, n. 20, p. 2470, 2021.
[41] RATHORE, B. Cloaked in code: AI & machine learning advancements in fashion marketing. Development, v. 6, n. 2, 2017.
[42] ARAÚJO, M. A. P. de. Zalando: An Equity Research Analysis of the E-Commerce Giant. Dissertação (Mestrado) — Universidade NOVA de Lisboa (Portugal), 2023.
[43] XIAO, H.; RASUL, K.; VOLLGRAF, R. Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
[44] TALAVERA, G. L. Image classification using embedded spaces generated by siamese networks. 2019.
[45] KEELE, S. et al. Guidelines for performing systematic literature reviews in software engineering. [S.l.], 2007.
[46] LEITHARDT, V. Classifying garments from Fashion-MNIST dataset through CNNs. Advances in Science, Technology and Engineering Systems Journal, v. 6, n. 1, p. 989–994, 2021.
[47] NOCENTINI, O. et al. Image classification using multiple convolutional neural networks on the Fashion-MNIST dataset. Sensors, MDPI, v. 22, n. 23, p. 9544, 2022.
*[48] ERKOÇ, T.; ESKIL, M. T. A novel similarity based unsupervised technique for training convolutional filters. IEEE Access, IEEE, v. 11, p. 49393–49408, 2023.
[49] SWAIN, D. et al. An intelligent fashion object classification using CNN. EAI Endorsed Transactions on Industrial Networks and Intelligent Systems, v. 10, n. 4, p. e2, 2023.
[50] YU, F. et al. FFENet: frequency-spatial feature enhancement network for clothing classification. PeerJ Computer Science, PeerJ Inc., v. 9, p. e1555, 2023.
[51] WAN, G.; YAO, L. LMFRNet: A lightweight convolutional neural network model for image analysis. Electronics, MDPI, v. 13, n. 1, p. 129, 2023.
[52] SUN, Y. et al. MADPL-Net: Multi-layer attention dictionary pair learning network for image classification. Journal of Visual Communication and Image Representation, Elsevier, v. 90, p. 103728, 2023.
[53] SHIN, S.-Y.; JO, G.; WANG, G. A novel method for fashion clothing image classification based on deep learning. Journal of Information and Communication Technology, v. 22, n. 1, p. 127–148, 2023.
[54] ÇETINER, H.; METLEK, S. CNNTuner: Image classification with a novel CNN model optimized hyperparameters. Bitlis Eren Üniversitesi Fen Bilimleri Dergisi, Bitlis Eren University, v. 12, n. 3, p. 746–763, 2023.
[55] HAJI, L. M. et al. Enhanced convolutional neural network for fashion classification. Engineering, Technology & Applied Science Research, v. 14, n. 5, p. 16534–16538, 2024.
[56] VENKATARAVANAPPA, V. et al. Conquering Fashion-MNIST with CNNs using computer vision by pretrained models: VGG19 and ResNet50. In: AIP Conference Proceedings. [S.l.]: AIP Publishing, 2024. v. 3131, n. 1.
[57] CHHABRA, S.; VENKATESWARA, H.; LI, B. Patchswap: A regularization technique for vision transformers. In: BMVC. [S.l.: s.n.], 2022. p. 996.
[58] ALAZIZ, H. M. A. et al. Enhancing fashion classification with vision transformer (ViT) and developing recommendation fashion systems using DINOv2. Electronics, MDPI, v. 12, n. 20, p. 4263, 2023.
[59] RODRIGUEZ, D.; KRISHNAN, R. Learnable image transformations for privacy enhanced deep neural networks. In: 2023 5th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems and Applications (TPS-ISA). [S.l.]: IEEE, 2023. p. 64–73.
[60] SHAH, S. M. A. H. et al. A hybrid neuro-fuzzy approach for heterogeneous patch encoding in ViTs using contrastive embeddings and deep knowledge dispersion. IEEE Access, IEEE, v. 11, p. 83171–83186, 2023.
[61] XU, C. et al. Transformer in optronic neural networks for image classification. Optics & Laser Technology, Elsevier, v. 165, p. 109627, 2023.
[62] LIN, Z. et al. RViT: Robust fusion vision transformer with variational hierarchical denoising process for image classification. Guidance, Navigation and Control, World Scientific, v. 4, n. 03, p. 2441007, 2024.
[63] SHAO, R.; BI, X.-J. Transformers meet small datasets. IEEE Access, IEEE, v. 10, p. 118454–118464, 2022.
[64] COOLS, A.; MAHMOUDI, S. A.; BELARBI, M. A. Carenet: A novel architecture for low data regime mixing convolutions and attention. In: Proceedings of the 20th International Conference on Content-Based Multimedia Indexing (CBMI). [S.l.]: IEEE, 2023. p. 1–6.
[65] MENG, Y. et al. MixMobileNet: A mixed mobile network for edge vision applications. Electronics, MDPI, v. 13, n. 3, p. 519, 2024.
[66] XU, C. et al. HSViT: Horizontally Scalable Vision Transformer. arXiv preprint arXiv:2404.05196, 2024. ⟨https://arxiv.org/abs/2404.05196⟩. Accessed: August 2025.
[67] YUAN, M.; XU, C. A novel approach for clothing classification: Integrating CNN and transformer in RST-Net. In: 2024 5th International Seminar on Artificial Intelligence, Networking and Information Technology (AINIT). [S.l.]: IEEE, 2024. p. 1652–1656.
[68] SIKDER, A. S.; ROLFE, S. The power of e-commerce in the global trade industry: A realistic approach to expedite virtual market place and online shopping from anywhere in the world: E-commerce in the global trade industry. International Journal of Imminent Science & Technology, v. 1, n. 1, p. 79–100, 2023.
[69] PACAL, I. et al. A novel CNN-ViT-based deep learning model for early skin cancer diagnosis. Biomedical Signal Processing and Control, Elsevier, v. 104, p. 107627, 2025.
[70] ABUALKEBASH, H.; SALEH, R. A.; ERTUNC, H. M. Automated explainable deep learning framework for multiclass skin cancer detection and classification using hybrid YOLOv8 and vision transformer (ViT). Biomedical Signal Processing and Control, Elsevier, v. 108, p. 107934, 2025.
[71] LU, X.; SUGANUMA, M.; OKATANI, T. SBCFormer: lightweight network capable of full-size ImageNet classification at 1 fps on single board computers. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. [S.l.: s.n.], 2024. p. 1123–1133.
[72] GOU, Q.; REN, Y. Research on multi-scale CNN and transformer-based multi-level multi-classification method for images. IEEE Access, IEEE, 2024.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Sonia Bouzidi, Ghazala Hcini, Imen Jdey, Fadoua Drira

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.Accepted 2025-09-23
Published 2026-01-30













