Improving Image Classification through Active Learning and Pseudo-Labeling in Vision-Language Models
DOI:
https://doi.org/10.22456/2175-2745.150941Keywords:
visual language models, active learning, pseudo-labelingAbstract
Visual Language Models (VLMs) combine natural language processing and computer vision to interpret multimodal data, such as images and text, showing great potential in image classification applications. This paper investigates the integration of Active Learning (AL) and pseudo-labeling techniques with VLMs to improve image classification in various domains. To achieve this, five AL strategies (Random Sampling, Uncertainty Sampling, Margin Sampling, Entropy Sampling, and Query-by-Committee) and three pseudo-labeling approaches (Direct, Confidence Threshold, and Feature Similarity) were evaluated iteratively. The results demonstrate that the combination of active learning and pseudo-labeling can achieve promising results, in addition to full class coverage in a few iterations. We conclude that the integration of AL with feature similarity-based pseudo-labeling offers a robust and efficient solution for image classification in limited-data scenarios, promoting high accuracy, class representativeness, and the reduction of propagation errors, with potential for applications in critical domains like healthcare and industry.
Downloads
References
[1] RADFORD, A. et al. Learning transferable visual models from natural language supervision. International Conference on Machine Learning (ICML), PMLR, p. 8748–8763, 2021.
[2] ESTEVA, A. et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature, Nature Publishing Group, v. 542, n. 7639, p. 115–118, 2017.
[3] WINDSOR, R. et al. Vision-language modelling for radiological imaging and reports in the low data regime. In: Proceedings of Machine Learning Research. [S.l.: s.n.], 2023. v. 227, p. 53–73.
[4] NGUYEN, D. M. H. et al. LVM-Med: Learning large-scale self-supervised vision models for medical imaging via second-order graph matching. In: Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023. [S.l.: s.n.], 2023.
[5] ZHU, X. Semi-supervised learning literature survey. Computer Sciences Technical Report 1530, University of Wisconsin-Madison, 2005.
[6] LEE, D.-H. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. ICML Workshop on Challenges in Representation Learning, 2013.
[7] COHN, D. A.; GHAHRAMANI, Z.; JORDAN, M. I. Active learning with statistical models. Journal of Artificial Intelligence Research, v. 4, p. 129–145, 1996.
[8] ZHANG, J. et al. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, v. 46, p. 5625–5644, 2023.
[9] BANG, J.; AHN, S.; LEE, J.-G. Active prompt learning in vision language models. In: International Conference on Learning Representations. [S.l.: s.n.], 2023.
[10] LI, L. H. et al. VisualBERT: A simple and performant baseline for vision and language. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020. p. 1492–1502.
[11] LI, G. et al. Unicoder-VL: A universal encoder for vision and language by cross-modal pre-training. In: Proceedings of the AAAI Conference on Artificial Intelligence. [S.l.: s.n.], 2020. v. 34, n. 07, p. 11316–11323.
[12] CHEN, Y.-C. et al. UNITER: Universal image-text representation learning. In: Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX. Berlin, Heidelberg: Springer-Verlag, 2020. p. 104–120. ISBN 978-3-030-58576-1.
[13] LU, J. et al. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In: Advances in Neural Information Processing Systems. [S.l.: s.n.], 2019. p. 13–23.
[14] TAN, H. H.; BANSAL, M. LXMERT: Learning cross-modality encoder representations from transformers. In: Conference on Empirical Methods in Natural Language Processing. [S.l.: s.n.], 2019.
[15] LECUN, Y. et al. Gradient-based learning applied to document recognition. Proceedings of the IEEE, v. 86, n. 11, p. 2278–2324, 1998.
[16] KRIZHEVSKY, A.; HINTON, G. Learning multiple layers of features from tiny images. [S.l.], 2009.
[17] XIAO, H.; RASUL, K.; VOLLGRAF, R. Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms. In: ICLR 2017 Workshop on Challenges in Representation Learning. [S.l.: s.n.], 2017.
[18] HELBER, P. et al. EuroSAT: A novel dataset for deep learning in remote sensing. IEEE Transactions on Geoscience and Remote Sensing, v. 57, n. 11, p. 9414–9426, 2019.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Willian P. Amorim, Santos, C. A. N., Saito, P. T. M.

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.













