Two-Stage Fine-Tuning of Object Detectors for Industrial Environments

Authors

DOI:

https://doi.org/10.22456/2175-2745.150904

Keywords:

industrial object detection, automated quality control, two-stage fine-tuning, computer vision

Abstract

Automated quality control in industrial production has the potential to reduce errors and provide real-time information. However, the inspection of secondary packaging, such as counting boxes in crates, still represents a challenge, since manual methods are slow and error-prone, while automatic methods are limited by the scarcity of domain-specific datasets and the high cost of annotation. This paper proposes an efficient and low-cost two-stage object detection workflow for automatic box counting. The central novelty lies in the integration of the Segment Anything Model (SAM) to accelerate the creation of a high-quality dataset from production line videos, making model specialization economically feasible. Initially, a generalist model is trained on heterogeneous datasets to capture general visual features. Then, the model is fine-tuned with the domain-specific dataset of approximately 750 images. Experiments with YOLOv11x and Faster R-CNN achieved mAP@0.5 above 98%, with YOLOv11x showing higher accuracy and faster inference. These results demonstrate the efficiency of the proposed approach, establishing it as a replicable and low-cost solution for monitoring secondary packaging in industrial environments.

Downloads

Download data is not yet available.

Author Biographies

Michel M. Silva, Universidade Federal de Viçosa (UFV)

Michel Silva is currently an Assistant Professor at Universidade Federal de Viçosa (UFV) in the Department of Informatics (DPI), and a researcher at Vision and Robotics Laboratory (VeRLab) in the Universidade Federal de Minas Gerais (UFMG), Brazil, from where he earned the Ph.D. degree in the area of computer vision focusing in Fast-Forward of Egocentric Videos emphasizing semantic information. He holds M.S. and B.S. degrees in Computer Science from the Universidade Federal de Lavras (UFLA). During the bachelor, he was a visiting student at James Hutton Institute (JHI), Scotland UK. His topics of interest include computer vision, egocentric videos, semantic analysis, and three-dimensional reconstruction.

Thiago L. Gomes, Universidade Federal de Viçosa (UFV)

Thiago L. Gomes is currently an Assistant Professor at Universidade Federal de Viçosa (UFV) in the Department of Informatics (DPI). He received his B.sC. and M.sC. degrees in Computer Science from Universidade Federal de Vicosa, Brasil. He received his Ph.D. degree in Computer Science from Universidade Federal de Minas Gerais (UFMG), Brazil. His research interests focus on computer graphics, computer vision, human motion analysis, and the applications of computer vision to the visual arts.

References

[1] GHOBAKHLOO, M. Industry 4.0, digitization, and opportunities for sustainability. Journal of Cleaner Production, v. 252, p. 119869, 2020. ISSN 0959-6526. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S0959652619347390⟩.

[2] ZRIBI, M.; PAGLIUCA, P.; PITOLLI, F. A computer vision-based quality assessment technique for the automatic control of consumables for analytical laboratories. Expert Systems with Applications, v. 256, p. 124892, 2024. ISSN 0957-4174. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S0957417424017597⟩.

[3] KHAN, M. G. et al. Perfcam: Digital twinning for production lines using 3d gaussian splatting and vision models. IEEE Access, Institute of Electrical and Electronics Engineers (IEEE), v. 13, p. 85390–85406, 2025. ISSN 2169-3536. Disponível em: ⟨http://dx.doi.org/10.1109/ACCESS.2025.3567702⟩.

[4] A Review of Vision Based Defect Detection Using Image Processing Techniques for Beverage Manufacturing Industry. Jurnal Teknologi (Sciences & Engineering), v. 81, n. 3, 2019. Disponível em: ⟨https://doi.org/10.11113/jt.v81.12505⟩.

[5] KIRILLOV, A. et al. Segment Anything. 2023. ArXiv preprint arXiv:2304.02643; Accessed: 2025-08-21. Disponível em: ⟨https://segment-anything.com⟩.

[6] JOCHER, G.; QIU, J. Ultralytics YOLO11. 2024. Disponível em: ⟨https://github.com/ultralytics/ultralytics⟩.

[7] REN, S. et al. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. 2016. Disponível em: ⟨https://arxiv.org/abs/1506.01497⟩.

[8] DWYER, B. et al. Roboflow (Version 1.0) [Software]. 2024. Computer vision. Disponível em: ⟨https://roboflow.com⟩.

[9] ANGELOLMG. Open Source Dataset, milk-detector-v2 Dataset. Roboflow, 2022. Accessed: 2025-07-03. Disponível em: ⟨https://universe.roboflow.com/angelolmg/milk-detector-v2⟩.

[10] 1, S. test. Open Source Dataset, sort the trash V4 Dataset. Roboflow, 2023. Accessed: 2025-07-04. Disponível em: ⟨https://universe.roboflow.com/sampahsem5-test-1/sort-the-trash-v4⟩.

[11] JAHANGIR, S. Open Source Dataset, Slice Juice Dataset. Roboflow, 2024. Accessed: 2025-07-04. Disponível em: ⟨https://universe.roboflow.com/sarmad-jahangir-ph67v/slice-juice⟩.

[12] LONKYKONG. Open Source Dataset, juice-detection-4-items Dataset. Roboflow, 2023. Accessed: 2025-07-04. Disponível em: ⟨https://universe.roboflow.com/lonkykong/juice-detection-4-items⟩.

[13] LEARNING, M. Open Source Dataset, Grocery Items Dataset. Roboflow, 2024. Accessed: 2025-07-04. Disponível em: ⟨https://universe.roboflow.com/machine-learning-chipg/grocery-items-re4qf⟩.

[14] University. Open Source Dataset, RetinaNet Dataset. Roboflow, 2024. Accessed: 2025-07-04. Disponível em: ⟨https://universe.roboflow.com/university-d2l2p/retinanet-besb5⟩.

[15] YANG, J. et al. Scd: A stacked carton dataset for detection and segmentation. Sensors, v. 22, n. 10, 2022. ISSN 1424-8220. Disponível em: ⟨https://www.mdpi.com/1424-8220/22/10/3617⟩.

Downloads

Published

2026-03-10

How to Cite

Moreira, G., Otoni, I., M. Silva, M., & L. Gomes, T. (2026). Two-Stage Fine-Tuning of Object Detectors for Industrial Environments. Revista De Informática Teórica E Aplicada, 33(2), 115–122. https://doi.org/10.22456/2175-2745.150904

Issue

Section

WVC2025

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.