A Luminous Pen for Air-Writing: Hybrid Approach with Computer Vision and Vision–Language Models

Authors

DOI:

https://doi.org/10.22456/2175-2745.150973

Keywords:

VLM, air writing, human-computer Interaction (HCI), image recognition, word recognition

Abstract

Air writing is a form of human–computer interaction that enables users to write in the air through different technological approaches, such as radio waves, wearable sensors, dedicated devices, and computer vision. In this work, we present a hybrid solution that combines the use of a luminous-tip pen with computer vision techniques to support pen detection and the air-writing process. The pen was designed and fabricated with a 3D printer and integrates an ESP32 microcontroller board, which provides Bluetooth connectivity for real-time communication with the computer. In addition to tracking movements for writing, the device can transmit specific commands to the system, such as starting or stopping the writing process and changing colors. This approach bridges the gap between wearable devices and vision-based methods, offering a practical and accessible alternative for air-writing applications. The system integrates vision–language models capable of recognizing both words and images. To determine which models were best suited for our system, we evaluated the performance of several candidates. Five participants contributed to the creation of the dataset. For drawing recognition, each participant produced images in the following classes: Tree, Moon, Cat, Heart, and Pen. For word recognition, the dataset included the following eight words: Purple, Window, Jungle, Pillow, Team, Doctor, Words, and Science. The results showed that the best-performing models achieved 76.00% accuracy with Kosmos-2 for image detection and 97.39% accuracy with Gemini 2.5 Flash for word detection, demonstrating satisfactory outcomes for the system implementation. The proposed system demonstrates the potential of combining electronic devices and advanced machine learning techniques for air-writing, with applications in education and assistive technologies.

Downloads

Download data is not yet available.

References

[1] YANAY, T.; SHMUELI, E. Air-writing recognition using smart-bands. Pervasive and Mobile Computing, v. 66, p. 101183, 2020. ISSN 1574-1192. Disponível em: https://www.sciencedirect.com/science/article/pii/S1574119220300602.

[2] CHEN, M.; ALREGIB, G.; JUANG, B.-H. Air-writing recognition—part i: Modeling and recognition of characters, words, and connecting motions. IEEE Transactions on Human-Machine Systems, v. 46, n. 3, p. 403–413, 2016.

[3] VAIDYA, V.; PRAVANTH, T.; VIJI, D. Air writing recognition application for dyslexic people. In: 2022 International Mobile and Embedded Technology Conference (MECON). [S.l.: s.n.], 2022. p. 553–558.

[4] ELSHENAWAY, A. R.; GUIRGUIS, S. K. On-air hand-drawn doodles for iot devices authentication during covid-19. IEEE Access, v. 9, p. 161723–161744, 2021.

[5] BARBOSA, C. E. et al. Reconhecimento de texto para sistemas air writing: Um estudo experimental. In: SBC. Escola Regional de Informática do Espírito Santo (ERI-ES). [S.l.], 2024. p. 21–30.

[6] VLOISON, V.; XIWEI, H. Deep Learning framework for Line-level Handwritten Text Recognition. 2021. https://github.com/vloison/Handwritten_Text_Recognition.

[7] LI, M. et al. TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models. 2021. https://github.com/microsoft/unilm/tree/master/trocr.

[8] ALAM, M. S.; KWON, K.-C.; KIM, N. Trajectory-based air-writing character recognition using convolutional neural network. In: 2019 4th International Conference on Control, Robotics and Cybernetics (CRC). [S.l.: s.n.], 2019. p. 86–90.

[9] CHEN, Y.-H.; SU, P.-C.; CHIEN, F.-T. Air-writing for smart glasses by effective fingertip detection. In: 2019 IEEE 8th Global Conference on Consumer Electronics (GCCE). [S.l.: s.n.], 2019. p. 381–382.

[10] KARBHARI, J.; MUKHERJI, P. Alphabet recognition using air written trajectories. In: 2023 International Conference on Emerging Smart Computing and Informatics (ESCI). [S.l.: s.n.], 2023. p. 1–6.

[11] PAL, A. Micapen: A pen to write in air using mica motes. In: 2020 16th International Conference on Distributed Computing in Sensor Systems (DCOSS). [S.l.: s.n.], 2020. p. 151–154.

[12] GHOSH, A. et al. Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions. 2025. Disponível em: https://arxiv.org/abs/2404.07214.

[13] PENG, Z. et al. Kosmos-2: Grounding multimodal large language models to the world. ArXiv, abs/2306.14824, 2023.

[14] LI, J. et al. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. arXiv, 2022. Disponível em: https://arxiv.org/abs/2201.12086.

Downloads

Published

2026-03-10

How to Cite

T. L. de Souza, L., A. D. Caldeira, R., D. C. Leal, S., Clara P. de Souza, M., M. Paixão, T., & J. M. G. Tello, R. (2026). A Luminous Pen for Air-Writing: Hybrid Approach with Computer Vision and Vision–Language Models. Revista De Informática Teórica E Aplicada, 33(2), 52–57. https://doi.org/10.22456/2175-2745.150973

Issue

Section

WVC2025

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.