Sign Language Segmentation Models and Their Impact on Automatic Translation Systems
DOI:
https://doi.org/10.22456/2175-2745.150939Keywords:
sign Language segmentation, sentence boundary detection, automatic sign language translation, brazilian sign languageAbstract
This work investigates automatic segmentation of Brazilian Sign Language videos for translation systems, addressing challenges of the visual-spatial modality of signed languages. We introduce the JW-Bible-Libras dataset, the largest resource for this task, and evaluate two segmentation approaches: Optical Flow-based models and Spatio-Temporal Graph Convolutional Networks (ST-GCN). Segmentation performance is analyzed both intrinsically and in relation to downstream translation using the gloss-free Sign2GPT architecture. Results show that the nine-layer ST-GCN with bidirectional LSTM achieves the best segmentation results (F1: 0.7358, IoU: 0.5820), while the unidirectional variant yields the strongest translation scores (BLEU1: 9.31, ROUGE: 9.49). Notably, a simple heuristic based on average sentence duration performs competitively, highlighting the gap between segmentation accuracy and translation quality. Our findings demonstrate the importance of segmentation strategies while revealing opportunities for integrating linguistic cues and boundary-aware learning to advance sign language translation.
Downloads
References
[1] World Federation of the Deaf. Frequently Asked Questions. 2025. ⟨https://wfdeaf.org/contact/faqs/⟩. Accessed: 2025-08-25.
[2] TANZER, G. et al. Reconsidering sentence-level sign language translation. Conference on Empirical Methods in Natural Language Processing, 2024. Disponível em: ⟨https://aclanthology.org/2024.emnlp-main.360/⟩.
[3] ORMEL, E.; CRASBORN, O. Prosodic correlates of sentences in signed languages: a literature review and suggestions for new types of studies. Sign Language Studies, v. 12, p. 279–315, 01 2011. Disponível em: ⟨https://doi.org/10.1353/sls.2011.0019⟩.
[4] FENLON, J. et al. Seeing sentence boundaries. Sign Language Linguistics, v. 10, p. 177–200, 09 2008. Disponível em: ⟨https://doi.org/10.1075/sll.10.2.06fen⟩.
[5] KHAN, S. Segmentation of continuous sign language. Tese (PhD thesis) — Massey University, 2014. Disponível em: ⟨https://mro.massey.ac.nz/items/f97a8223-7144-46ea-83ca-d38eed5f7fc2⟩.
[6] SISTO, M. D. et al. Defining meaningful units: challenges in sign segmentation and segment-meaning mapping. In: Proceedings of the 1st International Workshop on Automatic Translation for Signed and Spoken Languages (AT4SSL). [s.n.], 2021. p. 98–103. Disponível em: ⟨https://aclanthology.org/2021.mtsummit-at4ssl.11/⟩.
[7] GABARRÓ-LÓPEZ, S.; MEURANT, L. When nonmanuals meet semantics and syntax: a practical guide for the segmentation of sign language discourse. In: . [s.n.], 2014. Disponível em: ⟨https://www.sign-lang.uni-hamburg.de/lrec/pub/14018.pdf⟩.
[8] MORYOSSEF, A. et al. Real-time sign language detection using human pose estimation. In: BARTOLI, A.; FUSIELLO, A. (Ed.). Computer Vision – ECCV 2020 Workshops. Cham: Springer International Publishing, 2020. p. 237–248. ISBN 978-3-030-66096-3. Disponível em: ⟨https://www.slrtp.com/papers/full papers/SLRTP.FP.04.017.paper.pdf⟩.
[9] CAO, Z. et al. OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. 2019. Disponível em: ⟨https://arxiv.org/abs/1812.08008⟩.
[10] KONRAD, R. et al. languageresource, MEINE DGS – annotiert. Öffentliches Korpus der Deutschen Gebärdensprache, 3. Release / MY DGS – annotated. Public Corpus of German Sign Language, 3rd release. Universität Hamburg, 2020. Disponível em: ⟨https://doi.org/10.25592/dgs.corpus-3.0⟩.
[11] MORYOSSEF, A. et al. Linguistically motivated sign language segmentation. In: Findings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). [s.n.], 2023. p. 12703–12724. Disponível em: ⟨https://doi.org/10.18653/v1/2023.findings-emnlp.846⟩.
[12] BULL, H.; GOUIFFÈS, M.; BRAFFORT, A. Automatic segmentation of sign language into subtitle-units. In: BARTOLI, A.; FUSIELLO, A. (Ed.). Computer Vision – ECCV 2020 Workshops. Cham: Springer International Publishing, 2020. p. 186–198. ISBN 978-3-030-66096-3. Disponível em: ⟨https://slrtp.com/papers/full papers/SLRTP.FP.01.011.paper.pdf⟩.
[13] YAN, S.; XIONG, Y.; LIN, D. Spatial temporal graph convolutional networks for skeleton-based action recognition. AAAI Conference on Artificial Intelligence, 2018. Disponível em: ⟨https://arxiv.org/abs/1801.07455⟩.
[14] BULL, H.; BRAFFORT, A.; GOUIFFÈS, M. MEDIAPI-SKEL - a 2D-skeleton video database of French Sign Language with aligned French subtitles. In: CALZOLARI, N. et al. (Ed.). Proceedings of the Twelfth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association, 2020. p. 6063–6068. ISBN 979-10-95546-34-4. Disponível em: ⟨https://aclanthology.org/2020.lrec-1.743⟩.
[15] SILVA, D.; ESTEVAM, V.; MENOTTI, D. Towards a realistic libras to portuguese translation. In: Anais da XXXVI Conference on Graphics, Patterns and Images. Porto Alegre, RS, Brasil: SBC, 2023. p. 217–222. Disponível em: ⟨https://sol.sbc.org.br/index.php/sibgrapi/article/view/27372⟩.
[16] WONG, R.; CAMGOZ, N. C.; BOWDEN, R. Sign2GPT: Leveraging large language models for gloss-free sign language translation. In: The Twelfth International Conference on Learning Representations. [s.n.], 2024. Disponível em: ⟨https://openreview.net/forum?id=LqaEEs3UxU⟩.
[17] OQUAB, M. et al. Dinov2: Learning robust visual features without supervision. Transactions on Machine Learning Research, 2023. Disponível em: ⟨https://arxiv.org/abs/1812.08008⟩.
[18] LIN, X. V. et al. Few-shot learning with multilingual generative language models. Conference on Empirical Methods in Natural Language Processing, 2022. Disponível em: ⟨https://aclanthology.org/2022.emnlp-main.616/⟩.
[19] HONNIBAL, M. et al. spacy: Industrial-strength natural language processing in python. 2020. Disponível em: ⟨https://doi.org/10.5281/zenodo.1212303⟩.
[20] CHOLLET, F. et al. Keras. 2015. ⟨https://keras.io⟩.
[21] MORALES, F. keras-ocr. [S.l.]: PyPI, 2023. ⟨https://pypi.org/project/keras-ocr/⟩. Version 0.9.3.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Jéssica Luiza Ferreira Ramos, Michel Melo da Silva, Mario Fernando Montenegro Campos

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.













