Exploring Topic Modeling in Short Texts from Social Media: A Comparative Analysis of Algorithms
DOI:
https://doi.org/10.22456/2175-2745.145569Keywords:
Topic Modeling, Short Texts, Social Media, Performance ComparisonAbstract
Many texts are broadly disseminated on online social media platforms each day. Topic modeling is a Natural Language Processing and Unsupervised Learning technique used in this scenario. It identifies the topics in a collection of texts — that is, the most relevant groups of words in the context of all the texts are analyzed. The characteristics of the texts are a relevant factor for identifying topics. Unlike traditional sources, texts published on social media are usually short, because even when a character limit per publication (e.g., Twitter/X) is not imposed, users tend to be objective in the texts they write. This work evaluates the performance of four categories of topic modeling algorithms: traditional, Dirichlet Multinomial Mixture (DMM)-based, self-aggregation-based, and global co-occurrence-based. Real texts generated by social media users were used for the evaluation. Model performance was evaluated using quality metrics accepted in the literature. Finally, the results were analyzed such that the performance of each algorithm was weighted down, and to clarify whether there would be any detriment to the results of topic modeling using traditional algorithms on short texts.
Downloads
References
[1] LAUREATE, C. D. P.; BUNTINE, W.; LINGER, H. A systematic review of the use of topic models for short text social media analysis. Artificial Intelligence Review, Springer, v. 56, n. 12, p. 14223–14255, 2023.
[2] MCQUILLAN, L. et al. Cultural convergence: Insights into the behavior of misinformation networks on twitter. arXiv preprint arXiv:2007.03443, 2020.
[3] NATURE. The powers and perils of using digital data to understand human behaviour. 2021. ⟨https://www.nature.com/articles/d41586-021-01736-y⟩.
[4] BLEI, D. M.; NG, A. Y.; JORDAN, M. I. Latent dirichlet allocation. Journal of machine Learning research, v. 3, n. Jan, p. 993–1022, 2003.
[5] ALASHRI, S. et al. An analysis of sentiments on facebook during the 2016 us presidential election. In: IEEE. 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM). [S.l.], 2016. p. 795–802.
[6] ZHANG, Y. et al. iDoctor: Personalized and professionalized medical recommendations based on hybrid matrix factorization. Future Generation Computer Systems, Elsevier, v. 66, p. 30–35, 2017.
[7] CHEN, Y. et al. What we can do and cannot do with topic modeling: A systematic review. Communication Methods and Measures, Taylor & Francis, v. 17, n. 2, p. 111–130, 2023.
[8] WU, X.; LI, C. Short text topic modeling with flexible word patterns. In: IEEE. 2019 International Joint Conference on Neural Networks (IJCNN). [S.l.], 2019. p. 1–7.
[9] QIANG, J. et al. Short text topic modeling techniques, applications, and performance: a survey. IEEE Transactions on Knowledge and Data Engineering, IEEE, v. 34, n. 3, p. 1427–1445, 2020.
[10] G., A. M. A.; ROBLEDO, S.; ZULUAGA, M. Topic modeling: Perspectives from a literature review. IEEE Access, v. 11, p. 4066–4078, 2023.
[11] CHURCHILL, R.; SINGH, L. The evolution of topic modeling. ACM Comput. Surv., Association for Computing Machinery, New York, NY, USA, v. 54, n. 10s, 2022. ISSN 0360-0300.
[12] YIN, J.; WANG, J. A dirichlet multinomial mixture model-based approach for short text clustering. In: Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining. New York, NY, USA: ACM, 2014. p. 233–242.
[13] LI, C. et al. Enhancing topic modeling for short texts with auxiliary word embeddings. ACM Transactions on Information Systems (TOIS), ACM New York, NY, USA, v. 36, n. 2, p. 1–30, 2017.
[14] CHENG, X. et al. BTM: Topic modeling over short texts. IEEE Transactions on Knowledge and Data Engineering, IEEE, v. 26, n. 12, p. 2928–2941, 2014.
[15] HADI, M. A.; FARD, F. H. AOBTM: Adaptive online biterm topic modeling for version sensitive short-texts analysis. In: IEEE. 2020 IEEE international conference on software maintenance and evolution (ICSME). [S.l.], 2020. p. 593–604.
[16] GAO, W. et al. Incorporating word embeddings into topic modeling of short text. Knowledge and Information Systems, Springer, v. 61, p. 1123–1145, 2019.
[17] ZHAO, H. et al. Leveraging meta information in short text aggregation. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. p. 4042–4049.
[18] FANG, A. Analysing political events on twitter: Topic modelling and user community classification. SIGIR Forum, Association for Computing Machinery, New York, NY, USA, v. 53, n. 1, p. 38–39, 2021.
[19] CAMPAGNOLO, J. M.; DUARTE, D.; BIANCO, G. D. Topic coherence metrics: How sensitive are they? Journal of Information and Data Management, v. 13, n. 4, 2022.
[20] TOLEGEN, G. et al. A clustering-based approach for topic modeling via word network analysis. In: 2022 7th International Conference on Computer Science and Engineering (UBMK). Ankara, Turkey: IEEE, 2022. p. 192–197.
[21] WALLACH, H. M. et al. Evaluation methods for topic models. In: Proceedings of the 26th Annual International Conference on Machine Learning. Montreal, Quebec, Canada: ACM, 2009. p. 1105–1112.
[22] YUSOF, M. A.; SAEE, S. Code switching: exploring perplexity and coherence metrics for optimizing topic models of historical documents. International Journal of Systematic Innovation, v. 8, n. 4, p. 103–118, 2024.
[23] CAO, J. et al. A density-based method for adaptive LDA model selection. Neurocomputing, v. 72, n. 7-9, p. 1775–1781, 2009. Advances in Machine Learning and Computational Intelligence.
[24] ARUN, R. et al. On finding the natural number of topics with latent dirichlet allocation: Some observations. In: ZAKI, M. J. et al. (Ed.). Advances in Knowledge Discovery and Data Mining. Berlin, Heidelberg: Springer Berlin Heidelberg, 2010. p. 391–402.
[25] MIMNO, D. et al. Optimizing semantic coherence in topic models. In: Proceedings of the 2011 conference on empirical methods in natural language processing. Edinburgh, Scotland, UK.: Association for Computational Linguistics, 2011. p. 262–272.
[26] LIM, J. P.; LAUW, H. Large-scale correlation analysis of automated metrics for topic models. In: ROGERS, A.; BOYD-GRABER, J.; OKAZAKI, N. (Ed.). Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Toronto, Canada: Association for Computational Linguistics, 2023. p. 13874–13898.
[27] O’CALLAGHAN, D. et al. An analysis of the coherence of descriptors in topic modeling. Expert Systems with Applications, Elsevier, v. 42, n. 13, p. 5645–5657, 2015.
[28] BANDA, J. M. et al. A large-scale COVID-19 Twitter chatter dataset for open scientific research—an international collaboration. Epidemiologia, MDPI, v. 2, n. 3, p. 315–324, 2021. ISSN 2673-3986. Disponível em: ⟨https://www.mdpi.com/2673-3986/2/3/24⟩.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Ian Macedo Maiwald Santos, Luciana de Oliveira Rech, Ricardo Moraes

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.Accepted 2025-12-26
Published 2026-01-30













