Exploring Topic Modeling in Short Texts from Social Media: A Comparative Analysis of Algorithms

Authors

DOI:

https://doi.org/10.22456/2175-2745.145569

Keywords:

Topic Modeling, Short Texts, Social Media, Performance Comparison

Abstract

Many texts are broadly disseminated on online social media platforms each day. Topic modeling is a Natural Language Processing and Unsupervised Learning technique used in this scenario. It identifies the topics in a collection of texts — that is, the most relevant groups of words in the context of all the texts are analyzed. The characteristics of the texts are a relevant factor for identifying topics. Unlike traditional sources, texts published on social media are usually short, because even when a character limit per publication (e.g., Twitter/X) is not imposed, users tend to be objective in the texts they write. This work evaluates the performance of four categories of topic modeling algorithms: traditional, Dirichlet Multinomial Mixture (DMM)-based, self-aggregation-based, and global co-occurrence-based. Real texts generated by social media users were used for the evaluation. Model performance was evaluated using quality metrics accepted in the literature. Finally, the results were analyzed such that the performance of each algorithm was weighted down, and to clarify whether there would be any detriment to the results of topic modeling using traditional algorithms on short texts.

Downloads

Download data is not yet available.

Author Biographies

Luciana de Oliveira Rech, Universidade Federal de Santa Catarina (UFSC)

Departamento de Computação

Ricardo Moraes, Universidade Federal de Santa Catarina (UFSC)

Departamento de Ciência da Informação

References

[1] LAUREATE, C. D. P.; BUNTINE, W.; LINGER, H. A systematic review of the use of topic models for short text social media analysis. Artificial Intelligence Review, Springer, v. 56, n. 12, p. 14223–14255, 2023.

[2] MCQUILLAN, L. et al. Cultural convergence: Insights into the behavior of misinformation networks on twitter. arXiv preprint arXiv:2007.03443, 2020.

[3] NATURE. The powers and perils of using digital data to understand human behaviour. 2021. ⟨https://www.nature.com/articles/d41586-021-01736-y⟩.

[4] BLEI, D. M.; NG, A. Y.; JORDAN, M. I. Latent dirichlet allocation. Journal of machine Learning research, v. 3, n. Jan, p. 993–1022, 2003.

[5] ALASHRI, S. et al. An analysis of sentiments on facebook during the 2016 us presidential election. In: IEEE. 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM). [S.l.], 2016. p. 795–802.

[6] ZHANG, Y. et al. iDoctor: Personalized and professionalized medical recommendations based on hybrid matrix factorization. Future Generation Computer Systems, Elsevier, v. 66, p. 30–35, 2017.

[7] CHEN, Y. et al. What we can do and cannot do with topic modeling: A systematic review. Communication Methods and Measures, Taylor & Francis, v. 17, n. 2, p. 111–130, 2023.

[8] WU, X.; LI, C. Short text topic modeling with flexible word patterns. In: IEEE. 2019 International Joint Conference on Neural Networks (IJCNN). [S.l.], 2019. p. 1–7.

[9] QIANG, J. et al. Short text topic modeling techniques, applications, and performance: a survey. IEEE Transactions on Knowledge and Data Engineering, IEEE, v. 34, n. 3, p. 1427–1445, 2020.

[10] G., A. M. A.; ROBLEDO, S.; ZULUAGA, M. Topic modeling: Perspectives from a literature review. IEEE Access, v. 11, p. 4066–4078, 2023.

[11] CHURCHILL, R.; SINGH, L. The evolution of topic modeling. ACM Comput. Surv., Association for Computing Machinery, New York, NY, USA, v. 54, n. 10s, 2022. ISSN 0360-0300.

[12] YIN, J.; WANG, J. A dirichlet multinomial mixture model-based approach for short text clustering. In: Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining. New York, NY, USA: ACM, 2014. p. 233–242.

[13] LI, C. et al. Enhancing topic modeling for short texts with auxiliary word embeddings. ACM Transactions on Information Systems (TOIS), ACM New York, NY, USA, v. 36, n. 2, p. 1–30, 2017.

[14] CHENG, X. et al. BTM: Topic modeling over short texts. IEEE Transactions on Knowledge and Data Engineering, IEEE, v. 26, n. 12, p. 2928–2941, 2014.

[15] HADI, M. A.; FARD, F. H. AOBTM: Adaptive online biterm topic modeling for version sensitive short-texts analysis. In: IEEE. 2020 IEEE international conference on software maintenance and evolution (ICSME). [S.l.], 2020. p. 593–604.

[16] GAO, W. et al. Incorporating word embeddings into topic modeling of short text. Knowledge and Information Systems, Springer, v. 61, p. 1123–1145, 2019.

[17] ZHAO, H. et al. Leveraging meta information in short text aggregation. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. p. 4042–4049.

[18] FANG, A. Analysing political events on twitter: Topic modelling and user community classification. SIGIR Forum, Association for Computing Machinery, New York, NY, USA, v. 53, n. 1, p. 38–39, 2021.

[19] CAMPAGNOLO, J. M.; DUARTE, D.; BIANCO, G. D. Topic coherence metrics: How sensitive are they? Journal of Information and Data Management, v. 13, n. 4, 2022.

[20] TOLEGEN, G. et al. A clustering-based approach for topic modeling via word network analysis. In: 2022 7th International Conference on Computer Science and Engineering (UBMK). Ankara, Turkey: IEEE, 2022. p. 192–197.

[21] WALLACH, H. M. et al. Evaluation methods for topic models. In: Proceedings of the 26th Annual International Conference on Machine Learning. Montreal, Quebec, Canada: ACM, 2009. p. 1105–1112.

[22] YUSOF, M. A.; SAEE, S. Code switching: exploring perplexity and coherence metrics for optimizing topic models of historical documents. International Journal of Systematic Innovation, v. 8, n. 4, p. 103–118, 2024.

[23] CAO, J. et al. A density-based method for adaptive LDA model selection. Neurocomputing, v. 72, n. 7-9, p. 1775–1781, 2009. Advances in Machine Learning and Computational Intelligence.

[24] ARUN, R. et al. On finding the natural number of topics with latent dirichlet allocation: Some observations. In: ZAKI, M. J. et al. (Ed.). Advances in Knowledge Discovery and Data Mining. Berlin, Heidelberg: Springer Berlin Heidelberg, 2010. p. 391–402.

[25] MIMNO, D. et al. Optimizing semantic coherence in topic models. In: Proceedings of the 2011 conference on empirical methods in natural language processing. Edinburgh, Scotland, UK.: Association for Computational Linguistics, 2011. p. 262–272.

[26] LIM, J. P.; LAUW, H. Large-scale correlation analysis of automated metrics for topic models. In: ROGERS, A.; BOYD-GRABER, J.; OKAZAKI, N. (Ed.). Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Toronto, Canada: Association for Computational Linguistics, 2023. p. 13874–13898.

[27] O’CALLAGHAN, D. et al. An analysis of the coherence of descriptors in topic modeling. Expert Systems with Applications, Elsevier, v. 42, n. 13, p. 5645–5657, 2015.

[28] BANDA, J. M. et al. A large-scale COVID-19 Twitter chatter dataset for open scientific research—an international collaboration. Epidemiologia, MDPI, v. 2, n. 3, p. 315–324, 2021. ISSN 2673-3986. Disponível em: ⟨https://www.mdpi.com/2673-3986/2/3/24⟩.

Downloads

Published

2026-01-30

How to Cite

Santos, I. M. M., Rech, L. de O., & Moraes, R. (2026). Exploring Topic Modeling in Short Texts from Social Media: A Comparative Analysis of Algorithms. Revista De Informática Teórica E Aplicada, 33(1), 87–98. https://doi.org/10.22456/2175-2745.145569

Issue

Section

Regular Papers
Received 2025-02-02
Accepted 2025-12-26
Published 2026-01-30

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.