Efficiency Analysis of Fine-Tuning Compact Language Models for Biomedical Information Retrieval Systems

Authors

DOI:

https://doi.org/10.22456/2175-2745.153249

Keywords:

small language models (SLMs), ajuste fino eficiente em parâmetros (LoRA), biomedical embeddings, information retrieval (RAG)

Abstract

The rise of semantic search systems and Retrieval-Augmented Generation (RAG) architectures in the biomedical domain poses critical challenges regarding computational costs and data sovereignty. This paper investigates the effectiveness of specializing ultra-compact language models for clinical embedding generation, utilizing the EmbeddingGemma 300M as the base model. By applying Low-Rank Adaptation (LoRA) techniques and low-level optimizations via the Unsloth library, the study demonstrates the feasibility of performing fine-tuning on entry-level hardware (NVIDIA Tesla T4). Experimental results show a 79.5% reduction in Multiple Negatives Ranking Loss (MNRL) in only 44.7 minutes, consuming just 6.56 GB of VRAM. The achieved throughput of 3,809 tokens/s and the residual operational cost (~$0.26) validate the proposal as a sustainable and private alternative. The results demonstrate competitive performance, approaching the effectiveness of significantly larger models in specific biomedical retrieval tasks with a fraction of the computational cost, thus promoting technological democratization for healthcare institutions with limited resources.

Downloads

Download data is not yet available.

Author Biography

Martony Demes da Silva, Universidade Federal do Piauí

Graduated in Computer Science from UESPI and in Telecommunications Networks from CEFET-PI. Holds postgraduate degrees in Computer Networks and IT Governance. Currently works as a higher education professor, postgraduate lecturer, undergraduate thesis advisor, and Information Technology specialist.

References

[1] LEWIS, P. et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, v. 33, p. 9459–9474, 2020.

[2] THAKUR, N. et al. Beir: A heterogeneous benchmark for information retrieval. arXiv preprint arXiv:2104.08663, 2021.

[3] SCHWARTZ, R. et al. Green ai. Communications of the ACM, v. 63, n. 12, p. 54–63, 2020.

[4] HU, E. J. et al. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.

[5] HAN, D.; HAN, M. Unsloth: Lightweight and 2x faster LLM fine-tuning. 2024. Disponível em: ⟨https://github.com/unslothai/unsloth⟩.

[6] REIMERS, N.; GUREVYCH, I. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019.

[7] KAPLAN, J. et al. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.

[8] AGHAJANYAN, A. et al. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. arXiv preprint arXiv:2012.13255, 2020.

[9] KIRKPATRICK, J. et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, v. 114, n. 13, p. 3521–3526, 2017.

[10] DAO, T. et al. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 2022.

[11] DETTMERS, T. et al. Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems, 2024.

[12] HENDERSON, M. et al. Efficient natural language response suggestion for smart reply. arXiv preprint arXiv:1705.00652, 2017.

[13] JIN, Q. et al. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics, 2023.

[14] XU, B. et al. Bmretriever: Tuning large language models as better biomedical text retrievers. arXiv preprint arXiv:2404.18443, 2024.

[15] LABRAK, Y.; BAZOGE, A. et al. Biomistral: A collection of open-source pretrained large language models for medical applications. arXiv preprint arXiv:2402.10373, 2024.

[16] GU, Y. et al. Domain-specific language model pre-training for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), ACM New York, NY, USA, v. 3, n. 1, p. 1–23, 2021.

[17] YUAN, H. et al. Biobart: Pretraining and evaluation of a biomedical generative language model. arXiv preprint arXiv:2203.11123, 2022.

[18] HEVNER, A. R. et al. Design science in information systems research. MIS Quarterly, Management Information Systems Research Center, v. 28, n. 1, p. 75–105, 2004.

[19] SILVA, M. D. Notebook de Experimentos: Fine-tuning EmbeddingGemma 300M para RI Biomédica. [S.l.]: Google Colab, 2025. ⟨https://colab.research.google.com/drive/1KVt6Qd5IzdhkNIAALQ3JPl1iXGcllWNH⟩. Acessado em: 20/04/2026.

[20] DETTMERS, T. et al. 8-bit optimizers via block-wise quantization. arXiv preprint arXiv:2110.02861, 2022.

[21] Unsloth AI. Embedding Fine-tuning Documentation. 2025. Acessado em: 10 de jan. de 2026. Disponível em: ⟨https://unsloth.ai/docs/new/embedding-finetuning⟩.

Downloads

Published

2026-08-10

How to Cite

Demes da Silva, M. (2026). Efficiency Analysis of Fine-Tuning Compact Language Models for Biomedical Information Retrieval Systems. Revista De Informática Teórica E Aplicada, 33(4), 21–31. https://doi.org/10.22456/2175-2745.153249

Issue

Section

Regular Papers
Received 2026-01-28
Accepted 2026-06-04
Published 2026-08-10

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.