Efficiency Analysis of Fine-Tuning Compact Language Models for Biomedical Information Retrieval Systems
DOI:
https://doi.org/10.22456/2175-2745.153249Keywords:
small language models (SLMs), ajuste fino eficiente em parâmetros (LoRA), biomedical embeddings, information retrieval (RAG)Abstract
The rise of semantic search systems and Retrieval-Augmented Generation (RAG) architectures in the biomedical domain poses critical challenges regarding computational costs and data sovereignty. This paper investigates the effectiveness of specializing ultra-compact language models for clinical embedding generation, utilizing the EmbeddingGemma 300M as the base model. By applying Low-Rank Adaptation (LoRA) techniques and low-level optimizations via the Unsloth library, the study demonstrates the feasibility of performing fine-tuning on entry-level hardware (NVIDIA Tesla T4). Experimental results show a 79.5% reduction in Multiple Negatives Ranking Loss (MNRL) in only 44.7 minutes, consuming just 6.56 GB of VRAM. The achieved throughput of 3,809 tokens/s and the residual operational cost (~$0.26) validate the proposal as a sustainable and private alternative. The results demonstrate competitive performance, approaching the effectiveness of significantly larger models in specific biomedical retrieval tasks with a fraction of the computational cost, thus promoting technological democratization for healthcare institutions with limited resources.
Downloads
References
[1] LEWIS, P. et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, v. 33, p. 9459–9474, 2020.
[2] THAKUR, N. et al. Beir: A heterogeneous benchmark for information retrieval. arXiv preprint arXiv:2104.08663, 2021.
[3] SCHWARTZ, R. et al. Green ai. Communications of the ACM, v. 63, n. 12, p. 54–63, 2020.
[4] HU, E. J. et al. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
[5] HAN, D.; HAN, M. Unsloth: Lightweight and 2x faster LLM fine-tuning. 2024. Disponível em: ⟨https://github.com/unslothai/unsloth⟩.
[6] REIMERS, N.; GUREVYCH, I. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019.
[7] KAPLAN, J. et al. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
[8] AGHAJANYAN, A. et al. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. arXiv preprint arXiv:2012.13255, 2020.
[9] KIRKPATRICK, J. et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, v. 114, n. 13, p. 3521–3526, 2017.
[10] DAO, T. et al. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 2022.
[11] DETTMERS, T. et al. Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems, 2024.
[12] HENDERSON, M. et al. Efficient natural language response suggestion for smart reply. arXiv preprint arXiv:1705.00652, 2017.
[13] JIN, Q. et al. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics, 2023.
[14] XU, B. et al. Bmretriever: Tuning large language models as better biomedical text retrievers. arXiv preprint arXiv:2404.18443, 2024.
[15] LABRAK, Y.; BAZOGE, A. et al. Biomistral: A collection of open-source pretrained large language models for medical applications. arXiv preprint arXiv:2402.10373, 2024.
[16] GU, Y. et al. Domain-specific language model pre-training for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), ACM New York, NY, USA, v. 3, n. 1, p. 1–23, 2021.
[17] YUAN, H. et al. Biobart: Pretraining and evaluation of a biomedical generative language model. arXiv preprint arXiv:2203.11123, 2022.
[18] HEVNER, A. R. et al. Design science in information systems research. MIS Quarterly, Management Information Systems Research Center, v. 28, n. 1, p. 75–105, 2004.
[19] SILVA, M. D. Notebook de Experimentos: Fine-tuning EmbeddingGemma 300M para RI Biomédica. [S.l.]: Google Colab, 2025. ⟨https://colab.research.google.com/drive/1KVt6Qd5IzdhkNIAALQ3JPl1iXGcllWNH⟩. Acessado em: 20/04/2026.
[20] DETTMERS, T. et al. 8-bit optimizers via block-wise quantization. arXiv preprint arXiv:2110.02861, 2022.
[21] Unsloth AI. Embedding Fine-tuning Documentation. 2025. Acessado em: 10 de jan. de 2026. Disponível em: ⟨https://unsloth.ai/docs/new/embedding-finetuning⟩.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Martony Demes da Silva

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.Accepted 2026-06-04
Published 2026-08-10













