Analyzing Correctness, Memory Consumption, CPU Usage, and Execution Time of the LLMs Llama-3.3-70B Versatile and Gemma2-9B Instruct
DOI:
https://doi.org/10.22456/2175-2745.153834Keywords:
LLM, code correctness, memory consumption, CPU usage, execution timeAbstract
Although the rapid advancement of Large Language Models (LLMs) has significantly transformed software development, enabling the automation of tasks such as code generation, testing, and documentation, concerns remain regarding the quality, efficiency, and security of automatically generated solutions, especially in relation to computational resource usage such as memory, CPU, and execution time. This scenario is intensified by the growing demand for digital sustainability and the need to minimize operational costs. This work investigates the efficiency and correctness of code generated by the LLMs Llama-3.3-70B Versatile and Gemma2-9B Instruct when solving Rosetta Code tasks in Python, comparing them with human implementations. The analyses address multiple aspects, including code correctness, clarity, memory consumption, CPU usage, and execution time, thus revealing the limitations and strengths of automatic solutions compared to human-developed code. The results suggest that, although LLMs are capable of quickly generating functional solutions, they still lack the algorithmic optimizations necessary to match the efficiency of manually written code, particularly regarding resource utilization. However, for routine or well-known tasks, automatic solutions can be highly competitive. This study highlights the importance of multidimensional evaluations and demonstrates that, if combined with human review and good engineering practices, the careful adoption of LLMs can enhance software development processes.
Downloads
References
[1] MASTROPAOLO, A. et al. On the robustness of code generation techniques: An empirical study on github copilot. arXiv preprint arXiv:2302.00438, 2023. Available at: ⟨https://arxiv.org/abs/2302.00438⟩.
[2] OPENAI. ChatGPT. 2024. Available at: ⟨https://openai.com/index/chatgpt/⟩.
[3] GOOGLE. Gemini. 2024. Available at: ⟨https://deepmind.google/technologies/gemini/⟩.
[4] Meta AI. Llama 3. 2024. Available at: ⟨https://llama.meta.com/⟩.
[5] ANTHROPIC. Claude. 2024. Available at: ⟨https://www.anthropic.com/claude/⟩.
[6] AUSTIN, J. et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021. Available at: ⟨https://arxiv.org/abs/2108.07732⟩.
[7] LIU, J. et al. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. 2023. Available at: ⟨https://arxiv.org/abs/2305.01210⟩.
[8] VARTZIOTIS, D.; MITROPOULOS, D. Green ai for code generation: Investigating the energy cost of large language models. arXiv preprint arXiv:2402.12513, 2024.
[9] PEREIRA, R. et al. Energy efficiency across programming languages: how do energy, time, and memory relate? In: ACM SIGPLAN International Conference on Software Language Engineering. [S.l.: s.n.], 2017.
[10] POP, M. Measuring the energy consumption and carbon footprint of encrypted databases using CodeCarbon. Tese (Doutorado) — University of Twente, 2025.
[11] STEENHOEK, B. et al. A comprehensive study of the capabilities of large language models for vulnerability detection. arXiv preprint arXiv:2403.17218, 2024. Available at: ⟨https://arxiv.org/html/2403.17218v1⟩.
[12] Rosetta Code. Rosetta Code. 2025. Available at: ⟨https://rosettacode.org/wiki/Rosetta_Code⟩.
[13] SALEWSKI, L. et al. In-context impersonation reveals large language models’ strengths and biases. In: Conference on Neural Information Processing Systems. [S.l.: s.n.], 2023. Available at: ⟨https://arxiv.org/abs/2305.14930⟩.
[14] Meta AI, via GroqCloud. Llama 3.3-70B Versatile [API documentation]. 2024. Available at: ⟨https://console.groq.com/docs/model/llama-3.3-70b-versatile/⟩.
[15] Google DeepMind. Gemma 2-9B IT [API documentation]. 2024. Available at: ⟨https://ai.google.dev/gemma/docs/⟩.
[16] Meta AI. Llama 3. 2024. ⟨https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_3⟩.
[17] TOUVRON, H. et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. Available at: ⟨https://arxiv.org/abs/2307.09288⟩.
[18] OUYANG, L. et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, v. 35, p. 27730–27744, 2022.
[19] Google. Gemma 2. 2024. ⟨https://ai.google.dev/gemma/docs⟩.
[20] Groq, Inc. GroqCloud API Documentation. 2024. Available at: ⟨https://console.groq.com/docs/⟩.
[21] WILCOXON, F. Individual comparisons by ranking methods. Biometrics bulletin, v. 1, n. 6, p. 80–83, 1945.
[22] VARGHA, A.; DELANEY, H. D. A critique and improvement of the CL common language effect size statistics of McGraw and Wong. Journal of Educational and Behavioral Statistics, v. 25, n. 2, p. 101–132, 2000.
[23] CHEN, M. et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021. Available at: ⟨https://arxiv.org/abs/2107.03374⟩.
[24] AMINI, A. et al. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. arXiv preprint arXiv:1905.13319, 2019. Available at: ⟨https://arxiv.org/abs/1905.13319⟩.
[25] The Algorithms. The Algorithms. 2025. Available at: ⟨https://the-algorithms.com/⟩.
[26] LeetCode. LeetCode Online Judge. 2025. Available at: ⟨https://leetcode.com/⟩.
[27] Codeforces. Codeforces: Competitive Programming Platform. 2025. Available at: ⟨https://codeforces.com/⟩.
[28] Beecrowd. Beecrowd: Programming Problems and Contests. 2025. Available at: ⟨https://www.beecrowd.com.br/⟩.
[29] LIMA, J.; ANDRADE, R. Analyzing Correctness, Memory Consumption, CPU Usage, and Execution Time of the LLMs Llama-3.3-70B Versatile and Gemma2-9B Instruct. 2026. Available at: ⟨https://zenodo.org/records/18776535⟩.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Rodrigo Andrade, Jackson Lima

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.Accepted 2026-05-18
Published 2026-06-21













