Behavioral Insights of SARSA-Based Reinforcement Learning for IPv6 Intrusion Classification
DOI:
https://doi.org/10.22456/2175-2745.148437Keywords:
epsilon greedy, IPv6, intrusion classification, policy behavior analysis, SARSA algoritmAbstract
Intrusion Detection Systems (IDS) are crucial for network security, but traditional models like signature-based and supervised learning approaches require frequent updates or retraining. This study introduces a reinforcement learning-based IDS using SARSA for IPv6 intrusion detection. The SARSA agent with epsilon 0.1 (E-0.1) performed best: optimal detection at step 11, 193,453 correct and 25,014 incorrect actions, total reward 842,195, average reward 3.855 per epoch, outperforming E-0.5 and E-0.9. Furthermore, a comparison of Q-Table exploitation between Q-Learning and SARSA indicates that both algorithms yield similar results. However, SARSA is preferable in security contexts because its more realistic approach results in a lower false alarm rate. This behavioral distinction is demonstrated by a -1.2109 value difference and a substantially lower mean state value for SARSA (4.903) compared to the more optimistic Q-Learning (16.310). To ensure robustness, the results were validated using different random seeds, and E-0.1 consistently produced stable total rewards with minimal variance. Additionally, Shannon entropy confirmed the agent's stable decision-making behavior, where E-0.1 showed faster convergence, indicating efficient learning. Altogether, these findings suggest that the SARSA-based IDS offers a more adaptive and reliable solution for intrusion detection, particularly in dynamic IPv6 environments.
Downloads
References
[1] M. Tajdini; H. Kolivand. IPv6 Detection Techniques and Solutions. In: 2023 16th International Conference on Developments in eSystems Engineering (DeSE). [S.l.: s.n.], 2023. p. 663–666. Journal Abbreviation: 2023 16th International Conference on Developments in eSystems Engineering (DeSE).
[2] SZYNKIEWICZ, P. Signature-Based Detection of Botnet DDoS Attacks. In: KOŁODZIEJ, J.; REPETTO, M.; DUZHA, A. (Ed.). Cybersecurity of Digital Service Chains: Challenges, Methodologies, and Tools. Cham: Springer International Publishing, 2022. p. 120–135. ISBN 978-3-031-04036-8. Disponível em: ⟨https://doi.org/10.1007/978-3-031-04036-8_6⟩.
[3] LIU, N. et al. A Survey on IPv6 Security Threats and Defense Mechanisms. In: SUN, X. et al. (Ed.). Artificial Intelligence and Security. Cham: Springer International Publishing, 2022. p. 583–598. ISBN 978-3-031-06794-5.
[4] T. Alsmadi; N. Alqudah. A Survey on malware detection techniques. In: 2021 International Conference on Information Technology (ICIT). [S.l.: s.n.], 2021. p. 371–376. Journal Abbreviation: 2021 International Conference on Information Technology (ICIT).
[5] KESKIN, O. F. et al. Cyber Third-Party Risk Management: A Comparison of Non-Intrusive Risk Scoring Reports. Electronics, v. 10, n. 10, 2021. ISSN 2079-9292.
[6] CREMER, F. et al. Cyber risk and cybersecurity: a systematic review of data availability. The Geneva papers on risk and insurance. Issues and practice, v. 47, n. 3, p. 698–736, 2022. ISSN 1468-0440 1018-5895. Place: England.
[7] ASLAN, O. et al. A Comprehensive Review of Cyber Security Vulnerabilities, Threats, Attacks, and Solutions. Electronics, v. 12, n. 6, 2023. ISSN 2079-9292. Disponível em: ⟨https://www.mdpi.com/2079-9292/12/6/1333⟩.
[8] HOI, S. C. et al. Online learning: A comprehensive survey. Neurocomputing, v. 459, p. 249–289, out. 2021. ISSN 0925-2312. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S0925231221006706⟩.
[9] GREENHOW, C. et al. Foundations of online learning: Challenges and opportunities. Educational Psychologist, v. 57, n. 3, p. 131–147, jul. 2022. ISSN 0046-1520. Publisher: Routledge. Disponível em: ⟨https://doi.org/10.1080/00461520.2022.2090364⟩.
[10] H. Kheddar et al. Reinforcement-Learning-Based Intrusion Detection in Communication Networks: A Review. IEEE Communications Surveys & Tutorials, v. 27, n. 4, p. 2420–2469, ago. 2025. ISSN 1553-877X.
[11] DARU, A. F.; HARTOMO, K. D.; PURNOMO, H. D. IPv6 flood attack detection based on epsilon greedy optimized Q learning in single board computer. International Journal of Electrical & Computer Engineering (2088-8708), v. 13, n. 5, p. 5782–5791, 2023. ISSN 2088-8708. Disponível em: ⟨https://ijece.iaescore.com/index.php/IJECE/article/view/30100⟩.
[12] DÍAZ-VERDEJO, J. et al. On the Detection Capabilities of Signature-Based Intrusion Detection Systems in the Context of Web Attacks. Applied Sciences, v. 12, n. 2, 2022. ISSN 2076-3417.
[13] ZHAO, R. et al. A Novel Intrusion Detection Method Based on Lightweight Neural Network for Internet of Things. IEEE Internet of Things Journal, v. 9, n. 12, p. 9960–9972, jun. 2022. ISSN 2327-4662.
[14] SCHRÖTTER, M.; NIEMANN, A.; SCHNOR, B. A Comparison of Neural-Network-Based Intrusion Detection against Signature-Based Detection in IoT Networks. Information, v. 15, n. 3, 2024. ISSN 2078-2489.
[15] KILINCER, I. F.; ERTAM, F.; SENGUR, A. Machine learning methods for cyber security intrusion detection: Datasets and comparative study. Computer Networks, v. 188, p. 107840, 2021. ISSN 1389-1286. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S1389128621000141⟩.
[16] DOUIBA, M. et al. Anomaly detection model based on gradient boosting and decision tree for IoT environments security. Journal of Reliable Intelligent Environments, jul. 2022. ISSN 2199-4676. Disponível em: ⟨https://doi.org/10.1007/s40860-022-00184-3⟩.
[17] UMAR, M. A. et al. Effects of feature selection and normalization on network intrusion detection. Data Science and Management, v. 8, n. 1, p. 23–39, mar. 2025. ISSN 2666-7649. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S2666764924000390⟩.
[18] ABDALLAH, E. E.; ELEISAH, W.; OTOOM, A. F. Intrusion Detection Systems using Supervised Machine Learning Techniques: A survey. The 13th International Conference on Ambient Systems, Networks and Technologies (ANT) / The 5th International Conference on Emerging Data and Industry 4.0 (EDI40), v. 201, p. 205–212, jan. 2022. ISSN 1877-0509. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S1877050922004422⟩.
[19] RAMACHANDRAN, S. et al. A survey on supervised machine learning algorithms. AIP Conference Proceedings, v. 3279, n. 1, p. 020003, abr. 2025. ISSN 0094-243X. Disponível em: ⟨https://doi.org/10.1063/5.0263126⟩.
[20] M. Agoramoorthy et al. An Analysis of Signature-Based Components in Hybrid Intrusion Detection Systems. In: 2023 Intelligent Computing and Control for Engineering and Business Systems (ICCEBS). [S.l.: s.n.], 2023. p. 1–5. Journal Abbreviation: 2023 Intelligent Computing and Control for Engineering and Business Systems (ICCEBS).
[21] THAKKAR, A.; LOHIYA, R. A survey on intrusion detection system: feature selection, model, performance measures, application perspective, challenges, and future research directions. Artificial Intelligence Review, v. 55, n. 1, p. 453–563, jan. 2022. ISSN 1573-7462. Disponível em: ⟨https://doi.org/10.1007/s10462-021-10037-9⟩.
[22] Van Hauser. The Hacker Choice’s IPv6 Attack Toolkit. 2020. Disponível em: ⟨https://github.com/vanhauser-thc/thc-ipv6/tree/v3.8⟩.
[23] Wireshark Foundation. Wireshark. 2025. Disponível em: ⟨https://gitlab.com/wireshark/wireshark/-/tree/release-4.2?ref_type=heads⟩.
[24] Farama Foundation. Gymnasium. 2024. Disponível em: ⟨https://pypi.org/project/gymnasium/⟩.
[25] CHOU, D.; JIANG, M. A Survey on Data-driven Network Intrusion Detection. ACM Comput. Surv., v. 54, n. 9, out. 2021. ISSN 0360-0300. Place: New York, NY, USA Publisher: Association for Computing Machinery. Disponível em: ⟨https://doi.org/10.1145/3472753⟩.
[26] SANTOS, K. C.; MIANI, R. S.; SILVA, F. de O. Evaluating the Impact of Data Preprocessing Techniques on the Performance of Intrusion Detection Systems. Journal of Network and Systems Management, v. 32, n. 2, p. 36, mar. 2024. ISSN 1573-7705. Disponível em: ⟨https://doi.org/10.1007/s10922-024-09813-z⟩.
[27] E. Chatzoglou et al. Pick Quality Over Quantity: Expert Feature Selection and Data Preprocessing for 802.11 Intrusion Detection Systems. IEEE Access, v. 10, p. 64761–64784, 2022. ISSN 2169-3536.
[28] SAVIĆ, I.; LIN, X. The Analysis and Implication of Data Deduplication in Digital Forensics. In: MENG, W.; CONTI, M. (Ed.). Cyberspace Safety and Security. Cham: Springer International Publishing, 2022. p. 198–215. ISBN 978-3-030-94029-4.
[29] N. Q. H. Othman; J. Li; Q. Yang. Multi-Agent Reinforcement Learning Aided Resource Allocation With SARSA in UAV Networks. In: 2023 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC). [S.l.: s.n.], 2023. p. 1–6. ISBN 2837-116X. Journal Abbreviation: 2023 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC).
[30] I. Zaghbani; R. Jarray; S. Bouallègue. Comparative Study of Q-Learning and SARSA Algorithms for UAV Path Planning in 3D Environments. In: 2024 IEEE 28th International Conference on Intelligent Engineering Systems (INES). [S.l.: s.n.], 2024. p. 000245–000250. ISBN 1543-9259. Journal Abbreviation: 2024 IEEE 28th International Conference on Intelligent Engineering Systems (INES).
[31] SIVAMAYIL, K. et al. A Systematic Study on Reinforcement Learning Based Applications. Energies, v. 16, n. 3, 2023. ISSN 1996-1073.
[32] CHAKRABORTY, J.; MAJUMDER, S.; MENZIES, T. Bias in machine learning software: why? how? what to do? In: Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. New York, NY, USA: Association for Computing Machinery, 2021. (ESEC/FSE 2021), p. 429–440. ISBN 978-1-4503-8562-6. Event-place: Athens, Greece. Disponível em: ⟨https://doi.org/10.1145/3468264.3468537⟩.
[33] SMITH, B.; KHOJANDI, A.; VASUDEVAN, R. Bias in Reinforcement Learning: A Review in Healthcare Applications. ACM Comput. Surv., v. 56, n. 2, set. 2023. ISSN 0360-0300. Place: New York, NY, USA Publisher: Association for Computing Machinery. Disponível em: ⟨https://doi.org/10.1145/3609502⟩.
[34] Q. Nguyen; N. Teku; T. Bose. Epsilon Greedy Strategy for Hyper Parameters Tuning of A Neural Network Equalizer. In: 2021 12th International Symposium on Image and Signal Processing and Analysis (ISPA). [S.l.: s.n.], 2021. p. 209–212. ISBN 1849-2266. Journal Abbreviation: 2021 12th International Symposium on Image and Signal Processing and Analysis (ISPA).
[35] BULUT, V. Optimal path planning method based on epsilon-greedy Q-learning algorithm. Journal of the Brazilian Society of Mechanical Sciences and Engineering, v. 44, n. 3, p. 106, mar. 2022. ISSN 1806-3691. Disponível em: ⟨https://doi.org/10.1007/s40430-022-03399-w⟩.
[36] CINCOTTA, P. M. et al. The Shannon entropy: An efficient indicator of dynamical stability. Physica D: Nonlinear Phenomena, v. 417, p. 132816, mar. 2021. ISSN 0167-2789. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S0167278920308174⟩.
[37] ALI, A.; ANAM, S.; AHMED, M. M. Shannon Entropy in Artificial Intelligence and Its Applications Based on Information Theory. Journal of Applied and Emerging Sciences; Vol 13, No 1 (2023), 2023. Disponível em: ⟨https://journal.buitms.edu.pk/j/index.php/bj/article/view/549⟩.
[38] SARAIVA, P. On Shannon entropy and its applications. Kuwait Journal of Science, v. 50, n. 3, p. 194–199, jul. 2023. ISSN 2307-4108. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S2307410823000433⟩.
[39] S. Dutta; A. Arunachalam; S. Misailovic. To Seed or Not to Seed? An Empirical Analysis of Usage of Seeds for Testing in Machine Learning Projects. In: 2022 IEEE Conference on Software Testing, Verification and Validation (ICST). [S.l.: s.n.], 2022. p. 151–161. ISBN 2159-4848. Journal Abbreviation: 2022 IEEE Conference on Software Testing, Verification and Validation (ICST).
[40] AHMED, H.; LOFSTEAD, J. Managing Randomness to Enable Reproducible Machine Learning. In: Proceedings of the 5th International Workshop on Practical Reproducible Evaluation of Computer Systems. New York, NY, USA: Association for Computing Machinery, 2022. (P-RECS ’22), p. 15–20. ISBN 978-1-4503-9313-3. Event-place: Minneapolis, MN, USA. Disponível em: ⟨https://doi.org/10.1145/3526062.3536353⟩.
[41] KACZMARCZYK, K.; MIAŁKOWSKA, K. Backtesting comparison of machine learning algorithms with different random seed. Knowledge-Based and Intelligent Information & Engineering Systems: Proceedings of the 26th International Conference KES2022, v. 207, p. 1901–1910, jan. 2022. ISSN 1877-0509. Disponível em: ⟨https://www.sciencedirect.com/science/article/pii/S1877050922011346⟩.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 April Firman Daru, Alauddin Maulana Hirzan, Zainab Senan Mahmod Attar Bashi, Fajriannoor Fanani

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.Accepted 2026-08-02
Published 2026-08-10













