DOI QR코드

DOI QR Code

A Study on Improving Answer Accuracy in Medical Question Answering Using LLaMa 2 7B and Retrieval-Augmented Generation (RAG))

RAG를 적용한 LLaMa 2 7B 기반 의료 질의응답 정확도 향상에 관한 연구

  • 김나랑 (동아대학교 경영정보학과) ;
  • 정의범 (한신대학교 경영학과) ;
  • 김현종 (동아대학교 글로컬 네트워크 인공지능 연구소)
  • Received : 2025.03.26
  • Accepted : 2025.04.06
  • Published : 2025.04.30

Abstract

This study explores the feasibility of constructing a medical question-answering (QA) system using LLama 2 7B, an open-source small large language model (sLLM), and aims to enhance its performance by applying the Retrieval-Augmented Generation (RAG) framework. The system was designed to retrieve semantically relevant documents from Wikipedia-based disease data and incorporate them into the model prompt for response generation. Experimental results showed that applying sLLM-RAG led to a notable improvement, with accuracy reaching 0.733 and BLEU score increasing to 0.500, compared to the non-RAG condition. Notably, the improvement in F1-score was largely attributed to a higher recall rate, indicating that RAG substantially improves the factual response capability of smaller models. This study offers three key contributions: First, it empirically demonstrates that small Large Language Model can deliver high-quality responses when supported by external knowledge retrieval. Second, it confirms that open data sources like Wikipedia can be effectively leveraged to build medical QA systems. Third, it presents a quantifiable and reproducible experimental framework for future RAG-based research. These findings contribute theoretically by providing evidence for sLLM-RAG integration in resource-constrained environments, and practically by offering a cost-effective and privacy-preserving alternative for healthcare and public institutions. Future work may involve expanding medical knowledge sources, supporting multi-turn dialogue, and optimizing RAG components for further system enhancement.

본 연구는 오픈소스 기반의 sLLM(small Large Language Model)인 LLaMA 2 7B를 이용하여 의료 질의응답 시스템 구축 가능성을 탐색하고, 그 성능을 향상시키기 위해 검색 증강 생성(RAG, Retrieval-Augmented Generation) 기법을 적용하는 실험을 실시하였다. 이를 위해 위키피디어의 질병 데이터를 외부 지식베이스로 활용하고, 사용자 질의에 의미 기반 검색을 한 후 관련 문서를 프롬프트에 포함하여 응답을 생성하는 구조를 실험하였다. 실험 결과, sLLM-RAG 적용 시 정확도는 0.733, BLEU 점수는 0.500으로 Non-RAG 대비 의미있는 성능 향상을 보였다. 특히 F1 점수의 향상은 재현율의 증대로 나타났으며, 이는 RAG 기법이 sLLM의 응답 능력을 실질적으로 향상시킬 수 있음을 시사한다. 본 연구는 첫째, RAG를 통해 소형 모델도 일정 수준의 응답이 가능함을 검증하였으며, 둘째, 위키피디어와 같은 오픈 데이터만으로도 의료 QA 시스템 구축이 가능함을 확인하였다. 셋째, 정량 평가 지표 기반의 실험 설계를 통해 후속 연구에 적용 가능한 RAG 구조와 분석 틀을 제시하였다. 이러한 결과는 학문적으로는 sLLM-RAG 구조의 이론 및 실증 연구 기반을 제공하며, 실무적으로는 비용 및 보안 측면에서 대형 모델 사용이 어려운 의료 및 공공기관에 효과적인 대안 기술로 활용될 수 있음을 보여준다. 향후 연구에서는 더 다양한 의료 지식베이스 확장, 멀티턴 대화 설계, RAG 구성 요소 최적화 등을 통해 시스템 고도화를 제안한다.

Keywords

Acknowledgement

이 논문은 정부(과학기술정보통신부)의 재원으로 한국연구재단의 지원을 받아 수행된 연구임(No.2022R1F1A1063537)

References

  1. Adejumo, P., Thangaraj, P., Shankar, S. V., Dhingra, L. S., Aminorroaya, A. and Khera, R. (2024). Retrieval-Augmented Generation for Extracting CHA2DS2-VASc Risk Factors from Unstructured Clinical Notes in Patients with Atrial Fibrillation, MedRxiv Preprint, https://doi.org/10.1101/2024.09.19.24313992
  2. Alsentzer, E., Murphy, J. R., Boag, W., Weng, W. H., Jin, D., Naumann, T. and McDermott, M. B. A. (2019). Publicly Available ClinicalBERT Embeddings, ArXiv Preprint, https://doi.org/10.48550/arXiv.1904.03323
  3. Amugongo, L. M., Mascheroni, P., Brooks, S. G., Doering, S. and Seidel, J. (2024). Retrieval Augmented Generation for Large Language Models in Healthcare: A Systematic Review, ArXiv Preprint, https://doi.org/10.20944/preprints202407.0876.v1
  4. Angels, B., Vinamra, B., Renato, C., Roberto, E., Hendry, T., Holstein, D., Marsman, J., Mecklenburg, N., Malvar, S., Nunes, L. O., Padilha, R., Sharp, M., Silva, B., Sharma, S., Aski, V. and Chandra, R. (2024). RAGVS Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture, ArXiv Preprint, https://doi.org/10.48550/arXiv.2401.08406
  5. Bae, J. G. (2024). A Study on the Construction of Financial-Specific Language Model Applicable to the Financial Institutions, Journal of Korea Society of Industrial Information Systems, 29(3), 79–87. https://doi.org/10.9723/JKSIIS.2024.29.3.079
  6. Choi, H. H., Cha, J. G. and Kim, S. W. (2024). LLM-Based Disease Prediction and Medical Service System for Marginalized Healthcare Groups, Proceedings of the Korean Institute of Electronics Engineers Conference, Nov. 22, Gangwon, Korea, pp. 1188-1191.
  7. Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazare, P. E., Lomeli, M., Hosseini, L. and Jegou, H. (2024). The FAISS Library, ArXiv Preprint, https://doi.org/10.48550/arXiv.2401.08281
  8. Ham, Y., Lee, S., Kim, M. and Kwak, C. (2024). Design of an Intelligent Korean Medical Service Recommendation System Based on KM-BERT and XGBoost, Proceedings of the KIIT Conference, Nov. 21-23, Jeju, Korea, pp. 1706-1709.
  9. Hao, B., Zhu, H. and Paschalidis, I. C. (2020). Enhancing Clinical BERT Embedding Using a Biomedical Knowledge Base, Proceedings of the 28th International Conference on Computational Linguistics (COLING 2020), Dec. 8-13, Barcelona, Spain, pp. 657-661.
  10. Hong, J., Ryu, E., Baek, J., Kim, S. and Oh, J. (2024). Infrastructure Proposal for the Safe Implementation of Private LLMs in SME: sLLM and Cloud Based Approach, Proceedings of the Annual Conference of the Korea Information Processing Society(KIPS), May. 23-25, Seoul, Korea, pp. 350-351.
  11. Huang, D., Hu, Z. and Wang, Z. (2024). Performance Analysis of LLaMA 2 Among Other LLMs, Proceedings of the 2024 IEEE Conference on Artificial Intelligence (CAI), Jun. 25-27, Singapore, pp. 1081-1085.
  12. Janssen, B. V., Kazemier, G. and Besselink, M. G. (2023). The Use of ChatGPT and Other Large Language Models in Surgical Science, BJS Open, 7(2), https://doi.org/10.1093/bjsopen/zrad032
  13. Jauhiainen, J. S. and Guerra, A. G. (2024). Evaluating Students' Open-Ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large, ArXiv Preprint, https://doi.org/10.48550/arXiv.2405.05444
  14. Jeon, J., Kim, K., Kim, J. and Park, S. (2024). Research on the Development of a Defense AI Platform Using a RAG-based Mil-sLLM, Proceedings of the Korean Institute of Information Scientists and Engineers Conference, Jun. 26, Jeju, Korea, pp. 44-46.
  15. Jeong, C. S. (2024). Domain-Specialized LLM: Financial Fine-Tuning and Utilization Method Using Mistral 7B, Journal of Intelligence and Information Systems, 30(1), 93–120. https://doi.org/10.13088/jiis.2024.30.1.093
  16. Jeong, J. H. (2025). A Study on the Evaluation of RAG using LLMs: Focusing on Inter-firm ESG Reports, Thesis, Graduate School of Sogang University, Seoul, Korea.
  17. Jung, H. C., Shin, K. S., Kim, H. D. and Park, S. B. (2024). Clinical Trials Utilizing LLM-Based Generative AI, Journal of the Korea Society of Computer and Information, 29(12), 169–180. https://doi.org/10.9708/jksci.2024.29.12.169
  18. Jung, J. K., Choi, S. K. and Kwon, H. C. (2023). Combining sLLM and Re-ranking Strategies for an Efficient GEC Model, Korean Institute of Information Scientists and Engineers, 2023(12), 362–364.
  19. Kang, K. S., Lee, Y. N. and Hong, A. R. (2024). Construction and Performance Analysis of an Automatic Classification Model for Domestic Academic Papers Using BERT and LLaMA2, Journal of Business Management Review, 15(4), 103–128.
  20. Ke, Y., Jin, L., Elangovan, K., Abdullah, H. R., Liu, N., Sia, A. T. H., Soh. C. R., Tung, J. Y. M., Ong, J. C. L. and Ting, D. S. W. (2024). Development and Testing of Retrieval-Augmented Generation in Large Language Models: A Case Study Report, ArXiv Preprint, https://doi.org/10.48550/arXiv.2402.01733
  21. Kim, J. Y., Kim, M. K., Hong, S. J. and Shin, J. W. (2024). Improving SLM Performance by Integrating Fine-Tuning and RAG, Proceedings of the KIIT Conference, May. 23-25, Jeju, pp. 156-158.
  22. Kim, W. S., Yu, S. C., Ju, C. Y., Kim, J. H., Ren, K. L. and Lee, D. H. (2024). A User-Friendly LLM-Based Medical Consultation System Exploiting a Dialogue Manager. Proceedings of KIISE Conference, Dec. 18-20, Yeosu, Korea, pp. 498-500.
  23. Kirchenbauer, J. and Barns, C. (2024). Hallucination Reduction in Large Language Models with Retrieval-Augmented Generation Using Wikipedia Knowledge, ArXiv Preprint, https://doi.org/10.31219/osf.io/pv7r5
  24. Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H. and Kang, J. (2020). BioBERT: A Pre-Trained Biomedical Language Representation Model for Biomedical Text Mining. Bioinformatics, ArXiv Preprint, 36(4), 1234–1240. https://doi.org/10.48550/arXiv.1901.08746
  25. Li, Y., Li, Z., Zhang, K., Dan, R., Jiang, S. and Zhang, Y. (2023). ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model (LLaMA) Using Medical Domain Knowledge, ArXiv Preprint, https://doi.org/10.48550/arXiv.2303.14070
  26. Long, C., Liu, Y., Ouyang, C. and Yu, Y. (2024). Bailicai: A Domain-Optimized Retrieval-Augmented Generation Framework for Medical Applications, ArXiv Preprint, https://doi.org/10.48550/arXiv.2407.21055
  27. Masoumi, S., Amirkhani, H., Sadeghian, N. and Shahraz, S. (2024). Natural Language Processing to Facilitate Abstract Review in Medical Research: The Application of BioBERT, Systematic Reviews, 13(1), 107. https://doi.org/10.1186/s13643-024-02470-y
  28. Moon, K., Kang, M., Yu, H. and Shin, Y. (2024). Research on the Application of Large Language Models (LLMs) in the Defense Domain, Journal of the Korea Academia-Industrial Cooperation Society, 25(8), 484–492. https://doi.org/10.5762/KAIS.2024.25.8.484
  29. Nam, Y. R., Choi, H. S., Choi, J. G., Kwon, H. J. and Lee, Y. H. (2025). Small Language Model and RAG-Based Military Maintenance Manual Curation System, Journal of Internet Computing & Services, 26(1), 181–188.
  30. Ng, K. K. Y., Matsuba, I. and Zhang, P. C. (2025). RAG in Health Care: A Novel Framework for Improving Communication and Decision-Making, NEJM AI, 2(1), https://doi.org/10.1056/AIra2400380
  31. Park, D. M. and Lee, H. J. (2024). Literature Review of AI Hallucination Research Since the Advent of ChatGPT: Focusing on Papers from ArXiv, Informatization Policy, 31(2), 3–38. https://doi.org/10.22693/NIAIP.2024.31.2.003
  32. Ranjit, M., Ganapathy, G., Manuel, R. and Ganu, T. (2023). Retrieval Augmented Chest X-ray Report Generation Using OpenAIGPT Models, Proceedings of the 8th Machine Learning for Healthcare Conference, Aug. 11-12, New York, USA, pp. 650-666.
  33. Smith, D. A. (2020). Situating Wikipedia as a Health Information Resource in Various Contexts: A Scoping Review, PLOS ONE, 15(2), https://doi.org/10.1371/journal.pone.0228786
  34. Son, Y. and Yang, G. (2024). Software Security Bug Report Template Generation and Prediction Method Using Text Similarity-Based RAG and a Smaller LLM, Proceedings of the Korean Institute of Information Scientists and Engineers Conference (KIISE), Dec. 18-20, Yeosu, Korea, pp. 387-389.
  35. Tjokro, V. C. and Sanjaya, S. A. (2024). Methods and Applications of Fine-Tuning LLaMA-2 and LLaMA-Based Models: A Systematic Literature Analysis, Journal of System and Management Sciences, 14(10), 254–266.
  36. Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Singh Koura, P., Lachaux, M. A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S. and Scialom, T. (2023). LLaMA 2: Open Foundation and Fine-Tuned Chat Models, ArXiv Preprint, https://doi.org/10.48550/arXiv.2307.09288
  37. Wang, C., Ong, J., Wang, C., Ong, H., Cheng, R. and Ong, D. (2024). Potential for GPT Technology to Optimize Future Clinical Decision-Making Using Retrieval-Augmented Generation, Annals of Biomedical Engineering, 52(5), 1115–1118. https://doi.org/10.1007/s10439-023-03327-6
  38. Woo, M., Si, J. and Kim, S. (2025). Document Summary-Based Chatbot System Utilizing Retrieval-Augmented Generation and Small Large Language Models, Journal of the Korean Institute of Information Technology,23(1), 13–20. https://doi.org/10.14801/jkiit.2025.23.1.13
  39. Xiong, G., Jin, Q., Lu, Z. and Zhang, A. (2024). Benchmarking Retrieval-Augmented Generation for Medicine, ArXiv Preprint, https://doi.org/10.48550/arXiv.2402.13178
  40. Yoon, Y. C. and Kim, S. G. (2025). Trends and Prospects of Retrieval-Augmented Generation (RAG) for Generative AI, The Journal of Korean Association of Computer Education, 28(2), 69–80. https://doi.org/10.32431/kace.2025.28.2.007
  41. Yu, Y. and Kim, H. (2023). Development of a Regulatory Q&A System for KAERI Utilizing Document Search Algorithms and Large Language Model, Journal of Korea Society of Industrial Information Systems, 28(5), 31-39.
  42. Zhang, Z., Yin, C. and Ouyang, C. (2024). A Study of Sentence Similarity Based on the All-MiniLM-L6-v2 Model with "Same Semantics, Different Structure" After Fine Tuning, Proceedings of ICIAAI 2024, pp. 677-684. https://doi.org/10.2991/978-94-6463-540-9_69