The Prediction of Survival of Breast Cancer Patients Based on Machine Learning Using Health Insurance Claim Data

Doeggyu Lee;Kyungkeun Byun;Hyungdong Lee;Sunhee Shin;

doi:10.9723/jksiis.2023.28.2.001

한국산업정보학회논문지 (Journal of Korea Society of Industrial Information Systems)

제28권2호
/
Pages.1-9
/
2023
/
1229-3741(pISSN)

한국산업정보학회 (Korea Society of Industrial Information Systems)

DOI QR Code

건강보험 청구 데이터를 활용한 머신러닝 기반유방암 환자의 생존 여부 예측

The Prediction of Survival of Breast Cancer Patients Based on Machine Learning Using Health Insurance Claim Data

이덕규 (숭실대학교 IT정책경영학과) ;
변경근 (숭실대학교 IT정책경영학과) ;
이형동 (숭실대학교 IT정책경영학과) ;
신선희 (강남대학교 교육학과)

투고 : 2023.02.12
심사 : 2023.03.13
발행 : 2023.04.30

https://doi.org/10.9723/jksiis.2023.28.2.001 인용 PDF

PDF 다운로드

⟨ 이전 논문 다음 논문 ⟩

초록

유방암 관련 기존 AI 연구는 보조적인 진단 예측이나 임상적 요인에 따른 진료 결과를 예측하는 주제가 많았다. 또한 연구기관의 코호트 자료나 일부 환자 자료를 이용하는 경우가 대부분이었다. 본 논문에서는 건강보험심사평가원이 보유하고 있는 전 국민 유방암 환자의 전수 데이터를 활용하여 유방암 환자의 40~50대와 다른 연령대 간의 생존 여부 예측과 생존 여부에 미치는 요인의 차이점을 분석했다. 그 결과, 환자들의 생존 여부 예측 정밀도는 40~50대가 평균 0.93으로 60~80대 0.86 보다 높았으며, 요인에 있어서도 40~50대는 치료횟수(46%)가, 60~80대는 나이(32%)의 변수 중요도가 제일 높았다. 기존 연구와 성능 비교 결과, 평균 정밀도가 0.90으로 기존 논문의 정밀도 0.81보다 높았다. 적용 알고리즘별 성능 비교 결과, 의사결정나무(Decision Tree), 랜덤포레스트(Random Forest) 및 그래디언트부스팅(Gradient Boosting)의 전체 평균 정밀도는 0.90, 재현율은 1.0으로 연령대 그룹 내에서 동일하였으며, 다층퍼셉트론(Multi-Layer Perceptron)의 정밀도는 0.89, 재현율은 1.0 이었다. 심평원의 전 국민 심사청구 빅데이터 가치 활용을 제고하기 위해 비전문가용 머신러닝 자동화(Auto ML) 도구를 사용한 더 많은 연구가 진행되기를 바란다.

Research using AI and big data is also being actively conducted in the health and medical fields such as disease diagnosis and treatment. Most of the existing research data used cohort data from research institutes or some patient data. In this paper, the difference in the prediction rate of survival and the factors affecting survival between breast cancer patients in their 40~50s and other age groups was revealed using health insurance review claim data held by the HIRA. As a result, the accuracy of predicting patients' survival was 0.93 on average in their 40~50s, higher than 0.86 in their 60~80s. In terms of that factor, the number of treatments was high for those in their 40~50s, and age was high for those in their 60~80s. Performance comparison with previous studies, the average precision was 0.90, which was higher than 0.81 of the existing paper. As a result of performance comparison by applied algorithm, the overall average precision of Decision Tree, Random Forest, and Gradient Boosting was 0.90, and the recall was 1.0, and the precision of multi-layer perceptrons was 0.89, and the recall was 1.0. I hope that more research will be conducted using machine learning automation(Auto ML) tools for non-professionals to enhance the use of the value for health insurance review claim data held by the HIRA.

키워드

참고문헌

Adam, Y., Constance, L., Tal, S., Tally, P. and Regina, B. (2019). A Deep Learning Mammography-based Model for Improved Breast Cancer Risk Prediction, Radiology, 292, 60-66. https://doi.org/ 10.1148/radiol.2019182716.
Ayelet, A. B. Michal, C., Yoel, S. ... Adam, S. (2019). Predicting Breast Cancer by Applying Deep Learning to Linked Health Records and Mammograms, Radiology, 292, 331-342. https://doi.org/10.1148/radiol.2019182622
Byun, K. K., Lee, D. G. and Shin, Y. T. (2022). A Study on the Prediction of Mortality Rate after Lung Cancer Diagnosis for the Elderly in their 80s and 90s Based on Deep Learning, Annual Spring Conference of KIPS, 29(1), 452-455.
Choi, H. N., Kim, M. R., Lee, J. H., Kim, H. J. and Kim J. E. (2021). AI development data research for breast cancer screening diagnosis, Proceedings of KIIT Conference 2021(11), 633-636.
Keping, Y. Liang, T. Long, L. Xiaofan, C. Zhang, Y. and Takuro, S. (2021). Deep-Learning-Empowered Breast Cancer Auxiliary Diagnosis for 5GB Remote E-Health, IEEE wireless communications, June, 54-61.
Kang, S. A., Kim, S. H., and Ryu, M. H. (2022). Analysis of Hypertension Risk Factors by Life Cycle Based on Machine Learning, Journal of the KIISR, 27(5), 73-82.
Kim, G. G., Lim, E. T. and Kim, J. H. (2022). Easily develop AI models with the AutoML platform WiseP ropet, Seoul, Cheongram.
Lee, M. S. (2020). Development of a prediction model for prognosis of triple-negative breast cancer based on deep learning using pathology images, Ph.D. Thesis, Graduate School of Ulsan University, Ulsan.
Ministry of Heath and Welfar. (2021). Cancer regstry statistics, Sejong
National Cancer Center. (2021). Cancer Monitoring Indicator.
Ryu, J. H. Hong, S. H. Park, H. G. Kim, D. M. Kim, S. J. and Park, S. J. (2017). Application of Machine Learning Techniques to Predict Stroke Diseases in Older Adults, The Spring Conference of KOSES, 42-42.
The Yakup. (2022). Analysis by age by cancer type in 2019, https://www.yakup.com (Accessed on March. 11th, 2022)
Yun, S. O., Jung, J. G., Wo, H. G. and Kim, J. E. (2022). Breast Cancer Survival Prediction: Model Comparison and Effect of Genetic Features, Database Research, 38(1), 3-15.

한국산업정보학회논문지 (Journal of Korea Society of Industrial Information Systems)

건강보험 청구 데이터를 활용한 머신러닝 기반유방암 환자의 생존 여부 예측

The Prediction of Survival of Breast Cancer Patients Based on Machine Learning Using Health Insurance Claim Data

초록

키워드

참고문헌

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)