• 제목/요약/키워드: Research dataset

검색결과 1,353건 처리시간 0.028초

Prediction Model for Gastric Cancer via Class Balancing Techniques

  • Danish, Jamil ;Sellappan, Palaniappan;Sanjoy Kumar, Debnath;Muhammad, Naseem;Susama, Bagchi ;Asiah, Lokman
    • International Journal of Computer Science & Network Security
    • /
    • 제23권1호
    • /
    • pp.53-63
    • /
    • 2023
  • Many researchers are trying hard to minimize the incidence of cancers, mainly Gastric Cancer (GC). For GC, the five-year survival rate is generally 5-25%, but for Early Gastric Cancer (EGC), it is almost 90%. Predicting the onset of stomach cancer based on risk factors will allow for an early diagnosis and more effective treatment. Although there are several models for predicting stomach cancer, most of these models are based on unbalanced datasets, which favours the majority class. However, it is imperative to correctly identify cancer patients who are in the minority class. This research aims to apply three class-balancing approaches to the NHS dataset before developing supervised learning strategies: Oversampling (Synthetic Minority Oversampling Technique or SMOTE), Undersampling (SpreadSubsample), and Hybrid System (SMOTE + SpreadSubsample). This study uses Naive Bayes, Bayesian Network, Random Forest, and Decision Tree (C4.5) methods. We measured these classifiers' efficacy using their Receiver Operating Characteristics (ROC) curves, sensitivity, and specificity. The validation data was used to test several ways of balancing the classifiers. The final prediction model was built on the one that did the best overall.

Proposing new models to predict pile set-up in cohesive soils

  • Sara Banaei Moghadam;Mohammadreza Khanmohammadi
    • Geomechanics and Engineering
    • /
    • 제33권3호
    • /
    • pp.231-242
    • /
    • 2023
  • This paper represents a comparative study in which Gene Expression Programming (GEP), Group Method of Data Handling (GMDH), and multiple linear regressions (MLR) were utilized to derive new equations for the prediction of time-dependent bearing capacity of pile foundations driven in cohesive soil, technically called pile set-up. This term means that many piles which are installed in cohesive soil experience a noticeable increase in bearing capacity after a specific time. Results of researches indicate that side resistance encounters more increase than toe resistance. The main reason leading to pile setup in saturated soil has been found to be the dissipation of excess pore water pressure generated in the process of pile installation, while in unsaturated conditions aging is the major justification. In this study, a comprehensive dataset containing information about 169 test piles was obtained from literature reviews used to develop the models. to prepare the data for further developments using intelligent algorithms, Data mining techniques were performed as a fundamental stage of the study. To verify the models, the data were randomly divided into training and testing datasets. The most striking difference between this study and the previous researches is that the dataset used in this study includes different piles driven in soil with varied geotechnical characterization; therefore, the proposed equations are more generalizable. According to the evaluation criteria, GEP was found to be the most effective method to predict set-up among the other approaches developed earlier for the pertinent research.

Identification of Combined Biomarker for Predicting Alzheimer's Disease Using Machine Learning

  • Ki-Yeol Kim
    • 생물정신의학
    • /
    • 제30권1호
    • /
    • pp.24-30
    • /
    • 2023
  • Objectives Alzheimer's disease (AD) is the most common form of dementia in older adults, damaging the brain and resulting in impaired memory, thinking, and behavior. The identification of differentially expressed genes and related pathways among affected brain regions can provide more information on the mechanisms of AD. The aim of our study was to identify differentially expressed genes associated with AD and combined biomarkers among them to improve AD risk prediction accuracy. Methods Machine learning methods were used to compare the performance of the identified combined biomarkers. In this study, three publicly available gene expression datasets from the hippocampal brain region were used. Results We detected 31 significant common genes from two different microarray datasets using the limma package. Some of them belonged to 11 biological pathways. Combined biomarkers were identified in two microarray datasets and were evaluated in a different dataset. The performance of the predictive models using the combined biomarkers was superior to those of models using a single gene. When two genes were combined, the most predictive gene set in the evaluation dataset was ATR and PRKCB when linear discriminant analysis was applied. Conclusions Combined biomarkers showed good performance in predicting the risk of AD. The constructed predictive nomogram using combined biomarkers could easily be used by clinicians to identify high-risk individuals so that more efficient trials could be designed to reduce the incidence of AD.

초거대 언어 모델로부터의 추론 데이터셋을 활용한 감정 분류 성능 향상 (Empowering Emotion Classification Performance Through Reasoning Dataset From Large-scale Language Model)

  • 박눈솔;이민호
    • 한국컴퓨터정보학회:학술대회논문집
    • /
    • 한국컴퓨터정보학회 2023년도 제68차 하계학술대회논문집 31권2호
    • /
    • pp.59-61
    • /
    • 2023
  • 본 논문에서는 감정 분류 성능 향상을 위한 초거대 언어모델로부터의 추론 데이터셋 활용 방안을 제안한다. 이 방안은 Google Research의 'Chain of Thought'에서 영감을 받아 이를 적용하였으며, 추론 데이터는 ChatGPT와 같은 초거대 언어 모델로 생성하였다. 본 논문의 목표는 머신러닝 모델이 추론 데이터를 이해하고 적용하는 능력을 활용하여, 감정 분류 작업의 성능을 향상시키는 것이다. 초거대 언어 모델(ChatGPT)로부터 추출한 추론 데이터셋을 활용하여 감정 분류 모델을 훈련하였으며, 이 모델은 감정 분류 작업에서 향상된 성능을 보였다. 이를 통해 추론 데이터셋이 감정 분류에 있어서 큰 가치를 가질 수 있음을 증명하였다. 또한, 이 연구는 기존에 감정 분류 작업에 사용되던 데이터셋만을 활용한 모델과 비교하였을 때, 추론 데이터를 활용한 모델이 더 높은 성능을 보였음을 증명한다. 이 연구를 통해, 적은 비용으로 초거대 언어모델로부터 생성된 추론 데이터셋의 활용 가능성을 보여주고, 감정 분류 작업 성능을 향상시키는 새로운 방법을 제시한다. 제시한 방안은 감정 분류뿐만 아니라 다른 자연어처리 분야에서도 활용될 수 있으며, 더욱 정교한 자연어 이해와 처리가 가능함을 시사한다.

  • PDF

재난안전관리를 위한 디지털 트윈 데이터셋 구조 연구 (A Study on the Dataset Structure of Digital Twin for Disaster and Safety Management)

  • 정기숙;정우석
    • 한국인터넷방송통신학회논문지
    • /
    • 제23권5호
    • /
    • pp.89-95
    • /
    • 2023
  • 지하공동구는 도시의 상하수, 전력, 통신 등과 같은 중요한 시설을 수용하여 관리하는 도시기반시설로 화재, 지진, 침수 등과 같은 재난으로부터 보호해야 하는 국가 시설이다. 예측, 예방, 대비, 대응, 복구 등의 재난안전 전주기 관리 체계를 구축함에 있어서 첨단 ICT 기술과 데이터가 융합된 디지털 트윈 기술을 활용하여 지하공동구의 재난안전관리 플랫폼을 개발 중에 있다. 이 논문에서는 재난안전 디지털 트윈에 대한 성숙도 모델을 살펴보고 각 성숙도 단계 별 재난안전 디지털 트윈 구현을 위해 필요한 데이터셋을 정의하였다. 정의된 데이터셋의 카테고리의 조합에 따라 성숙도 단계를 다르게 하여 디지털 트윈 구현이 가능하도록 구성하였다.

A Green Logistics Network Design to Increase Responsiveness to Eco-Friendly Consumers

  • Eungoo KANG
    • 산경연구논집
    • /
    • 제14권11호
    • /
    • pp.1-9
    • /
    • 2023
  • Purpose: The industrial sector, especially in developed countries, is seen as the primary threat to sustainability. As a result, contemporary organizations prioritize establishing sustainable business practices. This sustainability can be achieved by organizations being concerned with their external environments, which is referred to as going green. This study aims to provide a green logistics network design to explain how to attract green consumers. Research design, data and methodology: This study conducted a comprehensive process to obtain textual dataset in the current literature and finally the author could collect total 26 relevant prior studies to achieve the purpose of the study. All dataset was thoroughly screened and selected for the high-degree of validity. Results: Based on the intensive literature review, the author insists that the four findings presented in this study will be useful as they provide evidence of the importance of technology in achieving global sustainability.in the situation we face that technology has become an important part of human life. Conclusions: This study provides meaningful insights into the environmental strategies that organizations across the world can implement to achieve a green supply chain based on the solutions in this study. The strategies presented in this study are evidence-based and have been tested through different studies.

Human hand gesture identification framework using SIFT and knowledge-level technique

  • Muhammad Haroon;Saud Altaf;Zia-ur- Rehman;Muhammad Waseem Soomro;Sofia Iqbal
    • ETRI Journal
    • /
    • 제45권6호
    • /
    • pp.1022-1034
    • /
    • 2023
  • In this study, the impact of varying lighting conditions on recognition and decision-making was considered. The luminosity approach was presented to increase gesture recognition performance under varied lighting. An efficient framework was proposed for sensor-based sign language gesture identification, including picture acquisition, preparing data, obtaining features, and recognition. The depth images were collected using multiple Microsoft Kinect devices, and data were acquired by varying resolutions to demonstrate the idea. A case study was designed to attain acceptable accuracy in gesture recognition under variant lighting. Using American Sign Language (ASL), the dataset was created and analyzed under various lighting conditions. In ASL-based images, significant feature points were selected using the scale-invariant feature transformation (SIFT). Finally, an artificial neural network (ANN) classified hand gestures using specified characteristics for validation. The suggested method was successful across a variety of illumination conditions and different image sizes. The total effectiveness of NN architecture was shown by the 97.6% recognition accuracy rate of 26 alphabets dataset with just a 2.4% error rate.

Pile bearing capacity prediction in cold regions using a combination of ANN with metaheuristic algorithms

  • Zhou Jingting;Hossein Moayedi;Marieh Fatahizadeh;Narges Varamini
    • Steel and Composite Structures
    • /
    • 제51권4호
    • /
    • pp.417-440
    • /
    • 2024
  • Artificial neural networks (ANN) have been the focus of several studies when it comes to evaluating the pile's bearing capacity. Nonetheless, the principal drawbacks of employing this method are the sluggish rate of convergence and the constraints of ANN in locating global minima. The current work aimed to build four ANN-based prediction models enhanced with methods from the black hole algorithm (BHA), league championship algorithm (LCA), shuffled complex evolution (SCE), and symbiotic organisms search (SOS) to estimate the carrying capacity of piles in cold climates. To provide the crucial dataset required to build the model, fifty-eight concrete pile experiments were conducted. The pile geometrical properties, internal friction angle 𝛗 shaft, internal friction angle 𝛗 tip, pile length, pile area, and vertical effective stress were established as the network inputs, and the BHA, LCA, SCE, and SOS-based ANN models were set up to provide the pile bearing capacity as the output. Following a sensitivity analysis to determine the optimal BHA, LCA, SCE, and SOS parameters and a train and test procedure to determine the optimal network architecture or the number of hidden nodes, the best prediction approach was selected. The outcomes show a good agreement between the measured bearing capabilities and the pile bearing capacities forecasted by SCE-MLP. The testing dataset's respective mean square error and coefficient of determination, which are 0.91846 and 391.1539, indicate that using the SCE-MLP approach as a practical, efficient, and highly reliable technique to forecast the pile's bearing capacity is advantageous.

Exploring Public Opinion to Analyze the Consequences of Social Media on Students' Behaviors

  • Asif Nawaz;Tariq Ali;Saif Ur Rehman;Yaser Hafeez
    • International Journal of Computer Science & Network Security
    • /
    • 제24권8호
    • /
    • pp.159-168
    • /
    • 2024
  • Social media sites like as twitter, Facebook and flicker widely used by people, not only as a source of distributing information but also as for communication purpose, with the advancement of technology today. Now a day's one of the most frequently used communication methods are social networks. In various research studies, their use in different fields and the effects of social media on student's behaviors, chat sites and blogs caused by Facebook has been analyzed. In order to obtain the basic data, a general scanning model that is public opinion and views of parents and comments that are openly available across social media sites, used to perceive attitude of graduate students, instead of traditional methods like questionnaires and survey's conduction. A dataset of nearly 20000 reviews of parents was collected from different social media networks about their children's, while in another dataset in which 362 graduate school teachers who observe the students to use social media during classes, labs and in campus during free times, their comments about those students were chosen. As per this study, through different positive and negative factors the detailed analysis has been performed to show effect of social media on student's behavior.

다중 클래스 데이터셋의 메타특징이 판별 알고리즘의 성능에 미치는 영향 연구 (The Effect of Meta-Features of Multiclass Datasets on the Performance of Classification Algorithms)

  • 김정훈;김민용;권오병
    • 지능정보연구
    • /
    • 제26권1호
    • /
    • pp.23-45
    • /
    • 2020
  • 기업의 경쟁력 확보를 위해 판별 알고리즘을 활용한 의사결정 역량제고가 필요하다. 하지만 대부분 특정 문제영역에는 적합한 판별 알고리즘이 어떤 것인지에 대한 지식은 많지 않아 대부분 시행착오 형식으로 최적 알고리즘을 탐색한다. 즉, 데이터셋의 특성에 따라 어떠한 분류알고리즘을 채택하는 것이 적합한지를 판단하는 것은 전문성과 노력이 소요되는 과업이었다. 이는 메타특징(Meta-Feature)으로 불리는 데이터셋의 특성과 판별 알고리즘 성능과의 연관성에 대한 연구가 아직 충분히 이루어지지 않았기 때문이며, 더구나 다중 클래스(Multi-Class)의 특성을 반영하는 메타특징에 대한 연구 또한 거의 이루어진 바 없다. 이에 본 연구의 목적은 다중 클래스 데이터셋의 메타특징이 판별 알고리즘의 성능에 유의한 영향을 미치는지에 대한 실증 분석을 하는 것이다. 이를 위해 본 연구에서는 다중 클래스 데이터셋의 메타특징을 데이터셋의 구조와 데이터셋의 복잡도라는 두 요인으로 분류하고, 그 안에서 총 7가지 대표 메타특징을 선택하였다. 또한, 본 연구에서는 기존 연구에서 사용하던 IR(Imbalanced Ratio) 대신 시장집중도 측정 지표인 허핀달-허쉬만 지수(Herfindahl-Hirschman Index, HHI)를 메타특징에 포함하였으며, 역ReLU 실루엣 점수(Reverse ReLU Silhouette Score)도 새롭게 제안하였다. UCI Machine Learning Repository에서 제공하는 복수의 벤치마크 데이터셋으로 다양한 변환 데이터셋을 생성한 후에 대표적인 여러 판별 알고리즘에 적용하여 성능 비교 및 가설 검증을 수행하였다. 그 결과 대부분의 메타특징과 판별 성능 사이의 유의한 관련성이 확인되었으며, 일부 예외적인 부분에 대한 고찰을 하였다. 본 연구의 실험 결과는 향후 메타특징에 따른 분류알고리즘 추천 시스템에 활용할 것이다.