• Title/Summary/Keyword: Variable Clustering

Search Result 155, Processing Time 0.026 seconds

A Method of Object Identification from Procedural Programs (절차적 프로그램으로부터의 객체 추출 방법론)

  • Jin, Yun-Suk;Ma, Pyeong-Su;Sin, Gyu-Sang
    • The Transactions of the Korea Information Processing Society
    • /
    • v.6 no.10
    • /
    • pp.2693-2706
    • /
    • 1999
  • Reengineering to object-oriented system is needed to maintain the system and satisfy requirements of structure change. Target systems which should be reengineered to object-oriented system are difficult to change because these systems have no design document or their design document is inconsistent of source code. Using design document to identifying objects for these systems is improper. There are several researches which identify objects through procedural source code analysis. In this paper, we propose automatic object identification method based on clustering of VTFG(Variable-Type-Function Graph) which represents relations among variables, types, and functions. VTFG includes relations among variables, types, and functions that may be basis of objects, and weights of these relations. By clustering related variables, types, and functions using their weights, our method overcomes limit of existing researches which identify too big objects or objects excluding many functions. The method proposed in this paper minimizes user's interaction through automatic object identification and make it easy to reenginner procedural system to object-oriented system.

  • PDF

H.264/AVC to MPEG-2 Video Transcoding by using Motion Vector Clustering (움직임벡터 군집화를 이용한 H.264/AVC에서 MPEG-2로의 비디오 트랜스코딩)

  • Shin, Yoon-Jeong;Son, Nam-Rye;Nguyen, Dinh Toan;Lee, Guee-Sang
    • The Journal of the Korea institute of electronic communication sciences
    • /
    • v.5 no.1
    • /
    • pp.23-30
    • /
    • 2010
  • The H.264/AVC is increasingly used in broadcast video applications such as Internet Protocol television (IPTV), digital multimedia broadcasting (DMB) because of high compression performance. But the H.264/AVC coded video can be delivered to the widespread end-user equipment for MPEG-2 after transcoding between this video standards. This paper suggests a new transcoding algorithm for H.264/AVC to MPEG-2 transcoder that uses motion vector clustering in order to reduce the complexity without loss of video quality. The proposed method is exploiting the motion information gathered during h.264 decoding stage. To reduce the search space for the MPEG-2 motion estimation, the predictive motion vector is selected with a least distortion of the candidated motion vectors. These candidate motion vectors are considering the correlation of direction and distance of motion vectors of variable blocks in H.264/AVC. And then the best predictive motion vector is refined with full-search in ${\pm}2$ pixel search area. Compared with a cascaded decoder-encoder, the proposed transcoder achieves computational complexity savings up to 64% with a similar PSNR at the constant bitrate(CBR).

High-performance computing for SARS-CoV-2 RNAs clustering: a data science-based genomics approach

  • Oujja, Anas;Abid, Mohamed Riduan;Boumhidi, Jaouad;Bourhnane, Safae;Mourhir, Asmaa;Merchant, Fatima;Benhaddou, Driss
    • Genomics & Informatics
    • /
    • v.19 no.4
    • /
    • pp.49.1-49.11
    • /
    • 2021
  • Nowadays, Genomic data constitutes one of the fastest growing datasets in the world. As of 2025, it is supposed to become the fourth largest source of Big Data, and thus mandating adequate high-performance computing (HPC) platform for processing. With the latest unprecedented and unpredictable mutations in severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the research community is in crucial need for ICT tools to process SARS-CoV-2 RNA data, e.g., by classifying it (i.e., clustering) and thus assisting in tracking virus mutations and predict future ones. In this paper, we are presenting an HPC-based SARS-CoV-2 RNAs clustering tool. We are adopting a data science approach, from data collection, through analysis, to visualization. In the analysis step, we present how our clustering approach leverages on HPC and the longest common subsequence (LCS) algorithm. The approach uses the Hadoop MapReduce programming paradigm and adapts the LCS algorithm in order to efficiently compute the length of the LCS for each pair of SARS-CoV-2 RNA sequences. The latter are extracted from the U.S. National Center for Biotechnology Information (NCBI) Virus repository. The computed LCS lengths are used to measure the dissimilarities between RNA sequences in order to work out existing clusters. In addition to that, we present a comparative study of the LCS algorithm performance based on variable workloads and different numbers of Hadoop worker nodes.

An Empirical Comparative Study on the Clustering Measurement Using Fuzzy(Average Index Transformation) DEA and Cross-efficiency Models (퍼지(평균지수변환)DEA모형과 교차효율성모형을 이용한 클러스터링측정에 대한 실증적 비교연구)

  • Park, Ro-Kyung
    • Journal of Korea Port Economic Association
    • /
    • v.31 no.1
    • /
    • pp.85-110
    • /
    • 2015
  • The purpose of this paper is to show the clustering trend and the empirical comparison and to choose the clustering ports for 3 Korean ports(Busan, Incheon and Gwangyang Ports) by using the Fuzzy(Average Index Transformation) DEA and Cross-efficiency models for 38 Asian ports during 11 years(2001-2011) with 4 input variables(birth length, depth, total area, and number of crane) and 1 output variable(container TEU). The main empirical results of this paper are as follows. First, clustering results by using Fuzzy(AIT)DEA show that 3 Korean ports[Busan(56.29%), Incheon(57.96%), and Gwangyang(66.80%) each]can increase the efficiency. Second, according to Cross-efficiency model, Busan(Hongkong, Kobe, Manila, Singapore, and Kaosiung etc.), Incheon(Aquaba, Dammam, Karachi, Mohammad Byin Oasim and Davao), and Gwangyang(Damman, Yokohama, Nogoya, Keelong, Kaosiung, and Bangkok) should be clustered with those ports in parentheses. Third, when both Fuzzy(AIT)DEA and Cross-efficiency models are mixed, the empirical result shows that 3 Korean ports[Busan(71.38%), Incheon(103.89%), and Gwangyang(168.55%) each]can increase the efficiency. The efficiency ranking comparison among the three models by using Wilcoxon Signed-rank Test was matched with the average level of 66%-67%. The policy implication of this paper is that Korean port policy planner should introduce the Fuzzy(AIT)DEA, and Cross-efficiency models with the mixed two models when clustering is needed among the Asian ports for enhancing the efficiency of inputs and outputs. Also, the results of SWOT analysis among the clustering ports should be considered.

Spatial analysis of water shortage areas in South Korea considering spatial clustering characteristics (공간군집특성을 고려한 우리나라 물부족 핫스팟 지역 분석)

  • Lee, Dong Jin;Kim, Tae-Woong
    • Journal of Korea Water Resources Association
    • /
    • v.57 no.2
    • /
    • pp.87-97
    • /
    • 2024
  • This study analyzed the water shortage hotspot areas in South Korea using spatial clustering analysis for water shortage estimates in 2030 of the Master Plans for National Water Management. To identify the water shortage cluster areas, we used water shortage data from the past maximum drought (about 50-year return period) and performed spatial clustering analysis using Local Moran's I and Getis-Ord Gi*. The areas subject to spatial clusters of water shortage were selected using the cluster map, and the spatial characteristics of water shortage areas were verified based on the p-value and the Moran scatter plot. The results indicated that one cluster (lower Imjin River (#1023) and neighbor) in the Han River basin and two clusters (Daejeongcheon (#2403) and neighbor, Gahwacheon (#2501) and neighbor) in the Nakdong River basin were found to be the hotspot for water shortage, whereas one cluster (lower Namhan River (#1007) and neighbor) in the Han River Basin and one cluster (Byeongseongcheon (#2006) and neighbor) in the Nakdong River basin were found to be the HL area, which means the specific area have high water shortage and neighbor have low water shortage. When analyzing spatial clustering by standard watershed unit, the entire spatial clustering area satisfied 100% of the statistical criteria leading to statistically significant results. The overall results indicated that spatial clustering analysis performed using standard watersheds can resolve the variable spatial unit problem to some extent, which results in the relatively increased accuracy of spatial analysis.

Classification Methods for Fault Diagnosis of an Air Handling Unit (공조 시스템의 고장진단을 위한 분류기술 연구)

  • Lee, Won-Yong;Shin, Dong-Ryul;House, John M.
    • Proceedings of the KIEE Conference
    • /
    • 1998.07b
    • /
    • pp.420-422
    • /
    • 1998
  • All Fault Detection and Diagnosis(FDD) methods utilize classification techniques. The objective of this study was to demonstrate the application of classification techniques to the problem of diagnosing faults in data generated by a variable-air-volume(VAV) air-handling unit(AHU) simulation model and to describe the characteristics of the techniques considered. Artificial neural network classifier and fuzzy clustering classifier were considered for fault diagnostics.

  • PDF

A Study on the Variable Vocabulary Speech Recognition in the Vocabulary-Independent Environments (어휘독립 환경에서의 가변어휘 음성인식에 관한 연구)

  • 황병한
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • 1998.06e
    • /
    • pp.369-372
    • /
    • 1998
  • 본 논문은 어휘독립(Vocabulary-Independent) 환경에서 별도의 훈련과정 없이 인식대상 어휘를 추가 및 변경할 수 있는 가변어휘(Variable Vocabulary) 음성인식에 관한 연구를 다룬다. 가변어휘 인식은 처음에 대용량 음성 데이터베이스(DB)로 음소모델을 훈련하고 인식대상 어휘가 결정되면 발음사전에 의거하여 음소모델을 연결함으로써 별도의 훈련과정 없이 인식대상 어휘를 변경 및 추가할 수 있다. 문맥 종속형(Context-Dependent) 음소 모델인 triphone을 사용하여 인식실험을 하였고, 인식성능의 비교를 위해 어휘종속 모델을 별도로 구성하여 인식실험을 하였다. Unseen triphone 문제와 훈련 DB의 부족으로 인한 모델 파라메터의 신뢰성 저하를 방지하기 위해 state-tying 방법 중 음성학적 지식에 기반을 둔 tree-based clustering(TBC) 기법[1]을 도입하였다. Mel Frequency Cepstrum Coefficient(MFCC)와 대수에너지에 기반을 둔 3 가지 음성특징 벡터를 사용하여 인식 실험을 병행하였고, 연속 확률분포를 가지는 Hidden Markov Model(HMM) 기반의 고립단어 인식시스템을 구현하였다. 인식 실험에는 22 개 부서명 DB[3]를 사용하였다. 실험결과 어휘독립 환경에서 최고 98.4%의 인식률이 얻어졌으며, 어휘종속 환경에서의 인식률 99.7%에 근접한 성능을 보였다.

  • PDF

Visualizing Multi-Variable Prediction Functions by Segmented k-CPG's

  • Huh, Myung-Hoe
    • Communications for Statistical Applications and Methods
    • /
    • v.16 no.1
    • /
    • pp.185-193
    • /
    • 2009
  • Machine learning methods such as support vector machines and random forests yield nonparametric prediction functions of the form y = $f(x_1,{\ldots},x_p)$. As a sequel to the previous article (Huh and Lee, 2008) for visualizing nonparametric functions, I propose more sensible graphs for visualizing y = $f(x_1,{\ldots},x_p)$ herein which has two clear advantages over the previous simple graphs. New graphs will show a small number of prototype curves of $f(x_1,{\ldots},x_{j-1},x_j,x_{j+1}{\ldots},x_p)$, revealing statistically plausible portion over the interval of $x_j$ which changes with ($x_1,{\ldots},x_{j-1},x_{j+1},{\ldots},x_p$). To complement the visual display, matching importance measures for each of p predictor variables are produced. The proposed graphs and importance measures are validated in simulated settings and demonstrated for an environmental study.

Analysis of Virus Types by a Latent Variable Model (Latent variable model에 의한 바이러스 유형 분석)

  • Kim Soo-Jin;Joung Je-Gun;Tae Kang Soo;Zhang Byoung-Tak
    • Proceedings of the Korean Information Science Society Conference
    • /
    • 2005.11b
    • /
    • pp.262-264
    • /
    • 2005
  • 인유두종 바이러스(Human Papillomavirus: HPV)는 사마귀로부터 생식기 및 배설기의 침윤성 암에 이르기까지 여러 질병과 연관되어 있음이 알려져 있다. 현재 200종 이상이 알려져 있고, 이 중 85개는 전체 유전자가 밝혀져 있다. HPV 감염 시 만들어지는 단백질 중 E6. E7 단백질은 암 억제 유전자(p53, pRb)에 결합하여 세포의 암 억제 기능을 저하시키고 이로 인해 암을 발생시킨다. 본 논문은 암 발생과 밀접한 관련이 있는 HPV의 E6 단백질 서열과 HPV 유형(HPV Type)을 가지고, PLSA (Probabilistic Latent Semantic Analysis) 방법을 이용하여 HPV를 클러스터링(clustering) 해 보았다. 실험 결과, 특정 클러스터는 질병과 밀접하게 연관되어 있으며, 이와 관련된 주요 서열 분석이 가능함을 보여주고 있다.

  • PDF

A study on decision tree creation using intervening variable (매개 변수를 이용한 의사결정나무 생성에 관한 연구)

  • Cho, Kwang-Hyun;Park, Hee-Chang
    • Journal of the Korean Data and Information Science Society
    • /
    • v.22 no.4
    • /
    • pp.671-678
    • /
    • 2011
  • Data mining searches for interesting relationships among items in a given database. The methods of data mining are decision tree, association rules, clustering, neural network and so on. The decision tree approach is most useful in classification problems and to divide the search space into rectangular regions. Decision tree algorithms are used extensively for data mining in many domains such as retail target marketing, customer classification, etc. When create decision tree model, complicated model by standard of model creation and number of input variable is produced. Specially, there is difficulty in model creation and analysis in case of there are a lot of numbers of input variable. In this study, we study on decision tree using intervening variable. We apply to actuality data to suggest method that remove unnecessary input variable for created model and search the efficiency.