Search | Korea Science

Improving the Performance of Document Clustering with Distributional Similarities (분포유사도를 이용한 문헌클러스터링의 성능향상에 대한 연구)

Lee, Jae-Yun
- Journal of the Korean Society for information Management
- /
- v.24 no.4
- /
- pp.267-283
- /
- 2007
In this study, measures of distributional similarity such as KL-divergence are applied to cluster documents instead of traditional cosine measure, which is the most prevalent vector similarity measure for document clustering. Three variations of KL-divergence are investigated; Jansen-Shannon divergence, symmetric skew divergence, and minimum skew divergence. In order to verify the contribution of distributional similarities to document clustering, two experiments are designed and carried out on three test collections. In the first experiment the clustering performances of the three divergence measures are compared to that of cosine measure. The result showed that minimum skew divergence outperformed the other divergence measures as well as cosine measure. In the second experiment second-order distributional similarities are calculated with Pearson correlation coefficient from the first-order similarity matrixes. From the result of the second experiment, secondorder distributional similarities were found to improve the overall performance of document clustering. These results suggest that minimum skew divergence must be selected as document vector similarity measure when considering both time and accuracy, and second-order similarity is a good choice for considering clustering accuracy only.
https://doi.org/10.3743/KOSIM.2007.24.4.267 인용 PDF

A Study on Keyword Extraction From a Single Document Using Term Clustering (용어 클러스터링을 이용한 단일문서 키워드 추출에 관한 연구)

Han, Seung-Hee
- Journal of the Korean Society for Library and Information Science
- /
- v.44 no.3
- /
- pp.155-173
- /
- 2010
In this study, a new keyword extraction algorithm is applied to a single document with term clustering. A single document is divided by multiple passages, and two ways of calculating similarities between two terms are investigated; the first-order similarity and the second-order distributional similarity. In this experiment, the best cluster performance is achieved with a 50-term passage from the second-order distributional similarity. From the results of first experiment, the second-order distribution similarity was also applied to various keyword extraction methods using statistic information of terms. In the second experiment, pf(paragraph frequency) and $tf{\times}ipf$(term frequency by inverse paragraph frequency) were found to improve the overall performance of keyword extraction. Therefore, it showed that the algorithm fulfills the necessary conditions which good keywords should have.
https://doi.org/10.4275/KSLIS.2010.44.3.155 인용 PDF

Image Data Classification using a Similarity Function based on Second Order Tensor (2차 텐서 기반 유사도 함수를 이용한 영상 데이터 분류)

Yoon, Dong-Woo;Lee, Kwan-Yong;Park, Hye-Young
- Journal of KIISE:Software and Applications
- /
- v.36 no.8
- /
- pp.664-672
- /
- 2009
Recently, studies on utilizing tensor expression on image data analysis and processing have been attracting much interest. The purpose of this study is to develop an efficient system for classifying image patterns by using second order tensor expression. To achieve the goal, we propose a data generation model expressed by class factors and environment factors with second order tensor representation. Based on the data generation model, we define a function for measuring similarities between two images. The similarity function is obtained by estimating the probability density of environment factors using a matrix normal distribution. Through computational experiments on a number of benchmark data sets, we confirm that we can make improvement in classification rates by using second order tensor, and that the proposed similarity function is more appropriate for image data compared to conventional similarity measures.
PDF KSCI

SYMMETRY REDUCTIONS, VARIABLE TRANSFORMATIONS AND EXACT SOLUTIONS TO THE SECOND-ORDER PDES

Liu, Hanze;Liu, Lei
- Journal of applied mathematics & informatics
- /
- v.29 no.3_4
- /
- pp.563-572
- /
- 2011
In this paper, the Lie symmetry analysis is performed on the three mixed second-order PDEs, which arise in fluid dynamics, nonlinear wave theory and plasma physics, etc. The symmetries and similarity reductions of the equations are obtained, and the exact solutions to the equations are investigated by the dynamical system and power series methods. Then, the exact solutions to the general types of PDEs are considered through a variable transformation. At last, the symmetry and integration method is employed for reducing the nonlinear ODEs.
https://doi.org/10.14317/jami.2011.29.3_4.563 인용 PDF KSCI

The Relationships Between Children's Perceptions Toward Grandparents and Their Intimate Behavior

Jung, Min-Suk;Ko, Eun-Kyo;Rho, Joseph Y.;Lee, Seung-Hyun
- International Journal of Contents
- /
- v.5 no.4
- /
- pp.19-29
- /
- 2009
This study is focused on the causal relationship between children's intimate behavior and the level of perception towards their grandparents. Their perceptions are related to factors such as proximity, similarity, superiority, favorableness, and self-disclosure. We clarified the relation between intimate behavior and perception using effect factors of children's behavior regarding their grandparents so that this study could be used as an elementary material in developing a solution to improve grandparent-grandchild relationship where the grandparent actively encourages grandchildren's intimate behavior. Regression analysis was used as a hypothesis testing method. The results indicated the following three points. First, perception factors affect active intimate behavior in the order of favorableness, superiority, self-disclosure, and similarity. Second, perception factors affect intimate behavior will in the order of favorableness, superiority, and self-disclosure. Lastly, it was shown that a child's active intimate behavior has an influence on their intimate behavior will.
https://doi.org/10.5392/IJoC.2009.5.4.019 인용 PDF

Heat and mass transfer of a second grade magnetohydrodynamic fluid over a convectively heated stretching sheet

Das, Kalidas;Sharma, Ram Prakash;Sarkar, Amit
- Journal of Computational Design and Engineering
- /
- v.3 no.4
- /
- pp.330-336
- /
- 2016
The present work is concerned with heat and mass transfer of an electrically conducting second grade MHD fluid past a semi-infinite stretching sheet with convective surface heat flux. The analysis accounts for thermophoresis and thermal radiation. A similarity transformations is used to reduce the governing equations into a dimensionless form. The local similarity equations are derived and solved using Nachtsheim-Swigert shooting iteration technique together with Runge-Kutta sixth order integration scheme. Results for various flow characteristics are presented through graphs and tables delineating the effect of various parameters characterizing the flow. Our analysis explores that the rate of heat transfer enhances with increasing the values of the surface convection parameter. Also the fluid velocity and temperature in the boundary layer region rise significantly for increasing the values of thermal radiation parameter.
https://doi.org/10.1016/j.jcde.2016.06.001 인용 PDF KSCI

The Strength of the Relationship between Semantic Similarity and the Subcategorization Frames of the English Verbs: a Stochastic Test based on the ICE-GB and WordNet (영어 동사의 의미적 유사도와 논항 선택 사이의 연관성 : ICE-GB와 WordNet을 이용한 통계적 검증)

Song, Sang-Houn;Choe, Jae-Woong
- Language and Information
- /
- v.14 no.1
- /
- pp.113-144
- /
- 2010
The primary goal of this paper is to find a feasible way to answer the question: Does the similarity in meaning between verbs relate to the similarity in their subcategorization? In order to answer this question in a rather concrete way on the basis of a large set of English verbs, this study made use of various language resources, tools, and statistical methodologies. We first compiled a list of 678 verbs that were selected from the most and second most frequent word lists from the Colins Cobuild English Dictionary, which also appeared in WordNet 3.0. We calculated similarity measures between all the pairs of the words based on the 'jcn' algorithm (Jiang and Conrath, 1997) implemented in the WordNet::Similarity module (Pedersen, Patwardhan, and Michelizzi, 2004). The clustering process followed, first building similarity matrices out of the similarity measure values, next drawing dendrograms on the basis of the matricies, then finally getting 177 meaningful clusters (covering 437 verbs) that passed a certain level set by z-score. The subcategorization frames and their frequency values were taken from the ICE-GB. In order to calculate the Selectional Preference Strength (SPS) of the relationship between a verb and its subcategorizations, we relied on the Kullback-Leibler Divergence model (Resnik, 1996). The SPS values of the verbs in the same cluster were compared with each other, which served to give the statistical values that indicate how much the SPS values overlap between the subcategorization frames of the verbs. Our final analysis shows that the degree of overlap, or the relationship between semantic similarity and the subcategorization frames of the verbs in English, is equally spread out from the 'very strongly related' to the 'very weakly related'. Some semantically similar verbs share a lot in terms of their subcategorization frames, and some others indicate an average degree of strength in the relationship, while the others, though still semantically similar, tend to share little in their subcategorization frames.
PDF

THE SPACE-TIME FRACTIONAL DIFFUSION EQUATION WITH CAPUTO DERIVATIVES

HUANG F.;LIU F.
- Journal of applied mathematics & informatics
- /
- v.19 no.1_2
- /
- pp.179-190
- /
- 2005
We deal with the Cauchy problem for the space-time fractional diffusion equation, which is obtained from standard diffusion equation by replacing the second-order space derivative with a Caputo (or Riemann-Liouville) derivative of order ${\beta}{\in}$ (0, 2] and the first-order time derivative with Caputo derivative of order ${\beta}{\in}$ (0, 1]. The fundamental solution (Green function) for the Cauchy problem is investigated with respect to its scaling and similarity properties, starting from its Fourier-Laplace representation. We derive explicit expression of the Green function. The Green function also can be interpreted as a spatial probability density function evolving in time. We further explain the similarity property by discussing the scale-invariance of the space-time fractional diffusion equation.

A Study on Influence of Stroke Element Properties to find Hangul Typeface Similarity (한글 글꼴 유사성 판단을 위한 획 요소 속성의 영향력 분석)

Park, Dong-Yeon;Jeon, Ja-Yeon;Lim, Seo-Young;Lim, Soon-Bum
- Journal of Korea Multimedia Society
- /
- v.23 no.12
- /
- pp.1552-1564
- /
- 2020
As various styles of fonts were used, there were problems such as output errors due to uninstalled fonts and difficulty in font recognition. To solve these problems, research on font recognition and recommendation were actively conducted. However, Hangul font research remains at the basic level. Therefore, in order to automate the comparison on Hangul font similarity in the future, we analyze the influence of each stroke element property. First, we select seven representative properties based on Hangul stroke shape elements. Second, we design a calculation model to compare similarity between fonts. Third, we analyze the effect of each stroke element through the cosine similarity between the user's evaluation and the results of the model. As a result, there was no significant difference in the individual effect of each representative property. Also, the more accurate similarity comparison was possible when many representative properties were used.
https://doi.org/10.9717/kmms.2020.23.12.1552 인용 PDF KSCI HTML

Machine-Part Grouping with Alternative Process Plans (대체공정이 있는 기계-부품 그룹 형성)

Lee, Jong-Sub;Kang, Maing-Kyu
- Journal of Korean Institute of Industrial Engineers
- /
- v.31 no.1
- /
- pp.20-26
- /
- 2005
This paper proposes the heuristic algorithm for the generalized GT problems to consider the restrictions which are given the number of machine cell and maximum number of machines in machine cell as well as minimum number of machines in machine cell. This approach is split into two phase. In the first phase, we use the similarity coefficient which proposes and calculates the similarity values about each pair of all machines and sort these values descending order. If we have a machine pair which has the largest similarity coefficient and adheres strictly to the constraint about birds of a different feather (BODF) in a machine cell, then we assign the machine to the machine cell. In the second phase, we assign parts into machine cell with the smallest number of exceptional elements. The results give a machine-part grouping. The proposed algorithm is compared to the Modified p-median model for machine-part grouping.
PDF KSCI

Search Result 109, Processing Time 0.021 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)