• Title/Summary/Keyword: Second-order Similarity

Search Result 109, Processing Time 0.021 seconds

Improving the Performance of Document Clustering with Distributional Similarities (분포유사도를 이용한 문헌클러스터링의 성능향상에 대한 연구)

  • Lee, Jae-Yun
    • Journal of the Korean Society for information Management
    • /
    • v.24 no.4
    • /
    • pp.267-283
    • /
    • 2007
  • In this study, measures of distributional similarity such as KL-divergence are applied to cluster documents instead of traditional cosine measure, which is the most prevalent vector similarity measure for document clustering. Three variations of KL-divergence are investigated; Jansen-Shannon divergence, symmetric skew divergence, and minimum skew divergence. In order to verify the contribution of distributional similarities to document clustering, two experiments are designed and carried out on three test collections. In the first experiment the clustering performances of the three divergence measures are compared to that of cosine measure. The result showed that minimum skew divergence outperformed the other divergence measures as well as cosine measure. In the second experiment second-order distributional similarities are calculated with Pearson correlation coefficient from the first-order similarity matrixes. From the result of the second experiment, secondorder distributional similarities were found to improve the overall performance of document clustering. These results suggest that minimum skew divergence must be selected as document vector similarity measure when considering both time and accuracy, and second-order similarity is a good choice for considering clustering accuracy only.

A Study on Keyword Extraction From a Single Document Using Term Clustering (용어 클러스터링을 이용한 단일문서 키워드 추출에 관한 연구)

  • Han, Seung-Hee
    • Journal of the Korean Society for Library and Information Science
    • /
    • v.44 no.3
    • /
    • pp.155-173
    • /
    • 2010
  • In this study, a new keyword extraction algorithm is applied to a single document with term clustering. A single document is divided by multiple passages, and two ways of calculating similarities between two terms are investigated; the first-order similarity and the second-order distributional similarity. In this experiment, the best cluster performance is achieved with a 50-term passage from the second-order distributional similarity. From the results of first experiment, the second-order distribution similarity was also applied to various keyword extraction methods using statistic information of terms. In the second experiment, pf(paragraph frequency) and $tf{\times}ipf$(term frequency by inverse paragraph frequency) were found to improve the overall performance of keyword extraction. Therefore, it showed that the algorithm fulfills the necessary conditions which good keywords should have.

Image Data Classification using a Similarity Function based on Second Order Tensor (2차 텐서 기반 유사도 함수를 이용한 영상 데이터 분류)

  • Yoon, Dong-Woo;Lee, Kwan-Yong;Park, Hye-Young
    • Journal of KIISE:Software and Applications
    • /
    • v.36 no.8
    • /
    • pp.664-672
    • /
    • 2009
  • Recently, studies on utilizing tensor expression on image data analysis and processing have been attracting much interest. The purpose of this study is to develop an efficient system for classifying image patterns by using second order tensor expression. To achieve the goal, we propose a data generation model expressed by class factors and environment factors with second order tensor representation. Based on the data generation model, we define a function for measuring similarities between two images. The similarity function is obtained by estimating the probability density of environment factors using a matrix normal distribution. Through computational experiments on a number of benchmark data sets, we confirm that we can make improvement in classification rates by using second order tensor, and that the proposed similarity function is more appropriate for image data compared to conventional similarity measures.

SYMMETRY REDUCTIONS, VARIABLE TRANSFORMATIONS AND EXACT SOLUTIONS TO THE SECOND-ORDER PDES

  • Liu, Hanze;Liu, Lei
    • Journal of applied mathematics & informatics
    • /
    • v.29 no.3_4
    • /
    • pp.563-572
    • /
    • 2011
  • In this paper, the Lie symmetry analysis is performed on the three mixed second-order PDEs, which arise in fluid dynamics, nonlinear wave theory and plasma physics, etc. The symmetries and similarity reductions of the equations are obtained, and the exact solutions to the equations are investigated by the dynamical system and power series methods. Then, the exact solutions to the general types of PDEs are considered through a variable transformation. At last, the symmetry and integration method is employed for reducing the nonlinear ODEs.

The Relationships Between Children's Perceptions Toward Grandparents and Their Intimate Behavior

  • Jung, Min-Suk;Ko, Eun-Kyo;Rho, Joseph Y.;Lee, Seung-Hyun
    • International Journal of Contents
    • /
    • v.5 no.4
    • /
    • pp.19-29
    • /
    • 2009
  • This study is focused on the causal relationship between children's intimate behavior and the level of perception towards their grandparents. Their perceptions are related to factors such as proximity, similarity, superiority, favorableness, and self-disclosure. We clarified the relation between intimate behavior and perception using effect factors of children's behavior regarding their grandparents so that this study could be used as an elementary material in developing a solution to improve grandparent-grandchild relationship where the grandparent actively encourages grandchildren's intimate behavior. Regression analysis was used as a hypothesis testing method. The results indicated the following three points. First, perception factors affect active intimate behavior in the order of favorableness, superiority, self-disclosure, and similarity. Second, perception factors affect intimate behavior will in the order of favorableness, superiority, and self-disclosure. Lastly, it was shown that a child's active intimate behavior has an influence on their intimate behavior will.

Heat and mass transfer of a second grade magnetohydrodynamic fluid over a convectively heated stretching sheet

  • Das, Kalidas;Sharma, Ram Prakash;Sarkar, Amit
    • Journal of Computational Design and Engineering
    • /
    • v.3 no.4
    • /
    • pp.330-336
    • /
    • 2016
  • The present work is concerned with heat and mass transfer of an electrically conducting second grade MHD fluid past a semi-infinite stretching sheet with convective surface heat flux. The analysis accounts for thermophoresis and thermal radiation. A similarity transformations is used to reduce the governing equations into a dimensionless form. The local similarity equations are derived and solved using Nachtsheim-Swigert shooting iteration technique together with Runge-Kutta sixth order integration scheme. Results for various flow characteristics are presented through graphs and tables delineating the effect of various parameters characterizing the flow. Our analysis explores that the rate of heat transfer enhances with increasing the values of the surface convection parameter. Also the fluid velocity and temperature in the boundary layer region rise significantly for increasing the values of thermal radiation parameter.

The Strength of the Relationship between Semantic Similarity and the Subcategorization Frames of the English Verbs: a Stochastic Test based on the ICE-GB and WordNet (영어 동사의 의미적 유사도와 논항 선택 사이의 연관성 : ICE-GB와 WordNet을 이용한 통계적 검증)

  • Song, Sang-Houn;Choe, Jae-Woong
    • Language and Information
    • /
    • v.14 no.1
    • /
    • pp.113-144
    • /
    • 2010
  • The primary goal of this paper is to find a feasible way to answer the question: Does the similarity in meaning between verbs relate to the similarity in their subcategorization? In order to answer this question in a rather concrete way on the basis of a large set of English verbs, this study made use of various language resources, tools, and statistical methodologies. We first compiled a list of 678 verbs that were selected from the most and second most frequent word lists from the Colins Cobuild English Dictionary, which also appeared in WordNet 3.0. We calculated similarity measures between all the pairs of the words based on the 'jcn' algorithm (Jiang and Conrath, 1997) implemented in the WordNet::Similarity module (Pedersen, Patwardhan, and Michelizzi, 2004). The clustering process followed, first building similarity matrices out of the similarity measure values, next drawing dendrograms on the basis of the matricies, then finally getting 177 meaningful clusters (covering 437 verbs) that passed a certain level set by z-score. The subcategorization frames and their frequency values were taken from the ICE-GB. In order to calculate the Selectional Preference Strength (SPS) of the relationship between a verb and its subcategorizations, we relied on the Kullback-Leibler Divergence model (Resnik, 1996). The SPS values of the verbs in the same cluster were compared with each other, which served to give the statistical values that indicate how much the SPS values overlap between the subcategorization frames of the verbs. Our final analysis shows that the degree of overlap, or the relationship between semantic similarity and the subcategorization frames of the verbs in English, is equally spread out from the 'very strongly related' to the 'very weakly related'. Some semantically similar verbs share a lot in terms of their subcategorization frames, and some others indicate an average degree of strength in the relationship, while the others, though still semantically similar, tend to share little in their subcategorization frames.

  • PDF

THE SPACE-TIME FRACTIONAL DIFFUSION EQUATION WITH CAPUTO DERIVATIVES

  • HUANG F.;LIU F.
    • Journal of applied mathematics & informatics
    • /
    • v.19 no.1_2
    • /
    • pp.179-190
    • /
    • 2005
  • We deal with the Cauchy problem for the space-time fractional diffusion equation, which is obtained from standard diffusion equation by replacing the second-order space derivative with a Caputo (or Riemann-Liouville) derivative of order ${\beta}{\in}$ (0, 2] and the first-order time derivative with Caputo derivative of order ${\beta}{\in}$ (0, 1]. The fundamental solution (Green function) for the Cauchy problem is investigated with respect to its scaling and similarity properties, starting from its Fourier-Laplace representation. We derive explicit expression of the Green function. The Green function also can be interpreted as a spatial probability density function evolving in time. We further explain the similarity property by discussing the scale-invariance of the space-time fractional diffusion equation.

A Study on Influence of Stroke Element Properties to find Hangul Typeface Similarity (한글 글꼴 유사성 판단을 위한 획 요소 속성의 영향력 분석)

  • Park, Dong-Yeon;Jeon, Ja-Yeon;Lim, Seo-Young;Lim, Soon-Bum
    • Journal of Korea Multimedia Society
    • /
    • v.23 no.12
    • /
    • pp.1552-1564
    • /
    • 2020
  • As various styles of fonts were used, there were problems such as output errors due to uninstalled fonts and difficulty in font recognition. To solve these problems, research on font recognition and recommendation were actively conducted. However, Hangul font research remains at the basic level. Therefore, in order to automate the comparison on Hangul font similarity in the future, we analyze the influence of each stroke element property. First, we select seven representative properties based on Hangul stroke shape elements. Second, we design a calculation model to compare similarity between fonts. Third, we analyze the effect of each stroke element through the cosine similarity between the user's evaluation and the results of the model. As a result, there was no significant difference in the individual effect of each representative property. Also, the more accurate similarity comparison was possible when many representative properties were used.

Machine-Part Grouping with Alternative Process Plans (대체공정이 있는 기계-부품 그룹 형성)

  • Lee, Jong-Sub;Kang, Maing-Kyu
    • Journal of Korean Institute of Industrial Engineers
    • /
    • v.31 no.1
    • /
    • pp.20-26
    • /
    • 2005
  • This paper proposes the heuristic algorithm for the generalized GT problems to consider the restrictions which are given the number of machine cell and maximum number of machines in machine cell as well as minimum number of machines in machine cell. This approach is split into two phase. In the first phase, we use the similarity coefficient which proposes and calculates the similarity values about each pair of all machines and sort these values descending order. If we have a machine pair which has the largest similarity coefficient and adheres strictly to the constraint about birds of a different feather (BODF) in a machine cell, then we assign the machine to the machine cell. In the second phase, we assign parts into machine cell with the smallest number of exceptional elements. The results give a machine-part grouping. The proposed algorithm is compared to the Modified p-median model for machine-part grouping.