• 제목/요약/키워드: Statistical similarity

검색결과 311건 처리시간 0.033초

Statistical Fingerprint Recognition Matching Method with an Optimal Threshold and Confidence Interval

  • Hong, C.S.;Kim, C.H.
    • 응용통계연구
    • /
    • 제25권6호
    • /
    • pp.1027-1036
    • /
    • 2012
  • Among various biometrics recognition systems, statistical fingerprint recognition matching methods are considered using minutiae on fingerprints. We define similarity distance measures based on the coordinate and angle of the minutiae, and suggest a fingerprint recognition model following statistical distributions. We could obtain confidence intervals of similarity distance for the same and different persons, and optimal thresholds to minimize two kinds of error rates for distance distributions. It is found that the two confidence intervals of the same and different persons are not overlapped and that the optimal threshold locates between two confidence intervals. Hence an alternative statistical matching method can be suggested by using nonoverlapped confidence intervals and optimal thresholds obtained from the distributions of similarity distances.

Some new similarity based approaches in approximate reasoning and their applications to pattern recognition

  • Swapan Raha;Nikhil R. Pal;Ray, Kumar-Sankar
    • 한국지능시스템학회:학술대회논문집
    • /
    • 한국퍼지및지능시스템학회 1998년도 The Third Asian Fuzzy Systems Symposium
    • /
    • pp.719-724
    • /
    • 1998
  • This paper presents a systematic developement of a formal approach to inference in approximate reasoning. We introduce some measures of similarity and discuss their properties. Using the concept of similarity index we formulate two methods for inferring from vague knowledge. In order to illustrate the effectiveness of the proposed technique we use it to develop a vowel recognition system.

  • PDF

영어 동사의 의미적 유사도와 논항 선택 사이의 연관성 : ICE-GB와 WordNet을 이용한 통계적 검증 (The Strength of the Relationship between Semantic Similarity and the Subcategorization Frames of the English Verbs: a Stochastic Test based on the ICE-GB and WordNet)

  • 송상헌;최재웅
    • 한국언어정보학회지:언어와정보
    • /
    • 제14권1호
    • /
    • pp.113-144
    • /
    • 2010
  • The primary goal of this paper is to find a feasible way to answer the question: Does the similarity in meaning between verbs relate to the similarity in their subcategorization? In order to answer this question in a rather concrete way on the basis of a large set of English verbs, this study made use of various language resources, tools, and statistical methodologies. We first compiled a list of 678 verbs that were selected from the most and second most frequent word lists from the Colins Cobuild English Dictionary, which also appeared in WordNet 3.0. We calculated similarity measures between all the pairs of the words based on the 'jcn' algorithm (Jiang and Conrath, 1997) implemented in the WordNet::Similarity module (Pedersen, Patwardhan, and Michelizzi, 2004). The clustering process followed, first building similarity matrices out of the similarity measure values, next drawing dendrograms on the basis of the matricies, then finally getting 177 meaningful clusters (covering 437 verbs) that passed a certain level set by z-score. The subcategorization frames and their frequency values were taken from the ICE-GB. In order to calculate the Selectional Preference Strength (SPS) of the relationship between a verb and its subcategorizations, we relied on the Kullback-Leibler Divergence model (Resnik, 1996). The SPS values of the verbs in the same cluster were compared with each other, which served to give the statistical values that indicate how much the SPS values overlap between the subcategorization frames of the verbs. Our final analysis shows that the degree of overlap, or the relationship between semantic similarity and the subcategorization frames of the verbs in English, is equally spread out from the 'very strongly related' to the 'very weakly related'. Some semantically similar verbs share a lot in terms of their subcategorization frames, and some others indicate an average degree of strength in the relationship, while the others, though still semantically similar, tend to share little in their subcategorization frames.

  • PDF

문장구조 유사도와 단어 유사도를 이용한 클러스터링 기반의 통계기계번역 (Clustering-based Statistical Machine Translation Using Syntactic Structure and Word Similarity)

  • 김한경;나휘동;이금희;이종혁
    • 한국정보과학회논문지:소프트웨어및응용
    • /
    • 제37권4호
    • /
    • pp.297-304
    • /
    • 2010
  • 통계기계번역에서 번역성능의 향상을 위해서 문장의 유형이나 장르에 따라 클러스터링을 수행하여 도메인에 특화된 번역을 시도하는 방법이 있다. 그러나 기존의 연구 중 문장의 유형 정보와 장르에 따른 정보를 동시에 사용한 경우는 없었다. 본 논문에서는 각 문장의 문법적 구조 유사도에 따른 유형별분류 기법과, 단어 유사도 정보를 사용한 장르 구분법을 적용하여 기존의 두 기법을 통합하였다. 이렇게 분류된 말뭉치에서 추출한 도메인 특화 모델과 전체 말뭉치에서 추출된 모델에서 보간법(interpolation)을 사용하여 통계기계번역의 성능을 향상하였다. 문장구조 유사도와 단어 유사도의 계산 방법으로는 각각 커널과 코사인 유사도를 적용하였으며, 두 유사도를 적용하여 말뭉치를 분류하는 과정에서는 K-Means 알고리즘과 유사한 기계학습 기법을 사용하였다. 이를 일본어-영어의 특허문서에서 실험한 결과 최선의 경우 약 2.5%의 상대적인 성능 향상을 얻었다.

속용성 정제간의 용출유사성에 대한 통계학적 고찰 (Statistical Consideration on the Similarity in Dissolution Profile of Two Fast Releasing Tablets)

  • 조정환;이세희;김희선;오승열
    • Journal of Pharmaceutical Investigation
    • /
    • 제30권2호
    • /
    • pp.85-91
    • /
    • 2000
  • We have studied the dissolution kinetics of two fast releasing tablets in four media, and the similarity of dissolution profiles was compared using 3 methods. Two of the methods were introduced from statistical algorithm of distance methods, which are maximum distance and Mahalanobis distance. The dissolution kinetics were also analysed using FDA method for similarity evaluation, and the results were compared with those obtained using the distance methods.

  • PDF

유사측도를 이용한 신뢰성 있는 데이터의 추출 (Reliable Data Selection using Similarity Measure)

  • 류수록;이상혁
    • 한국지능시스템학회논문지
    • /
    • 제18권2호
    • /
    • pp.200-205
    • /
    • 2008
  • 데이터 분석을 위하여 데이터의 불확실성에 대한 측도로서 퍼지 집합에 대한 엔트로피를 소개하였고, 또한 데이터간의 유사도를 나타내는 유사측도를 구성하였다. 퍼지 소속 함수간의 유사측도는 거리측도를 이용하여 구성하였고, 제안한 유사측도를 증명을 통하여 확인하였다. 제안한 유사측도의 유용성을 확인하기 위하여 신뢰성 있는 데이터추출 예제에 적용하였다. 적용결과를 퍼지 엔트로피와 통계적 지식을 통하여 얻어진 이전의 결과와 비교하였다.

다중선택 시험에서 부정행위자 발견을 위한 새로운 통계적 측도 (A New Statistical Index for Detecting Cheaters on Multiple Choice Tests)

  • 한은수;임요한;이경은
    • 응용통계연구
    • /
    • 제26권1호
    • /
    • pp.81-92
    • /
    • 2013
  • 학문적 진실성(academic integrity)을 위반하는 잠재적 부적행위를 판단할 때, 잘못된 결정을 피하기 위해서는 확고한 근거를 마련하는 것이 중요하다. 교육학 연구자들은 부정행위를 발견 혹은 확신 할 수 있는 많은 통계적인 방법들을 발전시켰다. 그러나, 대부분의 방법들은 단순히 상관계수를 기초로한 방법들이어서 종종 응답자들의 패턴을 설명하기가 어렵다. 이 논문에서는, 이런 어려움을 해결하기 해결하기 위하여 표준화된 부호 엔트로피 유사성 점수(Standardized Signed Entropy Similarity Score)라는 새로운 통계적인 측도를 제안한다. 또한, 이 제안한 방법을 실제 시험 자료를 이용 부정행위자를 발견하는데 적용하였고, 다른 기존의 방법들과 비교하였다.

Recovery Levels of Clustering Algorithms Using Different Similarity Measures for Functional Data

  • Chae, Seong San;Kim, Chansoo;Warde, William D.
    • Communications for Statistical Applications and Methods
    • /
    • 제11권2호
    • /
    • pp.369-380
    • /
    • 2004
  • Clustering algorithms with different similarity measures are commonly used to find an optimal clustering or close to original clustering. The recovery level of using Euclidean distance and distances transformed from correlation coefficients is evaluated and compared using Rand's (1971) C statistic. The C values present how the resultant clustering is close to the original clustering. In simulation study, the recovery level is improved by applying the correlation coefficients between objects. Using the data set from Spellman et al. (1998), the recovery levels with different similarity measures are also presented. In general, the recovery level of true clusters was increased by using the correlation coefficients.

동물 및 임상 시험의 시계열 프로파일 데이터 비교를 위한 유사성 지수 개발 (Development of a New Similarity Index to Compare Time-series Profile Data for Animal and Human Experiments)

  • 이예경;이현정;장현애;신상문
    • 품질경영학회지
    • /
    • 제49권2호
    • /
    • pp.145-159
    • /
    • 2021
  • Purpose: A statistical similarity evaluation to compare pharmacokinetics(PK) profile data between nonclinical and clinical experiments has become a significant issue on many drug development processes. This study proposes a new similarity index by considering important parameters, such as the area under the curve(AUC) and the time-series profile of various PK data. Methods: In this study, a new profile similarity index(PSI) by using the concept of a process capability index(Cp) is proposed in order to investigate the most similar animal PK profile compared to the target(i.e., Human PK profile). The proposed PSI can be calculated geometric and arithmetic means of all short term similarity indices at all time points on time-series both animal and human PK data. Designed simulation approaches are demonstrated for a verification purpose. Results: Two different simulation studies are conducted by considering three variances(i.e., small, medium, and large variances) as well as three different characteristic types(smaller the better, larger the better, nominal the best). By using the proposed PSI, the most similar animal PK profile compare to the target human PK profile can be obtained in the simulation studies. In addition, a case study represents differentiated results compare to existing simple statistical analysis methods(i.e., root mean squared error and quality loss). Conclusion: The proposed PSI can effectively estimate the level of similarity between animal, human PK profiles. By using these PSI results, we can reduce the number of animal experiments because we only focus on the significant animal representing a high PSI value.

Semi-supervised learning using similarity and dissimilarity

  • Seok, Kyung-Ha
    • Journal of the Korean Data and Information Science Society
    • /
    • 제22권1호
    • /
    • pp.99-105
    • /
    • 2011
  • We propose a semi-supervised learning algorithm based on a form of regularization that incorporates similarity and dissimilarity penalty terms. Our approach uses a graph-based encoding of similarity and dissimilarity. We also present a model-selection method which employs cross-validation techniques to choose hyperparameters which affect the performance of the proposed method. Simulations using two types of dat sets demonstrate that the proposed method is promising.