• 제목/요약/키워드: LPC cepstrum coefficients

검색결과 33건 처리시간 0.028초

시간 정보와 VQ를 이용한 DDD 지역명 인식에 관한 연구 (A Study on the Speech Recognition for DDD Area - Name Using Vector Quantization with Time Information)

  • 이성권;이강성;안태옥;조형제;변용규;김순협
    • 한국음향학회지
    • /
    • 제8권5호
    • /
    • pp.102-112
    • /
    • 1989
  • 본 논문은 불특정 화자의 DDD 지역명 인식 실험에 관한 것으로 VQ(Vector Quantization) 방식을 이용하여 실험하였고 인식대상 어휘로는 다이얼링 시스템의 응용을 목적으로 전국 146재의 DDD 지역명을 선정하였다. 특징 파라메타로는 12차 LPC Cepstrum 계수를 사용하여 코우드북을 작성하였으며, 중심점을 찾는 방법으로는 MINSUM 방법과 MINIMAX 방법을 사용하였고 코우드북 작성에는 Splitting rule 3가지를 사용하였다. 코우드북도 Single Section 코우드북과 시간정보를 포함하는 Multi Section 코우드북으로 나누어 작성하였고 Section을 Overlapping 하여가면서 코우드북을 작성하여 실험하였다. 실험 결과 minsum 방법이 minimax 보다 인식률이 좋은 것으로 나타났으며 화자 독립의 경우 약 $90\%$의 인식율을 얻을 수 있었다.

  • PDF

Speaker-Dependent Emotion Recognition For Audio Document Indexing

  • Hung LE Xuan;QUENOT Georges;CASTELLI Eric
    • 대한전자공학회:학술대회논문집
    • /
    • 대한전자공학회 2004년도 ICEIC The International Conference on Electronics Informations and Communications
    • /
    • pp.92-96
    • /
    • 2004
  • The researches of the emotions are currently great interest in speech processing as well as in human-machine interaction domain. In the recent years, more and more of researches relating to emotion synthesis or emotion recognition are developed for the different purposes. Each approach uses its methods and its various parameters measured on the speech signal. In this paper, we proposed using a short-time parameter: MFCC coefficients (Mel­Frequency Cepstrum Coefficients) and a simple but efficient classifying method: Vector Quantification (VQ) for speaker-dependent emotion recognition. Many other features: energy, pitch, zero crossing, phonetic rate, LPC... and their derivatives are also tested and combined with MFCC coefficients in order to find the best combination. The other models: GMM and HMM (Discrete and Continuous Hidden Markov Model) are studied as well in the hope that the usage of continuous distribution and the temporal behaviour of this set of features will improve the quality of emotion recognition. The maximum accuracy recognizing five different emotions exceeds $88\%$ by using only MFCC coefficients with VQ model. This is a simple but efficient approach, the result is even much better than those obtained with the same database in human evaluation by listening and judging without returning permission nor comparison between sentences [8]; And this result is positively comparable with the other approaches.

  • PDF

변형된 Dynamic Averaging 방법을 이용한 단독어인식 (Isolated Word Recognition using Modified Dynamic Averaging Method)

  • 정의봉;고영혁;이종악
    • 한국음향학회지
    • /
    • 제10권2호
    • /
    • pp.23-28
    • /
    • 1991
  • 본 논문을 특정화자에 대한 단독어 음성 인식에 대한 연구이다. 우리는 표준패턴으로서 변형된 dynamic linear averaging 방법을 이용한 DTW 음성 인식 시스템을 제안한다. 57개의 모든 도시명이 인식 대상 어휘로 선정되었고 12차 LPC cepstram 계수를 특징계수로 사용하였다. 이 논문은 표준패턴으로 변형된 dynamic linear averaging 방법을 이용하여 인식 실험을 한것 이외에도 같은 데이터 같은 조건상에서 causal 방법과 dynamic averaging방법, linear averaging방법, clustering 방법을 이용하여 실험하였다. 실험결과로 변형시킨 dynamic linear averaging 방법을 이용한 DTW 음성인식이 97.6%로 가장 좋은 인식율을 보였다.

  • PDF

DMS 모델을 이용한 한국어 음성 인식 (Korean Speech Recognition using Dynamic Multisection Model)

  • 안태옥;변용규;김순협
    • 대한전자공학회논문지
    • /
    • 제27권12호
    • /
    • pp.1933-1939
    • /
    • 1990
  • In this paper, we proposed an algorithm which used backtracking method to get time information, and it be modelled DMS (Dynamic Multisection) by feature vectors and time information whic are represented to similiar feature in word patterns spoken during continuous time domain, for Korean Speech recognition by independent speaker using DMS. Each state of model is represented time sequence, and have time information and feature vector. Typical feature vector is determined as the feature vector of each state to minimize the distance between word patterns. DDD Area names are selected as recognition wcabulary and 12th LPC cepstrum coefficients are used as the feature parameter. State of model is made 8 multisection and is used 0.2 as weight for time information. Through the experiment result, recognition rate by DMS model is 94.8%, and it is shown that this is better than recognition rate (89.3%) by MSVQ(Multisection Vector Quantization) method.

  • PDF

다중 관측열을 토대로한 HMM에 의한 음성 인식에 관한 연구 (A study on the speech recognition by HMM based on multi-observation sequence)

  • 정의봉
    • 전자공학회논문지S
    • /
    • 제34S권4호
    • /
    • pp.57-65
    • /
    • 1997
  • The purpose of this paper is to propose the HMM (hidden markov model) based on multi-observation sequence for the isolated word recognition. The proosed model generates the codebook of MSVQ by dividing each word into several sections followed by dividing training data into several sections. Then, we are to obtain the sequential value of multi-observation per each section by weighting the vectors of distance form lower values to higher ones. Thereafter, this the sequential with high probability value while in recognition. 146 DDD area names are selected as the vocabularies for the target recognition, and 10LPC cepstrum coefficients are used as the feature parameters. Besides the speech recognition experiments by way of the proposed model, for the comparison with it, the experiments by DP, MSVQ, and genral HMM are made with the same data under the same condition. The experiment results have shown that HMM based on multi-observation sequence proposed in this paper is proved superior to any other methods such as the ones using DP, MSVQ and general HMM models in recognition rate and time.

  • PDF

음성인식을 위한 알고리즘에 관한 연구 (A study on the algorithm for speech recognition)

  • 김선철;이정우;조규옥;박재균;오용택
    • 대한전기학회:학술대회논문집
    • /
    • 대한전기학회 2008년도 제39회 하계학술대회
    • /
    • pp.2255-2256
    • /
    • 2008
  • 음성인식 시스템을 설계함에 있어서는 대표적으로 사람의 성도 특성을 모방한 LPC(Linear Predict Cording)방식과 청각 특성을 고려한 MFCC(Mel-Frequency Cepstral Coefficients)방식이 있다. 본 논문에서는 MFCC를 통해 특징파라미터를 추출하고 해당 영역에서의 수행된 작업을 매틀랩 알고리즘을 이용하여 그래프로 시현하였다. MFCC 방식의 추출과정은 최초의 음성신호로부터 전처리과정을 통해 아날로그 신호를 디지털 신호로 변환하고, 잡음부분을 최소화하며, 음성 부분을 강조한다. 이 신호는 다시 Windowing을 통해 음성의 불연속을 제거해 주고, FFT를 통해 시간의 영역을 주파수의 영역으로 변환한다. 이 변환된 신호는 Filter Bank를 거쳐 다수의 복잡한 신호를 몇 개의 간단한 신호로 간소화 할 수 있으며, 마지막으로 Mel-cepstrum을 통해 최종적으로 특징 파라미터를 얻고자 하였다.

  • PDF

Hidden LMS 적응 필터링 알고리즘을 이용한 경쟁학습 화자검증 (Speaker Verification Using Hidden LMS Adaptive Filtering Algorithm and Competitive Learning Neural Network)

  • 조성원;김재민
    • 대한전기학회논문지:시스템및제어부문D
    • /
    • 제51권2호
    • /
    • pp.69-77
    • /
    • 2002
  • Speaker verification can be classified in two categories, text-dependent speaker verification and text-independent speaker verification. In this paper, we discuss text-dependent speaker verification. Text-dependent speaker verification system determines whether the sound characteristics of the speaker are equal to those of the specific person or not. In this paper we obtain the speaker data using a sound card in various noisy conditions, apply a new Hidden LMS (Least Mean Square) adaptive algorithm to it, and extract LPC (Linear Predictive Coding)-cepstrum coefficients as feature vectors. Finally, we use a competitive learning neural network for speaker verification. The proposed hidden LMS adaptive filter using a neural network reduces noise and enhances features in various noisy conditions. We construct a separate neural network for each speaker, which makes it unnecessary to train the whole network for a new added speaker and makes the system expansion easy. We experimentally prove that the proposed method improves the speaker verification performance.

신경 회로망을 이용한 EMG신호 기능 인식에 관한 연구 (A Study on EMG functional Recognition Using Neural Network)

  • 조정호;최윤호;왕문성;박상희
    • 대한의용생체공학회:학술대회논문집
    • /
    • 대한의용생체공학회 1990년도 춘계학술대회
    • /
    • pp.73-78
    • /
    • 1990
  • In this study, LPC cepstrum coefficients are used as feature vector extracted from AR model of EMG signal, and a reduced-connection network which has reduced connection between nodes is constructed to classify and recognize EMG functional classes. The proposed network reduces learning time and improves system stability. Therefore it is shown that the proposed network is appropriate in recognizing the function of EMG signal.

  • PDF

피지에 기초를 둔 HMM을 이용한 음성 인식 (Speech Recognition Using HMM Based on Fuzzy)

  • 안태옥;김순협
    • 전자공학회논문지B
    • /
    • 제28B권12호
    • /
    • pp.68-74
    • /
    • 1991
  • This paper proposes a HMM model based on fuzzy, as a method on the speech recognition of speaker-independent. In this recognition method, multi-observation sequences which give proper probabilities by fuzzy rule according to order of short distance from VQ codebook are obtained. Thereafter, the HMM model using this multi-observation sequences is generated, and in case of recognition, a word that has the most highest probability is selected as a recognized word. The vocabularies for recognition experiment are 146 DDD are names, and the feature parameter is 10S0thT LPC cepstrum coefficients. Besides the speech recognition experiments of proposed model, for comparison with it, we perform the experiments by DP, MSVQ and general HMM under same condition and data. Through the experiment results, it is proved that HMM model using fuzzy proposed in this paper is superior to DP method, MSVQ and general HMM model in recognition rate and computational time.

  • PDF

개선된 MSVQ 인식 시스템을 이용한 단독어 인식에 관한 연구 (A Study on Isolated Word Recognition using Improved Multisection Vector Quantization Recognition System)

  • 안태옥;김남중;송철;김순협
    • 한국통신학회논문지
    • /
    • 제16권2호
    • /
    • pp.196-205
    • /
    • 1991
  • 본 논문은 화자 독립의 단독이 언직에 관한 연구로 기존의 MSVQ(multisection vector quantization) 일질시스템을 개선한 새로운 MSVQ 시스템을 제안한다. 제안된 내용은 기존의 시스템과는 달리 인식시 시험패턴의 구간 수를 표준패턴의 구간 수보다 한 구간 더 늘리는 것이다. 이 방법에 의한 실험시 인식 대상으로는 146개의 DDD 지역망을 선택했으며, 특징 파라베타로는 12사 LPC 스트럼(cepstrum) 계수를 사용했고 코드북 지정석 중심점 구하는 방법으로 MINSUM과 MINIMAX기법을 사용하였다. 실험 결과에 의하면 DTW(dynamic time warping) 패턴 매칭 방법, VQ(vector quantization)에 의한 방법은 물론 기존의 MSVQ 방법보다 계산량이 감소함과 동시에 더 높은 인식율을 얻을 수 있었다. 수 있었다.

  • PDF